I am trying to split one big file into individual entries. Each entry ends with the character “//”. So when I try to use
#!/usr/bin/python
import sys,os
uniprotFile=open("UNIPROT-data.txt") #read original alignment file
uniprotFileContent=uniprotFile.read()
uniprotFileList=uniprotFileContent.split("//")
for items in uniprotFileList:
seqInfoFile=open('%s.dat'%items[5:14],'w')
seqInfoFile.write(str(items))
But I realised that there is another string with “//“(http://www.uniprot.org/terms) hence it splits there as well and eventually I don’t get the result I want. I tried using regex but was not abler to figure it out.
Use a regex that only splits on // if it's not preceded by :
import re
myre = re.compile("(?<!:)//")
uniprotFileList = myre.split(uniprotFileContent)
I am using the code with modified split pattern and it works fine for me:
#!/usr/bin/python
import sys,os
uniprotFile = open("UNIPROT-data.txt")
uniprotFileContent = uniprotFile.read()
uniprotFileList = uniprotFileContent.split("//\n")
for items in uniprotFileList:
seqInfoFile = open('%s.dat' % items[5:17], 'w')
seqInfoFile.write(str(items))
If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!
Donate Us With