The application discloses a double-end sequencing sample mark weight method, comprising the following steps: aligning the double-end read length corresponding to the
DNA fragment in the double-end sequencing sample to the position corresponding to the whole
genome reference sequence respectively, the double-end read length of the fragment being at adjacent positions, obtaining the alignment information of the double-end read length, adding an
auxiliary field in the alignment information of each read length, and the
auxiliary field containing the starting position of the other end read length corresponding to the read length; sorting the read length according to the reference sequence site after the alignment is completed, and saving the ordered read length in a SAM file; sequentially reading the read length from the ordered SAM file, taking the read length at the same alignment position as a group of candidate read length repeats, screening the fragment
repeat sequence from the candidate read length
repeat sequence according to the
auxiliary field; deleting the
repeat sequence or marking the repeat sequence, and writing into a new SAM file, so that the application reduces the data reading
throughput in the double-end sequencing sample mark weight process, reduces the
running time, and improves the running efficiency.