A scaffolding method based on the statistical characteristics of the double-ended read insert size
A technique of statistical features and readings, applied in the field of bioinformatics, which can solve the problem of removing and ignoring the characteristics of long nodes and short nodes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Publication Date
- 2018-10-16
Smart Images

Figure 1 
Figure 2 
Figure 3
Abstract
Description
technical field
[0001] The invention relates to the field of bioinformatics, in particular to a scaffolding method based on the statistical characteristics of double-end read insert size. Background technique
[0002] Genome generally refers to all coding and non-coding deoxyribonucleic acid (DNA) sequences, which are composed of four bases: adenine (A), thymine (T), cytosine (C) and guanine (G) The sequence, that is, the genome sequence is a string, which only contains four characters A, T, G, and C. Another character N is also included in the actual genome sequence, representing that the base at this position cannot be determined. Through genome sequencing, short-segment base sequences (reads or reads) on a large number of genome sequences can be obtained. A collection of reads obtained from genome sequencing, generally with relatively short read lengths. Sequence assembly methods use these short reads to restore the complete original genome sequence. With the rapid de...
Examples
Embodiment Construction
[0093] Such as figure 1 Shown, the concrete realization process of the present invention is as follows:
[0094] 1. Pretreatment
[0095] In this method, it is assumed that in paired-end reads, the left and right reads are not in the same orientation, and that the right read is to the right of the left read relative to the 5' to 3' orientation of the left read.
[0096] The contig file and the comparison result file are used as input data. In the alignment result file, due to repeated regions and sequencing errors, a read often has multiple alignment position information. For each reading, this method only retains the alignment position information with the highest alignment score, and removes all the remaining non-optimal alignment position information. According to the alignment position information of all reads, the read coverage of each position on each contig can be obtained, that is, how many aligned reads cover a certain base position. At the same time, the average ...