Primer matching detection method

By constructing a molecular tag error correction method for legal lists and fault-tolerant lists, combining read-write separation strategies and primer dictionary, the problem of inefficiency in the pre-processing of sequencing libraries in the existing technology is solved, and efficient and unified sequencing library pre-processing is achieved.

CN120108503APending Publication Date: 2025-06-06HANGZHOU REPUGENE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510174440.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, there is a lack of problems such as quality control, primer matching detection and molecular label error correction in the preprocessing process of sequencing library at one time, resulting in complex data processing processes and inefficient efficiency.

Method used

Provide an efficient molecular tag error correction method, by constructing legal lists and fault-tolerant lists, correcting UMI sequences in sequencing data, and combining read-write separation strategies and primer dictionary for quality control and primer matching detection.

Benefits of technology

It significantly improves the speed and accuracy of sequencing library preprocessing, reduces the consumption of computing resources, and achieves the goal of solving quality control, primer matching detection and molecular label error correction at one time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108503A_ABST
    Figure CN120108503A_ABST
Patent Text Reader

Abstract

The invention provides a molecular tag error correction method and an efficient multiple amplicon capture sequencing library pretreatment method. In order to solve the problems that in the prior art, the molecular tag error correction processing speed is low, and a preprocessing method capable of achieving sequencing library quality control and / or primer matching detection and molecular tag error correction at a time is lacked, the invention provides a molecular tag error correction method based on base sequence digital storage. On the basis of the molecular tag error correction method, an efficient and uniform sequencing library preprocessing method is provided, and the method can significantly improve the preprocessing speed and reduce the consumption of computing resources on the basis of ensuring the accuracy, so that the increasing gene detection requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention is a divisional application of Chinese patent application 202411580041.2, with a filing date of November 7, 2024. The name of the invention of the original application is “Molecular tag error correction method, efficient multiple amplicon capture sequencing library pretreatment method”. Technical Field

[0002] The invention belongs to the field of bioinformatics, and in particular relates to a molecular tag error correction method and a high-efficiency multiple amplicon capture sequencing library preprocessing method based on the molecular tag error correction method. Background Art

[0003] With the widespread use of genetic testing technology, especially the in-depth application of second-generation sequencing technology that combines multiple primer probes and molecular tags (Unique Molecular Identifier, UMI), the accuracy and efficiency of sequencing library preprocessing have become crucial. This process involves multiple key steps, mainly including quality control of sequencing data, primer matching detection, and extraction and error correction of molecular tags.

[0004] However, for the preprocessing of multiple primer-probe second-generation sequencing libraries with molecular tags, there is currently no unified software that can solve quality control, and / or primer matching detection, and molecular tag extraction error correction at one time. Existing solutions usually rely on the tandem use of multiple third-party software, but these software each focus on solving specific problems. For example, quality control may use one software, while primer matching check and molecular tag error correction may use other software. This type of preprocessing method not only complicates the data processing process, but also reduces the overall efficiency, because the data between different software are incompatible, and frequent data input and output operations increase the computational burden.

[0005] Furthermore, in terms of primer matching detection, existing technologies often use a one-to-one comparison method, which is inefficient. For large-scale libraries, especially when the number of primers reaches thousands, traditional methods require a large number of matching operations, resulting in a significant increase in processing time. In addition, existing software has relatively limited support for primer length and mismatches, and cannot meet diverse experimental needs.

[0006] Furthermore, in terms of molecular tag error correction, traditional methods mainly rely on alignment strategies, which are slow to process, and few software can effectively support the matching of short sequences. The complexity of this process and the high requirements for computing resources further increase the difficulty of library preprocessing. Especially when dealing with molecular tags of fixed length and large number, it is difficult for existing technologies to achieve fast and accurate error correction.

[0007] Therefore, in response to the above technical problems, it is particularly important to provide an efficient and unified sequencing library preprocessing method. Summary of the invention

[0008] In order to solve the problems in the prior art of 1) slow molecular tag error correction processing speed and 2) lack of a preprocessing method that can solve sequencing library quality control, and / or primer matching detection, and molecular tag error correction at one time, the present invention provides an efficient molecular tag error correction method, and based on the molecular tag error correction method, provides an efficient and unified sequencing library preprocessing method, which can significantly improve the preprocessing speed and reduce the consumption of computing resources on the basis of ensuring accuracy, thereby meeting the growing demand for genetic testing.

[0009] The first aspect of the present invention provides a molecular tag error correction method, comprising: constructing a legal list and a fault-tolerant list according to a UMI sequence whitelist; correcting UMI sequences extracted from sequencing data according to the constructed legal list and fault-tolerant list; wherein the method for constructing the legal list comprises: digitally processing the legal UMI sequences existing in the UMI sequence whitelist, and adding legal indexes to the legal UMI sequences according to the digital processing information; storing the legal UMI sequences, the digital processing information corresponding to the legal UMI sequences, and the indexes corresponding to the legal UMI sequences in the legal list; the method for constructing the fault-tolerant list comprises: for each legal UMI sequence, performing full enumeration mutation according to a preset number of allowed error bases to obtain illegal mutation UMI sequences and legal mutation UMI sequences; digitally processing the legal UMI sequences, the legal mutation UMI sequences, and the illegal mutation UMI sequences, adding fault-tolerant indexes and corresponding legal indexes of the legal UMI sequences to the legal UMI sequences and the legal mutation UMI sequences according to the digital processing information, and setting illegal values ​​for the legal indexes of the illegal mutation UMI sequences; storing the legal UMI sequences, the legal mutation UMI sequences, and the legal mutation UMI sequences in the legal list; The digital processing information corresponding to the illegal UMI sequence, the fault tolerance index corresponding to the legal UMI sequence, the legal index corresponding to the legal UMI sequence, the legal mutation UMI sequence, the digital processing information corresponding to the legal mutation UMI sequence, the fault tolerance index corresponding to the legal mutation UMI sequence, the legal index corresponding to the legal mutation UMI sequence, the illegal mutation UMI sequence, the digital processing information corresponding to the illegal mutation UMI sequence, and the legal index corresponding to the illegal mutation UMI sequence are stored in the fault tolerance list; the method for correcting the extracted UMI sequence includes: digitally processing the extracted UMI sequence, adding an access index to the extracted UMI sequence according to the digital processing information; accessing the fault tolerance list according to the access index, if the return value is an illegal value, the extracted UMI sequence cannot match the UMI sequence whitelist, and is used as an erroneous UMI sequence; if the return value is not an illegal value, matching the access index to the corresponding fault tolerance index and legal index in the fault tolerance list in sequence, and then accessing the legal list according to the matched legal index and matching it to the corresponding legal UMI sequence as the correct UMI sequence of the extracted UMI sequence, and correcting the error through the UMI sequence.

[0010] In some embodiments, the digital processing method includes: converting the UMI sequence into a binary value for storage; wherein the bases A / a, C / c, G / g, and T / t of the UMI sequence are respectively mapped to binary values ​​with a storage space of 2 bits; preferably, the bases A / a, C / c, G / g, and T / t of the UMI sequence are respectively mapped to binary values ​​00, 01, 10, and 11.

[0011] In some implementations, the digitally processed information is any one of character information, RGB information, a binary value, an octal value, a hexadecimal value, and a decimal value corresponding to a binary value of a UMI sequence.

[0012] In some implementations, the method for correcting errors in the extracted UMI sequence further includes: Return the erroneous UMI sequence to the corresponding read segment in the quality control sequencing data, and obtain an offset UMI sequence of the same length as the erroneous UMI sequence according to the position information of the erroneous UMI sequence on the corresponding read segment and the preset offset distance; The offset UMI sequence is digitized, and an access index is added to the offset UMI sequence according to the digitized processing information; the fault tolerance list is accessed according to the access index. If the return value is an illegal value, the offset UMI sequence cannot match the UMI sequence whitelist and the offset UMI sequence is discarded; if the return value is not an illegal value, the access index is matched to the corresponding fault tolerance index and legal index in the fault tolerance list in sequence, and then the legal list is accessed according to the matched legal index and matched to the corresponding legal UMI sequence as the correct UMI sequence of the offset UMI sequence.

[0013] In some embodiments, the preset offset distance is 1 to 3 bases.

[0014] The second aspect of the present invention provides an efficient multiplex amplicon capture sequencing library pretreatment method, comprising: Acquire sequencing data with primer sequences and UMI sequences, perform quality control on the acquired sequencing data based on a read-write separation strategy, obtain quality control sequencing data, and extract UMI sequences from the quality control sequencing data; The molecular tag error correction method is used to correct the UMI sequence of the quality control sequencing data; The sequencing data after quality control and UMI sequence error correction is used as the preprocessing result, and the preprocessing result is output and reported.

[0015] In some embodiments, the efficient multiplex amplicon capture sequencing library preprocessing method comprises: Acquire sequencing data with primer sequences and UMI sequences, perform quality control on the acquired sequencing data based on a read-write separation strategy, obtain quality control sequencing data, and extract UMI sequences and primer sequences from the quality control sequencing data; A primer dictionary is constructed based on the primer whitelist and is graded according to the primer sequence length. The primer matching of the primer sequences in the quality control sequencing data is detected based on the primer dictionary. The molecular tag error correction method is used to correct the UMI sequence of the quality control sequencing data detected by primer matching; The sequencing data after quality control, primer matching detection, and UMI sequence error correction is used as the preprocessing result, and the preprocessing result is output and reported.

[0016] In some embodiments, the method of constructing a primer dictionary graded by primer sequence length according to the primer whitelist and performing primer matching detection on the extracted primer sequence according to the primer dictionary includes: The legal primer sequences existing in the primer whitelist and the legal mismatched primer sequences that have allowed base mismatches with the legal primers are classified according to the primer sequence length, and the legal primer sequences / legal mismatched primer sequences with the same primer sequence length are stored in the same hash table, and a primer dictionary with primer sequence length-legal primer sequence / legal mismatched primer sequence as key-value pairs is constructed; The extracted primer sequences are graded according to different primer sequence lengths in the primer dictionary to obtain primer subsequences of different primer sequence lengths, and are respectively consulted in a hash table of the corresponding primer sequence length; if there is a legal primer sequence / legal mismatched primer sequence that is the same as the primer subsequence in the hash table, the extracted primer sequence passes the primer matching test.

[0017] In some embodiments, the high-efficiency multiplex amplicon capture sequencing library preprocessing method further comprises: A space character and an extracted UMI sequence / correct UMI sequence are sequentially added to the original read name of the corresponding read segment of the extracted UMI sequence / correct UMI sequence as the output read name for outputting the preprocessing result.

[0018] In some embodiments, the method of outputting and reporting preprocessing results includes: Output preprocessing results in FASTQ and JSON formats, and report preprocessing results in the form of interactive dynamic web pages; The preprocessing result includes qualified data and unqualified data; the sequencing data that passes the quality control, the primer matching detection, and the UMI sequence error correction is regarded as qualified data; the sequencing data that fails the quality control and / or fails the primer matching detection and / or fails the UMI sequence error correction is regarded as unqualified data; The report content includes at least one of a quality control result, a primer matching detection result, and a UMI sequence error correction result.

[0019] Beneficial effects of the present invention: The efficient multiplex amplicon capture sequencing library preprocessing method provided by the present invention optimizes the full data calculation process of the multiplex primer probe second-generation sequencing library with molecular tags, and uses C / C++ high-performance implementation to ensure accuracy while bringing significant speed and performance improvements, and produces industrialized standard result files and user-friendly dynamic web pages.

[0020] In terms of sequencing data quality control, the pre-processing method provided by the present invention abandons the shortcomings of single-threaded reading and writing of traditional quality control software, adopts an efficient dual-threaded read-write processing separation mode design, and doubles the reading and writing speed. Combined with an efficient open source compression library, the speed can be further improved. At the same time, it also abandons the shortcomings of traditional quality control software, either with a single indicator or too many indicators that are not suitable for specific business scenarios. For this business scenario, the pre-processing method provided by the present invention only has built-in efficient basic quality control, filters low-quality fragments, automatically detects and removes joint sequences, extracts diversified molecular tags, and has sufficient calculation indicators, but does not abuse indicators. In addition, during the molecular tag extraction process, the pre-processing method provided by the present invention can add molecular tag information to the original read segment name, improves the subsequent processing speed, and supports the label option of the industrialized comparison result BAM.

[0021] In terms of primer matching detection, for variable segment sequence matching, the preprocessing method provided by the present invention creatively proposes a segmented primer dictionary, and constructs a primer dictionary with primer sequence length-legal primer sequence / legal mismatch primer sequence as key-value pairs according to the different lengths of whitelist primers. At the same time, the present invention also directly maps the error tolerance (i.e., legal mismatch primer sequence) to the true set sequence in the dictionary, achieving efficient non-matching matching.

[0022] In terms of molecular tag error correction, the molecular tag error correction method provided by the present invention abandons the algorithm of comparing strings character by character used in the traditional molecular tag error correction process, introduces all molecular tags within the digital coding error tolerance range at one time, and uses the coding information as a list index, and the list elements store the corresponding molecular tag information, such as the number of errors, correct sequence information, etc. In addition, the molecular tag error correction method provided by the present invention also introduces a rolling hash algorithm for base offset to avoid repeated construction of subsequences, and also provides support for the misalignment of the starting position.

[0023] In terms of preprocessing result output and reporting, the preprocessing method provided by the present invention abandons standard text content, introduces industrialized FASTQ and JSON formats for preprocessing result output, and uses HTML interactive dynamic web pages for reporting, which facilitates access to upstream and downstream work software flows and user-friendly viewing of results.

[0024] In general, compared with the traditional pretreatment method, the pretreatment method provided by the present invention has been significantly improved in analysis speed and also reduced resource consumption. The test results of 189 samples show that the pretreatment method provided by the present invention takes an average of 2 minutes, while the traditional process requires quality control, splitting, processing, and merging steps, which takes an average of more than 3 hours; the pretreatment method provided by the present invention consumes an average of 5.2G of memory, while the total memory consumption of the traditional process is more than 50G; at the same time, the temporary files of the pretreatment method provided by the present invention are only one-tenth of those of the traditional method. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic diagram of the process of the efficient multiplex amplicon capture sequencing library pretreatment method provided by the present invention. DETAILED DESCRIPTION

[0026] The following describes the embodiments of the present invention through specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0027] Molecular tags are often very short fixed-length sequences, basically within 10 bases in length. The traditional method of molecular tag matching and error correction based on the comparison strategy is not only slow but also rarely has comparison software that can support such short sequences. If the strategy of traditional primer detection is followed, on the one hand, the molecular tag is shorter than the primer, and on the other hand, the molecular tag has a fixed length and is numerous, which will inevitably require a larger dictionary, which brings great challenges to the query speed. Therefore, first of all, in order to solve the problem of slow molecular tag error correction processing speed in the prior art, an embodiment of the present invention provides a molecular tag error correction method that can be efficient.

[0028] In an embodiment of the present invention, molecular tag error correction is performed based on digital storage of base sequences, and the specific method includes: constructing a legal list and a fault-tolerant list according to a UMI sequence whitelist; and correcting UMI sequences extracted from sequencing data according to the constructed legal list and fault-tolerant list.

[0029] In order to improve the processing speed, the present invention first creatively adopts digital storage of base sequences, that is, each UMI sequence is converted into a number. In some embodiments, the digital processing method includes: converting the UMI sequence into a binary value for storage. Specifically, the bases A / a, C / c, G / g, and T / t of the UMI sequence can be mapped to binary values ​​with a storage space of 2 bits, for example, the bases A / a, C / c, G / g, and T / t of the UMI sequence are mapped to binary values ​​00, 01, 10, and 11, respectively. It can be understood that the present invention does not limit the specific mapping relationship between the bases A / a, C / c, G / g, and T / t and the binary values ​​00, 01, 10, and 11. In some embodiments, the bases A / a, C / c, G / g, and T / t can also be mapped to binary values ​​01, 00, 10, and 11, or other combinations, respectively. If a base is stored as a character (A / a, C / c, G / g, T / t), 8 bits are required, while storing a binary value such as 00, 01, 10, 11 only requires 2 bits. Therefore, for a 64-bit computer, the longest binary value that can be stored is 32 bases, and the length of most UMI sequences does not exceed 10 bases, so the digital storage of UMI sequences is not limited to hardware. For example, the UMI sequence ACGCTAGT is converted into a binary value 0001100111001011, and the corresponding decimal value is 6603. Therefore, for a nucleotide sequence with a length of 8 bases, only 16 bits of binary value are needed for storage. In addition, the prior art uses a judgment method when converting bases A, C, G, and T into numbers, such as 0, 1, 2, and 3. For example, if a base is A, it is converted into 0, if it is C, it is converted into 1, and so on; if the uppercase and lowercase are considered, a, c, g, and t also need to be judged. Therefore, each base needs to be judged 4 times on average to obtain the correct number. When the present invention performs digital processing on the base sequence, it utilizes the characteristic that a character determined by ASCII encoding is also an integer, and the range is 0-256, and constructs a list with a length of 256. The corresponding numbers can be stored in the list index positions corresponding to each base A, C, G, T, a, c, g, and t, such as binary values ​​00, 01, 10, and 11. In this way, when converting each base into a number, it only needs to be indexed by the number, without judgment, and the value can be directly obtained in one step. The algorithm complexity is 1, and the speed is increased by 4 times.

[0030] Based on the above principle of digital storage of base sequences, the legal UMI sequences in the UMI sequence whitelist and the illegal mutation UMI sequences and legal mutation UMI sequences obtained by full enumeration mutation according to the preset number of allowed error bases are digitally processed respectively, and then a legal list of legal UMI sequences, a fault tolerance list of legal UMI sequences, illegal mutation UMI sequences, and legal mutation UMI sequences are constructed according to the digital processing information. Here, the present invention creatively transforms the error correction problem of molecular tags into a fault tolerance problem, that is, according to the preset number of allowed error bases, the legal UMI sequences in the UMI sequence whitelist are automatically fully enumerated and mutated within a certain number of bases, and the illegal mutation UMI sequences and legal mutation UMI sequences obtained by mutation are also digitally processed and stored in the fault tolerance list.

[0031] In the constructed legal list, a unique legal index corresponding to its digital processing information is added to each legal UMI sequence. In the constructed fault-tolerant list, a fault-tolerant index and a corresponding legal index of the legal UMI sequence are added to each legal UMI sequence and legal mutation UMI sequence, and an illegal value is set for the legal index of the illegal mutation UMI sequence. On this basis, for any UMI sequence extracted from the sequencing data, it only needs to be digitized in the same way, and then the corresponding positions of the fault-tolerant list and the legal list are located in turn according to the access index determined by the digital processing information, and the extracted UMI sequence is judged according to the legal UMI sequence and the legal mutation UMI sequence whether the fault-tolerant matching is successful. If successful, the error is corrected to the corresponding legal UMI sequence, that is, the correct UMI sequence; otherwise, it is discarded. In some embodiments, the digital processing information is any one of character information, RGB information, binary value, octal value, hexadecimal value, and decimal value corresponding to the binary value of the UMI sequence, such as the decimal value corresponding to the binary value of the UMI sequence, and the decimal value is used as the index of the corresponding UMI sequence.

[0032] The traditional molecular tag error correction method compares each legal UMI sequence with the extracted UMI sequence of the same length as its read segment to determine whether the difference is within the allowed range. For example, for 1,000 legal UMI sequences, each extracted UMI sequence must be compared 500 times on average, while after digital storage, only two list index accesses are required, which increases the speed by hundreds of times. Specifically, the method for constructing a legal list includes: digitizing the legal UMI sequences existing in the UMI sequence whitelist, adding legal indexes to the legal UMI sequences based on the digital processing information; storing the legal UMI sequences, the digital processing information corresponding to the legal UMI sequences, and the indexes corresponding to the legal UMI sequences in the legal list. For example, the legal UMI sequence ACGCTAGT is first converted into a binary value 0001100111001011, then converted into a decimal value 6603, and then the decimal value 6603 is used as the legal index of the legal UMI sequence.

[0033] Specifically, the method for constructing a fault tolerance list includes: for each legal UMI sequence, performing full enumeration mutation according to a preset number of allowed error bases to obtain an illegal mutation UMI sequence and a legal mutation UMI sequence; performing digital processing on the legal UMI sequence, the legal mutation UMI sequence, and the illegal mutation UMI sequence, adding a fault tolerance index and a corresponding legal index of the legal UMI sequence to the legal UMI sequence and the legal mutation UMI sequence according to the digital processing information, and setting an illegal value for the legal index of the illegal mutation UMI sequence; storing the legal UMI sequence, the digital processing information corresponding to the legal UMI sequence, the fault tolerance index corresponding to the legal UMI sequence, the legal index corresponding to the legal UMI sequence, the legal mutation UMI sequence, the digital processing information corresponding to the legal mutation UMI sequence, the fault tolerance index corresponding to the legal mutation UMI sequence, the legal index corresponding to the legal mutation UMI sequence, the illegal mutation UMI sequence, the digital processing information corresponding to the illegal mutation UMI sequence, and the legal index corresponding to the illegal mutation UMI sequence in the fault tolerance list. For example, for a legal UMI sequence ACGCTAGT with a length of 8 bp, if the tolerance is allowed to be 1 bp, then while keeping other positions unchanged, a certain position can take any of the four nucleotides A, C, G, and T. A total of 8*3= 24 mutant sequences are allowed. Together with the original sequence, a total of 25 sequences are legal. For the legal UMI sequence ACGCTAGT, its tolerance index is 6603 and its legal index is 6603. The legal mutant UMI sequence CCGCTAGT obtained by the fault-tolerant mutation of the legal UMI sequence ACGCTAGT is first converted into a binary value 0101100111001011, and the corresponding decimal value 22987 is used as the fault-tolerant index of the legal mutant UMI sequence, and the legal index of the corresponding legal UMI sequence is 6603 as the legal index of the legal mutant UMI sequence. The illegal mutant UMI sequence TCGCTAGT obtained by the fault-tolerant mutation of the legal UMI sequence ACGCTAGT does not have a tolerance index, and its legal index is set to -1.

[0034] The method for correcting the extracted UMI sequence includes: digitizing the extracted UMI sequence, adding an access index to the extracted UMI sequence according to the digitized processing information; accessing the fault tolerance list according to the access index, if the return value is an illegal value, the extracted UMI sequence cannot match the UMI sequence whitelist, and is used as an erroneous UMI sequence; if the return value is not an illegal value, the access index is matched to the corresponding fault tolerance index and legal index in the fault tolerance list in sequence, and then the legal list is accessed according to the matched legal index and matched to the corresponding legal UMI sequence, as the correct UMI sequence of the extracted UMI sequence, and the error is corrected through the UMI sequence. For example, the extracted molecular tag is CCGCTAGT, and after digital processing, its decimal value 22987 is taken as the access index, the fault tolerance list is accessed, and the legal mutation UMI sequence CCGCTAGT with a fault tolerance index of 22987 is matched, and the legal list is accessed according to the legal index 6603 of the legal mutation UMI sequence, and the legal UMI sequence ACGCTAGT with a legal index of 6603 is matched, and the correct UMI sequence of the extracted molecular tag is corrected.

[0035] In addition, since sequencing will cause certain base errors at the starting position, the molecular tag error correction method provided by the present invention allows a certain starting position offset. If one base is stored as a character, 8 bits are required, and an array of length 8 is required to store a sequence of length 8 bp; and for the array, the elements in the array are independent of each other, and the rolling hash algorithm operation cannot be performed, so the array must be rebuilt for each sequence. The molecular tag error correction method based on digital storage of base sequences provided by the present invention can be calculated by the rolling hash algorithm, that is, when calculating the hash of all substrings of a string, it is often not necessary to calculate from scratch, only through bit operations, remove the bases that need to be offset, and add the subsequent bases, which greatly improves the calculation speed.

[0036] Therefore, in some embodiments, the method for correcting the extracted UMI sequence further includes: returning the erroneous UMI sequence to the corresponding read segment in the quality control sequencing data, and obtaining an offset UMI sequence of the same length as the erroneous UMI sequence according to the position information of the erroneous UMI sequence on the corresponding read segment and the preset offset distance; digitizing the offset UMI sequence, and adding an access index to the offset UMI sequence according to the digitized processing information; accessing the fault tolerance list according to the access index, if the return value is an illegal value, the offset UMI sequence cannot match the UMI sequence whitelist, and the offset UMI sequence is discarded; if the return value is not an illegal value, matching the access index to the corresponding fault tolerance index and legal index in the fault tolerance list in sequence, and then accessing the legal list according to the matched legal index and matching to the corresponding legal UMI sequence as the correct UMI sequence of the offset UMI sequence. Wherein, the preset offset distance can be 1 base, 2 bases, or 3 bases. For example, for the read end, an 8 bp UMI sequence is extracted from the 5' position, and the extracted UMI sequence is digitized. The access index is used to access the fault tolerance list. If it is found to be an erroneous UMI sequence, the bit position of 0-8 bp can be moved 2 bits to the left through bit operation, and then the 2-bit number of the 9th base is added to the rightmost, so that the access index of the offset UMI sequence is obtained and the error correction is performed again. At this time, the algorithm complexity is only 1. If the base sequence is not digitized, the 8 bp base sequence must be taken from the second base position again, and the subsequence must be reconstructed to compare all UMI sequences one by one. At this time, the algorithm complexity is 8.

[0037] On the basis of the above-mentioned molecular tag error correction method, the embodiment of the present invention also provides an efficient and unified sequencing library preprocessing method, which can significantly improve the preprocessing speed and reduce the consumption of computing resources while ensuring accuracy, and solves the problem that the prior art lacks a preprocessing method that can solve sequencing library quality control, and / or primer matching detection, and molecular tag error correction at one time, thereby meeting the growing demand for genetic testing.

[0038] The efficient multiplex amplicon capture sequencing library preprocessing method provided by the embodiment of the present invention mainly includes: obtaining sequencing data having primer sequences and UMI sequences, performing quality control on the obtained sequencing data based on a read-write separation strategy, obtaining quality control sequencing data, and extracting UMI sequences from the quality control sequencing data; using the molecular tag error correction method to correct the UMI sequence of the quality control sequencing data; taking the sequencing data after quality control and UMI sequence error correction as the preprocessing result, and outputting and reporting the preprocessing result.

[0039] In some embodiments, the efficient multiplex amplicon capture sequencing library preprocessing method further includes primer matching detection, and the specific method includes: obtaining sequencing data having primer sequences and UMI sequences, performing quality control on the obtained sequencing data based on a read-write separation strategy, obtaining quality control sequencing data, and extracting UMI sequences and primer sequences from the quality control sequencing data; constructing a primer dictionary graded by primer sequence length according to a primer whitelist, and performing primer matching detection on the primer sequences of the quality control sequencing data according to the primer dictionary; using the molecular tag error correction method to correct the UMI sequences of the quality control sequencing data that have passed the primer matching detection; using the sequencing data after quality control, primer matching detection, and UMI sequence error correction as the preprocessing result, and outputting and reporting the preprocessing result.

[0040] Please refer to Figure 1 As shown in the flowchart, in this embodiment, the efficient multiplex amplicon capture sequencing library preprocessing method includes four major steps, namely quality control, primer matching detection, molecular tag error correction, and report output.

[0041] In the quality control stage, sequencing data with primer sequences and UMI sequences are obtained from the double-end sequencing library, and quality control is performed on the obtained sequencing data based on the read-write separation strategy to obtain quality control sequencing data, and UMI sequences and primer sequences are extracted from the quality control sequencing data. Among them, the content of quality control includes filtering low-quality fragments, intelligent automatic detection, and excision of adapter sequences.

[0042] In the primer matching detection stage, firstly, a primer dictionary graded by primer sequence length is constructed according to the primer whitelist, and then the primer sequence of the quality control sequencing data is tested for primer matching according to the primer dictionary. In some embodiments, it can specifically include: grading the legal primer sequences present in the primer whitelist and the legal mismatched primer sequences with the legal primers according to the primer sequence length, storing the legal primer sequences / legal mismatched primer sequences with the same primer sequence length in the same hash table, and constructing a primer dictionary with primer sequence length-legal primer sequence / legal mismatched primer sequence as the key value pair; grading the extracted primer sequences according to the different primer sequence lengths in the primer dictionary, obtaining primer subsequences of different primer sequence lengths, and respectively looking up the corresponding primer sequence lengths in the hash table; if there is a legal primer sequence / legal mismatched primer sequence identical to the primer subsequence in the hash table, the extracted primer sequence passes the primer matching detection. For example, a primer whitelist has 2,000 legal primers, which are classified according to the length of the primer sequence as follows: 300 primer sequences with a length of 20 bp, 500 primer sequences with a length of 18 bp, 600 primer sequences with a length of 17 bp, 400 primer sequences with a length of 22 bp, and 200 primer sequences with a length of 21 bp. The traditional algorithm will match these 2,000 legal primer sequences one by one with each extracted primer sequence of different lengths, so that each extracted primer sequence needs to be matched 1,000 times on average. The primer matching detection method provided by the present invention stores legal primers of different lengths into different hash tables, for example, 300 legal primers of 20 bp in length are stored in a hash table h20, 500 legal primers of 18 bp in length are stored in a hash table h18, 600 legal primers of 17 bp in length are stored in a hash table h17, 400 legal primers of 22 bp in length are stored in a hash table h22, and 200 legal primers of 21 bp in length are stored in a hash table h21. In this way, for each extracted primer sequence, only five primer subsequences of 20 bp, 18 bp, 17 bp, 21 bp, and 22 bp need to be taken accordingly, and then the five primer subsequences can be respectively consulted in the hash tables h20, h18, h17, h21, and h22 corresponding to their respective lengths. The complexity of the hash table lookup algorithm is 1, so any extracted primer sequence only needs four matches at most (one match in each hash table), and only two matches on average, which increases the speed by an average of 1,000 times.

[0043] In the molecular tag error correction stage, a legal list and a fault-tolerant list are constructed based on the UMI sequence whitelist; based on the constructed legal list and fault-tolerant list, the UMI sequences of the quality control sequencing data that have passed the primer matching test are corrected. Please refer to the above for the specific method of molecular tag error correction, which will not be elaborated here.

[0044] In the report output stage, the sequencing data after quality control, primer matching detection, and UMI sequence error correction is used as the preprocessing result, and the preprocessing result is output and reported. In some embodiments, the preprocessing result includes qualified data and unqualified data; the sequencing data that passes the quality control, primer matching detection, and UMI sequence error correction is used as qualified data; the sequencing data that fails the quality control, and / or fails the primer matching detection, and / or fails the UMI sequence error correction is used as unqualified data. In one or more embodiments, it is also possible to choose to output the unqualified data that failed the preprocessing to a separate file. The unqualified data mainly includes the specific reasons for the failure of the sequencing data preprocessing, such as failure to pass the quality control, and / or failure to pass the primer matching detection, and / or failure to pass the UMI sequence error correction, to facilitate the retrospective review of the library and guide the R&D personnel to improve the experimental method. In some embodiments, it is preferred to output the preprocessing results in the industrialized FASTQ and JSON formats to facilitate the adaptation and use of upstream and downstream software. In some embodiments, the preprocessing results are preferably reported in the form of an interactive dynamic web page; the report content includes at least one of quality control results, primer matching detection results, and UMI sequence error correction results, so that users can view the results conveniently.

[0045] In some embodiments, the high-efficiency multiplex amplicon capture sequencing library preprocessing method further includes: sequentially adding a space character and an extracted UMI sequence / correct UMI sequence to the original read segment name of the read segment corresponding to the extracted UMI sequence / correct UMI sequence as an output read segment name for outputting the preprocessing result. The molecular tag is attached to the original read segment name in a special format and separated from the original read segment name by a space (indicated by ":"). For example, the original read name is followed by: @A02025:107:H2CYKDSX7:2:1101:10004:10551, and the extracted UMI sequence corresponding to the read is ACTCACG, then the output read name is: @A02025:107:H2CYKDSX7:2:1101:10004:10551:ACTCACG, which means that a molecular tag ACTCACG is extracted from the read; if the extracted UMI sequence corresponding to the read is CCTCACG, then the correct UMI sequence ACTCACG obtained after molecular tag error correction of the extracted UMI sequence is added to the original read name, and the output read name is: @A02025:107:H2CYKDSX7:2:1101:10004:10551:ACTCACG. The original read name of the sequence is not modified. Based on the labeled storage method, the alignment results can be directly stored in the form of labels in the subsequent alignment results for easy query. This is because the alignment software uses the space as the separator and stores the two parts separated by the space into a structure, which not only retains the original read name, but also realizes the convenience of UMI sequence acquisition, greatly improving the processing speed of subsequent alignment results. The currently popular fastp method for extracting molecular tags will directly connect a character to the read name. If the separator UMI is used, the above output read name is: @A02025:107:H2CYKDSX7:2:1101:10004:UMI:ACTCACG. This kind of read name destroys the original read name. If you need to find this read segment later, there is no corresponding original read name because the original read name has been modified. On the other hand, the subsequent acquisition of molecular labels requires a large number of string search operations. If the delimiter is used improperly, it may even lead to search errors. For example, if the delimiter is SX7 instead of UMI, the processed name is @A02025:107:H2CYKDSX7:2:1101:10004:SX7:ACTCACG. If you search for SX7, you will find that the first field "H2CYKDSX7" in the read segment name can be matched, while the following field: 2:1101:10004:UMI:ACTCACG is obviously not a pure UMI sequence.In fact, it is impossible to predict in advance what the library read segment names will be, so randomly selected delimiters are prone to conflict with existing characters in the read segment. Therefore, it is preferred to use space (indicated by ":") as the delimiter.

[0046] The method of the embodiment of this specification is performed by a computer device, and the computer device can be a terminal device or server with computing power. The terminal device can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, desktop computers, etc. In some embodiments, the computer device can be configured with sequencing data analysis software, and the molecular tag error correction method of the embodiment of this specification can be configured in a sequencing data analysis tool or a computing module. In some embodiments, the computer device is configured with sequencing library preprocessing software according to the above-mentioned efficient multiple amplicon capture sequencing library preprocessing method. After obtaining the double-end sequencing library, the sequencing data in the double-end sequencing library is preprocessed using the sequencing library preprocessing software configured in the computer device. The sequencing library preprocessing software is implemented based on C++, and only one command line run is required to complete the quality control of the multiple primer probe second-generation sequencing library with molecular tags, and / or primer matching detection, molecular tag error correction, no third-party software is called, and no temporary intermediate files are generated. In one embodiment, the software preprocessed 189 sequencing data. Test results showed that the software took an average of 2 minutes, while the traditional process required quality control, splitting, processing, and merging steps, which took an average of more than 3 hours; the software consumed an average of 5.2 GB of memory, while the total memory consumption of the traditional process was more than 50 GB; at the same time, the temporary files generated by the software were only one-tenth of those of the traditional method.

[0047] The embodiments described above are merely descriptions of preferred implementations of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope of the present invention.

Claims

1. A primer matching detection method, characterized in that: include: The legal primer sequences existing in the primer whitelist and the legal mismatched primer sequences that have allowed base mismatches with the legal primers are classified according to the primer sequence length, and the legal primer sequences / legal mismatched primer sequences with the same primer sequence length are stored in the same hash table, and a primer dictionary with primer sequence length-legal primer sequence / legal mismatched primer sequence as key-value pairs is constructed; The primer sequences are graded according to different primer sequence lengths in the primer dictionary, primer subsequences of different primer sequence lengths are obtained, and the primer subsequences are respectively searched in a hash table corresponding to the primer sequence length; If there is a legal primer sequence / legal mismatched primer sequence that is identical to the primer subsequence in the hash table, the primer sequence passes the primer matching test.

2. An efficient multiplex amplicon capture sequencing library pretreatment method, characterized in that: include: The primer sequence is tested for primer matching using the method described in claim 1.

3. The method according to claim 2, characterized in that include: Acquire sequencing data with primer sequences and UMI sequences, perform quality control on the acquired sequencing data based on a read-write separation strategy, obtain quality control sequencing data, and extract UMI sequences from the quality control sequencing data; Performing primer matching detection on the primer sequence using the method described in claim 1; Correct the UMI sequence of the quality control sequencing data that has passed the primer matching test; The sequencing data after quality control, primer matching detection, and UMI sequence error correction is used as the preprocessing result, and the preprocessing result is output and reported.

4. The method according to claim 3, characterized in that Also includes: A space character and an extracted UMI sequence / correct UMI sequence are sequentially added to the original read name of the corresponding read segment of the extracted UMI sequence / correct UMI sequence as the output read name for outputting the preprocessing result.

5. The method according to claim 3, characterized in that The method for outputting and reporting the preprocessing results includes: outputting the preprocessing results in FASTQ and JSON formats, and reporting the preprocessing results in the form of an interactive dynamic web page.

6. The method according to claim 3, characterized in that The preprocessing results include qualified data and unqualified data; sequencing data that has passed quality control, primer matching detection, and UMI sequence error correction is regarded as qualified data; Sequencing data that fails quality control, primer matching detection, and / or UMI sequence error correction are considered unqualified data; The method of claim 3, wherein the report content includes at least one of a quality control result, a primer matching detection result, and a UMI sequence error correction result.

7. A computer device, characterized in that: The computer device is configured with sequencing library preprocessing software according to the efficient multiplex amplicon capture sequencing library preprocessing method according to any one of claims 2 to 7.