Method for determining consensus sequence of nucleic acid molecules and use

The Levinston distance algorithm recognizes feature sequences in ONT sequencing technology, which solves the problem of long-term recognition of linker sequences, and generates highly accurate nucleic acid molecules consensus sequences, achieving convenient and efficient generation of consensus sequences.

WO2025137944A1PCT designated stage expired Publication Date: 2025-07-03SHENZHEN HUADA GENE INST
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/142432
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

When generating high-quality nucleic acid molecule consensus sequences, the prior art has the problem that recognition of linker sequences is not convenient and time-consuming enough. Especially in ONT sequencing technology, it is difficult to meet the needs of high accuracy and long read length.

Method used

The Levinston distance algorithm is used to identify the feature sequences of the feature fragments, and multi-copy sequences of nucleic acid molecules are generated through rolling loop amplification and sequencing. The sequencing information is divided using the feature sequence to generate consensus sequences.

Benefits of technology

The recognition time is significantly shortened and the recognition efficiency is improved. The generated consensus sequence accuracy reaches 99.5%, which is simpler and more convenient than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023142432_03072025_PF_FP_ABST
    Figure CN2023142432_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a method for determining a consensus sequence of nucleic acid molecules and the use. The method comprises the following steps: acquiring sequencing information of the nucleic acid molecules, wherein the sequencing information comprises multi-copy sequences of library molecules, and the library molecules comprise single-stranded circular molecules formed by ligation of the nucleic acid molecules and feature fragments; identifying feature sequences of the feature fragments in the sequencing information on the basis of a Levenshtein distance algorithm and by means of known sequences of the feature fragments; dividing the sequencing information according to the feature sequences, and determining repetitive nucleic acid sequences of the nucleic acid molecules; and generating the consensus sequence according to a set of the repetitive nucleic acid sequences of the nucleic acid molecules.
Need to check novelty before this filing date? Find Prior Art

Description

Method and application for determining consensus sequence of nucleic acid molecules Technical Field

[0001] The present application relates to the field of sequencing technology, and in particular to methods and applications for determining the consensus sequence of nucleic acid molecules. Background Art

[0002] A major challenge in sequencing technology currently lies in balancing read length and accuracy. Highly accurate sequences require shorter reads, while long reads are relatively inaccurate. In cancer and genetic disease research, accurately identifying mutations is crucial for understanding disease development, progression, and treatment. In clinical applications, high-precision sequences are also key to achieving precision medicine. Therefore, generating highly accurate, long-length sequences is essential in these scenarios.

[0003] Third-generation sequencing technologies, represented by platforms from companies like Oxford Nanopore Technologies (ONT), are increasingly popular in ecological research due to their portability and low-cost sequencing capabilities. While this technology excels at long-read sequencing, it can also be applied to sequencing amplicons. A drawback of ONT is the lower quality of raw reads, which doesn't meet the high-quality requirements of some scenarios. Therefore, generating high-quality consensus sequences remains a challenge.

[0004] For scenarios requiring high-quality sequences, constructing a single-stranded circular library is a common method. By amplifying the same nucleic acid sequence multiple times, multiple copies are achieved. Ultimately, a consensus sequence is generated using these multiple copies. Taking Pacific Biosciences (PB)'s circular consensus sequencing (CCS) process as an example, the steps for constructing a consensus sequence are as follows: (1) identifying the linker (hairpin) sequence between each repeated nucleic acid fragment based on alignment; (2) identifying each repeated nucleic acid (subread) sequence by the position of the linker sequence; and (3) constructing a consensus sequence using the identified nucleic acid sequences.

[0005] Currently, linker sequences are primarily identified using tools similar to the Basic Local Alignment Search Tool (BLAST). The BLAST algorithm is primarily used to compare biological sequences to identify similarities, homologies, and possible functional associations. BLAST focuses on similarity between longer sequences and lacks a simpler and more convenient method for identifying shorter fragments, such as linker sequences.

[0006] Summary of the Invention

[0007] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a method and application for determining the consensus sequence of nucleic acid molecules, which is simpler and more convenient than BLAST for identifying shorter fragments such as linker sequences, and does not require comparison with a priori known sequences.

[0008] In order to achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are:

[0009] A first aspect of an embodiment of the present invention provides a method for determining a consensus sequence of a nucleic acid molecule, comprising the following steps:

[0010] S100: Acquire sequencing information of the nucleic acid molecule, the sequencing information including multiple copy sequences of a library molecule, the library molecule including a single-stranded circular molecule formed by connecting the nucleic acid molecule and a characteristic fragment;

[0011] S200: identifying a characteristic sequence of the characteristic fragment in the sequencing information based on the known sequence of the characteristic fragment and a Levingston distance algorithm;

[0012] S300: Dividing the sequencing information according to the characteristic sequence to determine the repeated nucleic acid sequence of the nucleic acid molecule;

[0013] S400: generating the consensus sequence according to the set of repeated nucleic acid sequences of the nucleic acid molecule.

[0014] In some embodiments of the present invention, S200 includes calculating the Levingston distance between the sequence to be identified in the sequencing information and the known sequence. If the Levingston distance is not greater than a threshold, the sequence to be identified is determined to be a characteristic sequence.

[0015] In some embodiments of the present invention, the threshold is 0.3 times the length of the known sequence.

[0016] In some embodiments of the present invention, S200 further includes correcting the characteristic sequence.

[0017] In some embodiments of the present invention, correction is performed based on the position of the characteristic sequence in the sequencing information.

[0018] In some embodiments of the present invention, the method of correcting according to the position includes correcting a plurality of the characteristic sequences that are close in position into one.

[0019] In some embodiments of the present invention, the repetitive nucleic acid sequences in the collection in S400 are pre-screened.

[0020] In some embodiments of the present invention, the sequencing information includes the first to nth repeated nucleic acid sequences in the order of sequencing, and the screening includes screening based on the consistency of the kth repeated nucleic acid sequence with the first repeated nucleic acid sequence, and counting them into the set, and meeting 1 <k≤n。

[0021] In some embodiments of the present invention, the consistency is selected from at least one of length, similarity, base content, and quality score.

[0022] In some embodiments of the present invention, generating the consensus sequence according to the set of repeated nucleic acid sequences of the nucleic acid molecule in S400 includes performing a multiple sequence alignment on the repeated nucleic acid sequences in the set to generate the consensus sequence.

[0023] In some embodiments of the present invention, the multiple sequence alignment method is abpoa.

[0024] In some embodiments of the present invention, the method further includes performing error correction on the generated consensus sequence.

[0025] In some embodiments of the present invention, the error correction processing adopts at least one of gcpp, racon, pilon, and Nextpolish.

[0026] A second aspect of an embodiment of the present invention provides a method for sequencing a nucleic acid molecule, comprising connecting a nucleic acid molecule and a characteristic fragment to form a single-stranded circular molecule, performing rolling circle amplification and sequencing to obtain sequencing information of the nucleic acid molecule, and obtaining a consensus sequence of the nucleic acid molecule according to the aforementioned method.

[0027] According to a third aspect of the embodiments of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the aforementioned method.

[0028] According to a fourth aspect of an embodiment of the present invention, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the aforementioned method when running the computer program.

[0029] A fifth aspect of the embodiments of the present invention provides a sequencing data analysis system, comprising:

[0030] an acquisition module, the acquisition module being configured to acquire sequencing information of the nucleic acid molecule, the sequencing information including multi-copy sequences of a library molecule, the library molecule including a single-stranded circular molecule formed by connecting the nucleic acid molecule and a characteristic fragment;

[0031] an identification module, configured to identify a characteristic sequence of the characteristic fragment in the sequencing information based on the known sequence of the characteristic fragment and a Levingston distance algorithm;

[0032] a partitioning module, configured to partition the sequencing information according to the characteristic sequence and determine the repeated nucleic acid sequence of the nucleic acid molecule;

[0033] A generation module is used to generate the consensus sequence according to the set of repeated nucleic acid sequences of the nucleic acid molecule.

[0034] A sixth aspect of the embodiments of the present invention provides a sequencing system, comprising:

[0035] A sequencing module is used to connect the nucleic acid molecule and the characteristic fragment to form a single-stranded circular molecule, and obtain the sequencing information of the nucleic acid molecule through rolling circle amplification and sequencing;

[0036] An acquisition module, configured to acquire the sequencing information;

[0037] an identification module, configured to identify a characteristic sequence of the characteristic fragment in the sequencing information based on the known sequence of the characteristic fragment and a Levingston distance algorithm;

[0038] a partitioning module, configured to partition the sequencing information according to the characteristic sequence and determine the repeated nucleic acid sequence of the nucleic acid molecule;

[0039] A generation module is used to generate the consensus sequence according to the set of repeated nucleic acid sequences of the nucleic acid molecule.

[0040] The beneficial effects of the embodiments of the present invention are:

[0041] In the solution of the embodiment of the present invention, the edit distance (Levingston distance) method is used to identify the characteristic sequence. Compared with the traditional BLAST-based identification solution, the time required is shortened to only 1 / 10 or even shorter, and the whole process is simpler and more convenient. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] FIG1 is a schematic diagram of the sequencing process according to an embodiment of the present invention.

[0043] FIG2 is an example diagram of characteristic sequences identified in sequencing reads according to an embodiment of the present invention.

[0044] FIG3 is a flow chart of consensus sequence generation according to an embodiment of the present invention.

[0045] FIG4 is an Identity distribution diagram of each repetitive nucleic acid sequence (subread) in an embodiment of the present invention.

[0046] FIG5 is a diagram showing the length distribution of each repeated nucleic acid sequence in an embodiment of the present invention.

[0047] FIG6 is a diagram showing the number distribution of original repetitive nucleic acid sequences in an embodiment of the present invention.

[0048] FIG7 is a diagram showing the final distribution of the number of repeated nucleic acid sequences after screening in an embodiment of the present invention.

[0049] FIG8 is a diagram showing the evaluation results of consensus sequence recognition in an embodiment of the present invention.

[0050] FIG9 is a statistical result of relevant indicators in the process of constructing a consensus sequence according to an embodiment of the present invention. DETAILED DESCRIPTION

[0051] In the description of the embodiments of the present application, the terms "first", "second" and "third" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first", "second" and "third" may explicitly or implicitly include at least one of such features. In the description of the embodiments of the present application, "several" means more than one, "multiple" means at least two, such as two, three, etc., "greater than", "less than", "exceed" and the like are understood to exclude the number itself, and "above", "below", "within" and the like are understood to include the number itself, unless otherwise clearly and specifically defined.

[0052] In the description of the embodiments of the present application, the reference terms "one embodiment," "some embodiments," "illustrative embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the description of the embodiments of the present application, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples.

[0053] According to a first aspect of an embodiment of the present invention, a method for determining a consensus sequence of a nucleic acid molecule is provided, comprising the following steps S100 to S400.

[0054] Nucleic acid molecules include DNA, such as double-stranded DNA (which may be genomic DNA, plasmid DNA, mitochondrial DNA, chloroplast virus DNA, artificially synthesized double-stranded DNA, etc.), and single-stranded DNA (which may be viral DNA, cDNA, artificially synthesized single-stranded DNA, etc.). Nucleic acid molecules also include RNA, such as dsRNA, mRNA, tRNA, rRNA, miRNA, siRNA, snoRNA, etc.

[0055] In some embodiments, nucleic acid molecules are obtained from samples, such as environmental samples or biological samples. Environmental samples include water (e.g., domestic water, industrial water, medical water, agricultural water, and other types of sewage and wastewater, or river and ocean water), air, soil, compost, silt (e.g., river silt, wastewater sedimentation tank silt), volcanic ash, frozen soil, and food (e.g., solid food, liquid food, beverages, etc.). Biological samples include at least one of body fluids (such as blood, tissue fluid, lymph fluid, cerebrospinal fluid, urine, sweat, sputum, saliva, gastric juice, intestinal juice, pancreatic juice, bile, prostatic fluid, vaginal secretions, semen, serous cavity effusion, joint cavity effusion, bronchoalveolar lavage fluid, amniotic fluid, etc.), skin, feces, intestinal contents, swabs (such as nasal swabs, pharyngeal swabs, anal swabs, etc.), tissues (such as histological samples obtained by surgery, endoscopy or percutaneous puncture biopsy), cells, and microorganisms (such as bacteria, viruses, fungi, actinomycetes, rickettsia, mycoplasma, chlamydia, spirochetes, etc.).

[0056] In some embodiments, nucleic acid molecules are obtained from samples through a cleavage reaction, which can be at least one of a physical method, a chemical method, and a biological method. Among them, physical cleavage includes at least one of a boiling method, a glass bead method, an ultrasonic method, a grinding method, a freeze-thaw method, and a homogenization method; chemical cleavage includes at least one of a surfactant method (SDS method) and an alkaline lysis method; and biological cleavage includes an enzymatic cleavage method, such as cleavage by enzymes such as lysozyme and proteinase K. In some embodiments, before the sample is lysed, an extraction and purification step is also included to remove impurities such as salts and organic matter. In some embodiments, the extraction and purification also includes precipitating nucleic acids or adsorbing nucleic acids. In some embodiments, after precipitation or adsorption of nucleic acids, the nucleic acid is also eluted. In some embodiments, when the extracted and purified nucleic acid molecule is RNA, a step of reverse transcribing the RNA to obtain DNA is also included.

[0057] A consensus sequence is a sequence that is derived from a group of similar but not identical sequences and in which each position consists of the most likely base.

[0058] S100: Acquire sequencing information of nucleic acid molecules, where the sequencing information includes multiple copy sequences of library molecules, and the library molecules include single-stranded circular molecules formed by connecting nucleic acid molecules and characteristic fragments.

[0059] In some embodiments, the sequencing of nucleic acid molecules is performed by any of first-generation sequencing, second-generation sequencing, and third-generation sequencing. Among them, first-generation sequencing is such as Maxam-Gilbert sequencing technology, Sanger dideoxy sequencing technology, pyrophosphate sequencing technology, fluorescent automatic sequencing technology, and hybridization sequencing technology, second-generation sequencing is such as 454 Roche GS FLX, Illumina Solexa, SOLiD, Ion Torrent, BGISEQ, etc., and third-generation sequencing is such as HeliScope, PacBio HiFi, PacBio CLR, ONT, Cyclone WT, etc. Third-generation sequencing is also known as single-molecule sequencing. It can read nucleotide sequences at the single-molecule level and has advantages such as longer read length and faster sequencing speed. Therefore, in some embodiments of the present invention, nucleic acid molecules are sequenced by single-molecule sequencing to obtain sequencing information. The error rate of sequencing methods such as third-generation sequencing is high, and the error position is random. Therefore, error correction is performed by sequencing the nucleic acid molecules multiple times to reduce the sequencing error rate. For example, multiple copies can be obtained by amplification, and error correction is performed according to the sequencing results of these copies.

[0060] To this end, in some embodiments, referring to FIG1 , multiple copies of nucleic acid molecules are obtained by rolling circle amplification for sequencing. In some embodiments, nucleic acid molecules are linked to characteristic fragments to form single-stranded circular library molecules for rolling circle amplification, and the amplified products are sequenced, so that the sequencing information includes multiple copies of the library molecules. In some embodiments, the multiple copies of the sequencing information include sequence information of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, or 200 copies of the nucleic acid molecule. Due to sequencing errors, the sequence information of these copies is not completely consistent, and thus the sequencing results of these copies are used for error correction to reduce the error rate.

[0061] A characteristic segment refers to a nucleic acid segment of a certain length. By ligating the characteristic segment to the ends of the nucleic acid molecule, a single-stranded circular molecule is ultimately formed for rolling circle amplification and sequencing. In some embodiments, the characteristic segment can form a hairpin structure with complementary ends. In some embodiments, the nucleic acid molecule is double-stranded or single-stranded. When the nucleic acid molecule is single-stranded, in some embodiments, the single-stranded nucleic acid molecule is converted to a double-stranded structure and then ligated to the characteristic segment; in other embodiments, the single-stranded nucleic acid molecule is ligated to the characteristic segment and then extended to form a double-stranded structure. In some embodiments, the same or different characteristic segments are ligated to each end of the nucleic acid molecule. In some embodiments, the length of the characteristic segment is 5-500 nt, 10-200 nt, or 20-100 nt, for example, 5 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 40 nt, 50 nt, 60 nt, 70 nt, 80 nt, 90 nt, 100 nt, 200 nt, 300 nt, 400 nt, or 500 nt.

[0062] S200: Identify the characteristic sequence of the characteristic fragment in the sequencing information based on the known sequence of the characteristic fragment and the Levingston distance algorithm.

[0063] The characteristic fragment that is introduced on nucleic acid molecule is generally the nucleic acid fragment with part or all of known sequence, but in order-checking process, because the error of detection signal and other reasons cause order-checking error, the characteristic sequence that the characteristic fragment in multicopy sequence is measured deviates from its known sequence.Therefore, it is necessary to identify these characteristic sequences in order-checking information by a certain algorithm.Mainly utilize the principle that is similar to the search tool (Basic Local Alignment Search Tool, BLAST) based on local sequence alignment algorithm to identify characteristic sequence at present.But BLAST is more devoted to the comparison of the similarity of longer sequence, and is quite loaded down with trivial details for the identification of the shorter fragment of this length of characteristic sequence, and consuming longer time.For this reason, the embodiment of the present invention especially proposes to adopt the mode of Levingston distance algorithm to identify.

[0064] Among them, Levenshtein distance belongs to a type of edit distance algorithm, which refers to the minimum number of editing operations required to change one string into another. The allowed editing operations include replacement (replacing one character with another), insertion (inserting a character), and deletion (deleting a character). For strings a and b, the Levenshtein distance between the first i characters of string a and the first j characters of string b is lev a,b (i, j) satisfies the following formula:

[0065] in, is an indicator function, only when a i =b j In the calculation of min, the three from top to bottom represent deletion, insertion, and substitution respectively.

[0066] Therefore, in embodiments of the present invention, the Levingston distance between two sequences refers to the minimum number of editing operations required to transform one sequence into another through base substitution, insertion, or deletion. In some embodiments, the Levingston distance is calculated using at least one of dynamic programming and recursion. In some embodiments, sequences in the sequencing information whose Levingston distance with the known sequence of the signature segment is no greater than a threshold are identified as signature sequences. In some specific embodiments, a certain number of sequences to be identified can be provided from the sequencing information, and the Levingston distance between the sequences to be identified in the sequencing information and the known sequences is calculated. If the Levingston distance between the sequences to be identified and the known sequences is no greater than a threshold, the sequence to be identified is determined to be a signature sequence. In some specific embodiments, the sequence to be identified can be obtained from the sequencing information, for example, by defining the length of the sequence to be identified based on the length of the known sequence of the signature segment, and then determining possible sequences to be identified in the sequencing information based on the length. In some embodiments, the threshold for the Levingston distance used to determine whether a sequence to be identified is a signature sequence is a set multiple of the length of the known sequence of the signature segment, for example, 0.1 to 0.5 times, such as 0.1, 0.2, 0.3, 0.4, or 0.5 times. Taking the threshold value of 0.3 times the length of the known sequence of the characteristic fragment as an example, for a known sequence of 21 bp in length, the threshold value of the Levingston distance is 21×0.3≈7 bp.

[0067] The length of the characteristic sequence is relatively short, so there may be fragments in the sequence of the nucleic acid molecule that are slightly different from the characteristic sequence. In this way, there may be situations where some non-characteristic sequence positions are identified as characteristic sequences using the Levingston distance method. Therefore, in some embodiments, after identifying the characteristic sequence using the Levingston distance algorithm, the identified characteristic sequence is also corrected to eliminate the situation where these non-characteristic sequences are mistaken for characteristic sequences as much as possible. Based on the principle of this type of misidentification, in some embodiments, correction is performed based on the position of the characteristic sequence in the sequencing information. Typically, the nucleic acid molecule sequence and the characteristic sequence are sequentially spaced in the sequencing information obtained after the library molecules containing the nucleic acid molecule and the characteristic fragment are subjected to rolling circle amplification and sequencing. Therefore, adjacent characteristic sequences are separated by a base distance that varies slightly or remains essentially unchanged. For this reason, it can be determined whether the characteristic sequence belongs to a characteristic sequence or a characteristic sequence mistakenly identified in the nucleic acid molecule based on the approximate position of the characteristic sequence in the sequencing information. In some embodiments, the method of correcting based on position includes correcting multiple characteristic sequences that are located close to each other into one. Specifically, proximity of positions needs to be judged in conjunction with the length of the nucleic acid molecule. For example, in some sequencing methods, the length of the nucleic acid molecule fragment is more than 200bp. In theory, it is obviously impossible for adjacent characteristic fragments that are alternately connected to the nucleic acid molecule to appear twice consecutively within 200bp. Therefore, during the operation, two characteristic sequences identified as characteristic fragments that exist consecutively within 200bp are corrected to one, and usually the characteristic sequence of the previous characteristic fragment is retained.

[0068] S300: Divide sequencing information according to characteristic sequences to determine repeated nucleic acid sequences of nucleic acid molecules.

[0069] The multi-copy sequences in the sequencing information include sequences in which nucleic acid molecules and characteristic fragments are spaced apart from each other. Therefore, referring to FIG2 , the characteristic sequences can be used to divide the sequencing information. The remaining sequences between the characteristic sequences include multiple repeated nucleic acid sequences obtained by sequencing different copies of the nucleic acid molecules.

[0070] S400: Generate a consensus sequence based on a set of repeated nucleic acid sequences of nucleic acid molecules.

[0071] In some embodiments, the repeated nucleic acid sequences of the nucleic acid molecule obtained by dividing the characteristic sequence constitute a set of repeated nucleic acid sequences of the nucleic acid molecule, and a common sequence of the nucleic acid molecule is generated based on the set.

[0072] Due to problems such as sequencing errors, the lengths, qualities, etc. of each repeated nucleic acid sequence of a nucleic acid molecule may be inconsistent. Therefore, in some of these embodiments, the repeated nucleic acid sequences are pre-screened to obtain repeated nucleic acid sequences with better quality and improve the accuracy of the consensus sequence. In some specific embodiments, among the various repeated nucleic acid sequences of the nucleic acid molecule, the first repeated nucleic acid sequence to the nth repeated nucleic acid sequence are included in the order of sequencing after amplification. Based on the length and sequence information of the first repeated nucleic acid sequence, the remaining repeated nucleic acid sequences are screened. It should be noted that if the first two characteristic sequences are identified, the length of the first repeated nucleic acid sequence is relatively accurate; if the second repeated nucleic acid sequence or a later repeated nucleic acid sequence is used as a reference, it is necessary to require that no mis-identification problems occur in the first three or more characteristic sequences, and the risk is relatively high. Therefore, it is selected to screen the remaining repeated nucleic acid sequences based on the length and sequence information of the first repeated nucleic acid sequence. In some specific embodiments, screening is performed according to the consistency between the kth repeated nucleic acid sequence and the first repeated nucleic acid sequence. If the consistency requirement is met, the kth repeated nucleic acid sequence (1 < k ≤ n) is included in the set of repeated nucleic acid sequences. In some embodiments, the consistency between the kth repeated nucleic acid sequence and the first repeated nucleic acid sequence includes at least one, at least two, or all of length, similarity, base content, quality score, etc. For the consistency of length, for example, when the length difference between the kth repeated nucleic acid sequence and the first repeated nucleic acid sequence does not exceed ±0.1%, ±0.2%, ±0.3%, ±0.5%, ±1%, ±2%, ±3%, ±5%, ±10%, ±15%, ±20%, ±25%, ±30%, ±35%, ±40%, ±45%, ±50%, the consistency screening requirement is met. For similarity, for example, when the ratio of the number of matching bases in the cigar value of the length of the kth repeated nucleic acid sequence to the first repeated nucleic acid sequence to the total number of bases of the kth repeated nucleic acid sequence is greater than ⅆ, ⅆ, ⅆ, ⅆ, ⅆ, ⅆ, ⅆ, the consistency screening requirement is met. For base content, for example, when the difference in GC content between the kth repeated nucleic acid sequence and the first repeated nucleic acid sequence does not exceed ±0.01%, ±0.02%, ±0.03%, ±0.05%, ±0.1%, ±0.2%, ±0.3%, ±0.5%, ±1%, ±1.5%, ±2%, ±2.5%, ±3%, ±3.5%, ±4%, ±4.5%, ±5%, ±6%, ±7%, ±8%, ±9%, ±10%, the consistency screening requirement is met. In some embodiments, the quality score includes at least one of Phred score, Sanger score, Solexa score, etc.

[0073] It should be noted that there are some undefined values (ⅆ) in the translated text for similarity, which need to be filled in according to the specific content in the original Chinese text. If you can provide more accurate information, I can further optimize the translation.In some embodiments, generating a consensus sequence based on a set of repetitive nucleic acid sequences of nucleic acid molecules includes performing a multiple sequence alignment on the repetitive nucleic acid sequences in the set to generate a consensus sequence. In some sequencing scenarios, due to the lack of prior nucleic acid sequence information, in some specific embodiments, a multiple sequence alignment method without prior knowledge is used to generate a consensus sequence using the set of repetitive nucleic acid sequences, such as poa, spoa, abpoa, etc. In some embodiments, the multiple sequence alignment method without prior knowledge is abpoa.

[0074] Due to the high error rate of sequencing, in some embodiments, the generated consensus sequence is also subjected to error correction processing to further improve accuracy. In some embodiments, the error correction processing can adopt at least one of gcpp, racon, pilon, and Nextpolish.

[0075] A second aspect of an embodiment of the present invention provides a method for sequencing a nucleic acid molecule, comprising connecting a nucleic acid molecule and a characteristic fragment to form a single-stranded circular molecule, performing rolling circle amplification and sequencing to obtain sequencing information of the nucleic acid molecule, and obtaining a consensus sequence of the nucleic acid molecule according to the aforementioned method.

[0076] According to a third aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the aforementioned method.

[0077] According to a fourth aspect of an embodiment of the present invention, an electronic device is provided. The electronic device includes a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and the processor implements the aforementioned method when running the computer program.

[0078] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the aforementioned method described in the embodiments of the present invention. The processor determines the consensus sequence of the nucleic acid molecule by executing the non-transitory software program and instructions stored in the memory.

[0079] The memory may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function, while the data storage area may store and execute the aforementioned programs. Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0080] In some embodiments, the memory may include a memory remotely located relative to the processor, and the remote memory may be connected to the processor via a network. Examples of the aforementioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0081] The non-transitory software program and instructions required to implement the above method are stored in the memory, and when executed by one or more processors, the above method is performed.

[0082] A fifth aspect of the embodiments of the present invention provides a sequencing data analysis system, comprising:

[0083] An acquisition module is used to obtain sequencing information of nucleic acid molecules, wherein the sequencing information includes multiple copy sequences of library molecules, and the library molecules include single-stranded circular molecules formed by connecting nucleic acid molecules and characteristic fragments;

[0084] An identification module is used to identify the characteristic sequence of the characteristic fragment in the sequencing information based on the known sequence of the characteristic fragment and the Levingston distance algorithm;

[0085] A partitioning module is used to partition sequencing information according to characteristic sequences and determine the repeated nucleic acid sequences of nucleic acid molecules;

[0086] A generating module is used to generate the consensus sequence according to a set of repeated nucleic acid sequences of nucleic acid molecules.

[0087] In some embodiments, the identification module calculates the Levingston distance between the sequence to be identified and the known sequence in the sequencing information. If the Levingston distance is not greater than a threshold, the sequence to be identified is determined to be a signature sequence. In some specific embodiments, the threshold is a set multiple of the length of the known sequence of the signature segment, for example, 0.1 to 0.5 times, such as 0.1, 0.2, 0.3, 0.4, or 0.5 times. Taking the threshold of 0.3 times the length of the known sequence of the signature segment as an example, for a known sequence of 21 bp in length, the Levingston distance threshold is 21 × 0.3 ≈ 7 bp.

[0088] In some embodiments, the recognition module further comprises correcting the signature sequence. In some specific embodiments, the correction is performed based on the position of the signature sequence in the sequencing information. In some specific embodiments, the method of correcting based on position includes correcting multiple signature sequences that are closely located into one.

[0089] In some of these embodiments, the generation module pre - screens the repetitive nucleic acid sequences in the set. In some specific embodiments, the sequencing information includes the 1st to the nth repetitive nucleic acid sequences in the order of sequencing. The screening includes screening according to the consistency between the kth repetitive nucleic acid sequence (1 < k ≤ n) and the 1st repetitive nucleic acid sequence, and the kth repetitive nucleic acid sequence with the consistency meeting the requirements is included in the set. In some specific embodiments, the consistency is selected from at least one of length, base content, and mass fraction.

[0090] In some of these embodiments, the generation module performs multiple - sequence alignment on the repetitive nucleic acid sequences in the set to generate a consensus sequence. In some specific embodiments, the method of multiple - sequence alignment is abpoa. In some specific embodiments, the generation module further includes error - correction processing on the generated consensus sequence. In some specific embodiments, the error - correction processing uses at least one of gcpp, racon, pilon, and Nextpolish.

[0091] In a sixth aspect of the embodiments of the present invention, a sequencing system is provided, including:

[0092] A sequencing module, configured to ligate a nucleic acid molecule and a characteristic fragment to form a single - stranded circular molecule, and obtain sequencing information of the nucleic acid molecule through rolling - circle amplification and sequencing;

[0093] An acquisition module, configured to acquire the sequencing information;

[0094] An identification module, configured to identify the characteristic sequence of the characteristic fragment in the sequencing information based on the known sequence of the characteristic fragment and the Levenshtein distance algorithm;

[0095] A partitioning module, configured to partition the sequencing information according to the characteristic sequence and determine the repetitive nucleic acid sequences of the nucleic acid molecule;

[0096] A generation module, configured to generate the consensus sequence according to the set of repetitive nucleic acid sequences of the nucleic acid molecule.

[0097] The device or system implementation described above is only illustrative. The modules described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0098] It is understood that all or some of the steps disclosed above can be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). It is understood that computer storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include but are not limited to RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, disk storage or other magnetic storage device, or any other medium that can be used to store desired information and can be accessed by a computer.

[0099] Additionally, it will be appreciated that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0100] The present application is described below with reference to specific embodiments.

[0101] Example 1

[0102] Randomly fragmented nucleic acid molecules of E. coli DNA were linked to the characteristic hairpin fragment (sequence: CCTTGGCTCACAGAACGACAT) and subjected to rolling circle sequencing to obtain sequencing information for the multiple copies of these nucleic acid molecules. Based on this sequencing information, the consensus sequence of the nucleic acid molecules was identified. The steps are as follows, referring to Figure 3:

[0103] (1) Identification is performed through the known sequence of the characteristic fragment. The known sequence of the characteristic fragment is CCTTGGCTCACAGAACGACAT, and the Levenshtein distance threshold is 7bp. The Python library used in this method to calculate the Levenshtein distance is: python_Levenshtein (0.21.1) (https: / / maxbachmann.github.io / Levenshtein / levenshtein.html). The specific identification steps are: scanning from the beginning of fastq, the step length of each movement is 1, and the scanning window size is the length of the known sequence of the characteristic fragment, 21bp. Each step length can obtain the Levenshtein distance between the sequence to be identified and the known sequence in the current window. For positions that meet the threshold conditions, the sequence to be identified in the window is recorded as the characteristic sequence. The characteristic sequences screened in this step are shown in Figure 6. The read lengths are classified according to the number of characteristic sequences originally identified, and the proportion of each type of read length in the total data volume is obtained. This result can reflect the quality of library construction and sequencing.

[0104] (2) Correct the identified characteristic sequences. As shown in Figure 3, when two characteristic sequences appear consecutively within 200bp, the two characteristic sequences are merged (retained) as the previous characteristic sequence, and the latter characteristic sequence is discarded. Secondly, the specific position of characteristic sequence 1 is determined. Usually, a shorter nucleic acid fragment may appear before a stable inherent sequence. Therefore, after identifying the original characteristic sequence position, this method obtains the length information of each repeated nucleic acid sequence. This method determines: when the length of the first repeated nucleic acid sequence is less than 1 / 2 of the length of the second repeated nucleic acid sequence, the second repeated nucleic acid sequence is regarded as the first, and so on. If the length of the first repeated nucleic acid sequence is greater than 1 / 2 of the length of the second repeated nucleic acid sequence, no change is made. After the correction of the characteristic sequence position is completed, the characteristic sequence quantity distribution of all read lengths will be recorded, which is called the quantity distribution of the original characteristic sequence for downstream analysis.

[0105] (3) The repetitive nucleic acid sequences of the nucleic acid molecule are segmented according to the position of the characteristic sequence in the read length. After identifying N characteristic sequences, N+1 repetitive nucleic acid sequences can be obtained. The repetitive nucleic acid sequences obtained by segmentation according to the characteristic sequence are saved in files such as subread1, subread2, etc. according to their position distribution in the read length. This facilitates the evaluation of the quality of the sorted data. As shown in Figures 4 and 5, the recorded repetitive nucleic acid sequence (subread) information is analyzed for length and identity according to the position of the repetitive nucleic acid sequence to obtain the overall data quality.

[0106] (4) Screening of multiple repetitive nucleic acid sequences. Based on the length and sequence information of the first repetitive nucleic acid sequence, the remaining repetitive nucleic acid sequences were screened. The length screening criteria were to retain subreadk (k>1) with a length between 0.8 and 1.2 times the length of subread1, and discard other subreads with a length less than 0.8 times or greater than 1.2 times. The similarity screening criteria were to use the edlib library (1.3.9) in Python, and calculate the similarity by taking the ratio of the number of matched bases in the cigar value after the alignment between subreadk and subread1 to the total number of bases in subreadk. If the similarity between subreadk and subread1 is greater than 0.8, it is retained, otherwise it is discarded. The number of subreads finally retained based on the length screening and similarity screening is recorded to obtain a distribution of the number of subreads after screening. As shown in Figure 7, the proportion of each type of subread in the total number after screening can finally be obtained. Among them, ≥3D represents that the number of subreads finally retained in the read length of this category is ≥3, and so on.

[0107] (5) Using the screened set of repetitive nucleic acid sequences, a high-quality consensus sequence was generated using abpoa (1.2.5). As shown in Figure 8, the consensus sequences were finally classified based on the number of repetitive nucleic acid sequences, and the identity of each class was evaluated. ≥3D means that the number of subreads retained in this class is ≥3, and so on. Finally, the number of bases, number of reads, read N50 length, average length, and the proportion of reads in each step to the total data were summarized, as shown in Figure 9.

[0108] (6) Correct the generated consensus sequence using Nextpolish.

[0109] Comparative Example 1 constructs a consensus sequence for the rolling circle sequencing information in Example 1 according to the Pacific Biosciences circular consensus sequencing (CCS) process (see Nucleic Acids Res. 2010 Aug; 38(15):e159). The tool for identifying characteristic sequences based on the traditional BLAST recognition method is the Python BLAST library: Biopython (1.81) (https: / / doi.org / 10.1093 / bioinformatics / btp163).

[0110] For the same 10,000 sequences, the Levingston distance-based adapter sequence recognition method in Example 1 took 97 seconds, while the traditional BLAST-based recognition method in Comparative Example 1 required 20 minutes. Levingston distance-based adapter sequence recognition reduced the time to less than 1 / 10 of the traditional BLAST-based recognition scheme. Comparing the accuracy of the two methods, the consensus sequence provided by Example 1 had an accuracy of 99.5%, while the consensus sequence provided by the method in Comparative Example 1 had an accuracy of only 98.8%.

[0111] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A method for determining a consensus sequence of a nucleic acid molecule, characterized in that, It includes the following steps: S100: Obtain the sequencing information of the nucleic acid molecule, where the sequencing information includes multi-copy sequences of library molecules, and the library molecules include single-stranded circular molecules formed by ligating the nucleic acid molecule and a characteristic fragment; S200: Based on the known sequence of the characteristic fragment and the Levenshtein distance algorithm, identify the characteristic sequence of the characteristic fragment in the sequencing information; S300: Divide the sequencing information according to the characteristic sequence to determine the repetitive nucleic acid sequence of the nucleic acid molecule; S400: Generate the consensus sequence according to the set of repetitive nucleic acid sequences of the nucleic acid molecule.

2. The method according to claim 1, characterized in that, S200 includes calculating the Levenshtein distance between the sequence to be identified in the sequencing information and the known sequence. If the Levenshtein distance is not greater than the threshold, the sequence to be identified is determined as the characteristic sequence.

3. The method according to claim 2, wherein The threshold is 0.3 times the length of the known sequence.

4. The method according to claim 1, characterized in that S200 further includes correcting the characteristic sequence.

5. The method according to claim 4, wherein The correction is performed according to the position of the characteristic sequence in the sequencing information.

6. The method according to claim 5, characterized in that, The method of correction according to the position includes correcting multiple characteristic sequences with close positions into one.

7. The method according to claim 1, characterized in that The repetitive nucleic acid sequences in the set in S400 are pre-screened.

8. The method according to claim 7, wherein In the sequencing information, the 1st repetitive nucleic acid sequence to the nth repetitive nucleic acid sequence are included in the order of sequencing. The screening includes screening according to the identity between the kth repetitive nucleic acid sequence and the 1st repetitive nucleic acid sequence and including it in the set, where 1 < k ≤ n.

9. The method according to claim 8, wherein The identity is selected from at least one of length, similarity, base content, and quality score.

10. The method according to claim 1, wherein Generating the consensus sequence according to the set of repetitive nucleic acid sequences of the nucleic acid molecule in S400 includes performing multiple sequence alignment on the repetitive nucleic acid sequences in the set to generate the consensus sequence.

11. The method according to claim 10, wherein The method of multiple sequence alignment is abpoa.

12. The method according to claim 1, characterized in that, It further includes performing error correction processing on the generated consensus sequence.

13. The method according to claim 12, wherein The error correction processing uses at least one of gcpp, racon, pilon, and Nextpolish.

14. A method for sequencing a nucleic acid molecule, comprising the following steps: Ligate the nucleic acid molecule and the characteristic fragment to form a single-stranded circular molecule, perform rolling circle amplification and sequencing to obtain the sequencing information of the nucleic acid molecule, and obtain the consensus sequence of the nucleic acid molecule according to the method according to any one of claims 1 to 13.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the method according to any one of claims 1 to 14.

16. An electronic device, characterized in that, It includes a processor and a memory. A computer program that can run on the processor is stored on the memory, and the processor implements the method according to any one of claims 1 to 14 when running the computer program.

17. A sequencing data analysis system, characterized in that, It includes: An acquisition module, which is used to acquire the sequencing information of the nucleic acid molecule, where the sequencing information includes multi-copy sequences of library molecules, and the library molecules include single-stranded circular molecules formed by ligating the nucleic acid molecule and a characteristic fragment; An identification module, which is used to identify the characteristic sequence of the characteristic fragment in the sequencing information based on the known sequence of the characteristic fragment and the Levenshtein distance algorithm; A partitioning module, which is used to partition the sequencing information according to the characteristic sequence and determine the repetitive nucleic acid sequence of the nucleic acid molecule; A generating module, which is used to generate the consensus sequence according to the set of repetitive nucleic acid sequences of the nucleic acid molecule.

18. A sequencing system, characterized in that, Comprising: A sequencing module, which is used to ligate a nucleic acid molecule and a characteristic fragment to form a single-stranded circular molecule, and obtain the sequencing information of the nucleic acid molecule through rolling circle amplification and sequencing; An acquisition module, which is used to acquire the sequencing information; An identification module, which is used to identify the characteristic sequence of the characteristic fragment in the sequencing information based on the Levenshtein distance algorithm through the known sequence of the characteristic fragment; A partitioning module, which is used to partition the sequencing information according to the characteristic sequence and determine the repetitive nucleic acid sequence of the nucleic acid molecule; A generating module, which is used to generate the consensus sequence according to the set of repetitive nucleic acid sequences of the nucleic acid molecule.

Citation Information

Patent Citations

  • Methods for accurate sequence data and modified base position determination

    CN102076871A

  • Method for performing high-throughput sequencing on haplotype of genome subjected to two-time DNA fragmentation

    CN104357563A

  • Single stranded circular DNA libraries for circular consensus sequencing

    CN110062809A

  • Nucleic acid sequencing library and construction method thereof

    CN110219054A

  • Method for constructing sequencing library, obtained sequencing library and sequencing method

    CN112739829A