Method for obtaining and correcting biological sequence information

By using sequencing reagents with nucleotide monomers conjugated with different labels and mathematical analysis in high-throughput sequencing, the problem of sequencing errors in high-throughput sequencers was solved, and highly accurate sequence information acquisition was achieved.

CN108699599BActive Publication Date: 2026-01-06CYGNUS BIOSCI BEIJING CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN201680079417.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-10-14
Filing Date
2016-11-16
Publication Date
2026-01-06
Estimated Expiration
2036-11-16

AI Technical Summary

Technical Problem

Current high-throughput sequencers cannot completely eliminate sequencing errors, which mainly originate from accidental reaction errors, signal acquisition errors, and errors caused by signal correction. Current technology optimizations are mostly focused on chemical reactions and image signal processing, lacking innovation from the sequencing logic.

Method used

Sequencing reagents that provide nucleotide monomers conjugated with different labels in the presence of a polynucleotide replication catalyst are used. By detecting the fluorescence emission of nucleotide monomers after incorporation into polynucleotides, combined with mathematical analysis and algorithms to compare sequence information, sequencing errors can be reduced or eliminated.

Benefits of technology

It improves sequencing accuracy, enabling high-throughput sequencing to achieve up to 95% base pair accuracy and reducing or eliminating sequence errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0001735767580000031
    Figure BDA0001735767580000031
  • Figure BDA0001735767580000041
    Figure BDA0001735767580000041
  • Figure BDA0001735767580000361
    Figure BDA0001735767580000361
Patent Text Reader

Abstract

This disclosure provides methods for sequencing biomolecules such as nucleic acid molecules, and methods for detecting and / or correcting sequencing errors in sequencing results. Reagent kits and systems based on the methods disclosed herein are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates in some aspects to a high-throughput sequencing method, belonging to the field of gene sequencing. Background Technology

[0002] High-throughput sequencing is a technology that has developed rapidly in recent years. Compared to traditional Sanger sequencing, the biggest advantage of high-throughput sequencing is that it can read massive amounts of sequence information simultaneously. Although its accuracy is not as high as traditional sequencing methods, the analysis of massive amounts of data can yield information beyond the sequence itself, such as gene expression levels and copy number changes.

[0003] Most mainstream sequencers today use the SBS (sequencing-by-synthesis) method, such as Solexa / Illumina, 454, and IonTorrent. These sequencers share similar structures, including fluidic systems, optical systems, and chip systems. The sequencing reaction occurs within the chip. The sequencing process is also similar, involving: introducing the reaction solution into the chip, initiating the SBS reaction, acquiring the signal, and washing. Then, a new round of sequencing is performed. This is a cyclical process. With each cycle, consecutive single-base non-degenerate sequence information (such as ACTGACTG) is obtained. However, high-throughput sequencers cannot completely eliminate sequencing errors. Sequencing errors can originate from: accidental or cumulative reaction errors, signal acquisition errors, errors from signal correction, etc. In existing sequencers, these chemical, optical, or software errors can become noise, which cannot be identified at a single readout site. They can only be eliminated through deep sequencing, utilizing multiple readouts of the same sequence at different sites. More accurate readouts are an important direction for the development of high-throughput sequencing. However, current technologies for improving accuracy primarily focus on optimizing the chemical reaction itself and subsequent image signal processing, without innovating the sequencing logic. Therefore, there is a need for improved sequencing methods. Summary of the Invention

[0004] This application claims priority to the following Chinese patent applications: Chinese patent application CN201510822361.9, filed November 18, 2015, entitled "A method for sequencing nucleotide molecules with phosphate-modified fluorophores"; Chinese patent application CN201510815685.X, filed November 18, 2015, entitled "A method for sequencing nucleotide substrate molecules with fluorophores having fluorescence switching properties"; Chinese patent application CN201510944878.5, filed December 11, 2015, entitled "A method for detecting and correcting sequence data errors in sequencing results"; and Chinese patent application CN201610899880.X, filed October 14, 2016, entitled "A method for reading sequence information from raw signals of high-throughput DNA sequencing," the entire contents of which are incorporated herein by reference.

[0005] The summary of this invention is not intended to limit the scope of the claimed subject matter. Other features, details, utility, and advantages of the claimed subject matter will become apparent from the detailed description of those aspects disclosed in the drawings and appended claims.

[0006] On the one hand, this article provides a method for obtaining sequence information of a target polynucleotide, the method comprising: a) providing a first sequencing reagent to the target polynucleotide in the presence of a first polynucleotide replication catalyst, wherein the first sequencing reagent comprises at least two different nucleotide monomers each conjugated to a first label, and the nucleotide monomer / first label conjugate is substantially non-fluorescent until the nucleotide monomer is incorporated into the target polynucleotide according to complementarity with the target polynucleotide, wherein the first labels of the at least two different nucleotide monomers are the same or different; and b) providing a second sequencing reagent to the target polynucleotide in the presence of a second polynucleotide replication catalyst, wherein the second sequencing reagent comprises one or more nucleotide monomers each conjugated to a second label, and the nucleotide monomer / second label conjugate is substantially non-fluorescent until the nucleotide monomer is incorporated into the target polynucleotide according to complementarity with the target polynucleotide, wherein at least one of the one or more nucleotide monomers is different from the nucleotide monomers present in the first sequencing reagent, and wherein the second sequencing reagent is provided after the first sequencing reagent is provided; and c) obtaining at least a portion of the sequence information of the target polynucleotide by detecting fluorescence emission caused by the first label and the second label after the nucleotide monomer is incorporated into the polynucleotide in steps a) and b).

[0007] In one embodiment, the method is used to obtain sequence information of at least a partial single target polynucleotide. In another embodiment, the method is used to simultaneously obtain sequence information of at least a partial plurality of target polynucleotides.

[0008] In any of the foregoing embodiments, the first polynucleotide replication catalyst and the second polynucleotide replication catalyst may be the same polynucleotide replication catalyst or different polynucleotide replication catalysts.

[0009] In any of the foregoing embodiments, the sequence information can be obtained through one or more sequencing reactions, wherein optionally one or more sequencing reactions are performed in one or more reaction volumes (e.g., reaction chambers), such as about 1 × 10⁻⁶. 6 Approximately 5×10 8 One reaction volume, approximately 1 × 10 6 To approximately 1×10 8 One reaction volume or approximately 1 × 10⁻⁶ 6 Approximately 5×10 7 Sequencing takes place in individual reaction volumes, wherein optionally the reaction volumes are physically separated from each other and / or there is no or substantially no material exchange between reaction volumes, wherein optionally the reaction volumes are located in an array such as a chip, wherein optionally the reaction volumes are enclosed and / or isolated from each other by liquids immiscible with the liquids in the reaction volumes, such as oil. When there is substantially no material exchange between reaction volumes, some material exchange is permitted but this will not affect the sequencing results in any reaction volume to the point of causing cross-contamination.

[0010] In any of the foregoing embodiments, a reaction volume may be provided within a reaction chamber, and the target polynucleotide in each reaction chamber is immobilized on a solid support within the reaction chamber, wherein optionally the sequence information is obtained by high-throughput sequencing, for example, wherein at least about 10 3 10 4 10 5 10 6 10 7 10 8 Or 10 9 The sequences are read in parallel. In any of the foregoing embodiments, the first polynucleotide replication catalyst and / or the second polynucleotide replication catalyst is a polymerase, such as DNA polymerase, RNA polymerase or RNA-dependent RNA polymerase, ligase, reverse transcriptase or terminal deoxyribonucleoside transferase.

[0011] In any of the foregoing embodiments, the nucleotide monomers in the first and / or second sequencing reagents may be selected from the group consisting of: deoxyribonucleotides, modified deoxyribonucleotides, ribonucleotides, modified ribonucleotides, peptide nucleotides, modified peptide nucleotides, modified phosphate sugar backbone nucleotides, and mixtures thereof. In one embodiment, both the first and second sequencing reagents contain deoxyribonucleotides. In some embodiments, the nucleotide monomers are selected from the group consisting of: A, T / U, C, and G deoxyribonucleotides, and analogs thereof. In another embodiment, both the first and second sequencing reagents contain ribonucleotides. In a specific embodiment, the nucleotide monomers are selected from the group consisting of: A, U / T, C, and G ribonucleotides, and analogs thereof.

[0012] In any of the foregoing embodiments, the first and / or second label is releasably conjugated to the nucleotide monomer. In one embodiment, the first and / or second label is conjugated to the terminal phosphate group of the nucleotide monomer. In a specific embodiment, the nucleotide monomer / first-labeled conjugate in the first sequencing reagent and / or one or more nucleotide monomers / second-labeled conjugates in the second sequencing reagent have the structure of Formula I:

[0013]

[0014] Where n is 0-6, R is a nucleoside base, and X is H, OH, or OMe, or a salt thereof. In some embodiments, the first and / or second label is substantially non-fluorescent until after release from the terminal phosphate group of the nucleotide monomer. In another embodiment, the method further includes using an activating enzyme to release the first and / or second label from the terminal phosphate group of the nucleotide monomer. In one embodiment, the activating enzyme is an exonuclease, phosphotransferase, or phosphatase.

[0015] In any of the foregoing embodiments, the nucleotide monomer / first labeled conjugate in the first sequencing reagent and / or one or more nucleotide monomers / second labeled conjugates in the second sequencing reagent may have the structure of Formula II:

[0016]

[0017] In any of the foregoing embodiments, the first labels of at least two different nucleotide monomers may be the same or different from each other. In any of the foregoing embodiments, the method may further include a washing step between steps a) and b).

[0018] In any of the foregoing embodiments, the target polynucleotide is immobilized on a surface, such as a solid surface, a soft surface, a hydrogel surface, a microparticle surface, or a combination thereof. In one embodiment, the solid surface is part of a microreactor, and steps a) and b) are performed in the microreactor. In any of the foregoing embodiments, the method is performed at a temperature ranging from about 20°C to about 70°C.

[0019] In any of the foregoing embodiments, multiple rounds of steps a) and b) can be performed using different combinations of the first and second sequencing reagents.

[0020] In any of the foregoing embodiments, the sequence information obtained in step c) may be a degenerate sequence. In one embodiment, at least one additional round of steps a) and b) is performed using a combination of first and second sequencing reagents different from the combination of first and second sequencing reagents in one or more previous rounds of steps a) and b) to obtain at least one additional sequence, and the additional sequence is compared with the degenerate sequence to obtain a non-degenerate sequence.

[0021] In any of the foregoing embodiments, the initial sequence information obtained in step c) may be error-free or contain one or more errors. In one embodiment, at least one additional round of steps a) and b) is performed using a different combination of first and second sequencing reagents than in the previous rounds of steps a) and b) to obtain at least one additional sequence, and the additional sequence is compared with the initial sequence to reduce or eliminate sequence errors.

[0022] In any of the foregoing embodiments, mathematical analysis, algorithms, or methods are used for sequence comparison. In one embodiment, the mathematical analysis, algorithm, or method includes a Markov model or a maximum likelihood method based on a Bayesian scheme.

[0023] In any of the foregoing embodiments, the first sequencing reagent may comprise two different nucleotide monomers / first labeled conjugates, each containing a different nucleotide monomer. In any of the foregoing embodiments, the second sequencing reagent may comprise two different nucleotide monomers / second labeled conjugates, each containing a different nucleotide monomer. In any of the foregoing embodiments, the two nucleotide monomers in the first sequencing reagent may be different from the two nucleotide monomers in the second sequencing reagent.

[0024] In any of the foregoing embodiments, the two nucleotide monomers in the first sequencing reagent and the two nucleotide monomers in the second sequencing reagent may be selected from the group consisting of A, T / U, C, and G deoxyribonucleic acid, and analogues thereof. In one embodiment, the two nucleotide monomers in the first sequencing reagent and the two nucleotide monomers in the second sequencing reagent may be selected from the group consisting of: 1) A and T / U deoxyribonucleic acid in one sequencing reagent and C and G deoxyribonucleic acid in another sequencing reagent; 2) A and G deoxyribonucleic acid in one sequencing reagent and C and T / U deoxyribonucleic acid in another sequencing reagent; and 3) A and C deoxyribonucleic acid in one sequencing reagent and G and T / U deoxyribonucleic acid in another sequencing reagent. In another embodiment, one round of steps a) and b) or at least two rounds of steps a) and b) are performed, one combination of combinations 1)-3) is used in one round of steps a) and b), and another combination of combinations 1)-3) that is different from the combination used in the previous round of steps a) and b) is used in another round of steps a) and b). On the one hand, three rounds of steps a) and b) are performed, each round using a different combination selected from combinations 1)-3). In any of the foregoing embodiments, the sequences obtained from the multiple rounds of steps a) and b) can be compared to obtain non-degenerate sequences and / or reduce or eliminate sequence errors in the non-degenerate sequences.

[0025] In any of the foregoing embodiments, the two nucleotide monomers in the first sequencing reagent and the two nucleotide monomers in the second sequencing reagent may be selected from the group consisting of A, T / U, C, and G ribonucleotides, and analogues thereof. In one embodiment, the two nucleotide monomers in the first sequencing reagent and the two nucleotide monomers in the second sequencing reagent may be selected from the group consisting of: 1) A and T / U ribonucleotides in one sequencing reagent and C and G ribonucleotides in another sequencing reagent; 2) A and G ribonucleotides in one sequencing reagent and C and T / U ribonucleotides in another sequencing reagent; and 3) A and C ribonucleotides in one sequencing reagent and G and T / U ribonucleotides in another sequencing reagent. On one hand, one round of steps a) and b) or at least two rounds of steps a) and b) are performed, using one combination of combinations 1)-3) in one round of steps a) and b), and using another combination of combinations 1)-3) that is different from the combination used in the previous round of steps a) and b) in another round of steps a) and b). On the other hand, at least three rounds of steps a) and b) are performed, each round using different combinations from combinations 1)-3). In any of the foregoing embodiments, the sequences obtained from the multiple rounds of steps a) and b) can be compared to obtain non-degenerate sequences and / or reduce or eliminate sequence errors in the non-degenerate sequences.

[0026] In any of the foregoing embodiments, the first label of the two different nucleotide monomers may be the same, and the second label may be the same as the first label.

[0027] In any of the foregoing embodiments, the first label of the two different nucleotide monomers may be different, and the second label may be the same as the first label.

[0028] In any of the foregoing embodiments, one of the first and second sequencing reagents may contain three different nucleotide monomers / first label conjugates, each containing a different nucleotide monomer, while the other sequencing reagent may contain one nucleotide monomer / second label conjugate, and the three nucleotide monomers in one sequencing reagent may be different from the nucleotide monomers in the other sequencing reagent.

[0029] In any of the foregoing embodiments, the nucleotide monomers in the first and second sequencing reagents may be selected from the group consisting of A, T / U, C, and G deoxyribonucleotides, and analogues thereof. In a specific embodiment, the nucleotide monomers in the first and second sequencing reagents may be selected from the group consisting of: 1) C, G, and T / U deoxyribonucleotides in one sequencing reagent and A deoxyribonucleotides in another sequencing reagent; 2) A, G, and T / U deoxyribonucleotides in one sequencing reagent and C deoxyribonucleotides in another sequencing reagent; 3) A, C, and T / U deoxyribonucleotides in one sequencing reagent and G deoxyribonucleotides in another sequencing reagent; and 4) A, C, and G deoxyribonucleotides in one sequencing reagent and T / U deoxyribonucleotides in another sequencing reagent. In one implementation, steps a) and b) are performed in one round, or at least two rounds, using one combination from combinations 1)-4) in one round of steps a) and b), and using another combination from combinations 1)-4) that is different from the combination used in the previous round of steps a) and b) in another round of steps a) and b). In another implementation, steps a) and b) are performed in three rounds, each round using a different combination selected from combinations 1)-4). In yet another implementation, steps a) and b) are performed in four rounds, each round using a different combination selected from combinations 1)-4). In any of the foregoing implementations, sequences obtained from multiple rounds of steps a) and b) can be compared to obtain non-degenerate sequences and / or reduce or eliminate sequence errors in the non-degenerate sequences.

[0030] In any of the foregoing embodiments, the nucleotide monomers in the first and second sequencing reagents may be selected from the group consisting of A, T / U, C, and G ribonucleotides, and analogues thereof. In one embodiment, the nucleotide monomers in the first and second sequencing reagents may be selected from the group consisting of: 1) C, G, and T / U ribonucleotides in one sequencing reagent and A ribonucleotides in another sequencing reagent; 2) A, G, and T / U ribonucleotides in one sequencing reagent and C ribonucleotides in another sequencing reagent; 3) A, C, and T / U ribonucleotides in one sequencing reagent and G ribonucleotides in another sequencing reagent; and 4) A, C, and G ribonucleotides in one sequencing reagent and T / U ribonucleotides in another sequencing reagent. In one embodiment, one round of steps a) and b) or at least two rounds of steps a) and b) are performed, one combination of combinations 1)-4) is used in one round of steps a) and b), and another combination of combinations 1)-4) that is different from the combination used in the previous round of steps a) and b) is used in another round of steps a) and b). In one specific implementation, steps a) and b) are performed in at least three rounds, each round using a different combination from combinations 1)-4). In another implementation, steps a) and b) are performed in at least four rounds, each round using a different combination from combinations 1)-4). In any of the foregoing implementations, sequences obtained from multiple rounds of steps a) and b) can be compared to obtain non-degenerate sequences and / or reduce or eliminate sequence errors in the non-degenerate sequences.

[0031] In any of the foregoing embodiments, approximately 250bp, approximately 350bp, approximately 400bp, approximately 450bp, approximately 500bp, approximately 550bp, approximately 600bp, approximately 650bp, approximately 700bp, approximately 750bp, approximately 800bp, approximately 850bp, approximately 900bp, approximately 950bp, approximately 1000bp, approximately 1050bp, approximately 1100bp, approximately 1150bp, approximately 1200bp, approximately 1250bp, approximately 1300bp, and approximately 1350bp can be obtained. p, approximately 1400bp, approximately 1450bp, approximately 1500bp, approximately 1550bp, approximately 1600bp, approximately 1650bp, approximately 1700bp, approximately 1750bp, approximately 1800bp, approximately 1850bp, approximately 1900bp, approximately 1950bp, approximately 2000bp, approximately 2050bp, approximately 2100bp, approximately 2150bp, approximately 2200bp, approximately 2250bp, approximately 2300bp, approximately 2350bp, or approximately 2400 base pairs of read length.

[0032] In any of the foregoing embodiments, a code accuracy of at least approximately 95% can be obtained. In any of the foregoing embodiments, the target polynucleotide can be a single-stranded polynucleotide.

[0033] On the other hand, this document discloses a method for obtaining sequence information of a target polynucleotide, the method comprising: a) providing a first sequencing reagent to the target polynucleotide in the presence of a first polynucleotide replication catalyst, wherein the first sequencing reagent comprises two different nucleotide monomers each conjugated to a first label, and the nucleotide monomer / first label conjugate is substantially non-fluorescent until the nucleotide monomer is incorporated into the target polynucleotide according to its complementarity with the target polynucleotide; and b) providing a second sequencing reagent to the target polynucleotide in the presence of a second polynucleotide replication catalyst, wherein the second sequencing reagent comprises two different nucleotide monomers each conjugated to a second label, and the nucleotide monomer / second label conjugate is substantially non-fluorescent until the nucleotide monomer is incorporated into the target polynucleotide according to its complementarity with the target polynucleotide, and wherein the second sequencing reagent is provided subsequently after the first sequencing reagent is provided; and c) by steps a) and b). After incorporating a nucleotide monomer into a polynucleotide, the fluorescence emission caused by a first label and a second label is detected to obtain at least a portion of the target polynucleotide's sequence information. The nucleotide monomers in the first and second sequencing reagents are selected from the group consisting of: 1) an adenine (A) nucleotide monomer and a thymine (T) / uracil (U) nucleotide monomer in one sequencing reagent and a cytosine (C) nucleotide monomer and a guanine (G) nucleotide monomer in another sequencing reagent; 2) an adenine (A) nucleotide monomer and a guanine (G) nucleotide monomer in one sequencing reagent and a cytosine (C) nucleotide monomer and a thymine (T) / uracil (U) nucleotide monomer in another sequencing reagent; and 3) an adenine (A) nucleotide monomer and a cytosine (C) nucleotide monomer in one sequencing reagent and a guanine (G) nucleotide monomer and a thymine (T) / uracil (U) nucleotide monomer in another sequencing reagent. In one embodiment, the first label of the two different nucleotide monomers in step a) and the second label of the two different nucleotide monomers in step b) are the same label. In another embodiment, the first marker comprises two distinct markers, wherein one of the first markers is identical to one of the second markers, and the other of the first markers is identical to the other of the second markers. In any of the foregoing embodiments, multiple rounds of steps a) and b) are performed, each round using a combination selected from combinations 1)-3). In another embodiment, obtaining at least two or three sets of sequence information in step c) comprises: using combination 1) to perform multiple rounds of steps a) and b) in a first sequencing reaction volume to obtain a first set of sequence information, using combination 2) to perform multiple rounds of steps a) and b) in a second sequencing reaction volume to obtain a second set of sequence information, and / or using combination 3) to perform multiple rounds of steps a) and b) in a third sequencing reaction volume to obtain a third set of sequence information. In one embodiment, the first, second, and third sets of sequence information are obtained in parallel from separate sequencing reaction volumes.In another embodiment, the first, second, and third sets of sequence information are obtained sequentially from the same sequencing reaction volume, and the product of the previous sequencing reaction is removed before starting the next sequencing reaction. In any of the foregoing embodiments, the method further includes comparing at least two or three sets of sequence information to reduce or eliminate sequence errors. In one embodiment, the comparison indicates that there are no errors in the obtained target polynucleotide sequence when at least two or three sets of sequence information are consistent with each other. In another embodiment, the comparison indicates that there are errors in the obtained target polynucleotide sequence when at least two or three sets of sequence information contain differences in at least one nucleotide residue of the target polynucleotide sequence. In one embodiment, the method further includes correcting at least one nucleotide residue in the obtained target polynucleotide sequence such that, after correction, at least two or three sets of sequence information are consistent with each other.

[0034] In another aspect, this document discloses a method for obtaining sequence information of a target polynucleotide, the method comprising: a) providing a first sequencing reagent to the target polynucleotide in the presence of a first polynucleotide replication catalyst, wherein the first sequencing reagent comprises three different nucleotide monomers each conjugated to a first label, and the nucleotide monomer / first label conjugate is substantially non-fluorescent until the nucleotide monomer is incorporated into the target polynucleotide according to its complementarity with the target polynucleotide; and b) providing a second sequencing reagent to the target polynucleotide in the presence of a second polynucleotide replication catalyst, wherein the second sequencing reagent comprises a nucleotide monomer conjugated to a second label, and the nucleotide monomer / second label conjugate is substantially non-fluorescent until the nucleotide monomer is incorporated into the target polynucleotide according to its complementarity with the target polynucleotide, and wherein the second sequencing reagent is provided before or after the provision of the first sequencing reagent; and c) detecting the first label and the second label after incorporation of the nucleotide monomer into the polynucleotide in steps a) and b). The fluorescence emission of the sequencing reagent yields sequence information of at least a portion of the target polynucleotide, wherein the nucleotide monomers in the first and second sequencing reagents are selected from the group consisting of: 1) a sequencing reagent containing cytosine (C) nucleotide monomers, guanine (G) nucleotide monomers, thymine (T) / uracil (U) nucleotide monomers, and adenine (A) nucleotide monomers; 2) a sequencing reagent containing adenine (A) nucleotide monomers, guanine (G) nucleotide monomers, thymine (T) / uracil (U) nucleotide monomers, and cytosine (C) nucleotide monomers; and 3) a sequencing reagent containing adenine (A) nucleotide monomers, cytosine (C) nucleotide monomers, thymine (T) / uracil (U) nucleotide monomers, and guanine (G) nucleotide monomers; and 4) a sequencing reagent containing adenine (A) nucleotide monomers, cytosine (C) nucleotide monomers, and guanine (G) nucleotide monomers, and another sequencing reagent containing thymine (T) / uracil (U) nucleotide monomers. In one embodiment, the first label of the three different nucleotide monomers in step a) and the second label of one nucleotide monomer in step b) are the same label. In any of the foregoing embodiments, multiple rounds of steps a) and b) are performed, each round using a combination selected from combinations 1)-4). In one embodiment, obtaining at least two, three, or four sets of sequence information in step c) is achieved by: performing multiple rounds of steps a) and b) in a first sequencing reaction volume using combination 1) to obtain a first set of sequence information; performing multiple rounds of steps a) and b) in a second sequencing reaction volume using combination 2) to obtain a second set of sequence information; performing multiple rounds of steps a) and b) in a third sequencing reaction volume using combination 3) to obtain a third set of sequence information; and / or performing multiple rounds of steps a) and b) in a fourth sequencing reaction volume using combination 4) to obtain a fourth set of sequence information.In one implementation, the first, second, third, and fourth sets of sequence information are obtained in parallel from separate sequencing reaction volumes. In another implementation, the first, second, third, and fourth sets of sequence information are obtained sequentially from the same sequencing reaction volume, and the product of the previous sequencing reaction is removed before starting the next sequencing reaction. In any of the foregoing implementations, the method further includes comparing at least two, three, or four sets of sequence information to reduce or eliminate sequence errors. In one implementation, the comparison indicates that there are no errors in the obtained target polynucleotide sequence when at least two, three, or four sets of sequence information are consistent with each other. On the one hand, when using a single-color sequencing method, at least three sets of sequence information are required to monitor sequencing errors. On the other hand, when using a two-color sequencing method, only two sets of sequence information are required to detect sequencing errors because the information from the two fluorescent labels provides an additional set of information for sequence comparison.

[0035] In another embodiment, when at least two, three, or four sets of sequence information contain differences at at least one nucleotide residue in the target polynucleotide sequence, the comparison indicates an error in the obtained target polynucleotide sequence. In one embodiment, the method further includes correcting at least one nucleotide residue in the obtained target polynucleotide sequence such that, after correction, at least two, three, or four sets of sequence information are consistent with each other. On one hand, at least one nucleotide residue is corrected by deletion or insertion at the erroneous position to achieve the correct sequence. On the other hand, each insertion at the erroneous position extends the sequence by at least one nucleotide, and sequence information from one or more rounds of sequencing is compared with the extended sequence to achieve the corrected sequence. On the other hand, each deletion at the erroneous position shortens the sequence by at least one nucleotide, and sequence information from one or more rounds of sequencing is compared with the shortened sequence to achieve the corrected sequence.

[0036] In another aspect, this document discloses a kit or system for obtaining sequence information of a target polynucleotide, the kit or system comprising: a) a first sequencing reagent containing at least two different nucleotide monomers / first label conjugates, said at least two different nucleotide monomers / first label conjugates being substantially non-fluorescent until the nucleotide monomers are incorporated into the target polynucleotide according to their complementarity; and b) a second sequencing reagent containing one or more nucleotide monomers / second label conjugates, said one or more nucleotide monomers / second label conjugates being substantially non-fluorescent until the nucleotide monomers are incorporated into the polynucleotide according to their complementarity, at least one of the one or more nucleotide monomers being different from the nucleotide monomers present in the first sequencing reagent; and c) a detector for detecting fluorescence emission caused by the first label and the second label after the nucleotide monomers are incorporated into the polynucleotide. In one embodiment, the kit or system further comprises a first polynucleotide replication catalyst and / or a second polynucleotide replication catalyst. In any of the foregoing embodiments, the first and / or second labels are conjugated to the terminal phosphate group of the nucleotide monomer. In one embodiment, the kit or system further comprises an activating enzyme for releasing the first and / or second labels from the terminal phosphate group of the nucleotide monomer. In any of the foregoing embodiments, the kit or system may further include a solid surface on which the target polynucleotide is configured to be immobilized. In one embodiment, the solid surface is part of a microreactor.

[0037] In any of the foregoing embodiments, the kit or system further includes a tool for obtaining at least one sequence of the target polynucleotide based on fluorescence emission caused by a first label and a second label after incorporating a nucleotide monomer into the polynucleotide. In one embodiment, the tool includes a computer-readable medium containing executable instructions that, when executed, can obtain at least a portion of the target polynucleotide's sequence information based on fluorescence emission caused by a first label and a second label after incorporating a nucleotide monomer into the polynucleotide.

[0038] In any of the foregoing embodiments, the kit or system may further include tools for comparing multiple sequences to obtain non-degenerate sequences and / or reducing or eliminating sequence errors in the non-degenerate sequences. In one embodiment, the tool includes a computer-readable medium containing executable instructions that, when executed, can compare sequences to obtain non-degenerate sequences and / or reduce or eliminate sequence errors in the non-degenerate sequences.

[0039] On the one hand, the present disclosure provides a method for correcting sequencing information errors, which includes: (a) performing parameter estimation based on sequencing signals from one or more reference polynucleotides during a sequencing reaction and the known nucleic acid sequence of the reference polynucleotide, and obtaining information on leading and / or lagging out-of-phase phenomena of the sequencing reaction using the parameter estimation; (b) obtaining sequencing signals from a target polynucleotide during the sequencing reaction; (c) calculating a secondary leading amount of the target polynucleotide based on the information obtained from step (a) and the sequencing signals obtained from step (b); (d) calculating an out-of-phase amount of the target polynucleotide based on the sequencing signals obtained from step (b) and the secondary leading amount of step (c); (e) using the out-of-phase amount to correct the sequencing signals obtained from step (b) to generate predicted sequencing signals of the target polynucleotide; (f) repeating steps (c) to (e) for one or more rounds, wherein the predicted sequencing signals from the i-th round are used to calculate the secondary leading amount of the target polynucleotide in the (i + 1)-th round until the predicted sequencing signals of the target polynucleotide from the j-th round are mathematically convergent, where i and j are integers and 1 ≤ i < i + 1 ≤ j. In one embodiment, the secondary leading phenomenon refers to an unexpected nucleotide extension occurring at a residue of the target polynucleotide during sequencing, and the unexpected extension is further extended by a nucleotide other than the next residue. In another embodiment, the out-of-phase amount includes a change in the sequencing result due to leading and / or lagging out-of-phase phenomena during sequencing.

[0040] In any of the foregoing embodiments, the parameter estimation in step (a) may include obtaining an attenuation coefficient. In any of the foregoing embodiments, the parameter estimation in step (a) may further include obtaining an offset. In any of the foregoing embodiments, the parameter estimation in step (a) may include obtaining unit signal information. In any of the foregoing embodiments, the parameter estimation in step (a) may include obtaining a leading coefficient and / or a lagging coefficient for each nucleotide or nucleotide combination.

[0041] In any of the foregoing embodiments, the method includes obtaining information on leading and / or lagging out-of-phase phenomena for each round of the sequencing reaction when multiple rounds of sequencing reactions are performed.

[0042] On the other hand, the present disclosure provides a method for correcting sequencing information errors, which includes: (a) performing parameter estimation based on sequencing signals from one or more reference polynucleotides and the known nucleic acid sequence of the reference polynucleotide during a sequencing reaction; (b) obtaining sequencing signals from a target polynucleotide during the sequencing reaction; (c) calculating a secondary leading amount of the target polynucleotide based on the information obtained from the leading or lagging phase obtained by parameter estimation in step (a) and the sequencing signals obtained from step (b); (d) calculating the phase shift amount of the target polynucleotide based on the sequencing signals obtained from step (b) and the secondary leading amount of step (c); (e) using the phase shift amount to correct the sequencing signals obtained from step (b) to generate predicted sequencing signals of the target polynucleotide; (f) repeating steps (c) to (e) for one or more rounds, wherein the predicted sequencing signals from the i-th round are used to calculate the secondary leading amount of the target polynucleotide in the (i + 1)-th round until the predicted sequencing signals of the target polynucleotide from the j-th round are mathematically convergent, where i and j are integers and 1 ≤ i < i + 1 ≤ j. On the one hand, parameter estimation includes obtaining a leading amount, a lagging amount, an attenuation coefficient, and / or an offset based on the sequencing signals from the reference polynucleotide and the known nucleic acid sequence of the reference polynucleotide. On the other hand, the secondary leading phenomenon refers to an unexpected nucleotide extension occurring at a residue of the target polynucleotide during sequencing, and the unexpected extension is further extended by a nucleotide other than the next residue. On yet another hand, the phase shift amount includes a change in the sequencing result due to leading and / or lagging phase shift phenomena during sequencing.

[0043] On yet another hand, the present disclosure provides a method for correcting the leading amount during sequencing, which includes: obtaining sequencing signals from a target polynucleotide during the sequencing reaction, the sequencing signals corresponding to the sequence of the target polynucleotide; and optionally using parameter estimation to correct the sequencing signals from the target polynucleotide with the secondary leading amount caused by the secondary leading phenomenon. In one embodiment, the secondary leading phenomenon refers to an unexpected nucleotide extension occurring at a residue of the target polynucleotide during sequencing, and the unexpected extension is further extended by a nucleotide other than the next residue.

[0044] On the one hand, the sequencing signals from the target polynucleotide include a primary leading amount caused by the primary leading phenomenon, where the primary leading phenomenon refers to an unexpected nucleotide extension occurring at a residue of the target polynucleotide during sequencing.

[0045] In any of the foregoing embodiments, if the sequencing signal from a specific nucleotide residue of the target polynucleotide is close to the unit signal, the sequencing signal can be corrected using a secondary lead. In any of the foregoing embodiments, the deviation of the sequencing signal intensity from the unit signal intensity is within about 60%, about 50%, about 40%, about 30%, about 20%, about 10%, or about 5%.

[0046] In any of the foregoing embodiments, when the nth sequencing signal is obtained, the method may include: comparing the sequencing signal of a reference polynucleotide with a known sequence of the reference polynucleotide to identify errors during sequencing, and methods for correcting errors; using the sequencing signal of the target polynucleotide preceding n and the error correction methods to obtain a corrected sequencing signal, for example, by feeding the sequencing signal of the target polynucleotide preceding n into the error correction methods; and determining whether a secondary lead exists at residue n by comparing the sequencing signal of the target polynucleotide at residue n with the corrected sequencing signal.

[0047] In any of the foregoing embodiments, sequencing may include adding one or more sequencing reagents to a reaction solution, wherein said one or more sequencing reagents optionally comprise nucleotides and / or enzymes. In any of the foregoing embodiments, one, two, or three types of nucleotides may be added in each sequencing reaction. In any of the foregoing embodiments, the sequencing reaction involves the open or unclosed 3' end of a polynucleotide. In any of the foregoing embodiments, the added nucleotides may comprise one or more of A, G, C, and T, or one or more of A, G, C, and U. In any of the foregoing embodiments, the detected sequencing signal may include an electrical signal, a bioluminescent signal, a chemiluminescent signal, or any combination thereof.

[0048] In any of the foregoing embodiments, parameter estimation may include: inferring an ideal signal h from a reference polynucleotide sequence, calculating a phase-mismatched signal s and a predicted raw sequencing signal p based on preset parameters, and calculating a correlation coefficient c between p and the actual raw sequencing signal f. On one hand, the method also includes using an optimization method to find a set of parameters such that the correlation coefficient c reaches an optimal value. On the other hand, this set of parameters may include a lead coefficient or magnitude, a lag coefficient or magnitude, a decay coefficient, an offset, a unit signal, or any combination thereof.

[0049] In any of the foregoing embodiments, during sequencing, two sets of reaction solutions may be provided, each containing one or more nucleotides different from the other, and one set of reaction solutions may be provided for each sequencing reaction. Alternatively, the two sets of reaction solutions may be used alternately for the sequencing reaction. In any of the foregoing embodiments, sequencing of the target polynucleotide and the reference polynucleotide may be performed simultaneously.

[0050] In any of the foregoing embodiments, a reference polynucleotide can be used for parameter estimation to obtain one or more of the following parameters for the sequencing reaction: lead factor or amount, lag factor or amount, attenuation factor, offset, and unit signal. In any of the foregoing embodiments, one or more parameters of the sequencing reaction obtained through parameter estimation can be used to correct the signal of the target polynucleotide. In any of the foregoing embodiments, the target polynucleotide may comprise a label containing a known sequence and / or a known amount of nucleotides, and the known sequence and / or known amount of nucleotides are used to generate the unit signal of the sequencing reaction. In any of the foregoing embodiments, the unit signal may be different at each sampling point, for example, at each nucleotide residue of the target polynucleotide.

[0051] In another aspect, this document discloses a computer-readable medium including instructions for correcting sequencing information errors. In one aspect, the instructions include: a) receiving sequencing information of a target polynucleotide and a reference polynucleotide; and b) correcting the sequencing information of the target polynucleotide using any of the methods disclosed herein for correcting sequencing information.

[0052] On the other hand, a computer system for sequencing is provided, the system including the computer-readable medium disclosed herein. Attached Figure Description

[0053] Figure 1 The method for correcting sequence data errors is shown.

[0054] Figure 2 The data distribution for groups 1 through 5 is shown in violin and box plots. Black represents encoding accuracy, and gray represents decoding accuracy. Groups 1 through 5 are presented from left to right in the sequence.

[0055] Figure 3 A frequency distribution histogram is displayed, showing the number of signals modified during decoding for each of the 5000 sequence data.

[0056] Figure 4 This displays the correlation between the number of signals that are erroneous during encoding and the number of signals that are erroneously modified during decoding. The horizontal axis represents the number of signals that are erroneous during encoding, and the vertical axis represents the correlation between the number of signals that are erroneously modified during decoding. The grayscale of the color indicates the proportion of times that point is counted out of all sequences.

[0057] Figure 5A -C shows that the fluorogenic performance of TPLFN can be improved by altering the fluorophore structure.

[0058] Figure 6 The MALDI-TOF mass spectra of purified TPLFN are shown.

[0059] Figure 7 The excitation and emission spectra of TG (Tokyo Green) are shown.

[0060] Figure 8 The emission spectra of TG (Tokyo Green), Me-FAM, and Me-HCF are shown under the same conditions (2 μM, pH 8.3, TE buffer, calculated using area normalization).

[0061] Figure 9 The absorption spectra of TPLFN(TG-dA4P) before and after enzyme digestion are shown.

[0062] Figure 10 The emission spectra of TPLFN(TG-dA4P) before and after enzyme digestion are shown.

[0063] Figure 11 The dynamic mode is displayed.

[0064] Figure 12 The differences in reaction rates among the four substrates were shown.

[0065] Figure 13 This demonstrates substrate competition.

[0066] Figure 14 The linearity of the homopolymer length with respect to the signal is shown.

[0067] Figure 15 A shows a homopolymer consisting only of T. Figure 15 B shows a homopolymer consisting of four repeating TCs.

[0068] Figure 16 The temperature-dependent activity of Bst was demonstrated.

[0069] Figure 17 The synthesis of N-(5-(2-bromoacetamido)pentyl)acrylamide is shown.

[0070] Figure 18 Primer grafting is shown.

[0071] Figure 19 The difference in contact angle between the glass and the BPAM-coated surface is shown.

[0072] Figure 20 The design of the ECCS library is shown. Figure 20 a shows the ECCS library prior to solid-phase PCR. Figure 20 b shows the ECCS library prior to solid-phase PCR. Figure 20 c shows the ECCS library after annealing the sequencing primers.

[0073] Figure 21 The template preparation process is shown.

[0074] Figure 22 The gel electrophoresis results of the PCR products are shown. Lane 1 is the marker (Transgene, 100bp Plus II DNA Ladder); lanes 2 and 3 are two 200bp templates (L718-208 (330bp) and L10115-201 (323bp) respectively); lanes 4-6 are three 300bp templates (L718-308 (430bp), L4418-305 (427bp), and L10115-301 (423bp) respectively); lanes 7-9 are three 500bp templates (L501-500 (622bp), L30501-500 (622bp), and L46499-500 (622bp) respectively).

[0075] Figure 23 The solid-phase PCR process is shown.

[0076] Figure 24 The top image shows heatmaps of PCR product density across different lanes and positions. The x-axis of each graph represents the four different lanes of the chip; the y-axis represents the five different imaging positions within each lane. Colors from black to green indicate PCR product density from low to high. The bottom image shows PCR product density for different templates. The x-axis represents different experimental groups in solid-phase PCR; the y-axis represents the average density per lane of the chip.

[0077] Figure 25 According to one implementation, the top figure shows the sequencing instrument, the bottom left figure shows typical fluorescence reaction kinetics, and the kinetics for each reaction cycle throughout the sequencing process.

[0078] Figure 26 The phase loss process is shown.

[0079] Figure 27 The image shows simulated sequencing signals (left) and DNA concentration distribution at different locations (right). Color bars (grayscale bars): DNA proportion. Figure 27 (a and 27b) Impurities: 0; Reaction time: 300. Figure 27 Impurities (c and 27d): 0.003; Reaction time: 300. Figure 27 (e and 27f) Impurities: 0; Reaction time: 100.

[0080] Figure 28The top figure illustrates the One Pass, More Stop principle. The bottom figure shows the distribution and flux matrices and their relationship. The lead ε and lag λ coefficients are set to 2% and 1%, respectively. These coefficients are relatively large to show the obvious effect of phase misalignment, rather than as an estimate of experimental data.

[0081] Figure 29 A simplified flowchart of the correction algorithm is shown.

[0082] Figure 30 The application of the correction algorithm is shown.

[0083] Figure 31 The phase loss correction algorithm is shown.

[0084] Figure 32 This shows the effect of the phase loss coefficient on the condition number of (flux matrix) T.

[0085] Figure 33 This demonstrates the impact of phase loss coefficient deviation on signal correction.

[0086] Figure 34 It shows that global white noise reduces the accuracy of the correction signal and makes subsequent loops more error-prone.

[0087] Figure 35 This shows the number of error-free cycles after phase correction with a given phase decoupling coefficient and global white noise.

[0088] Figure 36 This demonstrates the effect of signal anomalies in certain loops.

[0089] Figure 37A The change trajectory of each coefficient in the phase loss coefficient estimation algorithm is shown. Figure 37B The phase loss coefficients in multiple rounds of sequencing were summarized. Figure 37C The relationship between phase loss coefficient and sequencing reaction time is shown.

[0090] Figure 38 This illustrates phase loss in high-throughput DNA sequencing. Squares represent nucleotides in the template DNA, and circles represent nucleotides that make up the nascent DNA strand. Diagonally lined patterns represent sequencing primer regions, and patterns filled with white or gray represent different types of nucleotides.

[0091] Figure 39 The primary and secondary leading phenomena are shown.

[0092] Figure 40 The display shows that Level 3 advance has no longer occurred.

[0093] Figure 41 The basic process of parameter estimation is shown.

[0094] Figure 42 The basic process of signal correction is shown.

[0095] Figure 43 The raw 2+2 monochromatic sequencing signal is shown.

[0096] Figure 44 The changes in each parameter during the parameter estimation process of the raw signal from monochrome 2+2 sequencing are shown.

[0097] Figure 45 The raw and phase-depleted signals from monochrome 2+2 sequencing are shown.

[0098] Figure 46 The iterative steps in signal correction of monochrome 2+2 sequencing signals are shown.

[0099] Figure 47 The raw signal from a single two-color 2+2 sequencing run is shown.

[0100] Figure 48 The changes in all parameters during the parameter estimation process of two-color 2+2 sequencing are shown.

[0101] Figure 49 The raw and phase-delayed signals from primary two-color 2+2 sequencing are shown.

[0102] Figure 50 The iterative steps in signal correction during two-color 2+2 sequencing are shown.

[0103] Figure 51 The statistical results of signal correction for multiple monochromatic 2+2 sequencing are shown.

[0104] Figure 52 According to one aspect of the present invention, the principle of degenerate base fluorescence gene sequencing is shown.

[0105] Figure 53 According to one aspect of the invention, degenerate base-calling results are shown.

[0106] Figure 54 According to one aspect of the present invention, an information communication model for ECC sequencing is shown.

[0107] Figure 55 According to one aspect of the invention, a sequence decoding result using dynamic programming is shown.

[0108] Figure 56 According to one aspect of the invention, display decoding improves the accuracy of ECC sequencing.

[0109] Figure 57According to one aspect of the invention, a cyclic range distribution of three base combinations is shown.

[0110] Figure 58 This shows an example of the layer and node traversal order of the rating matrix structure.

[0111] Figure 59 According to one aspect of the invention, a state transition network for a hidden Markov model of ECC decoding is shown.

[0112] Figure 60 The diagram shows a simulated distribution of accuracy before and after decoding, according to one aspect of the invention.

[0113] Figure 61 An example decoding result is shown. Detailed Implementation

[0114] The following provides a detailed description of one or more embodiments of the claimed subject matter, along with accompanying drawings illustrating the principles of the claimed subject matter. The claimed subject matter is described in conjunction with such embodiments, but is not limited to any specific embodiment. It should be understood that the claimed subject matter can be embodied in various forms and encompasses many alternatives, modifications, and equivalents. Therefore, the specific details disclosed herein should not be construed as limiting, but rather serve as the basis for the claims and as a representative basis for teaching those skilled in the art to adopt the claimed subject matter in virtually any suitable detailed system, structure, or manner. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. These details are provided for illustrative purposes only, and the claimed subject matter may be practiced according to the claims without some or all of these specific details. It should be understood that other embodiments may be used and structural changes may be made without departing from the scope of the claimed subject matter. It should be understood that the various features and functions described in one or more individual embodiments are not limited to their application to the specific embodiments in which they are described. Rather, they may be applied individually or in some combination to one or more other embodiments of this disclosure, whether or not such embodiments are described, and whether or not such features are presented as part of the described embodiments. For clarity, known technical materials in the relevant technical fields of the claimed subject matter have not been described in detail to avoid unnecessarily obscuring the claimed subject matter.

[0115] All technical terms, symbols, and other technical and scientific terms used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains, unless otherwise defined. In some instances, for clarity and / or convenience of reference, terms with their commonly understood meanings are defined herein, and such definitions are incorporated herein, but this should not necessarily be construed as indicating a material difference from the meanings commonly understood in the art. Many of the techniques and procedures described or referenced herein are known to those skilled in the art and are commonly used when employing conventional methods.

[0116] All publications referenced in this application, including patent documents, scientific papers, and databases, are incorporated herein by reference in their entirety for all purposes, just as each individual publication is incorporated individually by reference. Where a definition set forth herein is contrary to or otherwise inconsistent with a definition set forth in a patent, patent application, publication, or other publication incorporated herein by reference, the definition set forth herein shall prevail. References to publications or documents are not intended as an admission that any of them is applicable prior art, nor do they constitute any admission of the content or dates of such publications or documents.

[0117] Unless otherwise stated, all headings are for the reader's convenience and should not be used to limit the meaning of the text following them.

[0118] Unless otherwise indicated, the practices provided will employ conventional techniques and descriptions of organic chemistry, polymer technology, molecular biology (including recombinant technologies), cell biology, biochemistry, and sequencing technologies, which are understood by one of ordinary skill in the art. Such conventional techniques include peptide and protein synthesis and modification, polynucleotide synthesis and modification, polymer array synthesis, polynucleotide hybridization and ligation, and detection of hybridization using labels. Specific descriptions of suitable techniques can be obtained by referring to the examples herein. However, other equivalent conventional procedures may of course be used. Such routine techniques and descriptions can be found in standard laboratory manuals, such as Green et al., Genome Analysis: A Laboratory Manual Series (Volumes I-IV) (1999); Weiner, Gabriel, and Stephens, Genetic Variation: A Laboratory Manual (2007); Dieffenbach and Dveksler, PCR Primer: A Laboratory Manual (2003); Bowtell and Sambrook, DNA Microarrays: A Molecular Cloning Manual (2003); Mount, Bioinformatics: Sequence and Genome Analysis (2004); Sambrook and Russell, Condensed Protocols from Molecular Cloning: A Laboratory Manual (2006); and Sambrook and Russell, Molecular Cloning: A Laboratory Manual (2002) (all from ColdSpring Harbor Laboratory Press); Ausubel et al., Current Protocols in Molecular Biology (1987); and T. Brown, Essential Molecular Biology. Biology (1991), IRL Press; Goeddel, ed., Gene Expression Technology (1991), Academic Press; A. Bothwell et al., ed., Methods for Cloning and Analysis of Eukaryotic Genes (1990), Bartlett Publ.; M.Kriegler, Gene Transfer and Expression (1990), Stockton Press; R. Wu et al., eds., Recombinant DNAMethodology (1989), Academic Press; M. McPherson et al., PCR: A Practical Approach (1991), IRL Press at Oxford University Press; Stryer, Biochemistry (4th Edition) (1995), WH Freeman, New York NY; Gait, Oligonucleotide Synthesis: APractical Approach (2002), IRL Press, London; Nelson and Cox, Lehninger, Principles of Biochemistry (2000) 3rd ed., WHFreeman Pub., New York, NY; Berg, et al., Biochemistry (2002) 5th ed., WHFreeman Pub., New York, NY; D. Weir & C. Blackwell, ed., Handbook of Experimental Immunology(1996),Wiley-Blackwell; Cellular and Molecular Immunology (A. Abbas et al., WBSaunders Co. 1991, 1994); Current Protocols in Immunology (J. Coligan et al., eds., 1991). All references cited are incorporated herein by reference in their entirety for all purposes.

[0119] Throughout this disclosure, all aspects of the claimed subject matter are presented in a scope format. It should be understood that this scope format is for convenience and brevity only and should not be construed as a rigid limitation on the scope of the claimed subject matter. Therefore, the scope description should be considered to have specifically disclosed all possible sub-scopes, as well as individual numerical values ​​within that scope. For example, where a range of values ​​is provided, it should be understood that every intermediate value between the upper and lower limits of that range, as well as any other specified value or intermediate value within that specified range, is covered within the claimed subject matter. The upper and lower limits of these smaller ranges may be independently included within those smaller ranges and are also covered within the claimed subject matter, subject to any express exclusions within the specified scope. When the stated scope includes one or both of the stated limits, the scope extending beyond any or both of those included limits is also included within the claimed subject matter. This application is independent of the breadth of the scope. For example, a description of a range such as from 1 to 6 should be considered to specifically disclose subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as individual numerical values ​​within that range, such as 1, 2, 3, 4, 5, and 6.

[0120] I. Definition

[0121] Unless the context clearly indicates otherwise, as used herein, the singular forms “a / an” and “the” include plural indicators. For example, “a / an” means “at least one” or “one or more”. It should be understood that aspects and variations described herein include “consisting of” and / or “substantially composed of” aspects and variations.

[0122] As used herein, the term "about" refers to a common range of error for a corresponding value that is readily known to those skilled in the art. References to "about" a value or parameter herein include (and describe) implementations for said value or parameter itself. For example, a description of "about X" includes a description of "X" itself.

[0123] The terms “polynucleotide,” “oligonucleotide,” “nucleic acid,” and “nucleic acid molecule” are used interchangeably herein to refer to polymeric forms of nucleotides of any length, and include ribonucleotides, deoxyribonucleotides, and their analogues or mixtures. The terms include triple-stranded, double-stranded, and single-stranded deoxyribonucleic acid (“DNA”), and triple-stranded, double-stranded, and single-stranded ribonucleic acid (“RNA”). They also include polynucleotides in their unmodified forms, as well as those modified, for example by alkylation and / or by end-capping. More specifically, the terms “polynucleotide,” “oligonucleotide,” “nucleic acid,” and “nucleic acid molecule” include polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D-ribose), including tRNA, rRNA, hRNA, and mRNA (whether spliced ​​or unspliced), any other type of polynucleotide that is an N- or C-glycoside of a purine or pyrimidine base, and other polymers containing a non-nucleotide backbone, such as polyamides (e.g., peptide nucleic acids (“PNA”)) and polymorpholino polymers (commercially available from Anti-Virals, Inc., Corvallis, OR, as with Neugene), as well as other synthetic sequence-specific nucleic acid polymers, provided that the polymer contains nucleic acid bases in a structure that allows base pairing and base stacking, such as in the construction of DNA and RNA. Therefore, these terms include, for example, 3'-deoxy-2',5'-DNA, oligodeoxyribonucleotides N3' to P5' phosphoramidites, 2'-O-alkyl-substituted RNA, hybrids between DNA and RNA or between PNA and DNA or RNA; they also include modifications of known types, such as labeling, alkylation; "capping"; substitution of one or more nucleotides with analogs; internucleotide modifications, such as modifications with uncharged bonds (e.g., methylphosphonates, phosphate triesters, phosphoramidites, carbamates, etc.); and modifications with negatively charged bonds (e.g., thiophosphates, dithiophosphates, etc.). Modifications include: phosphodiester bonds; modifications with positively charged linkages (e.g., aminoalkylphosphamide esters, aminoalkylphosphotriesters); modifications containing side-linked portions, such as proteins (including enzymes (e.g., nucleases), toxins, antibodies, signal peptides, poly-L-lysine, etc.); modifications with intercalating agents (e.g., acridine, psoralen, etc.); modifications containing chelates (e.g., chelates of metals, radioactive metals, boron, metal oxides, etc.); modifications containing alkylating agents; modifications with modified linkages (e.g., α-anomeric nucleic acids, etc.); and unmodified forms of polynucleotides or oligonucleotides. Nucleic acids typically contain phosphodiester bonds, but in some cases may include nucleic acid analogs with an alternative backbone, such as phosphoramide, dithiophosphate, or methylphosphorimide linkages; or peptide nucleic acid backbones and linkages.Other nucleic acid analogs include those with a bicyclic structure, including locked nucleic acids, positively charged backbones, nonionic backbones, and nonribose backbones. Molecular stability can be increased by modifying the ribose-phosphate backbone; for example, PNA:DNA hybrids exhibit greater stability in certain environments. The terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" can include any suitable length, such as at least 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 1,000, or more nucleotides.

[0124] It should be understood that the terms “nucleoside” and “nucleotide” as used herein include not only known purine and pyrimidine bases but also other modified heterocyclic bases. Such modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, or other heterocycles. Modified nucleosides or nucleotides may also include modifications to the glycosyl moiety, for example, where one or more hydroxyl groups are substituted with halogens, aliphatic groups, or functionalized to ethers, amines, etc. The term “nucleotide unit” is intended to encompass both nucleosides and nucleotides.

[0125] The terms "complementary" and "substantially complementary" encompass hybridization or base pairing, or the formation of a double helix between nucleotides or nucleic acids (e.g., between the two strands of a double-stranded DNA molecule or between primer binding sites on oligonucleotide primers and single-stranded nucleic acids). Complementary nucleotides are typically A and T (or A and U) or C and G. Two single-stranded RNA or DNA molecules can be described as substantially complementary when the nucleotides of one strand (optimally arranged and contrasted, with appropriate nucleotide insertions or deletions) pair with at least about 80% of the other strands, typically at least about 90% to about 95% of the other strands, or even about 98% to about 100% of the other strands. On the one hand, the two complementary sequences of nucleotides are capable of hybridization with opposing nucleotides, preferably with less than 25% mismatch, more preferably less than 15% mismatch, even more preferably less than 5% mismatch, and most preferably no mismatch. Preferably, the two molecules will hybridize under highly stringent conditions.

[0126] As used herein, “hybridization” can refer to the process by which two single-stranded polynucleotides non-covalently combine to form a stable double-stranded polynucleotide. The resulting double-stranded polynucleotide can be a “hybrid” or a “double strand.” Typical “hybridization conditions” include salt concentrations of approximately less than 1 M, typically less than about 500 mM, and possibly less than about 200 mM. A “hybridization buffer” contains a buffered salt solution, such as 5% SSPE or other such buffers known in the art. Hybridization temperatures can be as low as 5 °C, but are typically above 22 °C, more typically above about 30 °C, and usually above 37 °C. Hybridization is often performed under stringent conditions, which are conditions under which the sequence will hybridize with its target sequence but not with other non-complementary sequences. Stringent conditions are sequence-dependent and vary under different conditions. For example, for specific hybridization, longer fragments may require higher hybridization temperatures than shorter fragments. Since other factors, including the base composition and length of the complementary strand, the presence of organic solvents, and the degree of base mismatch, can affect the stringency of hybridization, the combination of parameters is more important than any single parameter as an absolute measure. Typically, stringent conditions are chosen to be approximately 5 °C lower than the Tm of a specific sequence at defined ionic strengths and pH. The melting temperature Tm can be the temperature at which a group of double-stranded nucleic acid molecules begins to partially dissociate into single strands. Several equations for calculating the Tm of nucleic acids are known in the art. As shown in standard references, a simple estimate of the Tm value can be calculated using the equation Tm = 81.5 + 0.41 (%G + C) when the nucleic acid is in aqueous solution at 1 M NaCl (see, for example, Anderson and Young, Quantitative Filter Hybridization, in Nucleic Acid Hybridization (1985)). Other references (e.g., Allai and Santa Lucia, Jr., Biochemistry, 36:10581-94 (1997)) include alternative methods for calculation, where structural and environmental factors, as well as sequence characteristics, are considered in the calculation of Tm.

[0127] Typically, the stability of heterozygotes is a function of ion concentration and temperature. Hybridization reactions are usually carried out under lower stringency conditions, followed by washing with different but higher stringency conditions. Exemplary stringency conditions include a pH of about 7.0 to about 8.3 and a temperature of at least 25°C, with a salt concentration of at least 0.01 M to no more than 1 M sodium ion (or other salt). For example, 5×SSPE conditions (750 mM NaCl, 50 mM sodium phosphate, 5 mM EDTA at pH 7.4) and a temperature of about 30°C are suitable for allele-specific hybridization, although suitable temperatures are related to the length of the hybridization region and / or GC content. On one hand, the “stringency of hybridization” in the mismatch percentage can be determined as follows: 1) High stringency: 0.1×SSPE, 0.1% SDS, 65°C; 2) Moderate stringency: 0.2×SSPE, 0.1% SDS, 50°C (also known as moderate stringency); and 3) Low stringency: 1.0×SSPE, 0.1% SDS, 50°C. It should be understood that equivalent stringency can be achieved using alternative buffers, salts, and temperatures. For example, moderately stringent hybridization can refer to conditions that allow nucleic acid molecules, such as probes, to bind to complementary nucleic acid molecules. Hybridized nucleic acid molecules generally have at least 60% identity, including, for example, any of at least 70%, 75%, 80%, 85%, 90%, or 95% identity. Moderately stringent conditions can be equivalent to hybridization at 42°C in 50% formamide, 5×Denhardt's solution, 5×SSPE, and 0.2% SDS, followed by washing at 42°C in 0.2×SSPE and 0.2% SDS. Highly stringent conditions, for example, can be provided as follows: hybridization at 42°C in 50% formamide, 5×Denhardt's solution, 5×SSPE, and 0.2% SDS, followed by washing at 65°C in 0.1×SSPE and 0.1% SDS. Low-strictness hybridization can refer to conditions equivalent to hybridization at 22°C in 10% formamide, 5× Dunhardt solution, 6× SSPE, and 0.2% SDS, followed by washing at 37°C in 1× SSPE and 0.2% SDS. The Dunhardt solution contains 1% Ficoll, 1% polyvinylpyrrolidone, and 1% bovine serum albumin (BSA). 20× SSPE (sodium chloride, sodium phosphate, EDTA) contains 3M sodium chloride, 0.2M sodium phosphate, and 0.025M EDTA.Other suitable moderately stringent and highly stringent hybridization buffers and conditions are known to those skilled in the art and described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, Plainview, NY (1989); and Ausubel et al., Short Protocols in Molecular Biology, 4th ed., John Wiley & Sons (1999).

[0128] Alternatively, essential complementarity exists when the RNA or DNA strand hybridizes with its complement under selective hybridization conditions. Typically, selective hybridization occurs when there is at least about 65% complementarity, preferably at least about 75%, and more preferably at least about 90% complementarity, across a sequence segment of at least 14 to 25 nucleotides. See M. Kanehisa, Nucleic Acids Res. 12:203 (1984).

[0129] The "primers" used in this article can be natural or synthetic oligonucleotides that can act as initiation sites for nucleic acid synthesis after forming a double helix with a polynucleotide template, and can extend along the template from its 3' end to form an extended double helix. The sequence of nucleotides added during the extension process is determined by the sequence of the template polynucleotide. Primers are typically amplified using polymerases, such as DNA polymerase.

[0130] The “substantially non-fluorescent” portion refers to a portion that emits approximately or substantially no detectable fluorescence. For example, at approximately the same concentration of the fluorescent and substantially non-fluorescent portions, the ratio of detectable absolute fluorescence emission from the fluorescent portion to detectable absolute fluorescence emission from the substantially non-fluorescent portion is typically greater than or equal to about 500:1, more typically greater than or equal to about 1000:1, and even more typically greater than or equal to about 1500:1 (e.g., about 2000:1, about 2500:1, about 3000:1, about 3500:1, about 4000:1, about 4500:1, about 5000:1, about 10...). 4 1. Approximately 10 5 1. Approximately 10 6 1. Approximately 10 7 :1 or approximately 10 8 :1).

[0131] "Sequencing," such as nucleotide sequencing methods, involves the determination of information related to the nucleotide base sequence of nucleic acids. This information can include the confirmation or determination of partial and complete sequence information of the nucleic acid. Sequence information can be determined with varying degrees of statistical reliability or confidence. On one hand, the term includes determining the identity and order of multiple adjacent nucleotides in a nucleic acid. "High-throughput sequencing" or "next-generation sequencing" includes sequencing using methods that determine many (typically thousands to billions) nucleic acid sequences in an inherently parallel manner, i.e., where the DNA template is prepared for batch processing rather than sequencing one at a time, and preferably for parallel readout of many sequences, or using ultra-high-throughput serial processing that can be parallelized itself. Such methods include, but are not limited to, pyrosequencing (e.g., as commercialized by 454 Life Sciences, Inc., Branford, CT); ligation sequencing (e.g., as in SOLiD); and sequencing by linker (e.g., as in SOLiD). TM Technology, Life Technologies, Inc., Carlsbad, CA (commercialized); sequencing through the use of modified nucleotide synthesis (such as in TruSeq) TM and HiSeq TM The technology is commercialized by Illumina, Inc., San Diego, CA; in HeliScope TM Commercialized in China by Helicos Biosciences Corporation, Cambridge, MA; and in PacBio RS by Pacific Biosciences of California, Inc., Menlo Park, CA, using ion detection technologies such as Ion Torrent TM Sequencing technologies include: sequencing of DNA nanospheres (Complete Genomics, Inc., Mountain View, CA); sequencing technologies based on nanopores (e.g., developed by Oxford Nanopore Technologies, LTD, Oxford, UK); and highly parallelized sequencing methods, for example.

[0132] In any of the embodiments disclosed herein, the method for obtaining sequence information of the target polynucleotide can be performed in a multiplexing assay. As used herein, “multiplexing” or “multiplex assay” can refer to an assay or other analytical method in which the presence and / or amount of multiple targets (e.g., multiple nucleic acid sequences) can be determined simultaneously, wherein each target has at least one distinct detection property, such as fluorescence properties (e.g., excitation wavelength, emission wavelength, emission intensity, FWHM (full width at half maximum peak) or fluorescence lifetime) or unique nucleic acid or protein sequence properties.

[0133] In any of the embodiments disclosed herein, the sequencing reaction of the target polynucleotide can be performed on an array, such as a microchip. The array may include multiple reaction volumes, for example, created by multiple reaction chambers disposed on the array. The target nucleotide sequence or fragment thereof may be immobilized or immobilized in the reaction volume, such as by adsorption or specific binding to a trap molecule on a solid support in each reaction volume. After the reaction solution is provided in the reaction mixture and delivered to each reaction volume, each reaction volume may be sealed and / or separated from other reaction volumes on the array. Signals such as fluorescence information can then be detected and / or recorded by each reaction volume.

[0134] In any of the embodiments disclosed herein, the array may be addressable. On one hand, addressability includes the ability of the microchip to guide substances such as nucleic acids and enzymes, as well as other amplification components, from one location on the microchip to another (capture sites on the chip). On the other hand, addressability includes the ability to spatially encode sequencing reactions and / or their sequencing products on each arrayspot, such that after sequence readout, the sequencing reactions and / or their sequencing products can be mapped back to the specific arrayspot and associated with additional identification information from that specific arrayspot. For example, spatially encoded tags may be conjugated to target polynucleotides such that when the conjugated target polynucleotides are sequenced, the tagged sequence reveals the location of the array target.

[0135] II. Sequencing Methods

[0136] On the one hand, this paper discloses a method for sequencing nucleotide molecules using phosphate-modified fluorophores. On the other hand, this paper discloses a method for sequencing nucleotide molecules modified with fluorescence-switching fluorophores.

[0137] On the one hand, this document discloses sequencing methods for mixed nucleotides. In a specific embodiment, this document discloses sequencing methods using phosphate-modified mixed nucleotide molecules with fluorophores. Furthermore, this disclosure also relates to sequencing methods based on fluorophores with fluorescence switching properties.

[0138] On the one hand, this paper discloses sequencing methods using mixed nucleotide molecules. In a specific embodiment, this paper discloses sequencing methods using modified mixed nucleotide molecules with fluorophores. Furthermore, this invention also relates to sequencing methods based on fluorophores with fluorescence switching properties. This invention combines fluorescence switching sequencing and mixed nucleotide molecule sequencing, achieving unexpected technical results. The unique signal acquisition method and efficiency make it promising for gene sequencing.

[0139] On one hand, this paper discloses a sequencing method using nucleotide substrate molecules, wherein sequencing is performed by modifying the 5' end or interphosphate of a nucleotide substrate molecule containing a fluorophore; each round of sequencing uses a set of reaction solutions, each set comprising two reaction solutions, each containing two nucleotides with different bases. In one embodiment, the nucleotides in one reaction solution are complementary to two bases on the nucleotide sequence to be tested, and the nucleotides in the other reaction solution are complementary to another two bases on the nucleotide sequence to be tested. In one embodiment, the method includes first providing a fragment of the nucleotide sequence to be tested (e.g., by immobilizing the nucleotide sequence on a solid support), and then providing a first reaction solution from the set of reaction solutions to begin a first round of sequencing. In one embodiment, the method includes detecting and recording a fluorescence signal from the first round of sequencing. In another embodiment, the method includes providing a second reaction solution from the same set of reaction solutions to continue the first round of sequencing. The fluorescence signal is detected and recorded again. On the other hand, the above steps are repeated, and the first and second reaction solutions can be provided sequentially in any suitable order to obtain the coding information of the nucleotide sequence to be tested by analyzing the fluorescence signal.

[0140] In one embodiment, each reaction solution contains two nucleotides with different bases, which can be labeled with two different or the same fluorophores.

[0141] In any of the foregoing embodiments, sequencing can be performed by modifying the 5' end or middle phosphate of a nucleotide substrate molecule with a fluorescence-switching property. On the one hand, fluorescence-switching property refers to a significant change in the fluorescence signal after sequencing compared to before the sequencing reaction.

[0142] In any of the foregoing embodiments, fluorescence switching property can refer to a significant enhancement (or increase) in fluorescence signal after sequencing compared to before the sequencing reaction.

[0143] On one hand, this paper also discloses a sequencing method using nucleotide substrate molecules with fluorescent switching properties. On one hand, sequencing is performed by modifying the 5' end or middle phosphate of the nucleotide substrate molecule with fluorescent switching properties. On one hand, fluorescent switching properties refer to a significant enhancement of the fluorescence signal intensity after sequencing compared to the state before the sequencing reaction. Each round of sequencing uses a set of reaction solutions, each set of reaction solutions including two reaction solutions, each reaction solution containing nucleotide substrate molecules with two different bases. On one hand, the nucleotide substrate molecules in one reaction solution are complementary to two bases on the nucleotide sequence to be tested, and the nucleotide substrate molecules in the other reaction solution are complementary to two other bases on the nucleic acid sequence to be tested. On one hand, the method includes immobilizing the nucleotide sequence fragment to be tested in a reaction chamber and then introducing a first reaction solution from the set of reaction solutions. On one hand, the method includes using an enzyme to release the fluorescent group on the nucleotide substrate with fluorescent switching properties, thereby causing fluorescence switching. On one hand, the method includes introducing a second reaction solution from the same set of reaction solutions. On one hand, the method includes using an enzyme to release a fluorophore on a nucleotide substrate having a fluorescence switching property, thereby causing fluorescence switching. On the other hand, the method includes adding two portions of reaction solution alternately and obtaining the encoding information of the nucleotide substrate to be tested through fluorescence information.

[0144] On the other hand, this paper discloses a sequencing method using nucleotide substrate molecules with fluorophores exhibiting fluorescence switching properties. Sequencing is performed by modifying the 5' end or middle phosphate of the nucleotide substrate molecule with a fluorophore exhibiting fluorescence switching properties. Fluorescence switching properties refer to a significant increase in fluorescence signal intensity after sequencing compared to before the sequencing reaction. Each sequencing run uses a set of reaction solutions, each set comprising at least two portions, each containing at least one of A, G, C, or T nucleotide substrate molecules or at least one of A, G, C, or U nucleotide substrate molecules. First, the nucleotide sequence fragment to be tested is immobilized in a reaction chamber, and reaction solutions from the set of reaction solutions are added to the chamber. The sequencing reaction can be initiated under appropriate conditions, and fluorescence signals are recorded. Then, an additional portion of reaction solution is provided each time, such that other portions of the same set of reaction solutions are successively provided during the sequencing reaction. Simultaneously, one or more fluorescence signals from each portion of reaction solution are recorded. At least one portion of reaction solution is included in the set containing two or three nucleotide molecules.

[0145] On the other hand, this paper discloses a sequencing method using nucleotide substrate molecules with fluorescence switching properties, achieved by modifying the 5' end or middle phosphate of the nucleotide substrate molecule with fluorescence switching properties. Fluorescence switching properties refer to a significant enhancement of the fluorescence signal intensity after sequencing compared to before the sequencing reaction. Each sequencing run uses a set of reaction solutions, each set comprising at least two portions, each containing any one of A, G, C, or T nucleotide substrate molecules, or each containing any one of A, G, C, or U nucleotide substrate molecules. The method includes first immobilizing the nucleotide sequence fragment to be tested in a reaction chamber, and then introducing one portion of the reaction solution from the set. The method includes testing and recording fluorescence information. The method includes adding one portion of the reaction solution at a time, followed by the subsequent addition of other reaction solutions from the same set. Fluorescence information from each sequencing reaction is recorded.

[0146] On the other hand, this paper discloses a sequencing method using nucleotide substrate molecules with fluorescence-switching properties. Sequencing is performed by modifying the 5' end or middle phosphate of the nucleotide substrate molecule with fluorescence-switching properties. Fluorescence-switching properties refer to a significant increase in fluorescence signal intensity after sequencing compared to before the sequencing reaction. In one aspect, each round of sequencing uses a set of reaction solutions containing either four nucleotide substrate molecules (A, G, C, T) or four nucleotide substrate molecules (A, G, C, U). In another aspect, the method includes immobilizing the nucleotide sequence fragment to be tested in a reaction chamber, introducing the reaction solution, and recording fluorescence information.

[0147] In any of the foregoing embodiments, the method may further include removing residual reaction solution and fluorescent molecules with a washing solution before proceeding to the next round of sequencing. In any of the foregoing embodiments, the reaction solution may be added at low temperature and then heated to the enzyme reaction temperature, wherein the fluorescence signal is detected. In any of the foregoing embodiments, after the reaction solution is added to the reaction mixture, the reaction chamber may be sealed, and fluorescence information may be detected and / or recorded.

[0148] In any of the foregoing embodiments, after the reaction solution is added, the space outside the reaction chamber can be filled with oil to isolate and seal the reaction chamber. In any of the foregoing embodiments, the polyphosphate nucleotide substrate molecule can refer to a nucleotide having 4 to 8 phosphate molecules. In any of the foregoing embodiments, the modified fluorophoretic nucleotide substrate molecule can be labeled with one fluorescent group for monochrome sequencing; or labeled with different fluorescent groups for multicolor sequencing.

[0149] In any of the foregoing embodiments, the method may include using an enzyme to release a fluorophore on a nucleotide substrate having fluorescence switching properties, wherein the enzyme may optionally include DNA polymerase and / or alkaline phosphatase.

[0150] In any of the foregoing embodiments, the two bases on the nucleotide sequence to be tested may include any two of A, G, C and T bases or A, G, C and U bases; wherein base C is methylated C or unmethylated C.

[0151] In any of the foregoing embodiments, the reaction solution may contain an enzyme that, when passed into the reaction region containing the gene fragment to be tested, releases a fluorophore on a nucleotide substrate with a fluorescence switching property.

[0152] In any of the foregoing embodiments, the reaction solution and the enzyme may not be added simultaneously; that is, the first reaction solution from a group of reaction solutions is first introduced, followed by the enzyme solution; then, the second reaction solution from the same group of reaction solutions is introduced, followed by the enzyme solution.

[0153] In any of the foregoing embodiments, one set of reaction solutions can be used for one round of sequencing, or two sets of reaction solutions can be used for two rounds of sequencing, or three sets of reaction solutions can be used for three rounds of sequencing.

[0154] In any of the foregoing embodiments, the method may include performing a round of sequencing using a set of reaction solutions and obtaining degenerate code results.

[0155] In any of the foregoing embodiments, the method may include performing two rounds of sequencing using two sets of reaction solutions to obtain base sequence information.

[0156] In any of the foregoing embodiments, the method may include performing three rounds of sequencing using three sets of reaction solutions, and performing error checking and correction based on the results of (comparison) of (any) two rounds of sequencing from the mutual information between the three rounds of sequencing.

[0157] In any of the foregoing embodiments, the fluorophore with fluorescence switching properties may include fluorophores having structures such as methylfluorescein, halomethylfluorescein, DDAO, or resorufin.

[0158] In any of the foregoing embodiments, the method may include using an enzyme to release a fluorophore on a nucleotide substrate having a fluorescence switching property, wherein optimally optionally, this includes first using a DNA polymerase to release the polyphosphate-substituted fluorophore, and then using a phosphatase to cleave the substituted polyphosphate, thereby releasing the fluorophore.

[0159] In any of the foregoing embodiments, the reaction solution may contain two or more nucleotides with different bases, and the reaction solution may be simply decomposed into two or more portions of the reaction solution, such that each portion of the reaction solution contains one or more nucleotides; and at least one portion of the reaction solution may contain two or three nucleotides with different bases.

[0160] This document also discloses a high-throughput sequencing method according to any of the foregoing embodiments, wherein the sequencing reaction is performed on a chip having multiple reaction chambers. The method may optionally include immobilizing the nucleotide sequence fragment to be tested within the reaction chambers.

[0161] On the other hand, this paper discloses a sequencing method using nucleotide substrate molecules with fluorophores exhibiting fluorescence switching properties, and sequencing of nucleotide substrate molecules with fluorophores exhibiting fluorescence switching properties modified by using 5' polyphosphate. The method provided in this paper involves first immobilizing the nucleotide sequence fragment to be tested, and then adding a reaction solution containing nucleotide substrate molecules. Then, an enzyme can be used to release the fluorophore on the nucleotide substrate, thereby causing fluorescence switching.

[0162] In one embodiment, the sequencing method further includes using a washing buffer to remove residual reaction solution and fluorescent molecules before proceeding to the next sequencing reaction. In any of the foregoing embodiments, the sequencing method may include a reaction solution at a low temperature, which is then heated to the enzyme reaction temperature. Fluorescent signals can then be detected and / or recorded.

[0163] In any of the foregoing embodiments, the nucleotide substrate molecule may include a nucleotide molecule containing A, G, C, and T bases or a nucleotide molecule containing A, G, C, and U bases; wherein C is methylated C or unmethylated C. In any of the foregoing embodiments, the nucleotide substrate molecule may include a fluorophore with fluorescence switching properties modified with a 5' polyphosphate. In any of the foregoing embodiments, the nucleotide substrate molecule may include a fluorophore with fluorescence switching properties modified with a 5' phosphate.

[0164] This document also discloses methods according to any of the foregoing embodiments, wherein different nucleotide substrate molecules can be linked to a single fluorophore for monochromatic sequencing, or linked to multiple fluorophores for multicolor sequencing, depending on the bases present.

[0165] This document discloses the method according to any of the foregoing embodiments, wherein the fluorescence switching property refers to a significant increase or decrease in fluorescence signal after each sequencing reaction compared to before the sequencing reaction, or a significant change in the emission light frequency range.

[0166] This document discloses the method according to any of the foregoing embodiments, wherein the fluorescence switching property refers to a significant enhancement of the fluorescence signal after each sequencing reaction compared to the state before the sequencing reaction.

[0167] This document discloses a method according to any of the foregoing embodiments, wherein a reaction solution containing a nucleotide substrate molecule is used for sequencing. The nucleotide substrate molecule refers to any two or three mixtures of A, G, C, and T nucleotide substrate molecules; or any two or three mixtures of A, G, C, and U nucleotide substrate molecules.

[0168] This document discloses a method according to any of the foregoing embodiments, wherein a reaction solution containing a nucleotide substrate molecule is used for sequencing. The nucleotide substrate molecule refers to any one of A, G, C, or T nucleotide substrate molecules; or any one of A, G, C, or U nucleotide substrate molecules.

[0169] This document discloses a sequencing method for nucleotide substrate molecules using fluorophores with fluorescence switching properties according to any of the foregoing embodiments, wherein each sequencing run uses a set of reaction solutions, each set of reaction solutions comprising at least two portions of reaction solution, each portion of reaction solution containing at least one of A, G, C, and T nucleotide substrate molecules, or each portion of reaction solution containing at least one of A, G, C, and U nucleotide substrate molecules. In one aspect, the method includes immobilizing the nucleotide sequence fragment to be tested, introducing one portion of the reaction solution from the set of reaction solutions, and recording fluorescence information. In another aspect, the method includes introducing one portion of the reaction solution at a time, followed by the introduction of another portion of the same set of reaction solutions. In yet another aspect, the reaction solution set contains at least one portion of the reaction solution containing two or three nucleotide molecules.

[0170] This document discloses a sequencing method for nucleotide substrate molecules using fluorophores with fluorescence switching properties according to any of the foregoing embodiments, wherein each sequencing run uses a set of reaction solutions, each set comprising two reaction solutions, each containing two nucleotides with different bases. On one hand, the nucleotides in one reaction solution are complementary to two bases on the target nucleotide sequence, and the nucleotides in the other reaction solution are complementary to another two bases on the target nucleic acid sequence. On the other hand, the method includes immobilizing the target nucleotide sequence fragment and introducing a first reaction solution into the set of reaction solutions. Then, a second reaction solution from the same set of reaction solutions is added. The two reaction solutions can be added alternately to obtain the coding information of the target nucleotide substrate through fluorescence information.

[0171] In any of the foregoing embodiments, after the reaction solution is added to the sequencing reaction, the reaction chamber can be sealed, and then the fluorescence signal can be recorded.

[0172] In any of the foregoing embodiments, after the reaction solution is added to the sequencing reaction, the space outside the reaction chamber is filled with oil or oily substances that can isolate and seal the reaction chamber.

[0173] In any of the foregoing embodiments, the polyphosphate substrate may be a nucleotide having about 4 to about 8 phosphate molecules.

[0174] In any of the foregoing embodiments, one set of reaction solutions can be used for one round of sequencing, or two sets of reaction solutions can be used for two rounds of sequencing, or three sets of reaction solutions can be used for three rounds of sequencing.

[0175] In any of the foregoing embodiments, the method may include using an enzyme to release a fluorophore from a nucleotide substrate having fluorescence switching properties. The enzyme may include DNA polymerase and / or alkaline phosphatase.

[0176] In any of the foregoing embodiments, the method may include performing a round of sequencing using a set of reaction solutions and obtaining degenerate code results.

[0177] In any of the foregoing embodiments, the method may include performing two rounds of sequencing using two sets of reaction solutions, and obtaining base sequence information.

[0178] In any of the foregoing embodiments, the method may include performing three rounds of sequencing using three reaction solutions, and performing error checking and correction based on mutual information from the results of any two of the three rounds of sequencing.

[0179] In any of the foregoing embodiments, the reaction solution may contain an enzyme. When the reaction solution is passed into the reaction region containing the gene fragment to be tested, the contained enzyme can release the fluorophore on the nucleotide substrate with fluorescence switching properties.

[0180] In any of the foregoing embodiments, the reaction solution and enzyme can be added at different times. On one hand, a first reaction solution from the same set of reaction solutions is added to the reaction first, followed by the enzyme solution. Next, a second reaction solution from the same set of reaction solutions is added, followed by the enzyme solution.

[0181] In any of the foregoing embodiments, the fluorophore having fluorescence switching properties may include a fluorophore containing groups such as methylfluorescein, halogenated methylfluorescein, DDAO (7-hydroxy-9H-(1,3-dichloro-9,9-dimethylacridin-2-one)) and / or halogen.

[0182] In any of the foregoing embodiments, the release of the fluorophore on the nucleotide substrate having the fluorescence switching property can be optimized, for example, using an enzyme. One aspect of optimization involves first using a DNA polymerase to release the polyphosphate-substituted fluorophore, and then using a phosphatase to cleave the substituted polyphosphate to release the fluorophore.

[0183] In any of the foregoing embodiments, the reaction solution may contain two or more nucleotides with different bases. On one hand, two or more portions of the reaction solution may be used, such that each portion of the reaction solution contains one or more nucleotides. The order in which the reaction solutions are added during the reaction can be appropriately adjusted; on the other hand, at least one portion of the reaction solution contains two or three nucleotides with different bases.

[0184] This document also provides a high-throughput sequencing method according to any of the foregoing embodiments, wherein the sequencing reaction is performed on a chip having multiple reaction chambers. One aspect of the method involves immobilizing the nucleotide sequence fragments to be tested in each reaction chamber.

[0185] On one hand, the present invention relates to sequencing methods, for example, using mixed nucleotide molecules. More specifically, this sequencing method uses mixed nucleotide molecules with fluorophores that are modified (e.g., phosphate-modified). Furthermore, the present invention also relates to sequencing methods based on fluorophores with fluorescence switching properties. Fluorophores with fluorescence switching properties are sequenced using nucleotide substrates labeled with terminal phosphates. The substrate of the fluorophore with fluorescence switching properties is a fluorophore with fluorescence switching properties modified by a 5' polyphosphate or interphosphate, characterized in that the fluorophore with fluorescence switching properties is modified with 4, 5, 6 or more phosphate-deoxyribonucleotides (including A, C, G, T, U and other nucleotides) at the terminal or interphosphate, and is not labeled at the bases and 3'-hydroxyl groups. The absorption and / or emission spectra of the phosphate-modified fluorophore differ from those of the fluorophore without phosphate. The sequencing reaction typically comprises successive and similar cycles. Each cycle may include steps such as sample injection / coating, reaction, signal acquisition, and washing of unreacted reactant molecules. In the previously reported method, when a substrate molecule with a base pair enters, no reaction occurs if it is not properly paired; and the polymerase links the substrate molecule to its 3' end, releasing a polyphosphate-modified fluorescent molecule, which alters the fluorescence spectrum. If it pairs continuously with a homopolymer, the spectrum will change multiple times. In practice, as a modification label for substrate molecules such as methylfluorescein, halomethylfluorescein, DDAO, halogenated fluorescein, and fluorescent molecules as described in CN104844674, fluorophores with fluorescence switching properties that do not absorb in the terminal phosphate ester and whose release state has a high quantum yield are frequently used. Four substrate molecules can be labeled with different fluorescent molecules. The sequencing process is performed via sample injection in ACGTACGT... or any cyclic or non-cyclic injection process, using a reaction solution containing the substrate molecule in limited phases to obtain extension information for each cycle, and then obtaining the DNA sequence.

[0186] On one hand, this invention relates to a sequencing method for various nucleotides. More specifically, this sequencing method uses phosphate to modify mixed nucleotide molecules containing fluorophores. Sequencing is performed by modifying the 5' end or middle phosphate of the fluorophore-containing nucleotide substrate molecule; each round of sequencing uses a set of reaction solutions, each set of reaction solutions comprising two reaction solutions, each reaction solution containing two nucleotides containing different bases; one reaction solution contains nucleotides that are complementary to two bases on the nucleotide sequence to be tested, and the other reaction solution contains nucleotides that are complementary to another two bases on the nucleic acid sequence to be tested; first, the nucleotide sequence fragment to be tested is fixed, and the first reaction solution in the set of reaction solutions is passed through; fluorescence information is tested and recorded; then the second reaction solution in the same set of reaction solutions is passed through; fluorescence information is tested and recorded; the two reaction solutions are added cyclically, and the coding information of the nucleotide substrate to be tested is obtained through fluorescence information.

[0187] In some embodiments, the reaction solution in this invention refers to a sequencing reaction solution in the general sense. An auxiliary solution, such as other washing or cleaning solutions, is introduced into the spaces between the reaction solutions. On one hand, each reaction solution contains two nucleotides with different bases, which can be labeled with different or the same fluorophores. On the other hand, sequencing is performed by modifying the 5' end or middle phosphate of a nucleotide substrate molecule with a fluorophore possessing fluorescence switching properties; fluorescence switching properties refer to a significant change in the fluorescence signal after sequencing compared to before the sequencing reaction.

[0188] On the one hand, fluorescence switching properties refer to a significant enhancement (improvement) in the fluorescence signal after sequencing compared to before the sequencing reaction. The frequency of emitted light may change, but the overall intensity of emitted light or the intensity of emitted light in certain frequency bands will be significantly enhanced.

[0189] On one hand, this invention relates to a method for sequencing nucleotide molecules using fluorophores with fluorescence switching properties, wherein sequencing is performed by modifying the 5' end or middle phosphate of the fluorophore-containing nucleotide substrate molecule; fluorescence switching property refers to a significant enhancement of the fluorescence signal intensity after sequencing compared to the state before the sequencing reaction; each round of sequencing uses a set of reaction solutions, each set of reaction solutions includes two reaction solutions, each reaction solution containing nucleotide substrate molecules with two different bases; one reaction solution contains nucleotide substrate molecules that are complementary to two bases on the nucleotide sequence to be tested, and the other reaction solution contains nucleotides that are complementary to the other two bases on the nucleotide sequence to be tested. First, the nucleotide sequence fragment to be tested can be fixed in a reaction chamber, and then one reaction solution from the set of reaction solutions is introduced; then an enzyme is used to release the fluorophore on the nucleotide substrate with fluorescence switching properties, thereby causing fluorescence switching; then a second reaction solution from the same set of reaction solutions is introduced; an enzyme is used to release the fluorophore on the nucleotide substrate with fluorescence switching properties, thereby causing fluorescence switching; the two reaction solutions are added in a cycle, and the coding information of the nucleotide substrate to be tested is obtained through fluorescence information.

[0190] On one hand, this invention relates to a method for sequencing nucleotide molecules using fluorophores with fluorescence switching properties, wherein sequencing is performed by modifying the 5' end or middle phosphate of the fluorophore-containing nucleotide substrate molecule; fluorescence switching property refers to a significant enhancement of the fluorescence signal intensity after sequencing compared to the state before the sequencing reaction; each round of sequencing uses a set of reaction solutions, each set of reaction solutions including at least two portions, each reaction solution containing at least one of A, G, C, or T nucleotide substrate molecules, or one of A, G, C, or U nucleotide substrate molecules. On the other hand, the nucleotide sequence fragment to be tested can be first immobilized in a reaction chamber, and one portion of the reaction solution from the set can be introduced; fluorescence information can be tested and recorded; one portion of the reaction solution can be introduced each time, followed by the other portion of the same set of reaction solutions. Simultaneously, fluorescence information can be checked and recorded after each portion of the reaction solution is introduced, wherein at least one portion of the reaction solution set contains two or three of the nucleotide molecules from the set of reaction solutions.

[0191] On one hand, this invention relates to a method for sequencing nucleotide molecules using fluorophores with fluorescence switching properties, wherein sequencing is performed by modifying the 5' end or middle phosphate of the fluorophore-containing nucleotide substrate molecule; fluorescence switching property refers to a significant enhancement of the fluorescence signal intensity after sequencing compared to the state before the sequencing reaction; each round of sequencing uses a set of reaction solutions, each set of reaction solutions including at least two portions, each portion containing any one of A, G, C, or T nucleotide substrate molecules, or any one of A, G, C, or U nucleotide substrate molecules. On the other hand, the nucleotide sequence fragment to be tested can be first immobilized in a reaction chamber, and one portion of the reaction solution from the set can be introduced; fluorescence information can be tested and recorded; one portion of the reaction solution is introduced each time, followed by the other portion of the same set of reaction solutions. Simultaneously, fluorescence information can be tested and recorded after each portion of the reaction solution is introduced.

[0192] On one hand, this invention relates to a method for sequencing nucleotide molecules using fluorophores with fluorescence switching properties, wherein sequencing is performed by modifying the 5' end or middle phosphate of the fluorophore-containing nucleotide substrate molecule; fluorescence switching property refers to a significant enhancement of the fluorescence signal intensity after sequencing compared to the state before the sequencing reaction; each round of sequencing uses a set of reaction solutions containing A, G, C, and T nucleotide substrate molecules, or A, G, C, and U nucleotide substrate molecules. On the other hand, the nucleotide sequence fragment to be tested can be immobilized in a reaction chamber, the reaction solution is introduced, and then the fluorescence information is tested and recorded.

[0193] On one hand, the method further includes using a washing solution to remove residual reaction solution and fluorescent molecules before proceeding to the next round of sequencing. On another hand, the method includes delivering the reaction solution at low temperature, then heating it to the enzyme reaction temperature and testing the fluorescence signal. On yet another hand, after introducing the reaction solution, the method includes sealing the reaction chamber and then testing and recording the fluorescence information.

[0194] On one hand, after the reaction solution is introduced, the method includes filling the space outside the reaction chamber with oil, thereby isolating and sealing the reaction chamber. On the other hand, polyphosphate nucleotide substrate molecules refer to nucleotides having 4-8 phosphate molecules. On the other hand, nucleotide substrate molecules modified with fluorophores can be labeled with one fluorescent group for single-color sequencing, or labeled with different fluorophores for multicolor sequencing, depending on the bases used.

[0195] On one hand, the method includes the following steps: using an enzyme (e.g., DNA polymerase and / or alkaline phosphatase) to release a fluorophore on a nucleotide substrate with fluorescence switching properties. On one hand, the two bases on the nucleotide sequence to be tested refer to any two of A, G, C, and T bases or any two of A, G, C, and U bases, wherein base C is methylated C or unmethylated C. On one hand, when the reaction solution is passed into the reaction region where the gene fragment to be tested is located, the enzyme in the reaction solution can release the fluorophore on the nucleotide substrate with fluorescence switching properties. On one hand, the method includes performing one round of sequencing using one set of reaction solutions and obtaining degenerate code results. On one hand, the method includes performing two rounds of sequencing using two sets of reaction solutions and obtaining base sequence information. On one hand, the method includes performing three rounds of sequencing using three sets of reaction solutions and performing error checking and correction based on the results of (any) two rounds of sequencing based on the interaction information between the three rounds of sequencing.

[0196] On one hand, this invention relates to a sequencing method for mixed nucleotide molecules. More specifically, this sequencing method uses phosphate to modify mixed nucleotide molecules containing fluorophores. Compared to sequencing methods for unphosphate-modified mixed nucleotides, this method is easier to hydrolyze, does not introduce other groups after the reaction, which is beneficial for extending the sequencing reaction, and the sequencing reaction is simple.

[0197] On one hand, the present invention relates to a sequencing method for mixed nucleotide molecules of nucleotide substrate molecules with fluorescence-switching properties modified by using 5' polyphosphate. On one hand, the method includes first immobilizing the nucleotide sequence fragment to be tested and then passing it through a reaction solution containing nucleotide substrate molecules. On another hand, the method includes using an enzyme to release the fluorophore on the nucleotide substrate with fluorescence-switching properties, thereby causing fluorescence switching. On yet another hand, the method further includes using a washing buffer to remove residual reaction solution and fluorescent molecules before proceeding to the next round of sequencing reactions.

[0198] In another embodiment, this invention combines fluorescence-switched sequencing and mixed nucleotide sequencing to achieve unexpected results. For example, fluorescence switching provides data redundancy and verification characteristics for mixed nucleotide sequencing, improving the accuracy of sequencing data. Furthermore, 3' end blocking sequencing eliminates the need for real-time data acquisition during the sequencing reaction, improving signal accuracy. It is independent of the sequencing chemistry principles themselves and can be used with different sequencing chemistry approaches. Furthermore, the 2+2 mode (sequencing with two bases at a time) of fluorescence switching offers significant advantages over other mixed nucleotide sequencing methods. For example, data parsing is relatively easy, and it also provides data redundancy and verification characteristics. The unique signal acquisition method and efficiency make it promising for gene sequencing. Fluorescence-switched multibase sequencing reduces the error rate and simplifies the reaction compared to non-fluorescence-switched nucleotide sequencing. The mixed nucleotide sequencing method using the fluorescence switching method of this disclosure achieves sequencing accuracy up to 99.99%, exceeding Illumina sequencing read lengths to 300 nt or more, and with very low raw material costs. It employs a reaction-before-scan approach, eliminating throughput limitations. Its single-round reaction time is relatively short, enabling rapid testing. Employing a strategy of fluorescence switching and mixed sequencing of multiple nucleotide molecules can extend the read length and information content of each reaction cycle. For example, Illumina sequencing yields a read length of 1 nt (1 base) and an information content of 2 bits per reaction cycle. 2+2 (using two different nucleotide molecules with different bases each time, employing two reaction solutions) single-color sequencing yields a read length of 2 nt and an information content of 2 bits per reaction cycle. Conversely, 2+2 dual-color sequencing yields a read length of 2 nt and an information content of 3.4 bits per reaction cycle.

[0199] In some respects, this article provides information on fluorescence generation and fluorophores. Some fluorophores exhibit a property where their fluorescence spectra (absorption and reflection spectra) change when substituents are altered; this is called fluorescence switching. On the other hand, when the intensity of the acquired signal increases under specific excitation and acquisition (emission) conditions, this is called fluorescence generation.

[0200] In some respects, this article provides information on nucleotides and nucleotide labeling. On one hand, a nucleotide molecule consists of a ribose backbone, a base at the glucoside position, and a polyphosphate chain attached to the 5-hydroxyl group on the ribose backbone. The 2C of the ribose ring may have a hydroxyl group attached (called a ribonucleotide), or only an H group attached (called a deoxyribonucleotide). Nucleotide molecules can be the four major bases ACGT, uracil, and modified bases such as methylated bases, hydroxymethylated bases, etc. The number of phosphate backbones can be 1-8. Molecular groups can be modified at multiple positions. At the bases, there can be one or more modification sites on the 3C hydroxyl group of the ribose backbone. For example, a fluorophore may be modified on the phosphate group, and an ethynyl group may be modified on the 3C group.

[0201] On the one hand, during polymerase chain reaction (PCR), the unmodified polyphosphate nucleotide substrate (with more than three phosphates) at the 3C has three active hydroxyl groups. On the other hand, the polymerase reaction continues as long as subsequent bases can still pair, until a pairing base is lacking or a nucleotide molecule with a non-hydroxyl group at the 3C is bound. In some respects, this article provides fluorescently generating nucleotides. On the one hand, nucleotide molecules with a phosphate terminus and labeled with a fluorophore that can be switched by phosphate hydrolysis are called fluorescently generating (or fluorescence-producing) nucleotides. The length of the phosphate chain can be 4-8.

[0202] On the one hand, the phosphate group can be at the terminal or on the side chain. There can be one or more labels. Multiple labels can be the same or different. More precisely, on the one hand, it is called a polymerase-fluorescent nucleotide. On the other hand, fluorescent nucleotides that are not labeled at the phosphate position and do not require polymerase fluorescence can also be used. The nucleotide molecule can be a ribonucleotide, a deoxyribonucleotide, or a (deoxy)ribonucleotide modified at the 3' C.

[0203] In some respects, this article provides a fluorescent nucleotide polymerase reaction. On one hand, the fluorescent nucleotide polymerase reaction uses a fluorescent nucleotide, a nucleic acid polymerase (DNA polymerase), a phosphatase, and a nucleic acid substrate. In some embodiments, firstly, the DNA polymerase polymerizes the fluorescent nucleotide into the nucleic acid substrate to release a phosphorylated fluorophore, which is then further hydrolyzed by the phosphatase to remove the phosphate and release a fluorophore with a altered fluorescence state.

[0204] In some respects, this article provides fluorescence-generating sequencing methods. One aspect of these methods is the use of fluorescence-generating nucleotide polymerase reactions to test changes in fluorescence (intensity and spectrum) of the fluorophore, thereby obtaining information about the polymerase reaction. In some respects, this article provides fluorescence-generating sequencing reaction solutions that may contain fluorescence-generating nucleotides, nucleic acid polymerase (DNA polymerase), and phosphatase.

[0205] As described herein, "fluorescent nucleotide" may comprise one or more fluorescent nucleotides. As described herein, "nucleotide" may comprise one or more nucleotides. In some embodiments, multiple nucleotides may be labeled with the same or different fluorescent substrates. In some aspects, this document provides a set of fluorescent sequencing reaction solutions that may contain two or more fluorescent sequencing reaction solutions, such as specific concentrations of A, C, G, and T reaction solutions, or specific concentrations of AC and GT reaction solutions.

[0206] In some aspects, this document provides a fluorescence sequencing reaction cycle, which may include performing a single fluorescence-generating polymerase reaction and testing the fluorescence signal using a sequencing reaction solution. In some aspects, this document provides a round of fluorescence-generating sequencing reactions, which may include using members of a set of fluorescence-generating sequencing reaction solutions in a predetermined order for sequencing reaction cycles. In some aspects, this document provides a set of fluorescence-generating sequencing reactions, which may include one or more rounds of fluorescence-generating sequencing.

[0207] In some respects, this paper provides single-base-resolved sequencing reactions. One approach is a (2+2 single-color two-set) reaction, where the first reaction solution is a mixture of two bases (e.g., AC), and the second reaction solution is a mixture of two other bases (GT), with the two solutions used alternately for sequencing. In this case, the number of extended bases increases in each cycle. After N rounds of sequencing, the number of extended bases is 2N nt, carrying 2N bits of information. There are three combinations that complete the above sequencing: AC / GT, AG / CT, and AT / CG; or, according to standard degenerate base (degenerate nucleotide) notation, written as M / K, R / Y, and W / S. These three combinations can be sequenced separately, or a new set of sequencing can be completed before re-sequencing. The i-th base determined on the DNA sequence will always pair and release a signal in a unique cycle of either of the two sequencing sets. In each sequencing set, the determined base sampling injection cycles include two types, so there are a total of 2×2=4 possible cases, corresponding to exactly four bases. The order of sequencing combinations does not affect base inference.

[0208] Table 1

[0209]

[0210] Table 2

[0211]

[0212]

[0213] Table 3

[0214]

[0215] In a further implementation, the method also includes sequencing using a third, different combination of reaction solutions after completing two different sequencing runs. The i-th base determined on the DNA sequence must pair and release a signal in a unique cycle among the three sequencing runs. In each sequencing run, the determined base sampling injection cycle includes two types, resulting in 2 × 2 × 2 = 8 possible cases, only four of which are valid, and the other four are invalid. Insertion or deletion errors are likely to occur in fluorescence switching sequencing. If a sequencing error occurs in one of the three sequencing runs for a specific base, the sequence cannot be correctly deduced, and it can be concluded that one or more of the three sequencing runs contain a sequencing error at that point.

[0216] Table 4

[0217]

[0218] This type of error can be corrected because when sequencing errors in a single dataset are corrected, a large number of subsequent errors will also be corrected.

[0219] Another specific implementation is a 2+2 two-color two-round mode. The first reaction solution is made from a mixture of two bases and carries different fluorescent labels (such as AX / CY), while the second reaction solution is made from a mixture of two other bases (GX / TY). In this case, the number of bases extended per cycle increases, averaging 2nt. The information carried is 2N bits.

[0220] III. Methods for detecting and / or correcting sequencing errors

[0221] On the one hand, this article relates to methods for detecting and / or correcting errors in one or more sequence data in sequencing results, which falls under the field of nucleic acid sequencing.

[0222] On the one hand, this paper provides a method for detecting and / or correcting sequence data errors in sequencing results. On the other hand, the sequencing reaction solution contains at least two types of nucleotide substrate molecules with different bases. On the other hand, degenerate gene coding information can be obtained. By comparing two or more degenerate coding information, it is possible to determine whether conflicting sequence information is present in one or more nucleotide residues. Using this method to correct sequence information, even small improvements that reduce the sequencing error rate in the original sequencing data can lead to a more significant reduction in the error rate of the corrected sequence information.

[0223] On the one hand, this paper discloses a method for detecting and / or correcting sequence data errors in sequencing results. On the one hand, the method includes sequencing a nucleic acid sequence to obtain sequence data of three or more orthogonal nucleotide degenerate sequences. On the other hand, the method further includes detecting errors in the sequence by comparing the three or more orthogonal nucleotide degenerate sequences. On the other hand, at the location where the error is found, a corrected sequence is obtained by modifying at least one sequence.

[0224] This document also discloses methods for detecting and / or correcting sequence data errors in sequencing results, wherein the methods include performing a sequencing reaction on nucleotide sequences to obtain three or more degenerate sequences represented by the letters M, K, R, Y, W, S, B, D, H, and V. On the one hand, according to the IUPAC nucleic acid symbols, the degenerate bases in this invention are represented by the letters in Table 5. For example, M represents A and / or C bases.

[0225] Table 5: Letters representing degenerate bases

[0226]

[0227]

[0228] In any of the foregoing embodiments, sequence errors can be detected by comparing three or more degenerate sequences. In any of the foregoing embodiments, the erroneous nucleotide position is identified during comparison, and a corrected sequence can be obtained by modifying at least one sequence; in any of the foregoing embodiments, the erroneous position identified during comparison can be the actual location where the sequencing error occurred.

[0229] On the other hand, this document discloses methods for detecting and / or correcting sequence data errors in sequencing results, wherein the methods include sequencing the same nucleic acid sequence to obtain two or more degenerate sequences represented by the letters M, K, R, Y, W, S, B, D, H, and V, to obtain sequence information represented by nucleic acid residues A, G, T, and C or nucleic acid residues A, G, U, and C. On another aspect, the methods also include detecting sequence errors by using light or electrical signals generated by one or more functional groups coupled to different bases in the sequencing reaction. For example, light or electrical signals from different fluorescent groups coupled to different bases in the sequencing reaction can be used as “redundant” information that distinguishes one base from another at a specific position in the sequence. In any of the foregoing embodiments, if an erroneous nucleotide position is found during comparison, a corrected sequence can be obtained by modifying at least one sequence; in any of the foregoing embodiments, the position of the error identified during comparison can be the actual location where the sequencing error occurred.

[0230] On the other hand, this paper discloses a method for detecting and / or correcting sequence errors in sequencing results using the memory property of nucleic acid sequences. One aspect of the method involves sequencing the same nucleic acid sequence to obtain data of three or more orthogonal degenerate nucleic acid sequences. Another aspect of the method includes comprehensively comparing degenerate sequences and utilizing the memory property of nucleic acid sequences to detect sequence errors. One aspect is that, at the location where an error occurs, a corrected sequence can be obtained by modifying at least one sequence. In some embodiments, each degenerate sequence represents only a portion of the sequence information of the actual polynucleotide template, and nucleotide identity at a position in one degenerate sequence does not indicate, or does not necessarily indicate, nucleotide identity at the same position in another degenerate sequence.

[0231] On the one hand, this document discloses a method for detecting and / or correcting sequence data errors in sequencing results. This method includes immobilizing a nucleic acid fragment to be sequenced onto a vector and providing a reaction solution to initiate a sequencing reaction, from which a degenerate nucleic acid sequence is obtained. The sequencing reaction can be repeated multiple times, such that a degenerate nucleic acid sequence is obtained from each sequencing round. After N sequencing rounds, N degenerate nucleic acid sequences are obtained. On the other hand, by comprehensively comparing the N degenerate sequences, the location of the sequence error can be detected. Furthermore, the method may also include, by comparing the location of the error, obtaining a corrected sequence by modifying at least one sequence. In any of the foregoing embodiments, the reaction solution may contain two or more types of nucleotide substrate molecules with different bases. In any of the foregoing embodiments, N may be a positive integer equal to or greater than 2.

[0232] In any of the foregoing embodiments, the method may include comparing N-1 degenerate nucleic acid sequences out of N to obtain nucleic acid sequence information encoded by A, G, T, and C, or nucleic acid sequence information encoded by A, G, U, and C. In one aspect, the method further includes comparing N degenerate nucleic acid sequences. In any of the foregoing embodiments, N may be a positive integer equal to or greater than 3.

[0233] In any of the foregoing embodiments, the method may include comparing N degenerate nucleic acid sequences to obtain nucleic acid sequence information encoded by A, G, T, and C, or nucleic acid sequence information encoded by A, G, U, and C. On one hand, the method further includes detecting the location of an error by using optical and / or electromagnetic information provided by two or more functional groups coupled to nucleotide residues. In any of the foregoing embodiments, N may be a positive integer equal to or greater than 2.

[0234] On the other hand, this document discloses a method for detecting and / or correcting sequence data errors in sequencing results, wherein the method includes immobilizing the nucleic acid fragment to be tested onto a vector. The method further includes providing a reaction solution to initiate a sequencing reaction, wherein the reaction solution contains nucleotide substrate molecules for sequencing and is divided into three groups according to different bases, each group comprising two different reaction solutions, and each sample of the reaction solution containing nucleotide substrate molecules with different bases. There is no intersection between the bases of the nucleotides in the two samples of the same reaction solution. Each round of sequencing uses one set of reaction solutions, providing two samples of each set to react sequentially with the nucleic acid template in any suitable order. Three rounds of sequencing are performed using the three sets of reaction solutions to obtain three degenerate sequences. The location of the sequence error can be detected by comprehensively comparing the three degenerate sequences. In one embodiment, by modifying at least one sequence at the location of the error, a corrected sequence can be obtained.

[0235] In any of the foregoing embodiments, the sequencing reaction can be carried out using a nucleotide substrate molecule modified with a fluorophore having fluorescence switching properties (such as dNTPs or ddNTPs), wherein the modification is on the 5'-terminal polyphosphate group of the nucleotide substrate molecule. On one hand, fluorescence switching properties can refer to a significant change in the fluorescence signal after sequencing compared to before the sequencing reaction. On the other hand, fluorescence switching occurs after the nucleotide substrate is polymerase-catalyzed and incorporated into an extension primer. On one hand, the nucleotide sequence fragment to be tested is immobilized on a vector, and then a reaction solution containing a nucleotide substrate molecule is provided to react with the template nucleotide sequence fragment. On the other hand, an enzyme is then used to release a fluorophore from the nucleotide substrate incorporated into the extension primer (and the duplex polymerase extension product), resulting in fluorescence switching.

[0236] On the one hand, after each sequencing reaction, the fluorescence signal may be significantly enhanced or weakened compared to before the sequencing reaction, or the frequency of emitted light may change significantly.

[0237] In any of the foregoing embodiments, sequence errors may include insertions and / or deletions. In any of the foregoing embodiments, a sequence data error can be considered to have occurred at a position when at least two degenerate nucleic acid sequences do not share a common base.

[0238] In any of the foregoing embodiments, correcting sequence errors may include correcting at least one nucleotide residue in a sequence such that the corrected sequence has the correct nucleotide residue at at least one position following the corrected nucleotide residue. On the one hand, the nucleotide residue is correct if the nucleic acid sequence information of any two rounds of sequencing determined at the same nucleotide residue position is not inconsistent with the nucleic acid sequence information of another round of sequencing.

[0239] In any of the foregoing embodiments, correcting sequence errors may include correcting errors in at least one sequence such that common nucleotide residues at at least one position in the sequence can be obtained by comparing sequence information from multiple rounds of sequencing.

[0240] In any of the foregoing embodiments, correcting sequence errors may include extending (e.g., by inserting nucleic acid residues at the location where an error is believed to have occurred) and / or shortening (e.g., by deleting nucleic acid residues at the location where an error is believed to have occurred) sequences representing nucleic acid sequence information from multiple rounds of sequencing. On one hand, by extending and / or shortening at least one sequence from multiple rounds of sequencing, the corrected sequence will be consistent with sequences from other rounds at at least one nucleotide residue location.

[0241] In any of the foregoing embodiments, the memory of a nucleic acid sequence can refer to the fact that, in the sequencing results, the nucleic acid sequence information at a specific position involves not only the nucleotide residues in the corresponding nucleic acid in the template, but also the sequence information preceding that sequence information.

[0242] In any of the foregoing embodiments, using sequencing signals from the other two rounds of sequencing, the sequence in the sequencing signal can be extended (e.g., by inserting nucleic acid residues at locations where errors are believed to have occurred) by a certain length to obtain a corrected nucleic acid sequence. In any of the foregoing embodiments, using sequencing signals from the other two rounds of sequencing, the sequence in the sequencing signal can be shortened (e.g., by deleting nucleic acid residues at locations where errors are believed to have occurred) by a certain length to obtain a corrected nucleic acid sequence.

[0243] In any of the foregoing embodiments, the reaction solution can be divided into three groups based on different bases, wherein the bases include A, G, C, and T bases or A, G, C, and U bases. In any of the foregoing embodiments, the bases can be methylated, hydroxymethylated, or modified with aldehyde or carboxyl groups, or be unmethylated, unhydroxymethylated, or not modified with aldehyde or carboxyl groups.

[0244] In any of the foregoing embodiments, the nucleotide substrate reaction solution may contain different bases and may be divided into two reaction solutions according to the different bases, for example, one reaction solution contains A+G and the other contains C+T; one reaction solution contains A+C and the other contains G+T; or one reaction solution contains A+T and the other contains C+G.

[0245] In any of the foregoing embodiments, the reaction solution may include multiple reaction solutions, one of which may be used for the sequencing reaction. On one hand, one or more reaction solutions are used per round of sequencing. On the other hand, at least one reaction solution contains two or more types of nucleotide substrate molecules with different bases. In any of the foregoing embodiments, the reaction solutions used in different rounds of sequencing contain different combinations of nucleotide substrate molecules.

[0246] In any of the foregoing embodiments, the nucleotide substrate molecule can be labeled with fluorescence. On one hand, a fluorescent group (or a functional group with fluorescence-switching properties acquired through a chemical reaction) is coupled to a base of the nucleotide residue. On the other hand, one of the fluorophores or functional groups can be used to modify the nucleotide substrate molecule, or multiple fluorophores or functional groups can be used to modify the nucleotide substrate molecule with different bases.

[0247] With the deepening understanding of genes in recent years, gene sequencing has brought about tremendous changes in pharmacy and biology. Conventional sequencing methods include Sanger DNA sequencing, restriction fragment length polymorphism (RFLP), single-strand conformation polymorphism (SCM), and allele-specific oligonucleotide hybridization sequencing based on gene chips. Due to various influencing factors during the sequencing process, such as inaccurate CCD luminescence, fluid movement, ambient light, contaminating DNA, signal correction system errors, or impurities in the sequencing reaction solution, errors are inevitable in the sequencing results. As genetic material, DNA stores the genetic information of an organism, a characteristic that makes it suitable as a storage medium for basic information. When DNA is used to store information, this information needs to be encoded into the DNA sequence, and then the information is read using gene sequencing methods. To avoid encoding and / or reading errors, redundant information is often introduced into the encoding process and used for signal correction during reading. For example, George Church et al., “Next-Generation Digital Information Storage in DNA,” Science, 2012, used Reed Solomon codes to encode information into DNA sequences and used the Illumina sequencing platform to read the information from the DNA sequences. DNA encoding-reading techniques are also used in combinatorial chemistry and other fields. In previous DNA encoding techniques, the type of each base was typically independent of bases at other positions (memoryless encoding) or only related to its neighboring bases. This paper presents a memory-based, distributed, orthogonal DNA encoding method where the type of each base is related to all bases in the preceding positions. Furthermore, the method can effectively improve the encoding-reading process and ultimately the decoding accuracy by comprehensively comparing multiple sets of orthogonal codes.

[0248] On one hand, the present invention provides a method for detecting and / or correcting coding errors in sequencing results, wherein the method includes sequencing the same nucleic acid sequence to obtain three or more orthogonal nucleotide degenerate sequence data, wherein errors in the sequence can be detected by comparing the three or more orthogonal nucleic acid degenerate sequences, and wherein a corrected sequence can be obtained by modifying at least one sequence at the location where the error was found during the comparison.

[0249] On the one hand, this invention provides a method for detecting and / or correcting code errors in sequencing results. The method involves sequencing the same nucleic acid sequence to obtain three or more degenerate sequence data represented by the letters M, K, R, Y, W, S, B, D, H, and V. Errors in the sequence can be detected by comparing the three or more degenerate sequences, and a corrected sequence can be obtained by modifying at least one sequence at the location where the error was found during the comparison. On the one hand, the method is applicable to routine sequencing. On the other hand, with proper sequencing substrate design, three or more coding results can be obtained through multiple rounds of sequencing, where redundancy of information can be used to detect and / or correct error codes.

[0250] On one hand, the present invention provides a method for detecting and / or correcting code errors using the memory of genetic code, wherein the method includes sequencing the same nucleic acid sequence to obtain two or more degenerate sequences represented by the letters M, K, R, Y, W, S, B, D, H, V, or obtaining nucleic acid sequence information encoded by A, G, T, C, or nucleic acid sequence information encoded by A, G, U, C, wherein the light or electrical signals caused by different functional groups linked on different bases in the sequencing reaction are used as redundant information to detect sequence errors, wherein a corrected sequence can be obtained by modifying at least one sequence at the location where the error is found during comparison.

[0251] On one hand, the present invention provides a method for detecting and / or correcting code errors using the memory of the gene code, wherein the method includes sequencing the same nucleic acid sequence to obtain three or more orthogonal degenerate nucleotide sequence data, and comprehensively comparing the degenerate sequences and using the memory of the nucleic acid sequences to detect sequence errors, wherein a corrected sequence can be obtained by modifying at least one sequence at the location where the error is found during the comparison, wherein in the degenerate sequences, each sequence signal represents a portion of gene sequence information, and the signal at the same position on another degenerate sequence cannot be inferred from the signal on one such degenerate sequence.

[0252] In any of the foregoing embodiments, the method may include immobilizing the nucleic acid fragment to be tested on a vector, providing a reaction solution to initiate a sequencing reaction, such that each round of sequencing yields a degenerate nucleic acid sequence; after at least N rounds of sequencing, N degenerate nucleic acid sequences are obtained, wherein by comprehensively comparing the N degenerate sequences, the location of an error in the sequence can be detected, wherein by modifying at least one sequence at the location of the error found during comparison, a corrected sequence can be obtained, wherein the reaction solution may contain two or more types of nucleotide substrate molecules with different bases, and wherein N is a positive integer equal to or greater than 2.

[0253] On the one hand, by comparing N-1 degenerate nucleic acid sequences, nucleic acid sequence information encoded by A, G, T, C, or nucleic acid sequence information encoded by A, G, U, C, can be obtained. On the other hand, by comparing N degenerate sequences, the location of sequence errors can be detected. N can be a positive integer equal to or greater than 3.

[0254] On the one hand, by comparing N degenerate nucleic acid sequences, nucleic acid sequence information encoded by A, G, T, and C, or nucleic acid sequence information encoded by A, G, U, and C, can be obtained. Furthermore, by comparing N degenerate sequences, the location of sequence errors can be detected. On the other hand, the location of errors can be detected using luminescent information provided by two or more functional groups linked to the bases, where N is a positive integer equal to or greater than 2. On the other hand, the method includes using changes in the bases themselves as redundant information for correction during the sequencing reaction that generates information from molecules such as phosphate and hydrogen ions released during the reaction.

[0255] On one hand, the present invention provides a method for detecting and / or correcting code errors in sequencing results, wherein the method includes immobilizing the nucleic acid fragment to be tested, providing a reaction solution to initiate a sequencing reaction, wherein the reaction solution for the nucleotide substrate molecules used for sequencing is divided into three groups according to different bases, each group comprising two different reaction solutions, each reaction solution containing nucleotide substrate molecules with different bases. On one hand, there is no overlap between the bases of the nucleotides in the two reaction solutions. On the other hand, one set of reaction solutions is used for each round of sequencing, with the two reaction solutions in each group provided alternately. On the other hand, the method includes performing three rounds of sequencing using the three sets of reaction solutions to obtain three degenerate sequences, the location of the error can be detected by comprehensively comparing the three degenerate sequences, and a corrected sequence can be obtained by modifying at least one sequence at the location of the error during comparison.

[0256] On the one hand, the reaction solution containing two different bases can be divided into two reaction solutions; other steps of the method can be adjusted accordingly.

[0257] On the one hand, the reaction solution may contain multiple reaction solutions, one for each sequencing run, wherein one or more reaction solutions are used per sequencing run, wherein at least one reaction solution contains two or more types of nucleotide substrate molecules with different bases, and wherein the reaction solutions used for different sequencing runs contain different combinations of nucleotide substrate molecules.

[0258] On one hand, the sequencing of the present invention includes sequencing by using a nucleotide substrate molecule with a fluorophore modified with a 5'-terminal polyphosphate to have fluorescence switching properties, wherein fluorescence switching properties refer to a significant change in the fluorescence signal after sequencing compared to the state before the sequencing reaction. In this process, the nucleotide sequence fragment to be tested is first immobilized on a vector, then a reaction solution containing nucleotide substrate molecules is provided, and then an enzyme is used to release the fluorophore on the nucleotide substrate, thereby causing fluorescence switching.

[0259] On the one hand, "significant changes in fluorescence signal after sequencing compared to before sequencing reaction" means that after each sequencing reaction, the fluorescence signal is significantly enhanced or significantly weakened compared to before the sequencing reaction, or the emission light frequency range is significantly changed.

[0260] On the one hand, sequence errors refer to insertion or deletion errors. On the other hand, sequence data errors refer to the situation where at least two nucleic acid sequence pieces do not represent the same base at the same position, which is considered an error. In yet another aspect, the method includes correcting errors in at least one sequence such that subsequent sequences at at least one position are correct, wherein a correct sequence means that the nucleic acid sequence information determined by any two rounds of sequences at the same position does not contradict the nucleic acid sequence information of another round of sequences, or in other words, the nucleic acid sequence information represented by any two rounds of sequences at the same position does not contradict the luminescence information provided by the functional group linked to the base or information from another sequencing process.

[0261] On one hand, the method includes correcting sequences by correcting errors in at least one sequence such that common bases can be obtained by comprehensively comparing sequences at at least one position.

[0262] On the one hand, by modifying at least one sequence, a corrected sequence can be obtained at the location where the error occurs by extending or shortening the sequence representing nucleic acid sequence information. Extension or shortening refers to increasing or decreasing the length of the same detection sequence. When encoding causes shortening or extension at that position, the sequence information represented by the code remains unchanged, resulting in the same code. For example, when the signal intensity of the degenerate code M is 2 (MM), it can be extended to 3 (MMM).

[0263] On the one hand, the memory property of nucleic acid sequences means that the nucleic acid sequence information at a certain position in the sequencing results is related not only to the sequence on the corresponding nucleic acid to be tested, but also to the sequence information preceding it.

[0264] On the one hand, by extending or shortening some sequencing signals at a certain position, the gene sequence represented by that position is extended or shortened in order to obtain a corrected nucleic acid sequence using two other rounds of sequencing signals. Extending the sequencing signals includes adding or inserting a specific length of the gene sequence represented by that position, while shortening some sequencing signals includes shortening or deleting a specific length of the gene sequence represented by that position, and obtaining a corrected nucleic acid sequence using two other rounds of sequencing signals.

[0265] On the one hand, the reaction solution is divided into three groups according to the different bases, where the base refers to A, G, C, T bases or A, G, C, U bases, and the base can be methylated, hydroxymethylated, or have an aldehyde or carboxyl group, or it can be unmethylated, unhydroxymethylated, or have no aldehyde or carboxyl group.

[0266] On the one hand, the reaction solution of nucleotide substrate containing two different bases can be divided into two reaction solutions according to the difference in bases.

[0267] On one hand, nucleotide substrate molecules can be labeled with fluorescence. On the other hand, the method includes modifying a fluorophore or functional group with fluorescence switching by a chemical reaction on the bases of the nucleotide substrate molecule. On the other hand, one of the fluorophores or functional groups can be used to modify the nucleotide substrate molecule, or multiple fluorophores or functional groups can be used to modify the nucleotide substrate molecule with different bases.

[0268] On the one hand, each round of sequencing yields a set of degenerate gene sequence information. On the other hand, degenerate gene sequence information refers to information containing possible gene sequences. For example, when the reaction solution contains nucleotide substrate molecules with A and G bases, the degenerate gene sequence information obtained from sequencing includes the gene sequence information of C and / or T bases in the nucleotide sequence to be tested. When the reaction solution contains nucleotide substrate molecules with T and G bases, the gene sequence information obtained by sequencing includes the gene sequence information of C and / or A bases in the nucleotide sequence to be tested.

[0269] On the one hand, in the comprehensive comparison of information from three rounds of sequencing, if the gene sequence information represented by the signal from one round of sequencing is an oversized erroneous sequencing signal, the gene sequence information represented by that sequence signal can be shortened so that the comparison result of at least one subsequent sequencing signal is correct.

[0270] On the one hand, in the comprehensive comparison of information from three rounds of sequencing, if the gene sequence information represented by the signal from one round of sequencing is a small erroneous sequence signal, then a gap can be added or the sequence can be extended to correct the comparison result of at least one subsequent sequencing signal. For example, when the signal intensity of degenerate code M is 2, i.e., MM, it can be extended to 3, i.e., MMM.

[0271] On the one hand, this paper provides methods for detecting and / or correcting errors in gene sequencing results, particularly sequencing methods using one or more reaction solutions containing nucleotide substrate molecules with two or more bases. In a specific aspect, this method is applicable to SBS (sequencing by synthesis) methods used for sequencing.

[0272] On the one hand, the degenerate gene sequence information in this paper includes possible gene sequence information for a given target (or template) sequence. For example, when the reaction solution contains nucleotide substrate molecules with A and G bases, the degenerate gene sequence information obtained by sequencing includes the gene sequence information of C and / or T bases in the nucleotide sequence to be tested. Assuming that the intensity information obtained from the sequencing reaction is 3, it means that the gene to be tested may contain three Cs and / or Ts, such as three Cs or three Ts, or one C and two Ts, or one T and two Cs, and the exact relative positions of Ts and / or Cs cannot be distinguished based on the degenerate sequence. Degenerate gene sequence information and degenerate codes are terms commonly used in this art.

[0273] On the one hand, the method described in this paper can detect and / or correct errors in sequencing, but it cannot completely eliminate sequence errors. It is possible that a specific location in the sequence signal that has been modified is not the actual location of a sequencing error, but the probability is extremely low. Ultimately, accuracy can be further improved. For example, if the modified signals of MK, RY, and WS are combined, and two modifications are found in N consecutive signals, it is considered highly likely that an error has occurred, and the corresponding sequence should be discarded. Here, N can be a positive integer equal to or greater than 2. The larger the value of N, the higher the probability that the sequence should be discarded, and the higher the final decoding ratio. In this paper, the optimized value of N is 3.

[0274] DNA sequences are copolymers; for example, DNA regions may include two different deoxyribonucleotides, such as AAC and GGTG.

[0275] On the one hand, methods for detecting and / or correcting sequence data errors can detect the location of errors and / or correct sequence errors.

[0276] On the one hand, in the actual sequencing process, the method includes first obtaining the relative intensity values ​​of optical or other signals through a cyclic sequencing reaction. These intensity values ​​can be represented in a specific form. For example, M represents positional information and the number of bases at that position (multiple bases are acceptable), or it can represent the degenerate gene coding result. By decoding the relative intensity values ​​of a sufficient amount of information, the gene sequence information to be tested can be obtained.

[0277] On one hand, delivering or providing reagents or reaction solutions means adding reagents or reaction solutions, such as the reaction mixture for a sequencing reaction, into a container. On the other hand, three or more rounds of sequencing can be used. Alternatively, two or more rounds of sequencing can be used. On the other hand, sequencing signals are counted by number of times. The signal intensity information for each sequencing run can be recorded, and in some embodiments, the intensity information is perfectly identical to the length of the corresponding copolymer.

[0278] Sequencing signals can be counted either by level or by the number of times a specific nucleotide is detected. For example, if the signal intensity is n and the number of nucleotides added to the reaction solution is X, then the sequencing result is represented as XXX...X, where the sequence length is n nucleotides. Figure 1 When the sequencing signal is counted by number, it can be converted into a level-counted sequencing signal MMMKKKKKMKKKMMK or written as (A / C, A / C, A / C, G / T, G / T, G / T, G / T, A / C, G / T, G / T, A / C, A / C and G / T).

[0279] For example, sequencing reaction solutions containing dA4P and dC4P (nucleotides with four phosphate groups and terminal phosphates labeled with fluorescent groups) can be used in odd-numbered cycles, while sequencing reaction solutions containing dG4P and dT4P can be used in even-numbered cycles. A set of fluorescence signal values ​​after multiple reactions can be found in Table 6 below.

[0280] Combinations of other fluorescently labeled nucleotides can be used to obtain fluorescence signal values ​​associated with the target DNA sequence. Examples of possible combinations are shown below:

[0281] M / K pattern: dA4P and dC4P are presented an odd number of times, and dG4P and dT4P are presented an even number of times; or vice versa.

[0282] R / Y mode: dA4P and dG4P are presented an odd number of times, and dC4P and dT4P are presented an even number of times; or vice versa; and

[0283] W / S mode: dA4P and dT4P are presented an odd number of times, and dC4P and dG4P are presented an even number of times; or vice versa.

[0284] Table 6

[0285]

[0286] Sequencing data obtained from three different nucleotide combinations can be combined as a horizontally counted signal. For each position, the next step is to analyze the intersection of the nucleotide types represented by the three horizontally counted sequencing signals at that position to obtain the target DNA sequence. This is, in part, the basic principle of signal decoding. For example, if the horizontally counted sequencing signals correspond to combinations of M / K, R / Y, and W / S such as (3, 5, 1, 3, 2, 1), (2, 4, 3, 2, 1, 3), and (2, 1, 3, 2, 3, 3, 1), the sequence can be summarized as AACTTTGGATTGCCT (SEQ ID NO:1).

[0287] On the one hand, the comprehensive comparison of the results of the three rounds of sequencing involves converting chemiluminescent signals or other forms of intensity signals into gene sequence information, and then comparing the results of the three rounds of sequencing at the same base position. If the results obtained from the three rounds of sequencing are consistent, the sequencing at that position is considered correct; if the gene sequence information represented by the results obtained from the three rounds of sequencing is inconsistent, the sequencing result at that base position is considered incorrect.

[0288] On the one hand, if factors such as inaccurate CCD luminescence, fluid movement, ambient light, contaminating DNA, errors in the signal correction system, or impurities in the sequencing reaction solution cause the sequence signal at a specific time point to be larger or smaller when counted in terms of number of iterations, this will result in an empty intersection of the nucleotide types represented by the corresponding or subsequent positions in the sequencing signal when counted horizontally, making it impossible to resolve the nucleotide types. Clearly, errors in the number-counted sequencing signal can cause the overall level-counted sequencing signal to shift from the position where the error occurred. Therefore, the level-counted sequencing signal is a signal with memory. Based on the memory characteristic of the level-counted sequencing signal, errors in the sequencing signal can be corrected.

[0289] On the one hand, this invention provides a method for detecting and / or correcting sequence data errors in sequencing results. The sequencing reaction solution contains at least two types of nucleotide substrate molecules with different bases; degenerate gene coding information can be obtained. Those skilled in the art can determine whether a conflict occurs in the code at a certain position by comparing two or more degenerate coding information. Compared to the same test substrate, using different primers or directly testing multiple rounds is easier, and the test can be completed with a single test design. On the other hand, the method provided herein is completely different from the method of testing the same gene in multiple rounds. In some aspects, the method provided herein lacks a correction basis if there are only two mutually orthogonal degenerate gene coding results (excluding cases where redundant information such as color is added). On the other hand, this invention first assumes that the detection and correction of errors in three or more mutually orthogonal degenerate codings leads to this sequencing type.

[0290] On the one hand, this article provides a method for detecting and / or correcting sequence data errors in sequencing results. Specifically, sequencing is performed on nucleotide substrate molecules with 5' polyphosphate-modified fluorophores exhibiting fluorescence switching properties; this method is also known as fluorescence switching sequencing. When using fluorescence switching sequencing in conjunction with a 2+2 sequencing method, the sequencing method itself offers many advantages, such as long readouts of 300 bp and sequencing accuracy up to 99.99%; all of these cannot be achieved by using either the 2+2 sequencing method or fluorescence switching sequencing alone. Furthermore, the combined method offers other advantages, such as higher throughput, simpler reaction, lower error rate, and no need for real-time information acquisition. Similarly, sequencing on other nucleotide substrate molecules with fluorescence switching properties exhibits the same characteristics. For example, fluorescence switching sequencing and 2+2 sequencing methods provide redundant information (luminescence or other detectable information) in addition to color information during three rounds of sequencing, which can be used for correction. This can also extend the effective reads without changing accuracy. The correction result depends on the accuracy of the sequencing method, and it can significantly improve the overall accuracy of the effective reads while keeping the sequencer accuracy constant. For example, the accuracy of sequencing a 400 bp nucleic acid fragment is up to 97.36%. The corrected accuracy reaches 99.17%. Therefore, if a sequencer employs this error detection and correction method, the effective reads can be extended accordingly. When using the methods presented in this paper for correction, a clear rule emerges: any small improvement in the sequencing method that can reduce the error rate can significantly reduce the error rate of the modified coding data.

[0291] IV. A method for extracting sequence information from raw signals of high-throughput DNA sequencing

[0292] On the one hand, this document relates to methods for reading nucleic acid sequence information from raw signals or original signals from sequencing reactions, such as high-throughput DNA sequencing reactions. In a particular aspect, this invention relates to methods for reading and / or correcting sequence information from raw signals or original signals from second-generation sequencing technologies (e.g., for gene or genome sequencing). On the one hand, this document considers the many reasons that cause deviations between raw signals and actual sequence information during nucleic acid sequencing to achieve comprehensive correction of the detected sequence information, thereby reading accurate DNA sequences from raw sequencing signals. On the other hand, the methods disclosed herein do not affect the normal process of the sequencing reaction. On the other hand, this document relates to the processing of both monochrome and multicolor sequencing signals. On the other hand, the processing of each type of signal includes parameter estimation and signal correction.

[0293] In high-throughput DNA sequencing, under ideal conditions, the intensity of the raw signal released in each sequencing reaction is directly proportional to the number of bases incorporated into the nascent DNA strand. However, in reality, this ratio does not always exist due to various reasons. For example, firstly, the intensity of the raw signal generally decreases due to fluid corrosion, DNA template hydrolysis, and / or base mismatches. Secondly, due to incomplete sequencing reactions, side reactions (e.g., unwanted reactions), and / or base mismatches, the length of the nascent DNA strand gradually becomes desynchronized as the sequencing reaction progresses (e.g., the length of the nascent DNA strand is inconsistent due to phase loss). The desynchronized length of the nascent DNA strand then leads to a deviation between the raw signal intensity and the actual target DNA sequence. Thirdly, the overall intensity of the raw signal will be high due to spontaneous hydrolysis of nucleotides and / or background fluorescence from the sequencing chip or substrate. All these factors make it difficult, and sometimes impossible, to directly read the target DNA sequence from the intensity of the raw sequencing signal based on the ideal ratio between the two.

[0294] Existing methods for reading sequence information from raw sequencing signals only consider some of the reasons mentioned above. For example, 454 sequencing technology only considers phase loss, the signal bias caused by phase loss during correction matrix transformation. In fact, since the above reasons coexist, considering only phase loss or merely separating phase loss from other factors such as attenuation and overall high values ​​will affect the accuracy of reading DNA sequence information. Furthermore, 454 sequencing technology only considers the primary lead of phase loss, ignoring the secondary lead, which also affects the accuracy of the final result. In addition, the effectiveness of 454 sequencing technology is also affected by many manually set parameters, making the technology inconvenient to use.

[0295] Ion Torrent sequencing technology attempts to mitigate signal bias caused by the aforementioned reasons by altering the order in which nucleotides are added during the sequencing reaction. However, on the one hand, this method only alleviates signal bias, rather than truly corrects it. On the other hand, changing the order of nucleotides added during the sequencing reaction reduces the average sequencing read length per sequencing reaction.

[0296] On the other hand, this paper discloses a sequencing method using nucleotide substrate molecules with fluorophores exhibiting fluorescence switching properties. Sequencing is performed by modifying the 5' end or middle phosphate of the nucleotide substrate molecule with a fluorophore exhibiting fluorescence switching properties. Fluorescence switching properties refer to a significant increase in fluorescence signal intensity after sequencing compared to before the sequencing reaction. Each sequencing run uses a set of reaction solutions, each set comprising at least two portions, each containing at least one of A, G, C, or T nucleotide substrate molecules or at least one of A, G, C, or U nucleotide substrate molecules. First, the nucleotide sequence fragment to be sequenced is immobilized in a reaction chamber, and reaction solutions from the set are added to the chamber. The sequencing reaction can be initiated under appropriate conditions, and fluorescence signals are recorded. Then, an additional portion of reaction solution is provided each time, such that other portions of the same set of reaction solutions are successively provided during the sequencing reaction. Simultaneously, one or more fluorescence signals from each portion of reaction solution are recorded. At least one portion of the reaction solution set contains two or three nucleotide molecules.

[0297] On the one hand, high-throughput sequencing aims to obtain the sequence information of the DNA to be tested by performing a series of enzymatic reactions and detecting the signals released during the reactions. Ideally, if some nascent DNA strands have extended to the nth base, and the nucleotides added to the current enzymatic reaction precisely pair with and are complementary to the (n+1)th and (n+m)th bases of the DNA template to be tested, then the nascent DNA strands in the enzymatic reaction will extend to the (n+m)th base. If the nascent DNA strands in the enzymatic reaction have actually extended beyond the (n+m)th base, then the nascent DNA strands in the enzymatic reaction have been "leaded". If the nascent DNA strands in the enzymatic reaction have not actually extended to the (n+m)th base, then the nascent DNA strands in the enzymatic reaction have been "lagged". "Leaded" and "lagged" phenomena are collectively referred to as phase loss. It should be noted that when the nascent DNA strands extend to the nth base, multiple "leaded" and "lagged" phenomena may have occurred in any possible order.

[0298] like Figure 38 As shown, all newly formed DNA strands have the same length of 1 before the sequencing reaction. The slashed boxes, white boxes, or gray boxes represent nucleotides in the sequence to be tested, respectively. For example, if the slashed box represents A, the white box represents T, and the gray box represents C, then... Figure 38The template sequence shown is ATCCTT. After the sequencing reaction, DNA molecules 1, 3, and 5 were extended, which was normal, and their length was 2. In DNA molecule 2, for example, due to a side reaction (e.g., an undesirable one), "premature" extension occurred, and its length was 3 because the extension exceeded the expected length of 2 nucleotides. In DNA molecule 4, for example, due to incomplete reaction, "lagging" extension occurred, and its length was 1. On the one hand, the lengths of the newly formed DNA strands varied after the sequencing reaction. Figure 38 The five DNA molecules shown are for illustrative purposes only and do not represent five DNA molecules in actual sequencing. In fact, there can be multiple DNA molecules in actual sequencing.

[0299] like Figure 39 As shown, DNA template 1 may have the sequence ATCTTT, and DNA template 2 may have the sequence ATCCTT. After normal extension of polymer A (DNA template 1, normal extension, showing polymer A with the AT sequence), in the same sequencing reaction, polymer A (i.e., AT) can be further extended by a side reaction to generate polymer B (DNA template 1, primary advance, showing polymer B with the ATC sequence). Since only nucleotide T is provided in this sequencing reaction, and the polymer is expected to extend only to position 2 (i.e., with T at position 2), polymer B is "primary advance," having extended to position 3 and having the ATC sequence. It should be noted that in this sequencing reaction, only nucleotide T is provided, not nucleotide C, meaning that C at position 3 could be the result of contamination (e.g., from the previous sequencing reaction), a side reaction, or polymerase error. In this example, polymer B can be further extended to position 4 to generate polymer C (with the ATCT sequence) because nucleotide T is provided in the sequencing reaction; this phenomenon is called "secondary advance." Compare this to DNA template 2, which has C instead of T at position 4. When DNA template 2 is sequenced, the provision of nucleotide T can lead to primary extension of the polymer to position 3 (C) due to a side reaction. However, the probability of another side reaction occurring to add another C at position 4 is negligible. Therefore, DNA template 2 will not extend to position 4, and secondary extension will not occur in DNA template 2.

[0300] sequencing methods

[0301] In some respects, this paper employs DNA sequencing methods. In some embodiments, the method includes immobilizing the DNA to be tested on a solid surface, hybridizing it with one or more sequencing primers, and / or sequentially performing sequencing reactions and detecting the signals released by the reactions. Each reaction includes the following steps: adding a reaction solution containing nucleotides, enzymes, and other reagents necessary for the reaction to a reactor (e.g., a chip) to initiate a specific biochemical reaction; detecting the signals released by the reaction; and / or cleaning the reactor. The added nucleotides can be natural deoxynucleotides or nucleotides with chemically modified groups, but in one respect, they should have a hydroxyl group at their 3' end. The number of nucleotide types added in each reaction can be one, two, or three, but not four (referring to ACGT or ACGU). The union of the nucleotide types added in two adjacent reactions includes all four nucleotide types. For example, if A and G are added in the first reaction, C and T will be added in the second reaction. In another example, if ACG is added in the first reaction, T will be added in the second reaction.

[0302] If two types of nucleotides are added to a reaction, these two types of nucleotides can release the same or different types of signals. If three types of nucleotides are added to a reaction, these three types of nucleotides can release the same or different types of signals. Optionally, two types release the same signal, while the third type releases a different signal. The type of signal in this article refers to the form of the signal (e.g., electrical signal, bioluminescent signal, chemiluminescent signal, etc.), or the color of the optical signal (e.g., green fluorescence signal, red fluorescence signal, etc.), or a combination thereof. For simplicity, a signal in which all nucleotides in a reaction release the same type of signal is called a monochromatic signal; a signal in which all nucleotides in a reaction release more than one type of signal is called a polychromatic signal. The term "color" here is for simplicity only; the type of signal is not limited to different colors of optical signals (e.g., wavelength).

[0303] In some implementations, this document refers to three types of signals with different meanings:

[0304] 1. The ideal signal h refers to the sequencing signal that can be directly inferred under ideal conditions based on the sequence of the DNA to be tested and the order in which nucleotides are added. It directly reflects the sequence information of the DNA.

[0305] 2. A phase-delayed signal s refers to a signal formed by the deviation of an ideal signal h after it has been subjected to a phase-delay phenomenon;

[0306] 3. The predicted raw sequencing signal p refers to the signal formed by the phase mismatch (or phase-mismatch) s after considering the following factors: the number of extended bases, the fold relationship of the sequencing signal intensity, signal attenuation, and overall shift. The predicted raw sequencing signal p is a prediction of the actual raw sequencing signal based on preset parameters;

[0307] 4. The actual raw sequencing signal f refers to the signal directly measured by the instrument in high-throughput DNA sequencing.

[0308] Parameter estimation

[0309] The process of inferring relevant parameters of the sequencing reaction based on one or more reference DNA molecules with known sequences and the actual raw sequencing signal is called parameter estimation. The basic process of parameter estimation is as follows: Figure 41 As shown. Parameter estimation involves a set of parameters that describe the relevant properties of the sequencing response, such as phase loss coefficient, unit signal intensity, attenuation coefficient, and overall offset coefficient.

[0310] First, the method includes inferring an ideal signal h based on the sequence of a reference DNA molecule, and then calculating a phase-mismatched signal s and a predicted raw sequencing signal p based on preset parameters. On one hand, the method includes calculating the correlation coefficient c between p and the actual raw sequencing signal f. On the other hand, the method includes using an optimization method to find a set of parameters that makes the correlation coefficient c reach an optimal value. The correlation coefficient c in this paper includes, but is not limited to, Pearson correlation coefficient, Spearman correlation coefficient, average mutual information, Euclidean distance, Hamming distance, Chebyshev distance, Mahalanobis distance, Manhattan distance, Minkowski distance, and the maximum or minimum absolute value of the corresponding signal difference, etc. The optimization methods discussed here include, but are not limited to, grid search, exhaustive search, gradient descent, Newton's method, Hessian matrix method, and heuristic search. Heuristic search methods include, but are not limited to, genetic algorithms, simulated annealing, ant colony optimization, harmonic optimization, spark optimization, particle swarm optimization, and immune optimization. The correlation coefficients and optimization methods mentioned here are all standard mathematical concepts.

[0311] On one hand, the ideal signal h and the actual raw sequencing signal f can be transformed based on the effects of lead, lag, and / or offset on the sequencing signal. On the other hand, these parameters (e.g., lead, lag, and / or offset) can also be obtained during parameter estimation in the process of inferring the relationship between the ideal signal h and the actual raw sequencing signal f (e.g., based on signals measured from reference sequences with known nucleotide sequences). In some aspects, the estimation process involves the use of matrices (e.g., transformation matrix T) and / or functions (e.g., transformation functions). ).

[0312] If the sequencing data is a monochromatic signal, the calculation is performed directly as described above. If the sequencing data is a multicolor signal, each type of signal is separated from the multicolor signal, and the calculation is performed separately using the method described above.

[0313] On one hand, the method for calculating s using h includes constructing a transformation matrix T based on the characteristics and relevant parameters of h, and then using T to transform h into s. On the other hand, the method for calculating p using s includes constructing a transformation function based on relevant parameters. And d is used to transform s into p. The specific implementation method will be detailed below.

[0314] signal correction

[0315] On the one hand, signal correction includes the process of inferring the sequence information of the DNA to be tested based on the parameters obtained from (1) parameter estimation and (2) the actual raw sequencing signal of the DNA to be tested with an unknown sequence. On the other hand, the basic process of signal correction is as follows: Figure 42 As shown, it can basically be regarded as the inverse process of parameter estimation.

[0316] In a first aspect, the process includes using a transformation function based on the parameters obtained from parameter estimation. The inverse function transforms the actual raw sequencing signal f into a phase-disconnected signal (or phase-mismatched) s. On one hand, the process includes treating s as a zero-order phase-disconnected signal s0, constructing a transformation matrix T1 based on s0 and relevant parameters, and using the generalized inverse matrix of T1 to transform s0 into a first-order phase-disconnected signal s1. On the other hand, the process also includes constructing a transformation matrix T2 based on s1 and relevant parameters, and using the generalized inverse matrix of T2 to transform s1 into a second-order phase-disconnected signal s2. Yet another aspect of the process includes... i Construct the transformation matrix T with relevant parameters i+1 and using T i+1 The generalized inverse matrix will s i Transformed into an (i+1)th order dephase signal s i+1, where i is an integer of 2 or greater. On one hand, the process involves calculating a series of phase-delayed signals s0, s1, s2, ..., s... i+1 ... s j On the one hand, if two adjacent out-of-phase signals s are found during the calculation... i and s i+1 If they are equal, stop the calculation and return s. i As a result of signal correction.

[0317] On the one hand, the aforementioned generalized inverse matrix can also be replaced by Tikhonov regularization.

[0318] If the sequencing data is a monochromatic signal, the calculation is performed directly as described above. If the sequencing data is a multicolor signal, each type of signal is separated from the multicolor signal, and the calculation is performed separately using the method described above.

[0319] The above utilizes transformation functions The process of transforming f into s using the inverse function of T, and the process of using the generalized inverse matrix of T to transform s i Transform into s i+1 The process will be detailed below.

[0320] Construction method of transformation matrix T

[0321] On one hand, the construction of the transformation matrix T depends on a sequencing-related signal X and phase-loss parameters. In parameter estimation, signal a is the ideal signal h; in signal correction, signal x is the phase-loss signal s of various orders. i To improve correction accuracy, signal x can be lengthened by adding several 1s after it; in preferred embodiments, typically 1-100 1s are added. In specific embodiments, 5-10 1s are added. On one hand, the phase decoupling parameters include the lead coefficient ε and the lag coefficient λ.

[0322] On one hand, the construction of the transformation matrix T also includes the construction of the secondary matrix D. On the other hand, assuming the signal x has m values ​​and the sequencing reaction is actually performed n times, then both the transformation matrix T and the auxiliary matrix D have n rows and m columns. For example, in the first row of the auxiliary matrix D, only the first column has an element of 1, and all other elements are 0.

[0323] On one hand, the method includes using the k-th row of the auxiliary matrix D to calculate the k-th row of the transformation matrix T. For the first element of the k-th row of the transformation matrix T:

[0324] 1. If k is odd, then the lag phenomenon should be considered, and the element should be designated as (1-λ)D. 1i ;

[0325] 2. If k is even, then the element is set to 0.

[0326] For the i-th element in the k-th row of the transformation matrix T (excluding the 1st element):

[0327] 1. If k and i have the same parity, then lag should be considered, and the element should be designated as (1-λ)D. ki ;

[0328] 2. If k and i have different parity, then the primary lead phenomenon should be considered, and the element should be designated as ε(1-λ)D. k,i-1 ;

[0329] 3. If the (i-1)th element of signal x is less than 2, then the secondary lead phenomenon should be considered. Based on the calculation results of steps 1 and 2 above, this element should be supplemented by the (i-1)th element T in the same row of the transformation matrix T. k,i-1 .

[0330] On one hand, the method involves using the k-th row of the transformation matrix T to calculate the (k+1)-th row of the auxiliary matrix. In the first row of the auxiliary matrix D, only the element in the first column is 1, and all other elements are 0. For the k-th row of the auxiliary matrix (excluding the first row):

[0331] 1. The first element is the element D in the same column of the previous row of the auxiliary matrix. k-1,i And the elements in the row above and in the same column of the transformation matrix T. k-1,i The difference;

[0332] 2. The element D of the i-th element (excluding the 1st element) in the same row and column of the auxiliary matrix. k-1,i And the elements in the row above and in the same column of the transformation matrix T. k-1,i Based on the difference, add the elements of the previous row and column of the corresponding element in the transformation matrix T. k-1,i-1 .

[0333] Therefore, on the one hand, this paper first defines the value of the first row of the auxiliary matrix D, and then calculates the first row of the transformation matrix based on the first row of the auxiliary matrix D. On the other hand, the method also includes using the first row of the transformation matrix T to calculate the second row of the auxiliary matrix; and using the second row of the auxiliary matrix D to calculate the second row of the transformation matrix T. The values ​​of all elements of the auxiliary matrix and the transformation matrix are obtained in the same way.

[0334] On the one hand, the auxiliary matrix D is introduced only for computational convenience and can be eliminated through conventional mathematical transformations, thus allowing direct calculation of the transformation matrix T.

[0335] In the above calculations, the phase loss parameter is related to the nucleotide type, as well as the row number k and column number i of the element being calculated. In actual calculations, for simplicity, the phase loss coefficients ε and / or λ can be kept constant, or they can be varied with the nucleotide type, row number k, and / or column number i.

[0336] On the one hand, in parameter estimation, the transformation matrix T is obtained according to the preset phase loss coefficient and the ideal signal h, using the calculation method described above. On the other hand, the phase loss signal (or phase mismatch) s is the product of the transformation matrix T and the ideal signal h. If the ideal signal h is represented as a column vector, then s is T multiplied by h; if the ideal signal is represented as a row vector, then s is the transpose of h multiplied by T.

[0337] During parameter correction, the preset phase loss coefficient and the i-th order phase loss signal s can be used as the basis. i The transformation matrix T is obtained according to the above calculation method. On the one hand, the (i+1)th order dephase signal s is the generalized inverse matrix T of the transformation matrix T. + The product of the i-th order phase-depleted signal and the i-th order phase-depleted signal. If s i If s is represented as a column vector, then i+1 For T + Multiply by s i If s i If s is represented as a row vector, then i+1 For s i Multiply by T + The transpose of . The (i+1)th order phase-depleted signal s i+1 After calculating using the above method, further rounding can be performed. Rounding methods include, but are not limited to:

[0338] 1. Rounding: Take the nearest integer value;

[0339] 2. Round up: round to the nearest integer greater than s. i+1 The smallest integer;

[0340] 3. Round down: round to the nearest integer less than s. i+1 The largest integer

[0341] 4. Rounding to zero: If s i+1 If s is greater than 0, round down; if s i+1 If the value is less than 0, it will be rounded up.

[0342] 5. Positive rounding: Round down using any of the methods described above, and then change all non-positive numbers to 1.

[0343] Construction method of transformation function

[0344] On the one hand, transformation function This involves several parameters, including unit signal a (the number of extended bases is proportional to the sequencing signal intensity), attenuation coefficient b, and overall offset c. Parameters a, b, and c in this paper can be single coefficients or a set of coefficients. For example, unit signal a is related to the type of nucleotide and the number of sequencing reactions. In calculations, single values ​​of these parameters can be used for simplicity, or these parameters can be varied with relevant factors for accuracy. Alternatively, some parameters can use single values ​​while others vary with relevant factors.

[0345] Transformation function The forms include, but are not limited to:

[0346]

[0347]

[0348]

[0349]

[0350] In the above function, where and Mathematical functions relating to a, b, and c, including but not limited to constant functions, power functions, exponential functions, logarithmic functions, trigonometric functions, inverse trigonometric functions, floor functions, special functions, and functions generated by the interaction, composition, iteration, or piecewise division of the above functions. In some implementations, special functions include but are not limited to elliptic functions, gamma functions, Bessel functions, and beta functions.

[0351] On the one hand, transformation function The phase-mismatched signal s is transformed into the predicted raw sequencing signal p, i.e. On the one hand, transformation function inverse function The actual raw sequencing signal f is transformed into a phase-disconnected signal (or phase-mismatched signal) s, i.e. The inverse function in this article will be interpreted using the conventional meaning in mathematics.

[0352] Compared to existing methods (e.g., the 454 patent method, such as the one disclosed in US 2011 / 0213563 A1, "System and method to correct out of phase errors in DNA sequencing data by use of arecursive algorithm," published as US 8,364,417), this paper mainly makes the following three improvements. First, the method in this paper involves constructing a transformation matrix that simultaneously considers primary lead, secondary lead, and lag in the phase loss phenomenon, and using this transformation matrix to correct sequencing errors caused by phase loss. Second, the method in this paper addresses signal bias caused by attenuation, phase loss, or overall shift as a whole. This method neither corrects signal bias caused by individual problems nor simply solves problems one by one. Third, the signal correction method is improved, avoiding the introduction of parameter settings that require subjective human judgment, thus improving the robustness and reproducibility of the method. Fourth, the method disclosed in this paper can correct both monochromatic and two-color signals.

[0353] On the one hand, this paper does not consider three levels of advance ( Figure 40 ).

[0354] On the one hand, compared with the methods mentioned in the background section, the method disclosed in this paper has the following effects and advantages:

[0355] 1. In the 2+2 sequencing method, secondary lead is very significant, and the resulting bias cannot be corrected by the 454 patent, which does not consider secondary lead. In this paper, secondary lead is taken into account, thus effectively correcting the signal bias caused by this phenomenon.

[0356] 2. In practical applications, if only a simple linear fitting method is used to read sequence information from the raw sequencing signal, the accuracy of the read will typically be at most about 100 bp. If the method described in this paper is used on the same data, an accurate read of about 350 bp can be achieved, greatly improving sequencing read length and sequencing accuracy. In some implementations, the read accuracy can reach approximately 400 bp, approximately 450 bp, approximately 500 bp, approximately 550 bp, approximately 600 bp, approximately 650 bp, approximately 700 bp, approximately 750 bp, approximately 800 bp, approximately 850 bp, approximately 900 bp, approximately 950 bp, approximately 1000 bp, approximately 1050 bp, approximately 1100 bp, approximately 1150 bp, approximately 1200 bp, approximately 1250 bp, approximately 1300 bp, approximately 1350 bp, and approximately... 1400bp, approximately 1450bp, approximately 1500bp, approximately 1550bp, approximately 1600bp, approximately 1650bp, approximately 1700bp, approximately 1750bp, approximately 1800bp, approximately 1850bp, approximately 1900bp, approximately 1950bp, approximately 2000bp, approximately 2050bp, approximately 2100bp, approximately 2150bp, approximately 2200bp, approximately 2250bp, approximately 2300bp, approximately 2350bp, or approximately 2400bp.

[0357] 3. On the one hand, this paper can correct both monochrome and two-color signals.

[0358] 4. On the other hand, compared to certain methods in the art, such as the Ion Torrent sequencing method (alternative nucleotide streams in sequencing-by-synthesis method) disclosed in US 2014 / 0031238A1 and US Patent No. 9,416,413, this does not affect the normal sequence of samples and / or reagents (e.g., dNTPs or ddNTPs) added for sequencing.

[0359] On the one hand, this paper discloses a method for iteratively generated errors in feedback template molecular sequence data, including: a) detecting multiple signals corresponding to nucleic acid sequences, which are generated due to the introduction of multiple nucleotides into the sequencing reaction; b) using the detected signals to generate quantitative (normalized or digitized) information; c) using parameter estimation to obtain a series of lead and / or lag information; d) using the amount of newly generated nucleotides and the accumulation of secondary lead to obtain phase mismatch; e) using phase mismatch to calculate the amount of newly generated nucleotides in each reaction; and f) repeating steps d) and e) until the amount of newly generated nucleotides in each reaction converges, wherein the parameter estimation refers to inferring lead and / or lag based on a reference sequence and its sequencing signal; wherein secondary lead refers to the occurrence of an extension in the sequencing reaction that does not match the nucleotide substrate of the sequencing reaction, and on this basis, an extension that matches the nucleotide substrate of the sequencing reaction; wherein phase mismatch is due to changes in the sequencing results of lead and / or lag, and wherein the amount of newly generated nucleotides is the extension length of the sequence after the addition of the sequencing reaction solution.

[0360] On one hand, the method further includes obtaining an attenuation coefficient in parameter estimation. On the other hand, the method further includes obtaining an offset in parameter estimation. On yet another hand, the method further includes obtaining unit signal information in parameter estimation.

[0361] On the other hand, this paper discloses a method for iteratively generating errors in feedback template molecular sequence data, including: a) detecting multiple signals corresponding to nucleic acid sequences, which are generated due to the introduction of multiple nucleotides into the sequencing reaction; b) using the detected signals to generate quantitative (normalized or digitized) information; c) using parameter estimation to obtain a series of lead and / or lag amounts, attenuation coefficients, and offsets; d) using the amount of newly generated nucleotides and the accumulation of secondary lead amounts to obtain phase mismatches; e) using phase mismatches to calculate the amount of newly generated nucleotides in each reaction; and f) repeating steps d) and e) until the amount of newly generated nucleotides in each reaction converges, wherein parameter estimation refers to inferring lead and / or lag amounts based on a reference sequence and its sequencing signal; wherein secondary lead amounts refer to extensions that do not match the nucleotide substrate of the sequencing reaction, and extensions that match the nucleotide substrate of the sequencing reaction, on top of which occur; wherein phase mismatches are due to changes in sequencing results caused by lead and / or lag amounts; and wherein the amount of newly generated nucleotides refers to the extension length of the sequence after the addition of the sequencing reaction solution.

[0362] On the one hand, this paper discloses a method for correcting lead in sequencing results using secondary lead, wherein if the signal obtained from a particular reaction in the sequencing results is similar to a unit signal, the method includes correcting the signal using secondary lead; wherein secondary lead refers to the occurrence of an extension in the sequencing reaction that does not match the nucleotide substrate of the sequencing reaction, and then the occurrence of an extension that matches the nucleotide substrate of the sequencing reaction.

[0363] On the one hand, sequencing results include primary lead, which refers to the mismatch between the extension and the nucleotide substrate in the sequencing reaction.

[0364] On the one hand, the impact of subsequent advances includes the impact of secondary advances. Primary advances other than the primary advance will accumulate in subsequent sequencing reactions.

[0365] In any of the foregoing embodiments, the signal obtained by the reaction being close to the unit signal means that the signal obtained by the reaction is close to the unit signal; the preferred reaction can obtain a signal intensity information with a deviation of less than about 60% from the unit information, the further preferred reaction can obtain a deviation of less than about 50% from the above two, the further preferred reaction can obtain a deviation of less than about 40% from the above two, the further preferred reaction can obtain a deviation of less than about 30% from the above two, the further preferred reaction can obtain a deviation of less than about 20% from the above two, the further preferred reaction can obtain a deviation of less than about 10% from the above two, and the further preferred reaction can obtain a deviation of less than about 5% from the above two.

[0366] On the one hand, in the sequencing reaction, the method includes obtaining a corrected sequencing signal using sequencing signals prior to n by feeding back errors generated in the template molecular sequence data when the nth sequencing signal is obtained; then, determining whether there is a secondary lead at that position according to the judgment rule described above.

[0367] In any of the foregoing embodiments, sequencing can be a process of adding sequencing reagents, such as a reaction solution of nucleotides and enzymes, to the nucleic acid sequence to be tested.

[0368] In any of the foregoing embodiments, during sequencing, one, two, three, or four types of nucleotides may be added in each reaction.

[0369] In any of the foregoing embodiments, sequencing can be a three-ends-open sequencing process. One, two, or three types of nucleotides can be added to the sequencing reaction. In any of the foregoing embodiments, the added nucleotides during sequencing can be one or more of A, G, C, and T, or one or more of A, G, C, and U.

[0370] In any of the foregoing embodiments, during sequencing, the detection signal can be an electrical signal, a bioluminescent signal, a chemiluminescent signal, or a combination thereof.

[0371] In any of the foregoing embodiments, the method in parameter estimation may include first inferring an ideal signal h from a reference DNA molecule, then calculating a phase-disconnected signal (or phase mismatch) s and predicting the raw sequencing signal p based on preset parameters, and calculating the correlation coefficient c between p and the actual raw sequencing signal f.

[0372] In any of the foregoing embodiments, the method may include using an optimization method to find a set of parameters such that the correlation coefficient c reaches an optimal value. The found parameters may include lead and / or lag, and may also include one or more of attenuation coefficients, offset, and unit signal.

[0373] In any of the foregoing embodiments, lead and / or lag can refer to the degree of phase loss caused by lead and / or lag in the sequencing reaction.

[0374] In any of the foregoing embodiments, during sequencing, nucleotides can be separated into two groups, and the method may include adding a sequencing reaction solution containing one group of nucleotide molecules in each sequencing reaction.

[0375] Example

[0376] Example 1: Sequencing using the "2+2 monochromatic" method

[0377] To further describe this disclosure, specific embodiments are provided below. Unless otherwise specified, specific parameters, steps, etc., are conventional in the art. These specific embodiments are not intended to limit the scope of the invention.

[0378] For sequencing using the "2+2 monochromatic" method, three sets of reaction solutions were prepared. Each set consisted of two vials, each containing two bases labeled with the same fluorescent group X. For each set, the two vials contained exactly all four bases required for the sequencing reaction. The six vials (two vials per set) were all unique.

[0379] Table 7: Reaction solution in the "2+2 monochromatic" method

[0380] First bottle Second bottle First set AX+CX GX+TX Second set AX+GX CX+TX The third set AX+TX CX+GX

[0381] The complete sequencing process consists of three rounds of sequencing, which can be performed sequentially in any suitable order. Each round of sequencing uses one of the three sets of reaction solutions listed in Table 7. For example, the order of the three rounds could be set 1 → set 2 → set 3, or set 2 → set 3 → set 1, etc. Each round of sequencing uses the three sets of reaction solutions mentioned above, and all other conditions are exactly the same (e.g., the same sequencing primers and reaction conditions are used in all three rounds). Two bottles from the same set of reaction solutions can also be used in any suitable order; for example, the first bottle can be used before or after the second bottle.

[0382] Each round of sequencing includes:

[0383] 1. Hybridize the sequencing primers onto the prepared DNA array.

[0384] 2. Begin the sequencing reaction. Steps 2.1-2.4 can be repeated multiple times.

[0385] 2.1. Add the first bottle of reaction solution (e.g., the first or second bottle of the first set) to the sequencing reaction mixture (e.g., in a flow cell) to allow the reaction to proceed and collect the fluorescence signal from the fluorescent group X.

[0386] 2.2. Clean all residual reaction solution and fluorescent molecules from the flow cell.

[0387] 2.3. Add a second bottle of reaction solution (e.g., the second bottle or the first bottle of the first set) to the sequencing reaction mixture, allow the reaction to proceed, and collect the fluorescence signal.

[0388] 2.4. Clean all residual reaction solution and fluorescent molecules from the flow cell.

[0389] 3. Unwind the extended sequencing primers.

[0390] At this point, a new round of sequencing reactions can begin.

[0391] The solution used in this embodiment can be prepared as follows. The washing solution for the sequencing reaction solution contains: 20 mM Tris-HCl pH 8.8; 10 mM (NH4)2SO4; 50 mM KCl; 2 mM MgSO4; and 0.1% 20. The master solution for the sequencing reaction contains: 20 mM Tris-HCl pH 8.8; 10 mM (NH4)2SO4; 50 mM KCl; 2 mM MgSO4; 0.1% 8000 units / mL Bst polymerase; and 100 units / mL CIP (alkaline phosphatase, bovine intestine).

[0392] The three sequencing reaction solutions were prepared as follows:

[0393] Set 1 (bottles 1A and 1B):

[0394] Bottle 1A: Main solution + 20 μM dA4P-TG + 20 μM dC4P-TG

[0395] Bottle 1B: Main solution + 20 μM dG4P-TG + 20 μM dT4P-TG

[0396] Set 2 (bottles 2A and 2B):

[0397] Bottle 2A: Main solution + 20 μM dA4P-TG + 20 μM dG4P-TG

[0398] Bottle 2B: Main solution + 20 μM dC4P-TG + 20 μM dT4P-TG

[0399] Set 3 (bottles 3A and 3B):

[0400] Bottle 3A: Main solution + 20 μM dA4P-TG + 20 μM dT4P-TG

[0401] Bottle 3B: Main solution + 20 μM dC4P-TG + 20 μM dG4P-TG

[0402] Place the prepared reaction solution and main solution in a 4°C refrigerator or on ice until needed.

[0403] To prepare the sequencing primers, the sequencing primer solution (10 μM primers in 1×SSC buffer) was injected into the sequencing chip, then heated to 90°C, and then cooled to 40°C at a rate of 5°C / min. The sequencing primer solution was then washed away with washing buffer.

[0404] To perform the sequencing reaction, place the sequencing chip on the sequencer. To perform sequencing using the first set of reaction solutions, follow these steps:

[0405] 1. Add 10 mL of cleaning solution to rinse the chip.

[0406] 2. Cool the chip down to 4°C.

[0407] 3. Add 100 μL of reaction solution 1A.

[0408] 4. Heat the chip to 65°C.

[0409] 5. Wait 1 minute.

[0410] 6. Take fluorescence images at an excitation laser wavelength of 473 nm.

[0411] 7. Add 10 mL of cleaning solution to rinse the chip.

[0412] 8. Cool the chip down to 4°C.

[0413] 9. Add 100 μL of reaction solution 1B.

[0414] 10. Heat the chip to 65°C.

[0415] 11. Wait 1 minute.

[0416] 12. Take fluorescence images at an excitation laser wavelength of 473 nm.

[0417] 13. Repeat steps 1-12 50 times to obtain 100 fluorescence signals.

[0418] The second round of sequencing can be performed as follows. First, cool the chip to room temperature. Then add 200 μL of 0.1 M NaOH solution to denature the extended DNA double strands from the first round of sequencing. Then add 10 ml of washing buffer to wash away residual NaOH and denatured DNA single strands.

[0419] Then, the sequencing primers were rehybridized to the DNA array, as described above. The sequencing reaction using the second set of reaction solutions was performed as follows:

[0420] 1. Add 10 mL of cleaning solution to rinse the chip.

[0421] 2. Cool the chip down to 4°C.

[0422] 3. Add 100 μL of reaction solution 2A.

[0423] 4. Heat the chip to 65°C.

[0424] 5. Wait 1 minute.

[0425] 6. Take fluorescence images at an excitation laser wavelength of 473 nm.

[0426] 7. Add 10 mL of cleaning solution to rinse the chip.

[0427] 8. Cool the chip down to 4°C.

[0428] 9. Add 100 μL of reaction solution 2B.

[0429] 10. Heat the chip to 65°C.

[0430] 11. Wait 1 minute.

[0431] 12. Take fluorescence images at an excitation laser wavelength of 473 nm.

[0432] 13. Repeat steps 1-12 50 times to obtain 100 fluorescence signals.

[0433] The third round of sequencing can be performed as follows. First, cool the chip to room temperature. Then add 200 μL of 0.1 M NaOH solution to denature the extended DNA double strands from the second round of sequencing. Then add 10 ml of washing buffer to wash away residual NaOH and denatured DNA single strands.

[0434] Then, the sequencing primers were rehybridized to the DNA array, as described above. The sequencing reaction using the third set of reaction solutions was performed as follows:

[0435] 1. Add 10 mL of cleaning solution to rinse the chip.

[0436] 2. Cool the chip down to 4°C.

[0437] 3. Add 100 μL of reaction solution 3A.

[0438] 4. Heat the chip to 65°C.

[0439] 5. Wait 1 minute.

[0440] 6. Take fluorescence images at an excitation laser wavelength of 473 nm.

[0441] 7. Add 10 mL of cleaning solution to rinse the chip.

[0442] 8. Cool the chip down to 4°C.

[0443] 9. Add 100 μL of reaction solution 3B.

[0444] 10. Heat the chip to 65°C.

[0445] 11. Wait 1 minute.

[0446] 12. Take fluorescence images at an excitation laser wavelength of 473 nm.

[0447] 13. Repeat steps 1-12 50 times to obtain 100 fluorescence signals.

[0448] At this point, the three rounds of sequencing were completed.

[0449] Example 2: Sequencing using the "2+2 two-color" method

[0450] In this embodiment, three sets of reaction solutions were prepared. Each set contained two vials, and each vial contained two types of nucleotide bases. The two nucleotide bases in each vial were labeled with two different fluorophores (so that their emission wavelengths were different) to distinguish the signals from the two nucleotide bases.

[0451] In this embodiment, the two types of fluorophores are X and Y. For each set, the two vials contain exactly all four bases of the sequencing reaction. The six vials (two vials per set) contain no duplicate solutions.

[0452] Table 8: Reaction solution in the "2+2 two-color" method

[0453] First bottle Second bottle First set AX+CY GX+TY Second set AX+GY CX+TY The third set AX+TY CX+GY

[0454] The complete sequencing process consists of three rounds of sequencing, which can be performed sequentially in any suitable order. Each round of sequencing uses one of the three sets of reaction solutions listed in Table 8. For example, the order of the three rounds could be set 1 → set 2 → set 3, or set 2 → set 3 → set 1, etc. Each round of sequencing uses the three sets of reaction solutions listed above, and all other conditions are exactly the same (e.g., the same sequencing primers and reaction conditions are used in all three rounds). Two bottles from the same set of reaction solutions can also be used in any suitable order; for example, the first bottle can be used before or after the second bottle.

[0455] Each round of sequencing includes:

[0456] 1. Hybridize the sequencing primers onto the prepared DNA array.

[0457] 2. Begin the sequencing reaction. Steps 2.1-2.4 can be repeated multiple times.

[0458] 2.1. Add a first bottle of reaction solution (e.g., the first or second bottle of the first set) to the sequencing reaction mixture (e.g., in a flow cell) to allow the reaction to proceed, and collect fluorescence signals from fluorescent group X and fluorescent group Y, respectively.

[0459] 2.2. Clean all residual reaction solution and fluorescent molecules from the flow cell.

[0460] 2.3. Add a second bottle of reaction solution (e.g., the second bottle or the first bottle of the first set) to the sequencing reaction mixture to allow the reaction to proceed, and collect fluorescence signals from fluorescent group X and fluorescent group Y, respectively.

[0461] 2.4. Clean all residual reaction solution and fluorescent molecules from the flow cell.

[0462] 3. Unwind the extended sequencing primers.

[0463] At this point, a new round of sequencing reactions can begin.

[0464] Example 3: Comparative Example

[0465] Comparative Example 1

[0466] In this comparative example, four 3'-blocked nucleotide molecules were used. The 3' blocking group prevents polymerase molecules from using this nucleotide molecule as a substrate for continuous elongation. The 3' blocking group can be cleaved under specific conditions to generate a terminal hydroxyl group. Each nucleotide molecule was labeled with a different fluorescent molecular group. The molecular group used here is not a fluorophore with fluorescence switching properties and can be cleaved under specific conditions. The fluorescent labels are W, X, Y, and Z. The labeled nucleotide monomers are WA, XC, YG, and ZT, respectively.

[0467] Reagent 1 is the main sequencing reaction solution, containing four 3'-terminal-blocked fluorescently labeled nucleotide molecules and a polymerase for polymerase-catalyzed extension using the labeled nucleotide molecules. Reagent 2 is the washing solution. Reagent 3 is the deblocking solution, containing reagents for removing the 3'-terminal blocking groups and fluorescent groups.

[0468] During sequencing, sequencing primers are first hybridized to the template strand. Reagent 1 is mixed with the hybridized template to initiate a polymerase reaction. After the reaction, reagent 2 is used to wash away any unreacted sequencing buffer. Fluorescence signals are collected to determine the nucleotide bases of the sequencing primers added to the polymerase extension reaction. Then, reagent 3 is used to remove all 3' end blocking groups and fluorescent groups. After washing, the template polynucleotides can be used in the next round of sequencing reactions. This sequencing method lacks data redundancy and quality control features.

[0469] Comparative Example 2

[0470] In this comparative example, sequencing reactions were performed using nucleotides with non-fluorescent switching properties. This example is similar to Example 1, except that the fluorescent label is not on the phosphate group. This example involves four types of nucleotide molecules, all of which can be freely extended by polymerase under complementary pairing conditions. The same fluorescent molecular group is labeled on the bases of each nucleotide molecule; this molecular group does not have fluorescent switching properties and can be cleaved under specific conditions. Three sets of reaction solutions are provided, two vials per set. For each set, the two vials contain exactly all four bases required for the sequencing reaction. The six vials (two vials per set) are all unique.

[0471] Table 9: Reaction solutions in Comparative Example 2

[0472] First bottle Second bottle First set AX+CX GX+TX Second set AX+GX CX+TX The third set AX+TX CX+GX

[0473] The complete sequencing process consists of three rounds of sequencing, which are performed sequentially in any suitable order. Each round of sequencing uses the three sets of reaction solutions described above, and all other conditions are exactly the same (e.g., the same sequencing primers and reaction conditions are used in all three rounds).

[0474] Each round of sequencing includes:

[0475] 1. Hybridize the sequencing primers onto the prepared DNA array.

[0476] 2. Begin the sequencing reaction. Steps 2.1-2.8 can be repeated multiple times.

[0477] 2.1. Add the first bottle of reaction solution to the sequencing reaction mixture to allow the reaction to proceed.

[0478] 2.2. Clean all residual reaction solution and fluorescent molecules from the flow cell.

[0479] 2.3. Acquire fluorescence signals from fluorescent groups.

[0480] 2.4. Add reagents to remove the fluorescent labeling group.

[0481] 2.5. Add the second reaction solution to the sequencing reaction mixture to allow the reaction to proceed.

[0482] 2.6. Clean all residual reaction solution and fluorescent molecules from the flow cell.

[0483] 2.7. Acquire fluorescence signals from fluorescent groups.

[0484] 2.8. Add reagents to remove the fluorescently labeled group.

[0485] 3. Unwind the extended sequencing primers.

[0486] Then, a new round of sequencing can begin. The sequencing experiment ends after three rounds of sequencing.

[0487] In this embodiment, a substrate (nucleotide molecule) with non-fluorescent switching properties is used, therefore a cleavage reagent needs to be introduced during the sequencing step to remove the fluorescent label, making the sequencing process longer. Furthermore, molecular scarring is generated and left on the resulting double-stranded DNA molecule, hindering further elongation.

[0488] Example 4: Detecting and / or correcting sequencing errors

[0489] In this embodiment, the single-stranded DNA molecule to be sequenced is immobilized on a solid surface. The immobilization method can be chemical cross-linking, molecular adsorption, etc. The 3' or 5' end of the DNA can be immobilized on the surface. The DNA to be tested contains a fixed fragment with a known sequence that can hybridize complementaryly with sequencing primers. The sequence from the 3' end of this fragment to the 3' end of the DNA to be tested is the region of the sequence to be tested. In this embodiment, the sequence to be tested is 5'-TGAACTTTAGCCACGGAGTA-3' (SEQ ID NO:2).

[0490] First, sequencing primers are hybridized to fragments containing known sequences of the target DNA. Each nucleotide substrate molecule has a fluorescently switching functional group attached to its bases; the number of phosphate molecules is four.

[0491] dG4P and dT4P, along with the corresponding reaction buffer, enzyme, and metal ions, are added to the reaction system to initiate a sequencing reaction that generates a fluorescent signal. The signal is acquired using a CCD (charge-coupled device). The values ​​of these fluorescent signals are recorded. This reaction is recorded as the first reaction.

[0492] The residual dG4P and dT4P from the reaction are washed away. Then, dA4P and dC4P are added to the reaction system to initiate the same sequencing reaction as above, and the fluorescence signal value is recorded. This reaction should be recorded as the second reaction. This method is also known as the monochromatic 2+2 sequencing method.

[0493] Repeat the above process. For odd-numbered reactions, add dG4P and dT4P; for even-numbered reactions, add dA4P and dC4P to obtain a set of sequencing signal values: x = (2, 3, 3, 1, 1, 3, 2, 1, 2, 1).

[0494] For example, the newly synthesized DNA strands from the sequencing reaction can be unwound and washed away using high temperatures or strongly hydrophilic substances (such as urea and formamide). Sequencing primers are then re-hybridized to the template DNA. For odd-numbered reactions, dC4P and dT4P are added; for even-numbered reactions, dA4P and dG4P are added to obtain a set of sequencing signal values: y = (1, 4, 4, 2, 2, 1, 1, 4, 1, 1).

[0495] For example, the newly synthesized DNA strands from the sequencing reaction can be unwound and washed away using high temperatures or strongly hydrophilic substances (such as urea and formamide). Sequencing primers are then re-hybridized to the template DNA. For odd-numbered reactions, dA4P and dT4P are added; for even-numbered reactions, dC4P and dG4P are added to obtain a set of sequencing signal values: z = (1, 1, 2, 1, 4, 3, 1, 3, 1, 1, 2).

[0496] The sequencing signal values ​​were then analyzed based on the type of nucleotide base represented by the signal to obtain sequencing information. For each residue of the target DNA, the common base in the three signals was identified and listed in the table below as the nucleotide residue at that position.

[0497] Table 10: Sequencing results before correction

[0498]

[0499] When analyzing the common bases at each position of the three sets of signals, several positions showed no common bases. This indicates that an error has occurred in the sequence. In this embodiment, changing the second value of signal Y from 4 to 3 and changing the sixth value of signal X from 3 to 4 will result in the signals shown in the table below.

[0500] Table 11: Corrected sequencing results

[0501]

[0502] In the table above, "changing the second value of signal y from 4 to 3" is represented by a strikethrough R, and "changing the sixth value of signal X from 3 to 4" is represented by adding an M (in italics and underline). After these two modifications, all positions of the three sets of signals have common bases, and the sequence composed of these common bases is the DNA sequence to be tested. This result shows that by "encoding" DNA with degenerate indicators (e.g., M, K, R, Y, W, S, B, and D), the method can effectively detect errors that occur during sequencing, while the method of "decoding" the sequence can effectively correct these errors. The short sequence of this embodiment can effectively explain the error correction method provided in this disclosure. The modification method used in this embodiment is the method with the least variation and is also the simplest method to achieve subsequent sequence matching. In practical applications, mathematical models can be constructed to achieve such variations. In practically feasible algorithms, all possible variations are statistically based on probability. After correction of probability parameters, the above variation is the most likely correct variation. On the one hand, this calculation is a simple application of the maximum likelihood method based on Bayesian probability. On the other hand, the calculation method is generally a conventional mathematical method.

[0503] By encoding and decoding DNA sequences, this method can effectively improve sequencing accuracy when applied to DNA sequencing signals. For decoding, the sequencing signal is represented as a weighted graph, such as... Figure 1 As shown. A weighted graph is denoted as G(V, E, W), where V are the nodes of the graph, E are the edges of the graph, and W is the weight of each edge (e.g., a real number). The encoding and decoding process is explained below, assuming the sequencing signal counted by number i is a. i .

[0504] 1) For each signal a i If the nucleotide provided in the i-th sequencing reaction is X, then plot node a of h. i Each node represents an X base.

[0505] 2) This a i The nodes are connected sequentially and in an orderly manner, meaning that the first point in this node points to the second point, the second point points to the third point, and so on.

[0506] 3) The last node of this node has a cycle pointing to itself.

[0507] 4) All nodes in the i-th iteration point to the first node representing the (i+1)-th iteration.

[0508] 5) Assign weights to all edges based on the statistical results of a large amount of sequencing data.

[0509] If a DNA sequence is sequenced once using M / K, R / Y, and W / S combinations respectively, three sequencing signals are obtained. These three sequencing signals can be represented graphically using the methods described above, as shown below. Figure 1 As shown.

[0510] The three sets of signals of sequence 5'-TGAACTTTAGCCACGGAGTA-3 (SEQ ID NO:2) are as follows (including errors):

[0511] M / K: 2, 3, 3, 1, 1, 3, 2, 1, 2, 1

[0512] R / Y: 1, 4, 4, 2, 2, 1, 1, 4, 1, 1

[0513] W / S: 1, 1, 2, 1, 4, 3, 1, 3, 1, 1, 2

[0514] A path in a directed weighted graph is defined as a set of nodes in the directed weighted graph, namely v1v2...v... n This set of nodes can be completely different, or some nodes can be the same (for example, v1 and v2 represent the same node). Furthermore, for any two adjacent nodes v in this set... i and v i+1 In this graph, there exists a directed edge from v. i Point to v i+1 The weight of a path is defined as the sum of the weights of all edges in that path. If each sequencing signal is represented as a weighted graph, then each path in the graph represents a possible DNA sequence. Signal decoding involves finding the maximum common path among all graphs. Specific implementation methods include exhaustive search, greedy search, dynamic programming, and heuristic search.

[0515] Example 5: Detecting and / or correcting sequencing errors

[0516] Following the sequencing method described in Example 4, 5000 DNA sequences of 400 bp in length were decoded; all DNA was divided into 5 groups, with 1000 DNA sequences in each group. The coding accuracy and decoding accuracy are summarized in the table below, based on the sequencing correction method described in Example 4.

[0517] Table 12: Sequencing accuracy

[0518] Group Code accuracy Decoding accuracy 1 0.9736 0.9917 2 0.9813 0.9951 3 0.9878 0.9977 4 0.9953 0.9997 5 0.9973 0.9999

[0519] Clearly, the encoder-decoder method presented in this paper can effectively improve sequencing accuracy. For example, when the error rate is 0.0364 (in other words, the accuracy is 0.9736), the corrected error rate becomes 0.0083 (in other words, the accuracy becomes 0.9917). When the error rate is 0.0047, the corrected error rate becomes 0.0003. By comparison, when the error rate before correction decreases by 7.74 times (0.0364 divided by 0.0047), the corrected error rate will decrease by 27.6 times (0.0083 divided by 0.0003). The overall data shows a clear trend: reducing the sequencing error rate leads to a further reduction in the error rate after correction. In other words, using the correction method disclosed in this paper, any small improvement to the sequencing method that can reduce the error rate can lead to a more significant reduction in the error rate of the corrected sequencing data.

[0520] The encoding accuracy and decoding accuracy of each group are calculated separately and represented using violin plots and box plots, such as... Figure 2 As shown.

[0521] Based on the characteristics of the modified signals during encoding, sequences with a higher probability of correct decoding can be selected, further improving decoding accuracy. The frequency distribution histogram of the number of modified signals in each sequence during decoding in the above data is shown below. Figure 3 As shown in the figure, this frequency distribution histogram has the following characteristics: there is a sharp peak on the left side of the image, while the frequency distribution to the right of the peak exhibits a long tail. If the sequences in the long-tailed distribution region in the figure below are discarded, and only the sequences in the peak region are used for analysis, the decoding accuracy can be further improved by 2-10 times.

[0522] Figure 4 This represents the relationship between the number of signals that occurred incorrectly during encoding and the number of signals that were incorrectly modified during decoding. The horizontal axis represents the number of signals that occurred incorrectly during encoding, and the vertical axis represents the correlation between the number of signals that were incorrectly modified during decoding. The grayscale value of the color indicates the proportion of times that point was counted out of all sequences. Figure 3 In most cases, even if an error occurs during decoding, the modified signal and the signal that actually erred are very close to each other. Therefore, this characteristic can be used to judge the quality of decoding. If a signal and its neighboring signals are not modified during decoding, the base type represented by this signal has extremely high reliability.

[0523] Example 6: Detecting and / or correcting sequencing errors

[0524] In this embodiment, the single-stranded DNA molecule to be sequenced is immobilized on a solid surface. The immobilization method can be chemical cross-linking, molecular adsorption, etc. The 3' or 5' end of the DNA can be immobilized on the surface. The DNA to be tested contains a fixed fragment with a known sequence that can hybridize complementaryly with sequencing primers. The sequence from the 3' end of this fragment to the 3' end of the DNA to be tested is the region of the sequence to be tested. In this embodiment, the sequence to be tested is 5'-TGAACTTTAGCCACGGAGTA-3' (SEQ ID NO:2).

[0525] First, sequencing primers are hybridized to a fragment containing a known sequence of the target DNA. Four types of dNTPs, along with corresponding reaction buffers, enzymes, and metal ions, are added to the reaction system. The 3' end of each type of dNTP is blocked by a chemical group. Furthermore, dGTP and dTTP are each labeled with a fluorescent group of the same color, while dATP and dCTP are each labeled with a fluorescent group of a different type of dNTP of the same color. During the reaction, dNTPs complementary to the bases at the extension site on the template DNA are incorporated into the nascent DNA strand by DNA polymerase. After the reaction, residual dNTPs are washed away, and the fluorescence signal is recorded using a CCD. The above reaction is repeated to obtain a set of sequencing signal values: x = KKMMMKKKMKMMMKKMKKM.

[0526] For example, the newly synthesized DNA strands from the sequencing reaction can be unwinded and washed away using high temperatures or strongly hydrophilic substances (such as urea and formamide). Sequencing primers are then rehybridized to the DNA template, and the sequencing process is repeated, but dCTP and dTTP are labeled with the same color fluorescent group, while dATP and dGTP are labeled with a different color fluorescent group. The sequenced signal value is obtained as: y = YRRRRYYYYRRYYRYRRRRYR.

[0527] For example, the newly synthesized DNA strands from the sequencing reaction can be unwound and washed away using high temperatures or strongly hydrophilic substances (such as urea and formamide). Sequencing primers are then re-hybridized to the DNA template, and the sequencing process is repeated, but dATP and dTTP are labeled with the same color fluorescent group, and dCTP and dGTP are labeled with a different color fluorescent group. The value of the sequencing signal is obtained: z = WSWWSWWWWSSSWSSSWSWW.

[0528] The sequencing signal values ​​were then analyzed based on the type of nucleotide base represented by the signal to obtain sequencing information. For each residue of the target DNA, the common base in the three signals was identified and listed in the table below as the nucleotide residue at that position.

[0529] Table 13: Sequencing results before correction

[0530] signal x K K M M M K K K M K M M M K K M K K M Signal y Y R R R R Y Y Y Y R R Y Y R Y R R R R Y R signal z W S W W S W W W W S S S W S S S W S W W common base T G A A ? T T T ? G ? C ? G ? ? ? G A ? ?

[0531] When analyzing the common bases at each position of the three sets of signals, several positions showed no common bases. This indicates that an error has occurred in the sequence. In this embodiment, changing the second value of signal Y from 4 to 3 and changing the sixth value of signal X from 3 to 4 will result in the signals shown in the table below.

[0532] Table 14: Corrected sequencing results

[0533]

[0534] In the table above, "changing the second value of signal y from 4 to 3" is represented by a strikethrough R, and "changing the sixth value of signal X from 3 to 4" is represented by adding an M (in italics and underline). After these two modifications, all positions in the three signal groups have common bases, and the sequence composed of these common bases is the DNA sequence to be tested. This result indicates that by "encoding" DNA with degenerate indicators (e.g., M, K, R, Y, W, S, B, and D), the method can effectively detect errors that occur during sequencing, while the method of "decoding" the sequence can effectively correct these errors.

[0535] Example 7: Detecting and / or correcting sequencing errors

[0536] In this embodiment, the DNA to be tested contains a fixed fragment with a known sequence that can hybridize complementaryly with the sequencing primers. The region from the 3' end of this fragment to the 3' end of the DNA to be tested is the region of the test sequence. In this embodiment, the test sequence is 5'-TGAACTTTAGCCACGGAGTA-3' (SEQ ID NO:2).

[0537] First, sequencing primers are hybridized to fragments containing known sequences of the target DNA. The reaction volume containing the template DNA molecule with the hybridization sequencing primers is divided into three portions, which can be measured in parallel or sequentially. Each portion contains four types of dNTPs, certain types of ddNTPs, and the enzymes and buffers necessary for DNA synthesis. In some respects, the added dNTPs are natural dNTPs, and the added ddNTPs have detectable labels (e.g., labels that can be detected by instruments), including but not limited to radioactive isotope labels, chemifluorescent group labels, etc. In the first portion, ddGTP and ddTTP have the same label, while ddATP and ddCTP have another identical label. In the second portion, ddCTP and ddTTP have the same label, and ddATP and ddGTP have another identical label. In the third portion, ddATP and ddTTP have the same label, and ddCTP and ddGTP have another identical label.

[0538] All three samples were reacted under suitable conditions for a period of time, during which DNA synthesis occurred. After the reaction was complete, the reaction products could be optionally washed or purified. Then, DNA electrophoresis was performed on the three reaction products. Based on the electrophoretic bands, three sequencing signals could be obtained:

[0539] x=KKMMMKKKMKMMMKKMKKM

[0540] y=YRRRRYYYYRRYYRYRRRRYR

[0541] z = WSWWSWWWWSSSWSSSWSWW

[0542] The sequencing signal values ​​are then analyzed based on the type of nucleotide base represented by the signal to obtain sequencing information. For each residue of the target DNA, the common base in the three signals is identified and listed in the table below as the nucleotide residue at that position.

[0543] Table 15: Sequencing results before correction

[0544] signal x K K M M M K K K M K M M M K K M K K M Signal y Y R R R R Y Y Y Y R R Y Y R Y R R R R Y R signal z W S W W S W W W W S S S W S S S W S W W common base T G A A ? T T T ? G ? C ? G ? ? ? G A ? ?

[0545] When analyzing the common bases at each position of the three sets of signals, several positions showed no common bases. This indicates that an error has occurred in the sequence. In this embodiment, changing the second value of signal Y from 4 to 3 and changing the sixth value of signal X from 3 to 4 will result in the signals shown in the table below.

[0546] Table 16: Corrected sequencing results

[0547]

[0548] In the table above, "changing the second value of signal y from 4 to 3" is represented by a strikethrough R, and "changing the sixth value of signal X from 3 to 4" is represented by adding an M (in italics and underline). After these two modifications, all positions of the three signal groups have common bases, and the sequence composed of these common bases is the DNA sequence to be tested. This result indicates that by "encoding" DNA with degenerate indicators (e.g., M, K, R, Y, W, S, B, and D), the method can effectively detect errors that occur during sequencing, while the method of "decoding" the sequence can effectively correct these errors.

[0549] Example 8: Sequencing was performed using a "2+2 two-color three-round" method.

[0550] In this embodiment, the single-stranded DNA molecule to be sequenced is immobilized on a solid surface. The immobilization method can be chemical cross-linking, molecular adsorption, etc. The 3' or 5' end of the DNA can be immobilized on the surface. The DNA to be tested contains a fixed fragment with a known sequence that can hybridize complementaryly with sequencing primers. The sequence from the 3' end of this fragment to the 3' end of the DNA to be tested is the region of the sequence to be tested. In this embodiment, the sequence to be tested is 5'-TGAACTTTAGCCACGGAGTA-3' (SEQ ID NO:2).

[0551] First, sequencing primers are hybridized to fragments containing known sequences of the target DNA. dG4P and dT4P (each labeled with a different fluorescent group, e.g., fluorescent group X and group Y), along with the appropriate reaction buffer, enzyme, and metal ions, are added to the reaction system to initiate a sequencing reaction that produces a fluorescent signal. The signal is acquired using a CCD. The values ​​of these fluorescence signals are recorded. This reaction is recorded as the first reaction.

[0552] Then, the residual dG4P and dT4P from the reaction are washed away. Next, dA4P and dC4P (each labeled with a fluorescent group emitting a different color, such as fluorescent group X and group Y) are added to the reaction system to initiate the same sequencing reaction as above, and the fluorescence signal values ​​are recorded. This reaction should be recorded as the second reaction.

[0553] Repeat the above process. For odd-numbered reactions, add dG4P and dT4P; for even-numbered reactions, add dA4P and dC4P. Each reaction adds two types of dN4P labeled with different colored fluorescent groups. A set of signal values ​​can be obtained: x = (1G+1T, 2A+1C, 0G+3T, 1A+0C, 1G+0T, 1A+2C, 2G+0T, 1A+0C, 1G+1T, 1A+0C).

[0554] For example, the newly synthesized DNA strands from the sequencing reaction can be unwound and washed away using high temperatures or strongly hydrophilic substances (such as urea and formamide). Sequencing primers are then re-hybridized to the template DNA. Two types of dN4P labeled with different colored fluorescent groups are added for each reaction. A set of signal values ​​can be obtained: y = (0C+1T, 3A+1G, 1C+3T, 1A+1G, 2C+0T, 1A+0G, 1C+0T, 1A+3G, 0C+1T, 1A+0G).

[0555] For example, the newly synthesized DNA strands from the sequencing reaction can be unwound and washed away using high temperatures or strongly hydrophilic substances (such as urea and formamide). Sequencing primers are then re-hybridized to the template DNA. For odd-numbered reactions, dG4P and dT4P are added; for even-numbered reactions, dA4P and dC4P are added. Each reaction uses two types of dN4P labeled with different colored fluorescent groups. This yields a sequencing signal: z = (0A+1T, 0C+1G, 2A+0T, 1C+0G, 1A+3T, 2C+1G, 1A+0T, 0C+1G, 1A+1T).

[0556] This method is called the "2+2 two-color" sequencing method. Sequence information can be obtained from the sequencing data of any two rounds of sequencing. It can be considered as orthogonal sequencing results.

[0557] The sequencing signal values ​​were then analyzed based on the type of nucleotide base represented by the signal to obtain sequencing information. For each residue of the target DNA, the common base in the three signals was identified and listed in the table below as the nucleotide residue at that position.

[0558] Table 17: Sequencing results before correction

[0559]

[0560] When analyzing the common bases at each position of the three sets of signals, no common bases were found at several positions, indicating an error in the sequence. Changing the second value of signal y (3A+1G) to (2A+1G) and the sixth value of signal X (1A+2C) to (1A+3C) will result in the signals shown in the table below.

[0561] Table 18: Corrected sequencing results

[0562]

[0563]

[0564] In the table above, "the second value of signal y (3A+1G) changed to (2A+1G)" is represented by a strikethrough A, and "the sixth value of signal x (1A+2C) changed to (1A+3C)" is represented by adding a C (in italics and underlined). After these two modifications, all positions of the three signal groups have common bases, and the sequence composed of these common bases is the DNA sequence to be tested. This result indicates that by "encoding" DNA with degenerate indicators (e.g., M, K, R, Y, W, S, B, and D), the method can effectively detect errors that occur during sequencing, while the method of "decoding" the sequence can effectively correct these errors.

[0565] Example 9: Sequencing using the "2+2 two-color two-round" method

[0566] In this embodiment, the single-stranded DNA molecule to be sequenced is immobilized on a solid surface. The immobilization method can be chemical cross-linking, molecular adsorption, etc. The 3' or 5' end of the DNA can be immobilized on the surface. The DNA to be tested contains a fixed fragment with a known sequence that can hybridize complementaryly with sequencing primers. The sequence from the 3' end of this fragment to the 3' end of the DNA to be tested is the region of the sequence to be tested. In this embodiment, the sequence to be tested is 5'-TGAACTTTAGCCACGGAGTA-3' (SEQ ID NO:2).

[0567] First, sequencing primers are hybridized to fragments containing known sequences of the target DNA. dG4P and dT4P (each labeled with a different fluorescent group, e.g., fluorescent group X and group Y), along with the appropriate reaction buffer, enzyme, and metal ions, are added to the reaction system to initiate a sequencing reaction that produces a fluorescent signal. The signal is acquired using a CCD. The values ​​of these fluorescence signals are recorded. This reaction is recorded as the first reaction.

[0568] Then, the residual dG4P and dT4P from the reaction are washed away. Next, dA4P and dC4P (each labeled with a fluorescent group emitting a different color, such as fluorescent group X and group Y) are added to the reaction system to initiate the same sequencing reaction as above, and the fluorescence signal values ​​are recorded. This reaction should be recorded as the second reaction.

[0569] Repeat the above process. For odd-numbered reactions, add dG4P and dT4P; for even-numbered reactions, add dA4P and dC4P. Each reaction adds two types of dN4P labeled with different colored fluorescent groups. A set of signal values ​​can be obtained: x = (1G+1T, 2A+1C, 0G+3T, 1A+0C, 1G+0T, 1A+2C, 2G+0T, 1A+0C, 1G+1T, 1A+0C).

[0570] For example, the newly synthesized DNA strands from the sequencing reaction can be unwound and washed away using high temperatures or strongly hydrophilic substances (such as urea and formamide). Sequencing primers are then re-hybridized to the template DNA. Two types of dN4P labeled with different colored fluorescent groups are added for each reaction. A set of signal values ​​can be obtained: y = (0C+1T, 3A+1G, 1C+3T, 1A+1G, 2C+0T, 1A+0G, 1C+0T, 1A+3G, 0C+1T, 1A+0G).

[0571] The sequencing signal values ​​were then analyzed based on the type of nucleotide base represented by the signal to obtain sequencing information. For each residue of the target DNA, the common base in the three signals was identified and listed in the table below as the nucleotide residue at that position.

[0572] Table 19: Sequencing results before correction

[0573]

[0574] When analyzing the common bases at each position of the two sets of signals, no common bases were found at several positions, indicating an error in the sequence. Changing the second value of signal y (3A+1G) to (2A+1G) and the sixth value of signal X (1A+2C) to (1A+3C) will result in the signals shown in the table below.

[0575] Table 20: Corrected sequencing results

[0576]

[0577]

[0578] In the table above, "the second value of signal y (3A+1G) changed to (2A+1G)" is represented by a strikethrough A, and "the sixth value of signal x (1A+2C) changed to (1A+3C)" is represented by adding a C (in italics and underlined). After these two modifications, all positions of the three signal groups have common bases, and the sequence composed of these common bases is the DNA sequence to be tested. This result indicates that by "encoding" DNA with degenerate indicators (e.g., M, K, R, Y, W, S, B, and D), the method can effectively detect errors that occur during sequencing, while the method of "decoding" the sequence can effectively correct these errors.

[0579] Example 10: Sequencing using the "1+3, monochromatic" method

[0580] In this embodiment, the single-stranded DNA molecule to be sequenced is immobilized on a solid surface. The immobilization method can be chemical cross-linking, molecular adsorption, etc. The 3' or 5' end of the DNA can be immobilized on the surface. The DNA to be tested contains a fixed fragment with a known sequence that can hybridize complementaryly with sequencing primers. The sequence from the 3' end of this fragment to the 3' end of the DNA to be tested is the region of the sequence to be tested. In this embodiment, the sequence to be tested is 5'-TGAACTTTAGCCACGGAGTA-3' (SEQ ID NO:2).

[0581] First, sequencing primers are hybridized to fragments containing known sequences of the target DNA. dC4P, dG4P, and dT4P, along with the corresponding reaction buffers, enzymes, and metal ions, are added to the reaction system to initiate a sequencing reaction that generates a fluorescent signal. The signal is acquired using a CCD. The values ​​of these fluorescence signals are recorded. This reaction is recorded as the first reaction.

[0582] Then, the residual dC4P, dG4P, and dT4P from the reaction are washed away. Next, dA4P is added to the reaction system to initiate the same sequencing reaction as described above, and the fluorescence signal value is recorded. This reaction should be recorded as the second reaction.

[0583] Repeat the above process. For odd-numbered reactions, add dC4P, dG4P, and dT4P; for even-numbered reactions, add dA4P. Obtain a set of signal values: x = (2, 2, 4, 1, 3, 1, 3, 1, 2, 1).

[0584] For example, the newly synthesized DNA strands from the sequencing reaction are unwound and washed away using high temperature or strongly hydrophilic substances (such as urea and formamide). Sequencing primers are then re-hybridized to the template DNA. dA4P, dG4P, and dT4P are added for odd-numbered reactions, and dC4P is added for even-numbered reactions. A set of signal values ​​is obtained: y = (4, 1, 6, 2, 1, 1, 6).

[0585] For example, the newly synthesized DNA strands from the sequencing reaction are unwound and washed away using high temperature or strongly hydrophilic substances (such as urea and formamide). Sequencing primers are then re-hybridized to the template DNA. dA4P, dC4P, and dT4P are added for odd-numbered reactions, and dG4P is added for even-numbered reactions. A set of signal values ​​is obtained: z = (1, 1, 7, 1, 4, 2, 1, 1, 2).

[0586] For example, the newly synthesized DNA strands from the sequencing reaction can be unwound and washed away using high temperature or a strongly hydrophilic substance (such as urea and formamide). Sequencing primers are then re-hybridized to the template DNA. dT4P is added for odd-numbered reactions, and dA4P, dC4P, and dG4P are added for even-numbered reactions. A set of signal values ​​is obtained: w = (1, 4, 3, 9, 1, 1).

[0587] The sequencing signal values ​​were then analyzed based on the type of nucleotide base represented by the signal to obtain sequencing information. For each residue of the target DNA, the common base in the three signals was identified and listed in the table below as the nucleotide residue at that position.

[0588] Table 21: Sequencing results before correction

[0589] signal x B B A A B B B B A B B B A B B B A B B A Signal y D D D D C D D D D D D C C D C D D D D D D signal z H G H H H H H H H G H H H H G G H G H H Signal w T V V V V T T T V V V V V V V V V T V common base T G A A C T T T A G ? C ? ? ? G A ? ? ? ?

[0590] When analyzing the common bases at each position of the two sets of signals, no common bases were found at several positions, indicating an error in the sequence. Changing the third value of signal y from 6 to 5 and the fourth value of signal w from 9 to 10 will alter the signals as shown in the table below.

[0591] Table 22: Corrected sequencing results

[0592]

[0593] In the table above, "changing the third value of signal y from 6 to 5" is represented by a strikethrough D, and "changing the fourth value of signal w from 9 to 10" is represented by adding a V (in italics and underline). After these two modifications, all positions in the four signal groups have common bases, and the sequence composed of these common bases is the target DNA sequence. This result indicates that by "encoding" DNA with degenerate indicators (e.g., M, K, R, Y, W, S, B, and D), the method can effectively detect errors occurring during sequencing, while the method of "decoding" the sequence can effectively correct these errors.

[0594] Example 11: Methods for detecting and / or correcting sequencing errors

[0595] Section 1: Substrate Synthesis and Spectral Properties

[0596] General: All anhydrous solvents were freshly distilled using standard procedures (Na or CaH2). Unless otherwise specified, reagents were used as received from the commercial supplier. Air and / or moisture sensitivity tests were performed under argon atmosphere. Mass spectrometry was performed using a Bruker APEX IV mass spectrometer and an AB Sciex MALDI-TOF5800 mass spectrometer. Reversed-phase HPLC was performed on a Shimadzu LC-20A HPLC system. Samples were dissolved in water and analyzed using an Inertsil ODS-3C18 column (250 × 4.6 mm, 5 μm) at a flow rate of 1 mL / min, with B (CH3CN) in A (50 mM TEAA pH 7.3) (0-20% B for 15 min, 20-30% B for 10 min).

[0597] 1.1 Synthesis of terminal phosphate-labeled fluorescent nucleotides (TPLFN)

[0598] Figure 5A -C shows how altering the fluorophore structure can improve the fluorescence performance of TPLFN. Figure 5A The previously developed Me-FAM-tagged nucleotides are shown. Figure 5B The previously developed Me-HCF-tagged nucleotides are shown. Figure 5C The TG-labeled nucleotides in this embodiment are shown.

[0599] For fluorescence sequencing purposes, the fluorophore used to label the terminal phosphate of nucleotides plays a crucial role. On the one hand, the phosphorylated fluorophore must be completely quenched, meaning that no fluorescence emission should be detected at a specific excitation wavelength. However, once the fluorophore is released, sufficient signal detection requires a strong fluorescence emission intensity. Based on this principle, Me-FAM was chosen as the labeling dye molecule previously reported. Figure 5A See Sims, PA; Greenleaf, WJ; Duan, H.; Xie, X. "Fluorogenic Pyrosequencing in PDMS Microreactors" Nature Methods 2011, 8, 575–580). Subsequently, the chlorinated form of Me-FAM, called Me-HCF, develops color with a significant red shift in excitation and emission wavelengths, suitable for multicolor sequencing purposes. Figure 5B Chen, Z.; Duan, H.; Qiao, S.; Zhou, W.; Qiu, H.; Kang, L.; Xie, X.; Huang, Y. Fluorogenic Sequencing using Halogen-Fluorescein Labeled Nucleotides. Chembiochem, 2015, DOI:10.1002 / cbic.201500117. Despite successful applications, the fluorescence properties of Me-FAM and Me-HCF (derived from FAM and HCF 3'-OH methylation) remain problematic, as shown by the parameters listed in Figure 5. 3'-OH methylation (or other protecting groups) is a prerequisite for the generation of fluorescent substrates, not only broadening the absorption and emission spectra but also significantly reducing the extinction coefficient and quantum yield, especially for Me-FAM. Therefore, there is still a strong need to develop fluorophores with better fluorescence properties.

[0600] TG (Tokyo Green) was developed by Nagano et al. (Y. Urano, M. Kamiya, K. Kanda, T. Ueno, K. Hirose, T. Nagano, Evolution of fluorescein as a platform for finely tunable fluorescence probes, J. Am. Chem. Soc., 2005, 127, 4888–4894). TG has shown excellent fluorescence properties. The unique structure of TG compared to 5(6)-FAM is the use of a methyl group instead of the carboxyl group in the benzene moiety to maintain the orthogonal relationship between the benzene ring and the fluorophore. Furthermore, phosphorylated TG has been shown to have excellent fluorescence properties. Another advantage is that, compared to the two phenolic groups in 5(6)-FAM or HCF, the single phenolic group in the TG structure promotes the synthesis of TPLFN because methylation is not required. The absence of this protective methyl group not only facilitates the synthesis of TPLFN, but also preserves its original high extinction coefficient and high quantum yield once the TG fluorophore is released via enzymatic digestion, resulting in a significantly higher fluorescence / background contrast. The detailed synthetic procedure is described below:

[0601] (I) Preparation of TG-monophosphate (S2)

[0602]

[0603] Based on the report, Tokyo Green S1 was synthesized using a procedural method [Y.Urano, M.Kamiya, K.Kanda, T.Ueno, K.Hirose, T.Nagano, Evolution of fluorescein as a platform for finely tunable fluorescence probes, J.Am.Chem.Soc., 2005, 127, 4888–4894].

[0604] S1 (332 mg, 1.00 mmol) was suspended in 15 mL of anhydrous CH2Cl2 in a flame-dried flask under an Ar atmosphere. A proton sponge (759 mg, 3.50 mmol) was added to the solution with stirring. After 10 minutes, the mixture was cooled to -10 °C, and phosphorus oxychloride (V) (275 μL, 3.00 mmol) was added. The reaction was maintained at the same temperature for 30 minutes. Then, TEAA buffer (20 mL of 1 M solution) was added to quench the reaction and hydrolyze the phosphoryl chloride intermediate at 0 °C for 1 hour. The two phases were then separated, and the aqueous solution was filtered under vacuum and concentrated for further purification by a reversed-phase rapid LC system. Conditions: AQ C-18 column (Agela 40 g), using 0-50% acetonitrile (pH 7.4) in 50 mM triethylammonium acetate buffer, flow rate 20 mL / min. The fraction containing the pure product was concentrated and co-evaporated twice with anhydrous DMF (2 mL). It was then dissolved in a specific amount of anhydrous DMF, and the resulting monophosphate S2 (100 mL DMF solution) was stored at -20°C for later use. MS (ESI): C 21 H 15 Calculated value of O7P(MH): 411.06. Measured value (m / z): 411.21.

[0605] (II) Synthesis of dN4P-δ-TG(TPLFN)

[0606]

[0607] 1) dA4P-δ-TG: Disodium 2′-deoxyadenosine-5′-triphosphate (dATP) (12.5 μL 100 mM solution, 12.5 μmol) was converted to tributylammonium salt by treatment with ion exchange resin (BioRad AG-50W-XB) and tributylamine. After removing water using an oil pump on a rotary evaporator, the obtained tributylammonium salt was co-evaporated twice with anhydrous DMF (1 mL), and then dissolved in 0.5 mL of anhydrous DMF under Ar. Carbonyl diimidazole (CDI, 10.1 mg, 63 μmol) was added to the solution, and the mixture was stirred at room temperature for 12 h. Then, MeOH (3.2 μL) was added, and the solution was stirred for 0.5 h. Then, using a syringe, 0.25 mL of the TG-tributylammonium monophosphate S2 (25 μmol) DMF solution from the previous step was transferred to the reaction mixture, followed by the addition of MgBr2 (25 mg, 100 μmol) in 0.5 mL of DMF. The mixture was stirred at room temperature for 30 h. The reaction mixture was then concentrated using an oil pump, diluted with water, and purified on a C18 reversed-phase HPLC system (Shimadzu) using a preparative Sepax Amethyst C18-H (21.2 x 150 mm) at a flow rate of 5 mL / min, with a gradient of B (CH3CN) in A (50 mM TEAA pH 7.3) (0–20% B for 15 min, 20–30% B for 10 min, 30–50% B for 10 min). The desired fractions were collected and concentrated using a Hi-Trap Q-HP 5 mL anion exchange column (GE Healthcare). The collected solution containing the desired product was purified again by HPLC under the same elution conditions and concentrated using a Hi-Trap Q-HP column. The product solution was stored at -20°C for later use. MS (MALDI-TOF): C 31 H 31 N5O 18 Calculated value of P4: 895.0615. Measured value: m / z: 884.1019 (MH). dC4P-δ-TG, dT4P-δ-TG, and dG4P-δ-TG were synthesized using the same procedure as dA4P-δ-TG. dC4P-δ-TG: MS (MALDI-TOF): C 30 H 31 N3O 19 Calculated value of P4: 861.0502. Measured value (m / z): 860.0732 (MH). dT4P-δ-TG: MS (MALDI-TOF): C 31 H 32 N2O 20 Calculated P4 value: 876.0499. Measured value: m / z: 875.0706 (MH). dG4P-δ-TG: MS (MALDI-TOF): C31 H 31 N5O 19 Calculated value of P4: 901.0564. Measured value: m / z: 900.0903 (MH). Figure 6 The MALDI-TOF mass spectra of purified TPLFN are shown.

[0608] 1.2. Spectral properties of fluorophores and TPLFN

[0609] The excitation / emission spectrum of TG(S1) is shown in Figure 7 Although Me-FAM has a similar maximum emission wavelength to TG, its extinction coefficient and quantum yield are much lower. Figure 8 Meanwhile, the broad emission spectrum of Me-FAM severely overlaps with other fluorophores such as Me-HCF, making it unsuitable for future multicolor sequencing applications. In contrast, the strong fluorescence and narrower spectrum of TG would more easily solve this problem. Figure 7 The excitation and emission spectra of TG (Tokyo Green) are shown. Figure 8 The emission spectra of TG (Tokyo Green), Me-FAM, and Me-HCF are shown under the same conditions (2 μM, pH 8.3, TE buffer, calculated using area normalization). The optical properties of TG, Me-FAM, and Me-HCF are listed in the table below (and...). Figure 5A -C (in Chinese).

[0610] Table 23

[0611] Excitation max(nm) Emission max (nm) Quantum yield (%) Extinction coefficient TG 490 513 82% <![CDATA[8×10 4 ]]> Me-FAM 463 514 55% <![CDATA[2×10 4 ]]> Me-HCF 544 567 57% <![CDATA[7×10 4 ]]>

[0612] In sequencing methods, the substrate (TPLFN) must be non-fluorescent before incorporation by DNA polymerase. After extension by polymerase primers, the dye-labeled triphosphates still attached to the substrate are released, and subsequently hydrolyzed in the presence of phosphatase to generate the fluorescent product triphosphate. Figure 9 and Figure 10 The differences in absorption and emission between TPLFN TG-dA4P and the released TG fluorophore were shown. Figure 9 As shown, TG-dA4P is not digested by CIP (bovine intestinal alkaline phosphatase) alone. However, once the polyphosphate chain of TG-dA4P is broken down by polymerase or PDE (phosphodiesterase), the remaining triphosphate chain labeled with TG is rapidly digested, yielding free TG molecules that have regained their strong uptake and emission strength.

[0613] Record the above spectra under the following conditions:

[0614] First, the spectrum of TG-dA4P was measured at room temperature. For emission measurements: the excitation wavelength was set to 460 nm, and the emission was scanned from 480 to 600 nm; for absorption measurements: the emission was scanned from 310 to 550 nm. Then, CIP and PDE were added, and the spectra were recorded sequentially under the same conditions.

[0615] The stability of TPLFN substrates under certain aqueous conditions must also be considered, as spontaneous hydrolysis of TPLFN increases fluorescence background during sequencing reactions, which can interfere with the desired signal and reduce sequencing accuracy. Fortunately, the hydrolysis rate of TPLFN substrates remains very low, approximately 2 ppm (substrate) / s, when measured at 65°C, which is negligible compared to the signal generated by polymerase incorporation. Nevertheless, in some respects, it is preferable to store the substrate solution in a 4°C cryogenic rack during sequencing and store it long-term in a -20°C freezer.

[0616] Section 2: Polymerase Kinetics Study

[0617] Polymerase kinetics assays were performed using a fluorometer to determine properties such as TPLFN incorporation / misincorporation ratio, homopolymer linearity, and temperature dependence. Figure 11 The proposed kinetic pathway for this sequencing-by-synthesis process is shown, where S is the matched substrate (TPLFN) and S* is the mismatched substrate; E is the enzyme (polymerase) and DN is the primer / template pair.

[0618] Although both TPLFN and the template are used as reaction substrates, this system can be simplified to a single-substrate reaction process because the concentration of one of the substrates, TPLFN (which is in significant excess compared to the primer / template), will remain almost constant. This makes the analysis of the process much easier. Figure 11 As shown, the polymerase-catalyzed reaction consists of three steps: a) DNA polymerase binds to the primer / template; b) incorporation of complementary nucleotides (TPLFNs); and c) nucleotide elongation along the template. The kinetic properties of the polymerase used in the sequencing process can be evaluated by varying reaction conditions such as primer / template concentration, the type of matching or mismatching TPLFNs, and temperature.

[0619] Figure 12The differences in polymerase (Bst) incorporation ratios among TPLFNs were shown. To examine and compare reaction rates, all four TPLFNs (TG-dA4P, TG-dG4P, TG-dC4P, and TG-dT4P) were adjusted to the same concentration (2.0 μM). Reactions were performed at 65 °C using Bst (120 nM), single-base extension primers / templates (T, C, G, and A relative to the four TPLFNs), CIP (0.01 U), and pH 8.3 buffer, triggered by Mn(II) (1 mM). Typically, the labeled polyphosphate moiety is released during Bst-mediated extension and requires hydrolysis by CIP to generate the fluorescent dye molecule. Excess CIP in the reaction was detected, confirming that the hydrolysis rate was very rapid and did not become a rate-determining step, thus not affecting the observation of the Bst reaction rate. Figure 12 In the study, the measured Bst incorporation ratios of the four labeled nucleotides were in the order of TG-dC4P > TG-dA4P > TG-dG4P > TG-dT4P.

[0620] Figure 12 The four curves can be fitted to the functions in the table below. The fitting results indicate that the reaction system can be considered a first-order reaction with respect to primer / template concentrations. However, unlike running the reaction in the sample cell of a fluorometer, the actual sequencing reaction on the chip will be slightly different because all primers / templates are grafted onto the chip surface. To ensure that each reaction cycle of the different TPLFNs is completed on the same time scale, the reaction rates of the four TPLFNs can be adjusted to the same level by increasing the concentration of the slower-running TPLFN.

[0621] Table 24

[0622] Substrate Fitting function <![CDATA[R 2 ]]> dA4P <![CDATA[9.319×10 5 (1-e -0.05242t )]]> 0.9976 dT4P <![CDATA[8.698×10 5 (1-e -0.02616t )]]> 0.9994 dC4P <![CDATA[8.977×10 5 (1-e -0.06189t )]]> 0.9959 dG4P <![CDATA[8.839×10 5 (1-e -0.0405t )]]> 0.9961

[0623] In 2+2 sequencing, two different nucleotides are added together to the reaction mixture. For example, "M" means dA4P and dC4P are added in the same cycle, and "K" means dG4P and dT4P are added in the same cycle. (See above.) Figure 11 As described above, one of the added nucleotides can be used as S*, which does not elongate the current template nucleotide but competes with the complementary substrate S for binding to Bst. Therefore, it is possible that S* can reduce the elongation rate of S. Thus, substrate competition is evaluated through competition experiments.

[0624] In this experiment, a 100 nM template-primer containing only one pair of nucleosides to be sequenced at the 3' end of the template, 2 μM complementary substrate, and 2 μM mismatched substrate, along with excess Bst and CIP enzyme, were mixed together. The reaction was carried out at 65 °C and pH 8.3, triggered by 1 mM Mn(II).

[0625] The results showed that the reaction rate did not decrease significantly when the substrate was added at the same concentration (see [link]). Figure 13 This can be explained as follows. Bst enzyme is polymerase I derived from Bacillus stearothermophilus cells. When Bst binds the primer-template with K... d Binds at 5 nM and combines the matching nucleotide with K. d When bound at 5 μM, the mismatched nucleotide binds to K. d Binds at 5 μM–10 μM. See, for example, Kornberg and Baker, DNA replication, 2nd ed., 2005, University Science Books, p. 126. Figure 11 Steps 1) and 2) in the process are considered as two thermodynamic equilibria, with their dissociation constant (K) d The values ​​are 5 nM and 30 μM (arithmetic mean), respectively. If no substrate competition occurs, the two equilibria can be merged into one, and the new equilibrium's K... d It equals 150 (nM)(μM).

[0626]

[0627]

[0628]

[0629] Therefore, D N ES and D N The concentrations of E were 25.6 nM and 63.9 nM, respectively. If competition occurs, then D... N The concentration of ES was 22.6 nM, D N ES * The concentration was 11.3 nM, D N The concentration of E is 56.5 nM. Calculations show that, with or without competition, D N The concentration of ES changed only slightly, so the reaction rate also changed only slightly.

[0630] In summary, in 2+2 sequencing, the reaction rates of the four substrates can be acceptablely different, but can be adjusted to be equal by varying the substrate concentrations. Competition between substrates does not significantly reduce the reaction rate. Therefore, in this method, the reaction rate for each cycle can be set to a specific value, as well as adjusted to optimized lead and lag values.

[0631] 100 nM of single-base extension primer / template poly-G was aliquoted into two PCR tubes. Two mismatched nucleotides, TG-dG4P (2 μM), along with excess Bst and CIP, were added to each tube. The mixtures were bubbled with argon for 2 minutes, and then the capped tubes were incubated at different temperatures: one at 4°C and the other at 65°C. After 1 hour, 2 μM of the matching nucleotide TG-dC4P was added to each tube, and the extension reaction was measured at 65°C using a fluorometer. If misincorporation occurred during incubation, different signal levels were expected in the two tubes, with the tube incubated at 65°C showing a lower signal level than the tube incubated at 4°C, because the misincorporation rate would be higher at 65°C. However, Figure 13 The results showed that the extension signals in the two tubes were almost identical, indicating that the misincorporation rate of Bst compared to TPLFN was undetectable under sequencing conditions.

[0632] One of the challenges of continuous fluorescence sequencing strategies is the need to accurately measure homopolymer or copolymer regions on the template using the generated fluorescence signals. Figure 14 Primer extensions of Bst polymerase for different homopolymer templates were demonstrated. Reactions were performed on a fluorometer under the following conditions: 100 nM / each template poly-T, poly-TT, poly-TTTT, and poly-TTTTTTTT, excess Bst and CIP, 2 μM TG-dA4P, pH 8.3 buffer, 65 °C, and triggered by Mn(II). Figure 14 The results showed that the generated fluorescence signal was proportional to the number of consecutive identical bases over a relatively wide range. Furthermore, Figure 15 The results show that by using a mixture of dA and dG instead of dA alone in this linear assay, the heteropolymer (or copolymer) sequence poly-TCTCTCTCTC can give the same signal level as poly-TTTTTTTT.

[0633] Besides reaction rate, polymerase fidelity is also a critical issue in the 2+2 sequencing strategy, especially considering the proofreading deficiencies of the polymerase used in this study. The incorporation of mismatched nucleotides not only reduces sequencing accuracy but also leads to signal attenuation in each sequencing cycle. Although fidelity is primarily an inherent capability of the polymerase, specific reaction conditions can still affect its ability to distinguish errors. To evaluate polymerase fidelity, a misincorporation experiment was designed, as described below:

[0634] Excess Bst and CIP, Mn(II), 100 nM primer-template (the template has G unpaired nucleosides except for the 3' end of the primers), and 2 μM dC4P were mixed at 65 °C and pH 8.3, generating a fluorescence signal of 4.5 × 10⁻⁶. 5 .

[0635] Next, a mixture of Bst, CIP, Mn(II), and primer-template at equal concentrations was mixed with 2 μM dG4P and bubbled with argon gas to prevent Mn(II) oxidation. One half of the mixture was incubated at 65 °C for 30 minutes, and the other half at 65 °C for 1 hour. After incubation, 2 μM dC4P was added to the mixture, resulting in a fluorescence signal of 4.6 × 10⁻⁶. 5 and 4.5×10 5 This indicates that mismatch extensions in the reaction system are virtually undetectable when using Bst polymerase. Minor signal differences are primarily due to inaccurate sample mixing. A very slow mismatch extension rate is highly preferred in sequencing reactions because once a primer-template is mismatched, a substitution mutation is generated at the current nucleotide site, altering the double-stranded structure and thus blocking further extension of that primer-template. In this way, mismatch extensions gradually reduce the effective concentration of the surface-grafted template array and lead to significant signal attenuation in each sequencing cycle. This study has ruled out the influence of mismatch extensions in the sequencing reaction and confirmed the high accuracy of the reaction system.

[0636] Figure 16 The results show that the extension rate of Bst is temperature-dependent, exhibiting optimal enzyme activity at 65°C and complete inactivity at 4°C. This temperature dependence is beneficial for sequencing performance because all reactions for high-throughput sequencing will be separated and confined to microreactors on the developed sequencing chip. Therefore, signal generation and diffusion are not critical requirements when loading substrate and enzyme at 4°C. However, once the temperature is raised to 65°C, the polymerase becomes fully active and rapidly generates a signal with a high signal-to-noise ratio.

[0637] The stability of the substrate TPLFN was also measured at different temperatures. The results showed that the higher the temperature, the greater the hydrolysis rate. However, the hydrolysis rate did not exceed 2 ppm / s, indicating that the background generated by autohydrolysis was still far below the polymerase extension signal. Even so, for better performance, the substrate would be preferably stored at low temperatures to prevent autohydrolysis before extension begins.

[0638] Section 3: Grafting on Sequencing Chip Surface

[0639] Between oligonucleotide grafting, the glass arrays used for sequencing were all modified with hydrogel. The modification method was based on a reported procedure, as described below. See, for example, U.S. Patent No. 8,247,177.

[0640] 3.1. Hydrogel polymer coating

[0641] 1) Synthesis of BRAPA:

[0642] The hydrogel monomer N-(5-(2-bromoacetamido)pentyl)acrylamide (BRAPA) was synthesized by the following method. Figure 17 )

[0643] 1,5-Diaminopentane (10.2 g, 0.1 mol) was dissolved in 300 mL of anhydrous methanol at 0 °C. Anhydrous THF solution of acryloyl chloride (0.9 g, 0.09 mol acryloyl chloride dissolved in 15 mL of anhydrous THF) was added dropwise with stirring. After addition, the reaction mixture was stirred for 10 h. 200 g of silica gel and 1% benzoquinone were added to the reaction mixture, and all solvent was removed using a vacuum evaporator. The silica gel powder adsorbed with the chemicals was loaded onto the top of a preparative silica gel column and eluted with DCM / methanol (10 / 1 to 1 / 1). The eluent containing the desired product was collected and concentrated to give 13 g of a grayish-white powder, which was used directly in the next step without further purification or prolonged storage to prevent polymerization.

[0644] The product was suspended in 150 mL of THF (20 mL of methanol could be added to increase solubility), and then an aqueous sodium bicarbonate solution (2 equivalents) was added at 0 °C. Bromoacetyl bromide (0.8 mol) was added dropwise to the mixture at °C, and the reaction was terminated after stirring for 10 h. Then, 50 mL of brine was added to the solution to separate the two phases, and the aqueous phase was extracted with 3 x 50 mL DCM. The combined organic phases were dried over Na₂SO₄, concentrated, and purified by silica gel column chromatography (eluting with EA / methanol) to give 13.5 g of BRAPA as a white solid. The product can be further purified by recrystallization in ethyl acetate. Mp 10²–10⁴ °C. HRMS C 10 H 18 The calculated value of BrN2O2(M+H) is 277.0541. The measured value (m / z) is 277.0546. 1 H NMR(500MHz,d6-DMSO)δ8.22(s,1H,NH),8.02(s,1H,NH),6.21(dd,J=15Hz,10Hz,1H,CH),6.07(dd,J=15Hz,5Hz,1H,CH) ,5.55(dd,J=10Hz,5Hz,1H,CH),3.82(s,2H,CH2),3.08(ddd,J=10Hz,5Hz,4H,CH2),1.43(m,4H,CH2),1.27(m,2H,CH2). 13 C NMR (126MHz, d6-DMSO) δ166.29,164.93,132.40,125.16,39.40,38.90,30.05,29.17,28.95,24.21.

[0645] 2) Chip surface cleaning:

[0646] Clean the channeled glass chip using the following procedure: rinse with chromic acid for 5 minutes, then wash thoroughly with milliQ H2O; after drying in a 120°C oven, treat the chip surface with oxygen-plasma for 3 minutes. Then immediately proceed with surface finishing.

[0647] 3) Hydrogel preparation:

[0648] Add BRAPA (70 mg in 700 μL DMF) to 10 mL of 2% acrylamide in milliQ H2O solution and mix thoroughly. Filter the mixture through a 0.22 μm filter and then bubble with argon for 15 minutes. Then, add 11.5 μL of TEMED, followed by 100 μL of potassium persulfate in milliQ H2O solution (50 mg / mL). Immediately load the thoroughly mixed solution into the channels of a clean chip and incubate under humid argon for 35 minutes. Then, wash the hydrogel-coated chip thoroughly with 200 mL of milliQ H2O.

[0649] 3.2. Primer grafting, template amplification, and hybridization

[0650] A solution of 10 μM PS-T10-P7 (5'-T*T*T*TTTTTTTCAAGCAGAAGACGGCATACGA-3', *=thiophosphate) in pH 8.0 PBS buffer was loaded into the coated channels and incubated at 50°C for 1 hour. The grafted chip surface was then blocked with 10 mM 2-mercaptoethanol in pH 8.0 PBS buffer for 40 minutes, followed by thorough washing with milliQ H2O. The grafted surface is shown in... Figure 18 .

[0651] 3.3. Preparation of DNA Template

[0652] ECCS Library Design:

[0653] A fragment of λ phage genomic DNA (approximately 300 bp) was used as a test DNA oligomer for preparing the sequencing template. The λ DNA was obtained from New England Biolabs, USA. The complete sequencing template included the reverse complementary strands of adapter 2 (43 bp) and P7 (21 bp) at the 5′ end of the ssDNA template, and adapter 1 (38 bp) and P5 (20 bp) at the 3′ end of the λ ssDNA. Except for a few bases, the sequences of P5, P7, adapter 1, and adapter 2 were identical to those of Illumina for compatibility.

[0654] Single-component library preparation (from phage λ):

[0655] A two-step PCR amplification method was used to prepare the sequencing template. In the first-step PCR, 50 μL of a mixture of λ genomic DNA (500 ng, NEB), the first-step PCR primers (200 nm each), and 1x Q5 high-fidelity 2x master mix (NEB) in H2O was treated with the following PCR thermal cycling spectrum: (i) initial heating at 95 °C for 90 seconds; (ii) 30 cycles, each cycle consisting of 30 seconds at 95 °C, 30 seconds at 65 °C, and 30 seconds at 72 °C. The amplified products were then purified using a PCR purification kit (Zymo, D4061) and transferred to Eppendorf tubes for the second-step PCR amplification. The conditions and thermal cycling spectrum for the second-step PCR were similar to those for the first step, but the primers used for the newly generated template were from above: P5-Adp1 (200 nM) and P7-Adp2 (200 nM).

[0656] The PCR products were purified by gel electrophoresis and verified by Sanger sequencing using primers P5, P7, and P5SeqP1. After measuring the final concentration, products containing the same DNA template were stored at -20°C for later use.

[0657] 3.4. Library Immobilization: Solid-Phase PCR in a Flow Cell

[0658] The same DNA template prepared above was mixed with PCR reagents and then loaded into a flow cell, where it was grafted onto the surface of primer P7 as described above. The mixture contained DNA template (1 nM), primer P5 (500 nM), primer P7 (62.5 nM), MgCl2 (6 mM), dNTPs (0.5 mM), platinum Taq polymerase (0.5 U / mL, Life Tech), BSA (0.2 mg / mL), and PCR buffer (200 mM Tris HCl, 500 mM KCl). The solid-phase amplification thermal cycling consisted of two phases with different temperature profiles. The first phase was an asymmetric pre-amplification process, consisting of (i) a 90-second hot start at 95 °C; and (ii) 15 cycles, each consisting of 30 seconds at 95 °C, 15 seconds at 65–60 °C (gradually decreasing), and 30 seconds at 72 °C. After asymmetric amplification, the primer P5-derived template strand was highly dominant in the PCR solution. Then, a second-stage solid-phase PCR thermal cycling was performed to primarily hybridize and extend the oligomer P7 grafted onto the flow cell surface. The thermal cycling pattern consisted of 30 cycles, including 30 seconds at 95°C and 300 seconds at 65°C. The sample was then denatured with formamide to remove the counterparts of the grafted oligomers, leaving only the P7-derived strands of the template on the flow cell surface.

[0659] After solid-phase PCR, the PCR solution was aspirated using a pipette. Formamide was injected into the flow cell to denature all remaining double-stranded DNA. Finally, the chip was washed with washing buffer (20 mM Tris-HCl buffer, pH 8.0, 50 mM KCl) to remove any remaining formamide.

[0660] Density measurement of solid-phase ssDNA template

[0661] First, 5 μM of oligonucleotide (FAM-T-SeqP1) with a fluorescent probe was injected into the flow cell, and the injection port was sealed. Then, the chip was placed on a heated plate at 80°C for 2 minutes, followed by cooling to room temperature (or below 30°C) over 30 minutes. The flow cell was thoroughly washed with washing buffer. Next, fluorescence images of the chip were captured using a fluorescence microscope with an automated stage. Images were taken at five different locations in each lane to check for uniformity and minimize random errors.

[0662] Previous experiments demonstrated a positive linear correlation between fluorescence values ​​and the number of FAM-modified primers. Therefore, a standard concentration curve was first established when calculating the PCR product concentration. A standard concentration curve was created by recording the fluorescence values ​​of lanes containing 0 nM (wash buffer without FAM-modified primers) and 100 nM TG solution. The average intensity of these images was fitted to the standard concentration curve to obtain the PCR product concentration.

[0663] Characterization of solid-phase PCR products showed that... Figure 24 . Figure 24 The image above shows a heatmap of PCR product density at different lanes and positions. Figure 24 The figure below shows the PCR product density for different templates.

[0664] Typically, the concentration of PCR products used in sequencing chips is approximately 50–150 nM (2.5–7.5 fmol / mm). 2 The average density of the four lanes of a single chip is approximately the same. Solid-phase PCR was performed on templates of different lengths, and there was no significant difference in density between the templates. To evaluate the uniformity of PCR product density, the coefficient of variation (CV) was measured by calculating the density values ​​at all imaging locations on the chip. The CV for all chips was 0.15 ± 0.13.

[0665] Prior to hybridization with sequencing primers (P5-SeqP1), the identified and qualified flow cell was denatured with formamide. The treated flow cell was then transferred to a microscope platform for sequencing.

[0666] 3.5. Sequencing

[0667] To conduct sequencing experiments, simple sequencing instruments were developed, such as... Figure 25 As shown. Figure 25 As shown in the image above, the sequencing chip (HiSeq 2000, for research purposes only) is placed on a temperature controller, beneath which is a 3D translation stage for 3D movement of the sequencing chip. Above the chip are a highly sensitive CCD and a 10× microscope. When blue light shines on the chip during the reaction, the emitted green light is captured by the microscope through the CCD. At one end of the chip, there is a thin tube connected to a valve and pump to introduce reaction buffer and wash buffer, while at the other end, the chip is attached to a tube to remove waste liquid.

[0668] For a sequential sequencing strategy, a mixture of two different nucleotides is added to the flow cell in each reaction cycle. Therefore, four nucleotide pairings are generated into three groups, each with two nucleotide pairs (AC / GT, AG / TC, or AT / GC). The six pairings are represented by M / K, R / Y, and W / S, respectively.

[0669] Before each sequencing run, reagents were premixed and stored in two separate vials within a cryo-support. Each vial contained Bst DNA polymerase (100 U / μL, McLab), bovine alkaline phosphatase (0.5 U / ml, NEB), MnCl2 (1 mM), and DTT (10 mM) in reaction buffer (40 mM Tris base, 40 mM HN4Cl, 100 mM KCl). One vial contained TG-dA4P (3 μM) / TG-dG4P (3 μM) for R, and the other contained TG-dC4P (2.5 μM) / TG-dT4P (5 μM) for Y. After each sequencing run, the vials were switched to W / S, and then M / K, with the same formulation as R / Y. These nucleotide groups do not necessarily need to be added in a specific order; any random sequence works in the same way.

[0670] The flow cell was mounted on a microscope platform, and the reagent vials were placed in a cryogenic holder. The automated sequencing process was performed using the following steps: (i) washing the flow cell and reagent input system (rotary valve, tubing between the flow cell and reagent vials) with wash buffer; (ii) washing the flow cell three times with wash buffer; (iii) cooling the flow cell to 4°C and loading one of the mixed nucleotides (TG-dA4P / TG-dG4P for R) via a syringe pump through the rotary valve; (iv) heating the flow cell to 15°C and capturing a background fluorescence image using a CCD camera (Hamamatsu); (v) heating the flow cell to 65°C to trigger polymerase-mediated nucleotide incorporation and primer extension, holding at 65°C for 1 minute; (vi) cooling the flow cell to 15°C, capturing an image to record the fluorescence signal, and then returning to step (ii). This process was automated until the entire template was sequenced or its sequencing limit was reached. The flow cell was then denatured with formamide to regenerate a single-stranded template. After primer annealing, the next round of sequencing with different reagent mixtures was performed in the same manner as described above.

[0671] Figure 25 The bottom left panel shows a typical fluorescence reaction kinetics curve, recording the fluorescence intensity every 5 seconds. When the chip was heated to 65°C, the fluorescence intensity significantly increased within approximately 20 seconds, reaching a plateau region, indicating that the reaction was nearing completion. The temperature controller was then cooled to 20°C to obtain the post-reaction fluorescence intensity, so the fluorescence intensity increased with decreasing temperature. However, the unit signal decreased throughout the sequencing process due to phase loss and template loss. Figure 25 The bottom right panel depicts the kinetic curves for each reaction cycle throughout the sequencing process.

[0672] Table 25: Oligonucleotide sequences used in this section

[0673]

[0674]

[0675] Note: "*" indicates a thiophosphate bond; FAM: 5,6-fluorescein phosphoramide

[0676] Table 26: Template sequences used in this section

[0677]

[0678]

[0679] Section 4: Phase Correction for Sequencing

[0680] 4.1. Signal Leading and Lagging

[0681] One unavoidable limiting factor for amplification-based sequencing-by-synthesis methods is phase loss, where extended molecules lose synchronization. This phenomenon is caused by accidental nucleotide addition (leading) or incomplete extension (lagging), and leads to increased noise and sequencing errors. Ideally, without phase loss, all nascent DNA molecules have the same extension length; however, when phase loss is taken into account, nascent DNA molecules can have different extension lengths. As the sequencing reaction progresses, the distribution of extension lengths becomes increasingly dispersed.

[0682] 4.2. Virtual Sequencing Instrument

[0683] 4.2.1. Virtual Sequencing Instrument Based on MATLAB

[0684] To monitor the distribution of nascent DNA extension length during sequencing reactions, a virtual sequencer program was developed using MATLAB to simulate all sequencing reactions. For a DNA sequence of length L, the chemical reactions considered and their corresponding kinetic constants are shown below:

[0685] Table 27: Chemical Reactions and Corresponding Kinetic Constants in Virtual Sequencing Program

[0686]

[0687] Where k = 1, 2, ..., L, and

[0688] Bst indicates Bst DNA polymerase,

[0689] DNA k-1 Indicates the (k-1)th position of the DNA to be sequenced.

[0690] dN k 4P indicates a terminal phosphate-labeled fluorescent nucleotide that can pair with the k-th position of DNA.

[0691] pFluorescein indicates non-fluorescein phosphoproteoside.

[0692] Phosphatase indicates alkaline phosphatase.

[0693] p indicates phosphoric acid,

[0694] Fluorescein indicates fluorescent, unphosphorylated fluorescein.

[0695] Bst-DNA k-1 Bst-DNA k-1 -dN k 4P indicates the corresponding complex.

[0696] The initial concentrations of the species used in the simulation are listed in the table below:

[0697] Table 28: Initial concentrations of various classes in the virtual sequencer program

[0698]

[0699]

[0700] The virtual sequencer program reads a given DNA sequence from a table and automatically generates a series of chemical reactions. These reactions are passed to the SimBiology toolbox in MATLAB to generate the corresponding ordinary differential equations (ODEs). All chemical kinetics used in the ODEs are mass actions. The ODEs are solved using the fourth-order Runge-Kutta method.

[0701] In the first sequencing cycle, the initial value of DNA0 was set to 0.05. k (k>0) is set to 0. DNA k The final value (k≥0) is set as the initial value for the next cycle. The concentrations of other species are reset to the values ​​listed in the table. A flowgram of the sequencing process is simulated by rotating the original value of dN4P in each cycle. The final value of Fluorescein is considered as the signal for each cycle.

[0702] In 2+2 sequencing simulated by a virtual sequencer program, if the concentration of the dominant dN4P species is sufficient and there are no impurities in the modified nucleotides, the signal given in each cycle is proportional to the length of each copolymer, and all nascent DNA molecules will have the exact same length. Figure 27 (ab). The sequence used in the simulation was L10115-301, with a base combination of M / K.

[0703] When impurities are present in the modified nucleotides or the reaction time is insufficient, a phase loss phenomenon occurs, and the sequencing signal is no longer proportional to the length of its corresponding copolymer. The effects of impurities and reaction time on the sequencing signal were evaluated using a virtual sequencer program, and the concentration distribution of nascent DNA molecules was monitored. When impurities were present but the reaction time was sufficient, a lead effect was observed. Figure 27 cd). When there are no impurities but the reaction time is insufficient, a hysteresis effect is observed ( Figure 27 ef).

[0704] 4.2.2. Principle of One-Time Pass, Multiple Terminations

[0705] To observe the effect of loss of molecule length on the distribution of elongation length in newly formed DNA molecules, a virtual sequencer program was used to simulate the sequencing reaction using ordinary differential equations (ODEs). In the simulation, the molecule to be sequenced was set as K(M).n In the KMM reaction, the dominant nucleotide species in the reaction solution were defined as K (G and T), and the impurities as M (A and C). Other parameters, such as reaction time and kinetic parameters, were set to estimated normal values. It was observed that after the first nucleotide K was extended by the dominant species, subsequent M nucleotides were partially extended by the impurities as expected, resulting in a lead effect. If n = 1, almost all K nucleotides following M would be extended by the dominant nucleotide species. However, if n > 1, this secondary lead would decrease rapidly. Figure 28 (See figure above). This single-pass, multiple-termination characteristic enables the prediction of DNA elongation length distribution and the development of the following correction algorithm (see below).

[0706] 4.3. Phase correction via flux matrix

[0707] In a 2+2 sequencing run, the parameters are defined as follows: N indicates the sequencing cycle number; M indicates the number of copolymers of the molecules to be sequenced; h is a column vector, and its elements h j Indicates the length of the j-th copolymer; s is a column vector whose elements s i The sequencing signal indicating cycle i; D N×M Indicator distribution matrix, its elements d ij Indicates the ratio of newly formed DNA molecules to the j-th copolymers that have extended in the i-th sequencing cycle; T N×M Indicator flux matrix, whose elements t ij The expression indicates the proportion of newly formed DNA molecules that extend (through) the j-th copolymer in the i-th sequencing cycle; λ indicates the lag factor, i.e., the proportion of newly formed DNA molecules of the same length that are not extended by the dominant nucleotide species in a given cycle; ε indicates the lead factor, i.e., the proportion of newly formed DNA molecules of the same length that are extended by the impurity nucleotide species in a given cycle; and h' is a column vector whose elements are...

[0708]

[0709] like Figure 27 As shown, phase loss leads to signal distortion and reduces sequencing accuracy. Algorithms have been developed to correct this distortion caused by phase loss, which will be discussed in detail below. Figure 28 The figure below provides a summary of the key concepts and an overview of the correction algorithm. Figure 28 The upper and lower parts of the figure below are the distribution matrix D. N×M and flux matrix T N×MA 3D representation of the matrix. Each entry in D and T is represented as a cube, its size along the sequence axis related to the length of its corresponding copolymer. Matrices D and T can be calculated interactively and iteratively, both being positive at or near their diagonals, otherwise zero. Ultimately, all nascent DNA strands extend beyond each copolymer; based on this fact, the accumulation of T along the cycle axis equals 1. The accumulation of T along the sequence axis is the measured dephasing sequencing signal. Matrices D and T, and their accumulations along both axes, can be categorized into three parts: primary, advanced, and lagging. The primary part is the diagonal of matrices D and T, representing nascent DNA strands of exactly the expected length. The advanced and lagging parts are the upper and lower triangular parts of matrices D and T, representing nascent DNA strands with lengths greater than or less than the expected values, respectively. Figure 28 As shown in the figure below, in the first few sequencing cycles, the primary fraction plays a dominant role in matrices D and T and their accumulation, contributing the vast majority of the sequencing signal. However, as the sequencing cycles continue, the primary fraction decreases while the leading and lagging fractions increase, indicating signal distortion.

[0710] 4.3.1. Distribution and Flux Matrix

[0711] The following assumptions are made: 1) No nucleotides were mis-introduced in the sequencing reaction, therefore this was not the cause of advance; 2) Advancement was caused by impurity nucleotides remaining from the previous cycle; 3) At most one base per molecule will be extended by the impurity nucleotide in a given cycle; 4) If the length extended by the copolymer of impurity nucleotides is 1, it will be further extended by the major nucleotide, a phenomenon known as secondary advance; 5) If the length of the copolymer extended by the impurity nucleotides is greater than 1, secondary advance will not occur; 6) The secondary advance chain will not be further extended by impurity nucleotides. Assumptions 3-6 are all based on the fact that the impurity nucleotides are present in trace amounts, consistent with the simulation results obtained using the virtual sequencer program in this paper (one-pass, multiple-termination principle).

[0712] Based on the above assumptions, for a given N, M, h, λ, and ε, calculate D and T as follows:

[0713]

[0714]

[0715] For example, considering sequencing using the combination M / K with the sequence AAGCTGTAGGAATCACT for 6 cycles, then h = (2,2,1,3,1,2,2,1,3,1). T Assuming both the lead and lag coefficients are 0.05, then matrices D and T are:

[0716]

[0717] The incorporation ratios of different nucleotides and impurity contents are also different. Considering this fact, different λ and ε are used for the two sequencing mixtures.

[0718] Dephasing correction algorithm

[0719] The relationship between h and s is as follows:

[0720] s = T(h’, ε, λ)h (4)

[0721] Since dim(s) < dim(h), this linear equation is indeterminate. Therefore, the Moore-Penrose pseudo-inverse and an iterative algorithm are used to obtain the minimum norm solution ( Figure 29 ):

[0722] 1. Set

[0723] 2. Calculate matrices D and T according to formulas (2) and (3).

[0724] 3. Set where is the pseudo-inverse of T.

[0725] 4. Compare [h2] and [h1], where [] is the rounding operation. If they are equal, return h2. Otherwise, go to step 5.

[0726] 5. Set h1 ← h2. Go to step 2.

[0727] Figure 29 shows a simplified flowchart of the dephasing correction algorithm. Briefly, the algorithm uses an iterative method to refine the sequencing signal until it converges. Usually, the iteration will terminate within 5 cycles. An example of applying it to real sequencing data is shown in Figure 30 . Figure 30 shows the refinement process during the iteration of the dephasing correction algorithm.

[0728] 4.3.2. General solution of the equation

[0729] The relationship between h and s is as follows:

[0730] s = T(h’, ε, λ)h (4)

[0731] Since dim(s) < dim(h), this linear equation is indeterminate and there are an infinite number of solutions that exactly satisfy the equation. The general form of these solutions is as follows:

[0732]

[0733] Where I is the identity matrix and w is an arbitrary vector. In the phase correction algorithm, w is set as the zero vector. (Check item) To observe its effect on h. The sequence was set to L10115-301, the base combination was set to M / K, the lead factor was set to 0.007, the lag factor was set to 0.005, and the sequencing cycle was set to 100. It was found that the entries in R between rows 1 and 99 and columns 1 and 99 were very close to zero (~10). -16 ), so much so that it can be considered a calculation error ( Figure 31 The matrix is ​​shown. The value of h is used to determine the actual determinant, except for the last element.

[0734] 4.3.3. Robustness of the Phase Loss Correction Algorithm - Condition Number

[0735] The Moore-Ponous pseudoinverse matrix is ​​used in the phase misalignment correction algorithm. For the flux matrix T, the condition number is defined as:

[0736]

[0737] A large condition number means that a small error in the elements of T can lead to a large error in the solution entries. The effect of the phase loss coefficient on the condition number of T is evaluated. The sequences used are poly(AG)(AGAGAG…), poly(AAGG)(AAGGAAGG…), L718-308, L4418-305, L9730-303, and L10115-301, with a base combination of M / K. The lead and lag coefficients used for evaluation are 0, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, and 0.1. For each sequence and phase loss coefficient, the flux matrix T is calculated according to Equation (3), and its condition number is calculated according to Equation (6). Figure 32 The logarithms of the condition number are shown for different dephasing coefficients. In all sequences except for AAGG, increasing the lead or lag coefficient leads to an increase in the condition number, indicating that the more dephasing molecules there are, the worse the correction. However, in AAGG sequences where the DPL is always 2, increasing the lead coefficient leads to a decrease in the condition number. This suggests that long DPLs (length > 2) have a significant hindering effect on dephasing.

[0738] 4.3.4. Algorithm Robustness

[0739] A) The impact of phase loss coefficient deviation on signal correction

[0740] The phase loss coefficient is obtained by fitting the signal of the reference sequence and is used to correct other unknown sequences. Ideally, the phase loss coefficients of the reference and unknown sequences are the same. However, due to random factors, slight differences between the two sets are inevitable. Therefore, if the coefficient is inaccurate, it is necessary to test how many errors it will produce in phase loss correction. One hundred 370bp DNA sequences were randomly generated, their phase loss signals at a given phase loss coefficient were calculated, and corrections were performed using different but very close coefficients. With the base combination set to M / K and the sequencing cycle number set to 150, the given phase loss coefficients tested were 0.001, 0.005, and 0.010. Since the correction algorithm will still produce errors in the last few cycles even if the phase loss coefficient is accurate, the difference in the number of errors between accurate and inaccurate phase loss coefficients will be used to characterize performance, and the average value is shown below. Figure 33 The graphs show the impact of phase deflection coefficient deviation on signal correction. Asterisks in each graph indicate the position of the accuracy coefficient, and the color bars are limited to the range of 0–5, so any error count greater than 5 is displayed in dark red. The results show that the greater the deviation of the phase deflection coefficient, the more errors it produces, and the tolerance to lead deviation is relatively greater than that to hysteresis deviation.

[0741] B) Tolerance to global noise

[0742] Sequencing signal noise can originate from defocused imaging, CCD imaging, fluid dynamics, instability, or anomalies. The impact...

Claims

1. A method for obtaining sequence information of a target polynucleotide, the method comprising: a) providing a first sequencing reagent to a target polynucleotide in the presence of a first polynucleotide replication catalyst, wherein the first sequencing reagent comprises at least two different nucleotide monomers each conjugated to a first label, and the nucleotide monomer / first label conjugate is non-fluorescent until the nucleotide monomer is incorporated into the target polynucleotide according to complementarity with the target polynucleotide, wherein the first label of the at least two different nucleotide monomers is the same or different; and b) providing a second sequencing reagent to the target polynucleotide in the presence of a second polynucleotide replication catalyst, wherein the second sequencing reagent comprises one or more nucleotide monomers each conjugated to a second label, and the nucleotide monomer / second label conjugate is non-fluorescent until the nucleotide monomer is incorporated into the target polynucleotide according to complementarity with the target polynucleotide, at least one of the one or more nucleotide monomers is different from the nucleotide monomers present in the first sequencing reagent, and wherein the second sequencing reagent is provided subsequent to the first sequencing reagent, and c) obtaining sequence information of at least a portion of the target polynucleotide by detecting fluorescence emission resulting from the first and second labels after the incorporation of the nucleotide monomers into the polynucleotide in steps a) and b); wherein: the initial sequence information obtained in step c) comprises one or more errors; at least one additional round of steps a) and b) using a combination of the first sequencing reagent and the second sequencing reagent; the combination is different from the combination of the first sequencing reagent in step a) and the second sequencing reagent in step b) to obtain at least one additional sequence, the additional sequence is compared to the initial sequence to reduce or eliminate sequence errors; further comprising a method of detecting and / or correcting sequence data errors in the sequence results, comprising obtaining sequence data of three or more orthogonal nucleotide degenerate sequences, detecting errors in the sequence by comparing the three or more orthogonal nucleotide degenerate sequences, and obtaining a corrected sequence by modifying at least one sequence at the position where the error is detected in the comparison; wherein the nucleotide monomer / first label conjugate in the first sequencing reagent and / or the one or more nucleotide monomers / second label conjugate in the second sequencing reagent has the following structure of Formula I: wherein n is 0-6, R is a nucleobase, X is H, OH, or OMe, or a salt thereof.

2. The method of claim 1 for obtaining sequence information of at least a portion of a single target polynucleotide, or for obtaining sequence information of at least a portion of a plurality of target polynucleotides simultaneously.

3. The method of claim 1 or 2, wherein the first polynucleotide replication catalyst and the second polynucleotide replication catalyst are the same polynucleotide replication catalyst, or wherein the first polynucleotide replication catalyst and the second polynucleotide replication catalyst are different polynucleotide replication catalysts.

4. The method of claim 1, wherein the sequence information is obtained by one or more sequencing reactions, wherein the one or more sequencing reactions are carried out in one or more reaction volumes, wherein the reaction volumes are physically separated from each other and / or there is no exchange of material between the reaction volumes, wherein the reaction volumes are located in an array chip, and wherein the reaction volumes are closed and / or isolated from each other by a liquid immiscible with the liquid in the reaction volumes.

5. The method of claim 4, wherein the reaction volumes are provided in reaction chambers, and the target polynucleotides in each of the reaction chambers are immobilized on a solid support in the reaction chamber, wherein the sequence information is obtained by high-throughput sequencing.

6. The method of any one of claims 1-5, wherein the first polynucleotide replication catalyst and / or the second polynucleotide replication catalyst is a polymerase selected from a DNA polymerase, an RNA polymerase or an RNA-dependent RNA polymerase, a ligase, a reverse transcriptase, or a terminal deoxynucleotidyl transferase.

7. The method of any one of claims 1-5, wherein the nucleotide monomers in the first and / or the second sequencing reagent are selected from the group consisting of deoxyribonucleotides, modified deoxyribonucleotides, ribonucleotides, modified ribonucleotides, peptide nucleotides, modified peptide nucleotides, modified phosphothio sugar backbone nucleotides, and mixtures thereof.

8. The method of claim 7, wherein the nucleotide monomers in the first sequencing reagent and the second sequencing reagent are both deoxyribonucleotides.

9. The method of claim 8, wherein the nucleotide monomers are selected from the group consisting of A, T / U, C, and G deoxyribonucleotides.

10. The method of claim 7, wherein the nucleotide monomers in the first sequencing reagent and the second sequencing reagent are both ribonucleotides.

11. The method of claim 10, wherein the nucleotide monomers are selected from the group consisting of A, U / T, C, and G ribonucleotides.

12. The method of any one of claims 1-5, wherein the first and / or the second label is releasably conjugated to the nucleotide monomer.

13. The method of claim 12, wherein the first and / or the second label is conjugated to a terminal phosphate group of the nucleotide monomer, or to the last, second to last, third to last, fourth to last, or fifth to last phosphate group of the nucleotide monomer.

14. The method of claim 13, wherein the nucleotide monomer / first label conjugate in the first sequencing reagent and / or the one or more nucleotide monomer / second label conjugates in the second sequencing reagent have the structure of Formula II:

15. The method of claim 13 or 14, wherein the first and / or second label is non-fluorescent until after release from the terminal phosphate group of the nucleotide monomer.

16. The method of claim 15, further comprising releasing the first and / or second label from the terminal phosphate group of the nucleotide monomer with an activating enzyme.

17. The method of claim 16, wherein the activating enzyme is an exonuclease, a phosphotransferase, or a phosphatase.

18. The method of any one of claims 1-5, wherein the first label of the at least two different nucleotide monomers is the same.

19. The method of any one of claims 1-5, wherein the first label of the at least two different nucleotide monomers is different.

20. The method of any one of claims 1-5, further comprising a washing step between steps a) and b).

21. The method of any one of claims 1-5, wherein the target polynucleotide is immobilized on a surface selected from the group consisting of a solid surface, a soft surface, a hydrogel surface, a microparticle surface, or a combination thereof.

22. The method of claim 21, wherein the solid surface is part of a microreactor, and steps a) and b) are performed in the microreactor.

23. The method of claim 22, performed at a temperature in the range of 20 °C to 70 °C.

24. The method of any one of claims 1-5, wherein multiple rounds of steps a) and b) are performed using different combinations of the first and second sequencing reagents.

25. The method of any one of claims 1-5, wherein the sequence information obtained in step c) is a degenerate sequence.

26. The method of claim 25, wherein at least one additional round of steps a) and b) is performed using a different combination of the first and second sequencing reagents than the combination of the first and second sequencing reagents in one or more previous rounds of steps a) and b) to obtain at least one additional sequence, and the additional sequence is compared to the degenerate sequence to obtain a non-degenerate sequence.

27. The method of any one of claims 1-5, wherein the initial sequence information obtained in step c) contains no errors or one or more errors.

28. The method of claim 27, wherein at least one additional round of steps a) and b) is performed using a different combination of the first and second sequencing reagents than the combination of the first and second sequencing reagents in one or more previous rounds of steps a) and b) to obtain at least one additional sequence, and the additional sequence is compared to the initial sequence to reduce or eliminate sequence errors.

29. The method of any one of claims 26 or 28, wherein the sequence comparison is performed using a mathematical analysis, algorithm, or method.

30. The method of claim 29, wherein the mathematical analysis, algorithm, or method comprises a Markov model or a maximum likelihood method based on a Bayesian profile.

31. The method of any one of claims 1-5, wherein the first sequencing reagent comprises two different nucleotide monomers / first label conjugates, each comprising a different nucleotide monomer, the second sequencing reagent comprises two different nucleotide monomers / second label conjugates, each comprising a different nucleotide monomer, and the two nucleotide monomers in the first sequencing reagent are different from the two nucleotide monomers in the second sequencing reagent.

32. The method of claim 31, wherein the two nucleotide monomers in the first sequencing reagent and the two nucleotide monomers in the second sequencing reagent are selected from the group consisting of A, T / U, C, and G deoxyribonucleotides.

33. The method of claim 32, wherein the two nucleotide monomers in the first sequencing reagent and the two nucleotide monomers in the second sequencing reagent are selected from the group consisting of the following combinations: 1) A and T / U deoxyribonucleotides in one sequencing reagent and C and G deoxyribonucleotides in the other sequencing reagent; 2) A and G deoxyribonucleotides in one sequencing reagent and C and T / U deoxyribonucleotides in the other sequencing reagent; and 3) A and C deoxyribonucleotides in one sequencing reagent and G and T / U deoxyribonucleotides in the other sequencing reagent.

34. The method of claim 33, wherein one round of steps a) and b) or at least two rounds of steps a) and b) are performed, one of the combinations 1)-3) is used for one round of steps a) and b), and another of the combinations 1)-3) different from the combination used in the previous round of steps a) and b) is used for another round of steps a) and b).

35. The method of claim 34, wherein three rounds of steps a) and b) are performed, each round using a different combination selected from the combinations 1)-3).

36. The method of claim 34 or 35, wherein sequences obtained from the multiple rounds of steps a) and b) are compared to obtain non-degenerate sequences and / or to reduce or eliminate sequence errors in the non-degenerate sequences.

37. The method of claim 30 or 31, wherein the two nucleotide monomers in the first sequencing reagent and the two nucleotide monomers in the second sequencing reagent are selected from the group consisting of A, T / U, C, and G ribonucleotides.

38. The method of claim 37, wherein the two nucleotide monomers in the first sequencing reagent and the two nucleotide monomers in the second sequencing reagent are selected from the group consisting of the following combinations: 1) A and T / U ribonucleotides in one sequencing reagent and C and G ribonucleotides in the other sequencing reagent; 2) A and G ribonucleotides in one sequencing reagent and C and T / U ribonucleotides in the other sequencing reagent; and 3) A and C ribonucleotides in one sequencing reagent and G and T / U ribonucleotides in the other sequencing reagent.

39. The method of claim 38, wherein one round of steps a) and b) or at least two rounds of steps a) and b) are performed, one of the combinations 1) - 3) is used for one round of steps a) and b), and another of the combinations 1) - 3) that is different from the combination used in the previous round of steps a) and b) is used for another round of steps a) and b).

40. The method of claim 39, wherein at least three rounds of steps a) and b) are performed, a different combination of the combinations 1) - 3) is used for each round.

41. The method of claim 39 or 40, wherein sequences obtained from multiple rounds of steps a) and b) are compared to obtain non-degenerate sequences and / or to reduce or eliminate sequence errors in the non-degenerate sequences.

42. The method of claim 41, wherein the first label of the two different nucleotide monomers is the same, and the second label is the same as the first label.

43. The method of claim 41, wherein the first label of the two different nucleotide monomers is different, and the second label is the same as the first label.

44. The method of any one of claims 1 - 5, wherein one of the first and second sequencing reagents comprises three different nucleotide monomer / first label conjugates, each comprising a different nucleotide monomer, one sequencing reagent comprises one nucleotide monomer / second label conjugate, and the three nucleotide monomers in one sequencing reagent are different from the nucleotide monomer in the other sequencing reagent.

45. The method of claim 44, wherein the nucleotide monomers in the first and second sequencing reagents are selected from the group consisting of A, T / U, C, and G deoxyribonucleotides.

46. The method of claim 45, wherein the nucleotide monomers in the first and second sequencing reagents are selected from the group consisting of the following combinations: 1) C, G, and T / U deoxyribonucleotides in one sequencing reagent and A deoxyribonucleotides in the other sequencing reagent; 2) A, G, and T / U deoxyribonucleotides in one sequencing reagent and C deoxyribonucleotides in the other sequencing reagent; 3) A, C, and T / U deoxyribonucleotides in one sequencing reagent and G deoxyribonucleotides in the other sequencing reagent; and 4) A, C, and G deoxyribonucleotides in one sequencing reagent and T / U deoxyribonucleotides in the other sequencing reagent.

47. The method of claim 46, wherein one round of steps a) and b) or at least two rounds of steps a) and b) are performed, one of the combinations 1) - 4) is used for one round of steps a) and b), and another of the combinations 1) - 4) that is different from the combination used in the previous round of steps a) and b) is used for another round of steps a) and b).

48. The method of claim 46, wherein three rounds of steps a) and b) are performed, a different combination selected from the combinations 1) - 4) is used for each round.

49. The method of claim 46, wherein four rounds of steps a) and b) are performed, each round using a different combination selected from the combinations 1) - 4).

50. The method of any one of claims 47-49, wherein sequences obtained from multiple rounds of steps a) and b) are compared to obtain non-degenerate sequences and / or to reduce or eliminate sequence errors in the non-degenerate sequences.

51. The method of claim 44, wherein the nucleotide monomers in the first and second sequencing reagents are selected from the group consisting of A, T / U, C, and G ribonucleotides.

52. The method of claim 51, wherein the nucleotide monomers in the first and second sequencing reagents are selected from the group consisting of the following combinations: 1) C, G, and T / U ribonucleotides in one sequencing reagent and A ribonucleotides in the other sequencing reagent; 2) A, G, and T / U ribonucleotides in one sequencing reagent and C ribonucleotides in the other sequencing reagent; 3) A, C, and T / U ribonucleotides in one sequencing reagent and G ribonucleotides in the other sequencing reagent; and 4) A, C, and G ribonucleotides in one sequencing reagent and T / U ribonucleotides in the other sequencing reagent.

53. The method of claim 52, wherein one round of steps a) and b) or at least two rounds of steps a) and b) are performed, one of the combinations 1) - 4) is used for one round of steps a) and b), and another of the combinations 1) - 4) that is different from the combination used in the previous round of steps a) and b) is used for another round of steps a) and b).

54. The method of claim 52, wherein at least three rounds of steps a) and b) are performed, a different combination of the combinations 1) - 4) is used for each round.

55. The method of claim 52, wherein at least four rounds of steps a) and b) are performed, a different combination of the combinations 1) - 4) is used for each round.

56. The method of any one of claims 52-55, wherein sequences obtained from multiple rounds of steps a) and b) are compared to obtain non-degenerate sequences and / or to reduce or eliminate sequence errors in the non-degenerate sequences.

57. The method of claim 56, wherein read lengths of 250, 350, 400, 500, 800, or 2400 base pairs are obtained.

58. The method of any one of claims 1-5, wherein at least 95% code accuracy is obtained.

59. The method of any one of claims 1-5, wherein the target polynucleotide is a single-stranded polynucleotide.

Citation Information

Patent Citations

  • System and method to correct out of phase errors in DNA sequencing data by use of a recursive algorithm

    US20110213563A1

  • Alternative nucleotide flows in sequencing-by-synthesis methods

    US20140031238A1

  • Method for sequencing a polynucleotide template

    US8247177B2

  • System and method to correct out of phase errors in DNA sequencing data by use of a recursive algorithm

    US8364417B2

  • Alternative nucleotide flows in sequencing-by-synthesis methods

    US9416413B2