A synthetic sequencing method based on nucleotide dimers as monomers

By employing the 3'-terminal hydroxyl-reversibly blocked nucleotide dimer synthesis sequencing method, two rounds of detection for each base were achieved. Four-color labeled nucleotide dimers were used for calibration, which solved the high error rate problem of second-generation high-throughput DNA sequencing platforms, improved sequencing accuracy and length, and expanded its application in bioscience research and clinical diagnosis.

CN115323043BActive Publication Date: 2026-03-10SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing second-generation high-throughput DNA sequencing platforms have a high error rate, which limits their accuracy in clinical applications and low-abundance mutation detection. In particular, SOLiD sequencing technology has a lower sequencing error rate but its application is limited.

Method used

The 3' hydroxyl-reversibly blocked nucleotide dimer synthesis sequencing method was adopted. Through two rounds of DNA sequencing, each base was detected twice. Four-color labeled nucleotide dimers were used for proofreading to reduce sequencing errors and improve accuracy.

Benefits of technology

It improves the accuracy and length of sequencing, expands the application of high-throughput DNA sequencing in bioscience research and early clinical diagnosis, and reduces errors in raw data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115323043B_ABST
    Figure CN115323043B_ABST
Patent Text Reader

Abstract

This invention discloses a synthetic sequencing method based on nucleotide dimers as monomers. The synthesized nucleotides for sequencing are nucleotide dimers. Under the action of polymerase, if the nucleotide dimer completely pairs with two bases in the DNA template hybridized to the sequencing primer, a polymerization reaction occurs, and the sequencing primer extends by two bases. If the two bases in the DNA template do not completely pair, the polymerization reaction does not occur. This method improves the length and accuracy of sequence determination, enabling high-throughput detection of nucleic acid sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, and in particular to a method for high-throughput synthesis and sequencing of nucleic acid sequences based on nucleotide dimers as monomers. Background Technology

[0002] DNA sequencing technology is the most commonly used technique in molecular biology research, and its emergence has greatly promoted the development of biology. The second-generation DNA sequencing technology developed in recent years has ushered in an era of high-throughput, low-cost DNA sequencing. High-throughput sequencing technology represents a revolutionary improvement over traditional sequencing technologies, capable of sequencing millions or even tens of millions of DNA sequences simultaneously. Currently, there are four main representatives of second-generation high-throughput sequencing technologies: Illumina's Solexa sequencing synthesis technology, Roche's 454 sequencing synthesis technology, ABI's IonTorrent sequencing synthesis technology, and SOLiD sequencing technology. However, current high-throughput DNA sequencing platforms, especially mainstream ones, still have a relatively high error rate. This high error rate makes the detection of single nucleotide polymorphisms (SNPs) or low-abundance mutations extremely difficult, limiting the clinical application of high-throughput DNA sequencing platforms, such as pharmacogenomics research primarily based on SNPs and early clinical diagnosis primarily based on low-abundance mutations. Currently, Sanger sequencing is still considered the gold standard due to its high accuracy (99.999%), and the results of high-throughput DNA sequencing platforms need to be validated by Sanger sequencing in clinical practice.

[0003] Among existing high-throughput DNA sequencing platforms, although the practical application of SOLiD ligation sequencing is limited by sequencing time and sequencing length, its raw sequencing error efficiency is 13, 16, 30, and 1 / 3 times lower than that of NextSeq, 454GS FLX, and Ion Torrent synthesis sequencing platforms, respectively, making it the highest among high-throughput DNA sequencing platforms. At a sequencing depth of 15×, its accuracy can reach 99.999%, matching the accuracy of Sanger sequencing. The low sequencing error efficiency of SOLiD sequencing technology is due to the fact that all bases are detected twice. This gives the method the characteristics of error correction and SNP / mutation site detection; that is, the information from the double base detection provides an inherent proofreading function, thereby reducing errors in the raw data.

[0004] To address the high error rate of existing second-generation high-throughput DNA synthesis and sequencing platforms, this invention uses nucleotide dimers as sequencing monomers. Through two rounds of DNA sequencing—that is, detecting all bases twice—the information provides a proofreading function, thereby reducing errors in the raw data and improving sequencing accuracy. This invention helps improve the accuracy of high-throughput DNA sequencing, increase sequencing length, expand the ability of high-throughput DNA sequencing technology to identify low-abundance sequences, and further expand the application of high-throughput DNA sequencing technology in bioscience research and early clinical diagnosis. Summary of the Invention

[0005] Purpose of the invention: The purpose of this invention is to provide a method for synthesizing and sequencing nucleotide dimers with reversible 3' hydroxyl blocking, which improves the length and accuracy of sequence determination and enables high-throughput detection of nucleic acid sequences.

[0006] Technical solution: 3' hydroxyl-reversibly blocked nucleotide dimer synthesis sequencing method. Under the action of polymerase, if the nucleotide dimer is completely paired with two bases in the DNA template of the sequencing primer, the nucleotide dimer polymerization reaction occurs, and the sequencing primer extends by two bases; if it is not completely paired with two bases in the DNA template, the nucleotide dimer polymerization reaction does not occur.

[0007] Furthermore, the basic structure of the nucleotide dimer consists of two identical or different nucleotides linked by a phosphate ester bond, wherein the 5' end also includes a triphosphate group. Nucleotide dimers include sixteen specific forms: AA, CC, GG, TT, AC, CA, GT, TG, AG, GA, CT, TC, AT, TA, CG, and GC.

[0008] Furthermore, the 3'-terminal hydroxyl group of the nucleotide dimer is blocked by groups including 3'-O-allyl, 3'-O-cyanoethyl, 3'-O-azidomethyl, and 3'-O-amino. These 3'-terminal modified nucleotide dimers can participate in the synthesis and sequencing reaction, and at most only one nucleotide dimer can be synthesized. Under suitable reaction conditions, the 3'-terminal blocking group can be deactivated, activating the 3'-terminal hydroxyl group of the nucleotide.

[0009] Furthermore, the nucleotide dimers are labeled with dyes, and the labeling position can be the 5' base position, the 5' phosphate terminus, or the 3' hydroxyl position. Specifically, when the 3' hydroxyl position is used, the dye can simultaneously replace the blocking group of claim 3. AA, CC, GG, and TT are labeled with the first dye; AC, CA, GT, and TG are labeled with the second dye; AG, GA, CT, and TC are labeled with the third dye; and AT, TA, CG, and GC are labeled with the fourth dye. The emission wavelengths of the four dyes do not interfere with each other, and these labeled dyes can be cleaved by chemical or photochemical methods without affecting the biochemical function of the DNA sequence.

[0010] Furthermore, the DNA template fragment to be tested refers to a single molecule, or a product of the same sequence amplified using a single molecule as a template. The DNA template to be tested can be a single sequence or an array of different DNA templates.

[0011] Furthermore, the sequencing information of the DNA template fragment to be tested is obtained by comparing the fluorescence intensity of four different dyes. That is, the dye with the highest fluorescence intensity among the four dyes is defined as the valid sequencing information of the DNA template, and the sequencing information and location coordinates of the DNA are recorded.

[0012] Furthermore, the specific base information of the DNA template to be tested is obtained through two rounds of sequencing, each round acquiring specific dye (or coding) information. The sequencing primers for the two rounds of sequencing differ by one base, with one round's sequencing primer corresponding to a known base in the sequencing template. The specific base information of the DNA template to be tested is obtained by sequentially decoding the coding information of two sets of base fragments from the known base in the sequencing template. In particular, for resequencing of genome sequencing with a reference sequence, the coding obtained from the real-time synthesis of DNA from the two sets of nucleotides can be directly used for alignment with the genome reference sequence without the need for decoding the coding, thus achieving resequencing of the genome sequence. Furthermore, the information of the DNA template to be tested is obtained by comparing the fluorescence intensity of the dye with the background value. When the strongest fluorescence intensity of the four labeling dyes containing a known base in the sequencing template is significantly higher than the background value, the DNA template can be determined as a valid sequencing template, and the fluorescence intensity and position coordinates of the DNA can be recorded. The steps are as follows:

[0013] A: Preparation of E. coli whole genome template: The target genome was sonicated into fragments of 100-1000 bp in size. These fragmented nucleic acid sequences were then ligated with a pair of known universal linkers using ligase. The sequence of linker 1 was: CTG CTG TAC CGT ACA GCC TTG GCC G, and the sequence of linker 2 was: CGC TTT CCT CTC TAT GGG CAG TCG GTGAT. Ten pre-amplification cycles were performed. Then, 200-800 bp DNA fragments were cut by gel electrophoresis and purified. These 200-800 bp DNA fragments were subjected to emulsion parallel PCR or bridge PCR to amplify the fragmented E. coli genome fragments and construct the E. coli genome sequencing DNA template chip.

[0014] B: First round of sequencing

[0015] a. Sequencing primer hybridization: The template fixed at the 5' end is hybridized with the first primer that is complementary to the 3' end linker. This hybridization primer serves as the sequencing primer for all *E. coli* genomic DNA templates (see [link to product]). Figure 2 );

[0016] b. Sequencing

[0017] (1) Sixteen 3'-O-azidomethyl-modified nucleotide dimers labeled with 100 μM tetrachromatograms (see...) Figure 1 The sequencing system (including 9° DNA polymerase) and Table 1 were added to the reaction chamber for the synthesis and sequencing reaction (60°C for 3 minutes). Then, the unreacted labeled nucleotide dimers were washed with 10mM disodium ethylenediaminetetraacetate (EDTA) buffer (pH = 7-8). The DNA template was identified as a valid sequencing template when the strongest fluorescence intensity of the four labeled dyes in a known base of the sequencing template was significantly higher than the background value. The fluorescence intensity and position coordinates of the DNA were also recorded.

[0018] (2) Add 100 mM tris(2-carboxyethyl)phosphine (pH 8.0) and react at 55 °C for 3 minutes, then wash with 10 mM EDTA buffer (pH = 7-8);

[0019] (3) Repeat steps (1) to (2) above to perform the synthesis and sequencing reaction in a cycle to obtain a set of sequencing information consisting of codes 1, 2, 3, and 4. Then perform the second round of sequencing reaction.

[0020] C: Second round of sequencing

[0021] a. Treat with 8M urea at 65℃ for 5 minutes twice to remove the sequencing primers and their synthetic strands from the first round of sequencing reaction and obtain a new single-stranded DNA template.

[0022] b. Sequencing primer hybridization: The 5' end-fixed template is hybridized with a second primer that is complementary to the 3' end linker. This hybridization primer serves as the sequencing primer for all *E. coli* genomic DNA templates (see [link to product]). Figure 3 );

[0023] c. Sequencing

[0024] (1) Add 16 3'-O-azidomethyl-modified nucleotide dimers labeled with 100 μM four-color and a sequencing system including 9° DNA polymerase to the reaction chamber for synthesis and sequencing reaction (react at 60°C for 3 minutes). Then wash the unreacted labeled nucleotide dimers with 10 mM disodium ethylenediaminetetraacetate (EDTA) buffer (pH = 7-8). When the strongest fluorescence intensity of the four labeling dyes containing a known base of the sequencing template is significantly higher than the background value, the DNA template is determined to be a valid sequencing template. Record the fluorescence intensity and position coordinates of the DNA.

[0025] (2) Add 100 mM tris(2-carboxyethyl)phosphine (pH 8.0) and react at 55 °C for 3 minutes, then wash with 10 mM EDTA buffer (pH = 7-8);

[0026] (3) Repeat steps (1) to (2) above to perform the synthesis and sequencing reaction in a cycle to obtain a set of sequencing information consisting of codes 1, 2, 3, and 4. Then perform the second round of sequencing reaction.

[0027] D. Decoding

[0028] Using the nucleotide dimer coding information obtained from two rounds of sequencing for each template, and using the nucleotide dimer coding information containing a known base from the sequencing template, the corresponding base sequence information for each template is decoded and assembled (see...). Figure 4 );

[0029] E. Sequence Assembly

[0030] Utilizing the base sequence information of all templates, and leveraging error correction and SNP recognition principles (see [link to documentation]). Figure 5 ), which are assembled into the E. coli genome sequence.

[0031] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0032] 1. The greatest advantage of using nucleotide dimers from sixteen four-color labeling methods in this invention as the raw material for sequencing is that all bases differing by only one base are sequenced twice using two sequencing primers. When there is an alignment sequence (such as reference sequence information or sequencing depth of 2× or higher), by comparison, if only one code changes between the sequencing information and the alignment information, this sequencing code is judged as a sequencing error; if two consecutive sequencing codes differ from the reference sequence, it is judged as a SNP. Because this base forms a nucleotide dimer with the preceding and following bases and is sequenced twice, the sequencing information has a proofreading function, thereby correcting errors in the original data and improving the accuracy of the sequencing information.

[0033] 2. Sixteen sequencing reaction methods using four-color labeling are employed, with nucleotide dimers used as the raw material for sequencing synthesis. This ensures that each DNA template meets the requirements for sequencing synthesis, resulting in fewer sequencing errors and reducing errors in the original sequencing. Using nucleotide dimers as the raw material for sequencing synthesis extends each sequencing reaction by two bases, significantly increasing the length of the synthesized sequence. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the structure of a nucleotide dimer monomer according to the method of the present invention;

[0035] Figure 2 This is a schematic diagram illustrating the first round of sequencing principle of the method of this invention;

[0036] Figure 3 This is a schematic diagram illustrating the principle of obtaining the second round of sequencing information using the method of this invention;

[0037] Figure 4 This is a schematic diagram illustrating the decoding principle of the method of the present invention;

[0038] Figure 5 This invention relates to the error correction method and its SNP identification principle. Detailed Implementation

[0039] This embodiment is based on the 3'-terminal hydroxyl-reversibly blocked nucleotide dimer synthesis sequencing method to determine the whole genome of Escherichia coli:

[0040] 1. Labeled nucleotide dimers: Synthesize or commercially purchase the following 3' hydroxyl-terminated reversibly blocked nucleotide dimers: AA, CC, GG, TT, AC, CA, GT, TG, AG, GA, CT, TC, AT, TA, CG, GC (see...) Figure 1 (and Table 1).

[0041] 2. Preparation of E. coli whole genome template: The target genome was fragmented into 100-1000 bp fragments using sonication. These fragmented nucleic acid sequences were then ligated using a pair of known universal linkers with the following sequence: Linker 1: CTG CTG TAC CGT ACA GCC TTG GCC G; Linker 2: CGC TTT CCT CTC TAT GGG CAG TCG GTGAT. Ten pre-amplification cycles were performed. Then, 200-800 bp DNA fragments were cut by gel electrophoresis and purified. These 200-800 bp DNA fragments were subjected to emulsion parallel PCR or bridge PCR to amplify the fragmented E. coli genome fragments and construct the E. coli genome sequencing DNA template chip.

[0042] B: First round of sequencing

[0043] a. Sequencing primer hybridization: The template fixed at the 5' end is hybridized with the first primer that is complementary to the 3' end linker. This hybridization primer serves as the sequencing primer for all *E. coli* genomic DNA templates (see [link to product]). Figure 2 );

[0044] b. Sequencing

[0045] (1) Sixteen 3'-O-azidomethyl-modified nucleotide dimers labeled with 100 μM tetrachromatograms (see...) Figure 1 The sequencing system (including 9° DNA polymerase) and Table 1 were added to the reaction chamber for the synthesis and sequencing reaction (60°C for 3 minutes). Then, the unreacted labeled nucleotide dimers were washed with 10mM disodium ethylenediaminetetraacetate (EDTA) buffer (pH = 7-8). The DNA template was identified as a valid sequencing template when the strongest fluorescence intensity of the four labeled dyes in a known base of the sequencing template was significantly higher than the background value. The dye coding record and position coordinate information corresponding to the strongest fluorescence intensity were recorded.

[0046] (2) Add 100 mM tris(2-carboxyethyl)phosphine (pH 8.0) and react at 55 °C for 3 minutes, then wash with 10 mM EDTA buffer (pH = 7-8);

[0047] (3) Repeat steps (1) to (2) above to perform the synthesis and sequencing reaction in a cycle to obtain a set of sequencing information consisting of codes 1, 2, 3, and 4. Then perform the second round of sequencing reaction.

[0048] C: Second round of sequencing

[0049] a. Treat with 8M urea at 65℃ for 5 minutes twice to remove the sequencing primers and their synthetic strands from the first round of sequencing reaction and obtain a new single-stranded DNA template.

[0050] b. Sequencing primer hybridization: The 5' end-fixed template is hybridized with a second primer that is complementary to the 3' end linker. This hybridization primer serves as the sequencing primer for all *E. coli* genomic DNA templates (see [link to product]). Figure 3 );

[0051] c. Sequencing

[0052] (1) 16 3'-O-azidomethyl-modified nucleotide dimers labeled with 100 μM tetrachromatograms and including 9 0 The DNA polymerase sequencing system was added to the reaction chamber for synthesis and sequencing reaction (reaction at 60°C for 3 minutes). Then, the unreacted labeled nucleotide dimers were washed with 10mM EDTA buffer (pH=7-8). The DNA template was identified as a valid sequencing template when the strongest fluorescence intensity of the four labeled dyes in a known base containing the sequencing template was significantly higher than the background value. The fluorescence intensity and position coordinates of the DNA were recorded.

[0053] (2) Add 100 mM tris(2-carboxyethyl)phosphine (pH 8.0) and react at 55 °C for 3 minutes, then wash with 10 mM EDTA buffer (pH = 7-8);

[0054] (3) Repeat steps (1) to (2) above to perform the synthesis and sequencing reaction in a cycle to obtain a set of sequencing information consisting of codes 1, 2, 3, and 4. Then perform the second round of sequencing reaction.

[0055] D. Decoding

[0056] Using the nucleotide dimer coding information obtained from two rounds of sequencing for each template, and using the nucleotide dimer coding information containing a known base from the sequencing template, the corresponding base sequence information for each template is decoded and assembled (see...). Figure 4 );

[0057] E. Sequence Assembly

[0058] Utilizing the base sequence information of all templates, and leveraging error correction and SNP recognition principles (see [link to documentation]). Figure 5 ), which are assembled into the E. coli genome sequence.

[0059] Table 1 Encoding of nucleotide dimers

[0060] coding Nucleotide dimers Marked dye 1 AA, CC, GG, TT FITC (fluorescein isothiocyanate): fluorescein isothiocyanate 2 AC, CA, GT, TG Cy3 (Cyanine 3): Anthocyanin 3 3 AG, GA, CT, TC Texas Red 4 AT, TA, CG, GC Cy5 (Cyanine 5): Anthocyanin 5

[0061] Table 2 shows the sequencing information obtained from the 3'-(T)TAATCAGGTCTG-5' sequence (where (T) in parentheses represents a known sequence) using a synthetic sequencing method based on nucleotide dimers as monomers.

[0062]

[0063] Explanation of the attached diagram:

[0064] Figure 1 This is a nucleotide dimer monomer structure based on a synthetic sequencing method using nucleotide dimers as monomers. The 3-terminal hydroxyl group is reversibly blocked with an azide group, and a cleavable fluorescent group is attached to the 2-base.

[0065] Figure 2 This is a sequencing method based on nucleotide dimers as monomers, using the first-round sequencing principle. 1 represents the DNA template, 1-1 and 1-2 are known common sequence fragments linked to both ends of the DNA template, 2 is the substrate, and 3 is the sequencing primer. The sequencing primers hybridize to fix the DNA template 1 on the substrate 3 to form a sequencing chip. Nucleotide dimers labeled with four different dyes are added, and a polymerization sequencing reaction occurs under DNA polymerase and its reaction system (1). Unreacted labeled nucleotide dimers are washed, and the fluorescence intensity of the four dyes is imaged and recorded. When the strongest fluorescence intensity among the four dyes is significantly higher than the background value, the DNA template is determined to be a valid sequencing template, and the dye code and position coordinates corresponding to the strongest fluorescence intensity are recorded. Then, a cleavage reagent is added to induce a cleavage reaction (2), cleaving the labeled dyes and activating the 3' hydroxyl groups to initiate the next polymerization sequencing reaction (1). (1) and (2) are repeated until the first round of sequencing is completed, obtaining the sequencing information of the DNA template from each sequencing reaction. Finally, treat with 8M urea at 65℃ for 5 minutes (3) twice in total to remove the sequencing primers and their sequencing primer synthesis strands from the first round of sequencing reaction, and obtain a new single-stranded DNA template for the second round of sequencing.

[0066] Figure 3This is a sequencing method based on nucleotide dimers as monomers, which is the principle for obtaining second-round sequencing information. 1 is the DNA template, 1-1 and 1-2 are known common sequence fragments linked to both ends of the DNA template, 2 is the substrate, and 3 is the sequencing primer. The sequencing primers hybridize to fix the DNA template 1 on the substrate 3 to form a sequencing chip. Nucleotide dimers labeled with four different dyes are added, and a polymerization sequencing reaction occurs under DNA polymerase and its reaction system (1). Unreacted labeled nucleotide dimers are washed, and the fluorescence intensity of the four dyes is imaged and recorded. When the strongest fluorescence intensity among the four dyes is significantly higher than the background value, the DNA template is determined to be a valid sequencing template, and the dye code and position coordinates corresponding to the strongest fluorescence intensity are recorded. Then, a cleavage reagent is added to induce a cleavage reaction (2), cleaving the labeled dyes and activating the 3' hydroxyl groups to initiate the next polymerization sequencing reaction (1). (1) and (2) are repeated until the second round of sequencing is completed, obtaining the sequencing information of the DNA template from each sequencing reaction.

[0067] Figure 4 This is a sequencing synthesis method based on nucleotide dimers as monomers, decoding the sequence. This method obtains sequencing encoding information from the 3'-(T)TAATCAGGTCTG-5' sequence to be tested (where (T) in parentheses represents a known sequence). In the left figure, the first row shows the dye-encoded information obtained from the first round of sequencing; the second row shows the nucleotide dimer base information decoded from the first round of sequencing; the third row shows the nucleotide dimer base information deduced from the first round of sequencing given the first base information; the fourth row shows the nucleotide dimer base information decoded from the second round of sequencing; the fifth row shows the dye-encoded information obtained from the second round of sequencing; and the sixth row shows the nucleotide dimer base information deduced from the second round of sequencing given the first round base information. The right figure shows the combined information from two rounds of sequencing: the first row shows the DNA template sequence to be tested; the second row shows the dye-encoded information obtained from both rounds of sequencing, where odd-numbered positions represent the first round and even-numbered positions represent the second round; and the third row shows the base information decoded from the dye-encoded information obtained from both rounds of sequencing given the first base information.

[0068] Figure 5This is a sequencing method based on nucleotide dimers as monomers, employing error correction and SNP identification principles. Reference information refers to information known in a database, or information obtained simultaneously from different DNA templates at the same position. The left image shows the comparison results between the sequencing information and the reference information, revealing a different code (shown by a triangle in the comparison results). This different code is identified as a sequencing error (at a sequencing depth of 2×, the presence of a sequencing error can be determined), and this incorrect code ② is corrected to the correct code ③. The right image shows the comparison results between the sequencing information and the reference information, revealing two consecutive different codes (shown by a triangle in the comparison results). These two consecutive different codes are identified as SNPs (at sequencing depths of 3× and above, sequencing errors can be corrected based on probability), and it is determined that this SNP has changed from the reference sequence base G to the sequencing sequence base C.

Claims

1. A method of sequencing by synthesis based on nucleotide dimers as monomers, characterized in that: The nucleotide for sequencing is a nucleotide dimer, under the action of polymerase, if the nucleotide dimer is completely paired with two bases in the DNA template hybridized with the sequencing primer, the polymerization of the nucleotide dimer occurs, and the sequencing primer is extended by two bases; if the two bases in the DNA template are not completely paired, the polymerization of the nucleotide dimer does not occur; The nucleotide dimer includes sixteen specific forms, namely AA, CC, GG, TT, AC, CA, GT, TG, AG, GA, CT, TC, AT, TA, CG, and GC; The 3' end hydroxyl of the nucleotide dimer is blocked by a group including 3'-O-allyl, 3'-O-cyanoethyl, 3'-O-azidomethyl, and 3'-O-amino; The specific base information of the DNA template to be tested is obtained through two rounds of sequencing, and each round of sequencing obtains specific dye or coding information. The sequencing primers for the two rounds of sequencing differ by one base, wherein the sequencing primer for one round of sequencing corresponds to one known base of the sequencing template, and the specific base information of the DNA template to be tested is obtained by decoding the coding information of the two groups of base fragments in sequence according to the one known base of the sequencing template. For resequencing of genomic sequencing with a reference sequence, the coding obtained by real-time synthesis DNA sequencing of the two groups of nucleotides can be directly used for alignment of the genomic reference sequence, without the need for decoding the coding, thereby achieving resequencing of the genomic sequence.

2. The nucleotide-dimer monomer-based sequencing-by-synthesis method according to claim 1, characterized in that: The basic structure of the nucleotide dimer is that two same or different nucleotides are connected by a phosphate bond, and a triphosphate group is further included at the 5' end.

3. The nucleotide-dimer monomer-based sequencing-by-synthesis method according to claim 1, wherein: The 3' end modified nucleotide dimer can participate in the synthesis sequencing reaction, and at most one nucleotide dimer synthesis can occur. The 3' end blocking group can be removed and activate the 3' end hydroxyl of the nucleotide under suitable reaction conditions.

4. The nucleotide-dimer monomer-based sequencing-by-synthesis method according to claim 1, wherein: The nucleotide dimer is labeled with a dye, and the labeling position can be the 5' end base position, the 5' end phosphate end, or the 3' end hydroxyl position.

5. The nucleotide dimer-based monomer synthetic sequencing method according to claim 4, characterized in that: When the labeling position is the 3' end hydroxyl position, the dye can replace the blocking group.

6. The nucleotide-dimer monomer-based sequencing-by-synthesis method of claim 1, wherein: AA, CC, GG, and TT are labeled with a first dye, AC, CA, GT, and TG are labeled with a second dye, AG, GA, CT, and TC are labeled with a third dye, and AT, TA, CG, and GC are labeled with a fourth dye. The emission wavelengths of the four dyes cannot interfere with each other, and the labeled dyes can be cut by chemical or photochemical methods, and the cutting does not affect the biochemical function of the DNA sequence.

7. The nucleotide-dimer monomer-based sequencing-by-synthesis method according to claim 1, wherein: The DNA template fragment to be tested refers to a single molecule or a product of the same sequence amplified by a single molecule as a template. The DNA template to be tested can be one sequence or an array of different DNA templates.

8. The nucleotide-dimer monomer-based sequencing-by-synthesis method of claim 1, wherein: The information of the DNA template to be tested is obtained by comparing the fluorescence intensity of the dye with the background value. When the strongest fluorescence intensity of the four labeling dyes in the DNA template containing a known base of the sequencing template is significantly higher than the background value, it can be determined that the DNA template is an effective sequencing template, and the fluorescence intensity and position coordinate information of the DNA are recorded.

9. The nucleotide dimer-based monomer synthetic sequencing method according to any one of claims 1 to 8, characterized in that: The method comprises the following steps: A: Preparation of E. coli genome-wide template: The target genome is broken into 100-1000 bp fragments by ultrasonic, and the fragmented nucleic acid sequences are ligated with a pair of sequence-known universal adapters, wherein the sequence of adapter 1 is CTG CTG TAC CGT ACA GCC TTG GCC G, and the sequence of adapter 2 is CGC TTT CCT CTC TAT GGG CAG TCG GTG AT; and pre-amplification is performed; then 200-800 bp DNA fragments are cut by gel electrophoresis and purified; the 200-800 bp DNA fragments are subjected to emulsion and parallel PCR reaction or bridge PCR to amplify the fragmented E. coli genomic fragments and construct E. coli genomic sequencing DNA template chip; B: First round of sequencing a, sequencing primer hybridization: the 5' end-fixed template is hybridized with the first primer complementary to the 3' end adapter, and the hybridized primer is used as the sequencing primer for all E. coli genomic DNA templates; b, sequencing (1) 16 3'-O-azidomethyl-modified nucleotide dimers labeled with four colors and a sequencing system including 9° DNA polymerase are added to the reaction pool for synthesis sequencing reaction, then uninvolved labeled nucleotide dimers are washed with ethylenediaminetetraacetic acid disodium buffer, imaging, and recording the fluorescence intensity of the four labeled dyes in the known base of the sequencing template is significantly higher than the background value, determining that the DNA template is an effective sequencing template, and recording the fluorescence intensity and position coordinate information of the DNA; (2) adding tri (2-carboxyethyl) phosphine, and then washing with 10 mM EDTA buffer; (3) repeating the synthesis sequencing reaction according to steps (1)-(2) above to obtain a set of sequencing information composed of 1, 2, 3, and 4, and then performing the second round of sequencing reaction; C: Second round of sequencing a, urea treatment to remove the sequencing primer and its synthesized strand in the first round of sequencing reaction to obtain single-stranded DNA template again; b, sequencing primer hybridization: the 5' end-fixed template is hybridized with the second primer complementary to the 3' end adapter, and the hybridized primer is used as the sequencing primer for all E. coli genomic DNA templates; c, sequencing (1) 16 3'-O-azidomethyl-modified nucleotide dimers labeled with four colors and a sequencing system including 9° DNA polymerase are added to the reaction pool for synthesis sequencing reaction, then uninvolved labeled nucleotide dimers are washed with ethylenediaminetetraacetic acid disodium buffer, imaging, and recording the fluorescence intensity of the four labeled dyes in the known base of the sequencing template is significantly higher than the background value, determining that the DNA template is an effective sequencing template, and recording the fluorescence intensity and position coordinate information of the DNA; (2) adding tri (2-carboxyethyl) phosphine, and then washing with 10 mM EDTA buffer; ​ (3) The synthesis sequencing reaction is cycled according to the above steps (1)-(2) to obtain a group of sequencing information composed of 1, 2, 3, 4, and then a second round of sequencing reaction is performed; D, decoding The nucleotide dimer encoding information obtained in the two rounds of sequencing of each template is used, and the nucleotide dimer encoding information containing a known base of the sequencing template is used to decode and assemble the corresponding base sequence information of each template; E, sequence assembly The base sequence information of all templates is used, and the error correction and SNP identification principle is used to assemble the E. coli genome sequence.

Citation Information

Patent Citations

  • DNA sequencing method capable of verifying base information for second time

    CN101575639A

  • 3'end reversible closed two-nucleotide real-time synthesizing and sequencing method

    CN106434866A