Method for determining sequence of template polynucleic acid
By performing multiple rounds of sequencing on template polynucleotides and combining the detection of a mixture of labeled and unlabeled nucleotides, the problem of signal crosstalk in four-color or two-color sequencing platforms was solved, achieving high-precision, low-cost single-color sequencing.
Patent Information
- Application Number
- CN202411157437.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-03
AI Technical Summary
In existing four-color or two-color sequencing platforms, the sequencing accuracy is affected by signal crosstalk between different wavelengths, and the optical system is complex and the hardware cost is high.
The method employs at least two rounds of sequencing of template polynucleotides, with each round containing multiple sequencing cycles. It uses a mixture of nucleotides containing and without detection tags, collects emission signals through a single imaging event, and combines signal combinations from different rounds to distinguish nucleotide types, thereby achieving single-wavelength fluorescence signal detection.
It reduces the requirements for the sequencer's optical system, simplifies the imaging system, reduces hardware costs, and improves sequencing accuracy and applicability.
Smart Images

Figure CN121592764A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of gene sequencing and relates to a method for determining the sequence of template polynucleotides. Background Technology
[0002] High-throughput sequencing technology has become a commonly used analytical method in life science research due to its high throughput, fast detection speed, flexibility, and low cost. This technology typically determines the types of bases incorporated into the sequencing reaction by detecting fluorescence signals in a chip, thereby obtaining the sequence information of the nucleic acid being tested. However, in existing four-color or two-color sequencing platforms, due to signal crosstalk between different wavelengths, when the optical channel for a specific base is photographed, the dyes for other bases will produce a certain amount of fluorescence signal. This is because different optical channels cannot be completely separated, and the excitation and emission wavelengths of different dyes overlap. This crosstalk phenomenon adversely affects sequencing accuracy. Furthermore, four-color or two-color imaging complicates the sequencer's optical system and significantly increases the hardware cost of the sequencer. Therefore, to overcome these problems, new sequencing methods need to be developed to improve sequencing accuracy and reduce costs. Summary of the Invention
[0003] This invention discloses a method for determining the sequence of template polynucleotides, which at least partially solves the problems existing in the prior art.
[0004] This invention provides a method for determining the sequence of a template polynucleotide, characterized by comprising: performing at least two rounds of sequencing on a template polynucleotide immobilized at a flow cell reaction site, each sequencing round comprising multiple sequencing cycles, each sequencing cycle comprising: introducing an incorporation mixture containing four types of nucleotide monomers, wherein at least one type and up to three types of nucleotide monomers contain nucleotides labeled with a detection tag, the incorporation mixture further comprising nucleotides without a detection tag; incorporating a single nucleotide from the incorporation mixture into a primer polynucleotide that binds to the template polynucleotide to generate an extended primer polynucleotide; performing a single imaging event and collecting the emission signal; wherein,
[0005] The combinations of at least one and at most three nucleotides containing the detection tag are different in at least two rounds of sequencing; the types of incorporated nucleotide monomers are distinguished based on the combination of emission signals obtained from at least two rounds of sequencing, thereby obtaining the initial sequence of the template polynucleotide.
[0006] In a specific embodiment of the present invention, two rounds of sequencing are performed on the template polynucleotide fixed at the reaction site in the flow cell. There are two types of nucleotide monomers containing nucleotides labeled with the detection tag. The combinations of the two types of nucleotides containing the detection tag are different in the two rounds of sequencing. One type of nucleotide is the same, and the other type of nucleotide is different.
[0007] In a specific embodiment of the present invention, the aforementioned method further includes:
[0008] An additional sequencing round is performed, in which the combination of two nucleotide substrates containing the detection tag used in this round of sequencing differs from the combination of two nucleotides containing the detection tag used in the previous two rounds of sequencing, to obtain at least one additional sequence, and the additional sequence is compared with the initial sequence to reduce or eliminate one or more errors contained in the initial sequence.
[0009] In a specific embodiment of the present invention, the relative signal intensity of the nucleotide monomer containing the detection tag is 0.5-1, preferably 0.8-1, and more preferably 0.95-1.
[0010] In a specific embodiment of the present invention, the relative signal intensities of the two nucleotide monomers containing the detection tag are both close to 1.
[0011] In a specific embodiment of the present invention, the structure of the nucleotide containing the detection tag is of formula (I):
[0012]
[0013] Wherein, Y is O or S, B is a heterocyclic base, and n is an integer from 0 to 6; R is selected from azidomethyl, amino, allyl, substituted dithioalkyl, substituted methoxymethyl, o-nitrobenzyl, coumarin, phosphate nitrile ethyl ester, trimethylsilyl, tetrahydropyranyl, azido, alkyl hydroxyamino, thiophosphate, malonyl, benzyl, acetal, thiocarbamate, and vinyl; Fluorogenic Dye is selected from anthracene, phenoxazine, acridine, and coumarin with fluorescence switching properties, wherein the anthracene with fluorescence switching properties includes oxanthracene-fluorescein, carbamate-Beijing orange, silanthracene, germananthracene, phosphoroxanthracene, and thioanthracene.
[0014] In a specific embodiment of the present invention, the reaction site comprises multiple reaction volumes created by multiple reaction chambers disposed on an array, wherein the template polynucleotide is immobilized in the reaction volume; after the incorporation mixture is delivered to each reaction volume, each reaction volume may be closed and / or separated from other reaction volumes on the array; then, the emission signal from the detection tag may be detected and / or recorded by each reaction volume.
[0015] In a specific embodiment of the present invention, the structure of the unlabeled nucleotide is as shown in formula (II):
[0016]
[0017] Where Y is O or S; X is O or a group that cannot be degraded by phosphatase, where GD is OH or any non-fluorescent organic group; n is an integer from 0 to 6; B is a heterocyclic base; and Cap1 is a reversible termination group.
[0018] Preferably, X is O, and GD is selected from optionally substituted C1-10 alkyl, C1-10 alkyloxy, amino, mono- or disubstituted amino, C6-14 aryl, C6-14 aryloxy, C6-14 arylamino, or C3-12 cycloalkyl, C3-12 cycloalkyloxy, C3-12 cycloalkylamino; C2-13 heteroaryl or C2-13 heterocyclic or combinations thereof; preferably, GD is selected from C1-4 alkyl, C1-4 alkoxy, C1-4 alkylamino or di(C1-4 alkyl)amino, or amino, or phenyl, phenoxy, aniline or combinations thereof;
[0019] Preferably, X is CH2, CF2, or NH, and GD is selected from OH, C1-10 alkyl, C1-10 alkyloxy, mono- or disubstituted amino, C6-14 aryl, C6-14 aryloxy, C6-14 arylamino, or C3-12 cycloalkyl, C3-12 cycloalkyloxy, C3-12 cycloalkylamino; C2-13 heteroaryl or C2-13 heterocyclic or combinations thereof; preferably, GD is OH.
[0020] In a specific embodiment of the present invention, the combination of the two types of nucleotide monomers containing the detection tag is selected from:
[0021] 1) Use A and T / U deoxyribonucleotides in one round of sequencing, or use C and G deoxyribonucleotides; or
[0022] 2) Use A and G deoxyribonucleic acid in one round of sequencing, or use C and T / U deoxyribonucleic acid; or
[0023] 3) Use A and C deoxyribonucleic acid in one round of sequencing, or use G and T / U deoxyribonucleic acid.
[0024] In an optional embodiment of the invention, the method further includes the aforementioned four types of nucleotide monomers having a reversible termination group at their 3' ends, and removing the reversible termination group at the 3' ends of the incorporated nucleotides before the next sequencing cycle, thereby making the nucleotides into an extendable state.
[0025] In a specific embodiment of the present invention, before performing a new round of sequencing reaction, the double strands from the previous sequencing are unwound into single strands, the extended strands of the primers from the previous round are washed away, and then hybridized with the new sequencing primers for the next round.
[0026] The method for determining the sequence of template polynucleotides disclosed in this invention has the following advantages: It requires less sophisticated optical systems from the sequencer, needing only a single wavelength of excitation light and a single wavelength of fluorescence signal reception. Compared to current mainstream four-color and two-color sequencers, the imaging system is simpler, and hardware costs are significantly reduced. Because only monochrome imaging is required, there is no crosstalk between optical channels, resulting in higher sequencing accuracy. Furthermore, the sequencing method disclosed in this invention does not have special requirements for the fluorescent labeling method of nucleotides; it only needs to meet the fluorescence signal detection requirements for gene sequencing, thus broadening its applicability. Attached Figure Description
[0027] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0028] Figure 1 A schematic diagram illustrating the sequencing method of the present invention is shown.
[0029] Figure 2 The figure shows the statistical results of the sequencing error rate for each sequencing cycle according to a specific embodiment of the present invention.
[0030] Figure 3 The figure shows the statistical results of the sequencing error rate for each sequencing cycle according to a specific embodiment of the present invention.
[0031] Figure 4 The figure shows the statistical results of the sequencing error rate for each sequencing cycle according to a specific embodiment of the present invention.
[0032] Figure 5 The figure shows the statistical results of the sequencing error rate for each sequencing cycle according to a specific embodiment of the present invention.
[0033] Figure 6 The figure shows the statistical results of the sequencing error rate for each sequencing cycle according to a specific embodiment of the present invention. Detailed Implementation
[0034] Terminology Explanation
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0036] The use of the term "comprising" is not restrictive. As used herein, whether in transitional phrases or in the body of the claims, the terms "comprising" and "comprising" are to be interpreted in an open-ended sense. That is, the terms should be interpreted synonymously with the phrases "at least having" or "at least including." For example, when used in the context of a process, the term "comprising" means that the process includes at least the listed steps, but may also include additional steps. When used in the context of a compound, composition, or device, the term "comprising" means that the compound, composition, or device contains at least the listed features or components, but may also contain additional features or components.
[0037] The term “or” means one or all of the listed elements, or any combination of two or more of the listed elements, unless the context otherwise requires.
[0038] The terms "preferred" and "ideally" refer to embodiments of the invention that provide certain benefits in certain circumstances. However, other embodiments may also be preferred in the same or other circumstances. Furthermore, the description of one or more preferred embodiments does not imply that other embodiments are unavailable, and is not intended to exclude other embodiments from the scope of the invention.
[0039] Unless otherwise stated, nucleic acids are written from left to right in the 5' to 3' direction; amino acid sequences are written from left to right in the direction from amino to carboxyl.
[0040] Unless otherwise stated, all headings are for the reader's convenience and should not be used to limit the meaning of the text following them. The headings provided herein are not intended to limit the various aspects or embodiments of the invention, which can be obtained by referring to the entire specification. Therefore, the terms defined below are defined more fully with reference to the entire specification.
[0041] As used herein, "polymerase" refers to an enzyme that catalyzes the polymerization of nucleotides (i.e., polymerase activity). Typically, this enzyme begins synthesis at the 3' end of a primer annealed to the polynucleotide template sequence and proceeds toward the 5' end of the template strand. "DNA polymerase" catalyzes the polymerization of deoxyribonucleotides. Deoxyribonucleotides can be native nucleotides or modified or labeled nucleotides.
[0042] As used in this invention, "nucleotide" comprises a nitrogenous heterocyclic base, a ribose sugar, and one or more phosphate groups. These are monomeric units of nucleic acid sequences. In RNA, the sugar is ribose, and in DNA, it is deoxyribose, i.e., a sugar lacking the hydroxyl group present in the ribose. The nitrogenous heterocyclic base may be a purine, denitropurine, or pyrimidine base. Purine bases include adenine (A) and guanine (G) and their modified derivatives or analogs, such as 7-denitroadenine or 7-denitroguanine. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U) and their modified derivatives or analogs. The C-1 atom of the deoxyribose sugar is bonded to the N-1 of a pyrimidine or the N-9 of a purine. In some cases, the term "nucleotide" may also cover nucleotide modifications or conjugates.
[0043] As used in this invention, the template polynucleotide is a single-stranded nucleotide or nucleotide cluster of any origin and of a length suitable for sequencing-by-synthesis; preferably, the template polynucleotide is immobilized on the surface of a vector.
[0044] As used in this invention, a sequencing reaction cycle refers to the use of a sequencing reaction solution to perform multiple sequencing reaction cycles in a predetermined order. For example, in one sequencing reaction cycle, the sequencing reaction solution used contains four substrate nucleotides: A, C, G, and T, where A and C are labeled with detection tags, and G and T are not labeled with detection tags. At least 50 sequencing reaction cycles are performed using this sequencing reaction solution. In another sequencing reaction cycle, the sequencing reaction solution used contains four substrate nucleotides: A, C, G, and T, where A and T are labeled with detection tags, and G and C are not labeled with detection tags. At least 50 sequencing reaction cycles are performed using this sequencing reaction solution.
[0045] As used in this invention, “extending” (or “incorporating”) into an oligonucleotide or polynucleotide means that a covalent bond is formed between the nucleotide and the oligonucleotide or polynucleotide. In some such embodiments, a phosphodiester bond is formed between the 3' hydroxyl group of the oligonucleotide or polynucleotide and the 5' phosphate group of the nucleotide.
[0046] Reversible termination nucleotide
[0047] The nucleic acid sequencing method of this invention requires that each sequencing cycle can only extend one base. Therefore, the sequencing method of this invention is based on the use of reversible termination nucleotides, which are well known in the fields of sequencing and nucleic acid synthesis, such as conventionally used 3′-modified reversible termination nucleotides (e.g., oxyamino groups). Those skilled in the art will understand how to attach a suitable protecting group to the 3′ position of the ribose to block the interaction between polymerase and the 3′-OH. The aforementioned protecting group can be directly attached to the 3′ position of the ribose. The protecting group attached to the 3′ position can be cleaved (or removed) to revert to an extendable 3′-OH. As used herein, the terms “blocking” or “closing” or “terminating” refer to the use of a specific group to protect the 3′-OH of a nucleoside or nucleotide to terminate the polymerization of a possible polymerase (e.g., DNA polymerase). The group used for blocking is called a “blocking group” or “closing group” or “terminating group” or “protecting group.” If the blocking group can be removed, thereby allowing the nucleotide to revert to an extendable state, such blocking is called “reversible blocking” or “reversible closing” or “reversible termination.” The group used for reversible termination is called the "reversible termination group." The process of removing the reversible termination group and changing the 3' position of the ribose back to an extendable 3'-OH is called "deprotection." Regarding reversible termination groups and corresponding deprotection reagents: The reversible termination group (Cap1) used in this article can be selected from the following groups: 3'-O-azidomethyl, 3'-O-amino, 3'-O-allyl, 3'-O-substituted dithioalkyl, 3'-O-substituted methoxymethyl, 3'-O-o-nitrobenzyl, 3'-O-coumarin, 3'-O-phosphocyanoethyl ester, 3'-O-trimethsilyl, 3'-O-THP, 3'-O-azido, 3'-O-alkylhydroxyamino, 3'-thiophosphate, 3'-O-malonyl, 3'-O-benzyl, 3'-O-acetal, 3'-O-thiocarbamate, 3'-vinyl, etc. Other available reversible terminating groups include, for example, those disclosed in international applications WO2014139596A1 and CN201780039312.5. This invention does not particularly limit the type of reversible terminating group; according to the principles of this invention, any group in the art that is used for reversible blocking of 3'-OH can be used as a reversible terminating group in this invention. For example, 3'-O reversible terminating groups that can be cleaved by a reducing agent (e.g., phosphine) include, but are not limited to, azidomethyl. 3'-O reversible terminating groups that can be cleaved by ultraviolet light include, but are not limited to, nitrobenzyl. 3'-O reversible terminating groups that can be cleaved by contact with an aqueous Pd solution include, but are not limited to, allyl. 3'-O reversible terminating groups that can be cleaved by acid include, but are not limited to, methoxymethyl.The 3'-O reversible terminating group that can be cleaved by contact with a sodium nitrite-buffered aqueous solution (pH = 5.5) includes, but is not limited to, aminoalkoxy groups. Correspondingly, commonly used deprotecting agents can be selected from: for example, for 3'-O-thiocarbamates, removal can be achieved under various conditions, such as, non-limiting exemplary conditions including NaIO4 and potassium persulfate; for 3'-O-acetal, cleavage can be achieved by a palladium catalyst; for 3'-O-azido or 3'-O-azidomethyl, the azide group can be converted to an amino group by contact with phosphine, such as, non-limiting exemplary conditions including phosphine, such as trialkylphosphine, non-limiting examples of which include tris(hydroxymethyl)phosphine (THP), tris(2-carboxyethyl)phosphine (TCEP), tris(hydroxymethyl)phosphine (THMP), or tris(hydroxyethyl)phosphine (THEP), etc. For 3'-vinyl, various tetrazine substances can be used to deprotect the 3'-vinyl protecting group. The deprotection mechanism of the 3'-vinyl protecting group may be similar to that of simple tetrazines, such as substituted tetrazines; each of v-tetrazine, as-tetrazine, and / or s-tetrazine can be used for the deprotection of the 3'-vinyl end-capped group. For 3'-O-allyl, the cleavage process employs a ligand of metallic palladium and trisodium triphenylphosphine tris(m-sulfonate), removing the blocking group via a palladium-catalyzed deallylation reaction. For 3'-O-THP, the THP protecting group is stable under basic conditions but unstable under mildly acidic conditions, for example, it can be cleaved using HOAc / THF / H2O (4:2:1) under slight heating.
[0048] 5'-terminal phosphorylated reversible termination nucleotide
[0049] In some preferred embodiments of the present invention, the nucleic acid sequencing method relies on a reversible 5′-terminal phosphate-labeled fluorescent nucleotide. The 5′-terminal phosphate-labeled nucleotide is described in CN104910229A. Hydroxyl-containing xanthracene, coumarin, and halogenated fluorescent molecules are introduced as detection tags via polyphosphate at the 5′ end of the nucleotide. This nucleotide structure can serve as a substrate for DNA or RNA polymerase, be recognized by the polymerase, and incorporated into the DNA or RNA chain, releasing a polyphosphate molecule containing a fluorescent group or a self-cleaving molecule containing a fluorescent group. The terminally phosphate-labeled polyphosphate molecule can be further decomposed as a substrate for an enzyme (e.g., alkaline phosphatase) until all phosphate groups are detached from the fluorescent molecule, thereby generating a fluorescent signal. The aforementioned fluorescent molecule is a fluorophore with fluorescence switching properties. In a preferred embodiment, the sequencing method may include using an enzyme to release the fluorophore with fluorescence switching properties, optionally including first using a DNA polymerase to release the polyphosphate-substituted fluorophore, and then using a phosphatase to cleave the substituted polyphosphate, thereby releasing the fluorophore. The dyes that can be used in the embodiments of this application include anthracene, phenoxazine, acridine, and coumarin dyes with fluorescence switching properties. The aforementioned anthracene dyes with fluorescence switching properties include: xanthracene-fluorescein, carbamate-Beijing Orange, silicane, germanium anthracene, phosphoroxane, and thioanthracene. Exemplary fluorophores include phenolic dyes such as fluorescein (e.g., methoxyfluorescein), phenoxazine (e.g., halogen), acridine (e.g., DDAO), and coumarins (e.g., 7-hydroxycoumarin, 3-carboxy-7-hydroxycoumarin, 4-acetic acid-7-hydroxycoumarin, 7-hydroxy-4-trifluoromethylcoumarin), Beijing Orange (PO), Tokyo Green (TG), and Tokyo Magenta (example structures of the above fluorophores can be found in Table 1). The chemical reactions of fluorescent nucleic acid substrates based on phenolic dyes are relatively simple because the phenolic oxygen is esterified into a phosphate group. For amine-containing dyes (e.g., rhodamine and its derivatives, cresol violet, etc.), nucleotide substrates can be introduced by introducing self-cleaving groups. Once DNA polymerase incorporates the labeled nucleotide substrate, it cleaves the nucleotide between its α- and β-phosphates. The released fluorophore fluoresces either directly from the nucleotide cleavage or after further enzymatic action by other enzymes (e.g., alkaline phosphatase). These new fluorescent molecules are then detected using standard fluorescence detection techniques (e.g., total internal reflection fluorescence, epifluorescence, or confocal microscopy). Fluorescence switching refers to a significant change in the fluorescence signal after sequencing compared to before the sequencing reaction; commonly, the fluorescence signal after sequencing is significantly enhanced (or increases) compared to before the sequencing reaction.
[0050] In a preferred embodiment of the present invention, the 5′-terminal phosphate-labeled reversible termination nucleotide used has the structure shown in formula (I):
[0051]
[0052] Wherein, Y is O or S, B is a heterocyclic base, and n is an integer from 0 to 6; R is selected from azidomethyl, amino, allyl, substituted dithioalkyl, substituted methoxymethyl, o-nitrobenzyl, coumarin, phosphate nitrile ethyl ester, trimethylsilyl, tetrahydropyranyl, azido, alkyl hydroxyamino, thiophosphate, malonyl, benzyl, acetal, thiocarbamate, and vinyl; Fluorogenic Dye is selected from the aforementioned fluorophores with fluorescence switching properties.
[0053] Table 1 summarizes exemplary alternative dyes. To reduce light damage during sequencing, dyes with excitation wavelengths greater than 400 nm are preferably selected from these options.
[0054] Table 1
[0055]
[0056]
[0057] The unlabeled reversible nucleotides (referred to as dark nucleotides) in this invention refer to those that do not produce a specific signal in a detection event. For example, the nucleotide substrate itself may lack a fluorescent label; these nucleotide substrates are referred to as dark nucleotides in this invention. This invention designs dark nucleotides with unique structures that can act as ordinary non-luminescent substrates, cooperating with fluorescently labeled substrates to achieve monochromatic sequencing reactions. For the most convenient implementation of the sequencing reaction, the reversible termination group contained in the dark nucleotide is the same as that contained in the 5′-terminal phosphate-labeled reversible nucleotide, or at least can be blocked in the same way.
[0058] In this invention, monochrome sequencing (or single-channel sequencing) refers to a sequencing method that uses a single excitation light source and a single optical channel for detection to identify four types of nucleotides. The four types of nucleotide analogs may contain the same type of fluorophore or different fluorophores with the same or similar excitation / emission spectra. The aforementioned use of a single excitation light source and a single optical channel for detection constitutes the execution of a single imaging event as described in this invention; therefore, this single imaging event is independent of the specific number of images captured.
[0059] In this invention, "relative signal intensity" refers to the ratio of the fluorescence signal intensity of each substrate nucleotide to the fluorescence signal intensity corresponding to the fluorescence signal produced when each substrate nucleotide molecule of that substrate nucleotide elongates (the value ranges from 0 to 1). For example, if substrate nucleotide A is 100% labeled with a fluorophore, meaning that each substrate nucleotide molecule elongates to produce a fluorescence signal, then the relative signal intensity of A is 1; if a portion of substrate nucleotide T is labeled with a fluorophore and the remaining portion is not labeled with a fluorophore, adjusting the ratio of the two can make the relative signal intensity of T a value between 0 and 1; substrate nucleotide C can be completely unlabeled with a fluorophore, then the relative signal intensity of C is 0. Invention Details
[0061] In existing high-throughput sequencing platforms, most employ four-color or two-color imaging to distinguish the four bases. This results in signal crosstalk between different wavelengths, affecting sequencing accuracy. Furthermore, four-color or two-color imaging makes the sequencer's optical system highly complex, significantly increasing hardware costs. To address these issues, patent CN202080003541.3 proposes a monochrome imaging sequencing method using nucleotides with photo-switching labels. These labels can reversibly change from an "on" to an "off" state under specific wavelengths of light, allowing control of the label's fluorescence state during sequencing, thus achieving monochrome sequencing. However, this monochrome sequencing method strictly relies on these photo-switching labels; ordinary fluorescent labels are not applicable. Moreover, the development and synthesis of photo-switching labels may involve complex chemical processes, potentially increasing production costs and difficulty. Therefore, this monochrome sequencing method has a narrow scope of application and has not been widely adopted.
[0062] This invention discloses a method for determining the sequence of a template polynucleotide, which at least partially solves the aforementioned problems. Specifically, the method includes:
[0063] At least two rounds of sequencing are performed on the template polynucleotide immobilized at the reaction site in the flow cell. Each sequencing round comprises multiple sequencing cycles. Each sequencing cycle includes: introducing an incorporation mixture containing four types of nucleotide monomers, wherein at least one type and up to three types of nucleotide monomers contain nucleotides labeled with a detection tag, and the incorporation mixture also includes nucleotides without a detection tag; incorporating a single nucleotide from the incorporation mixture into a primer polynucleotide that binds to the template polynucleotide to generate an extended primer polynucleotide; performing a single imaging event and collecting the emission signal; wherein...
[0064] The combinations of at least one and at most three nucleotides containing the detection tag are different in at least two rounds of sequencing; the types of incorporated nucleotide monomers are distinguished based on the combination of emission signals obtained from at least two rounds of sequencing, thereby obtaining the initial sequence of the template polynucleotide.
[0065] In this method, the incorporation mixture provided in each sequencing cycle contains four types of nucleotide monomers (i.e., A, G, C, T (or U)), and only one nucleotide is extended per cycle. The extended substrate nucleotide is a reversible termination nucleotide, which needs to be de-blocked and returned to an extendable state before the next sequencing cycle reaction. For example, its reversible end capping can be removed after performing an imaging event.
[0066] In some specific implementations, the template polynucleotide immobilized at the flow cell reaction site is sequenced in three rounds. Each sequencing round comprises multiple sequencing cycles. Each sequencing cycle includes: introducing an incorporation mixture containing four types of nucleotide monomers, three of which contain nucleotides labeled with a detection tag; incorporating a single nucleotide from the incorporation mixture into a primer polynucleotide that binds to the template polynucleotide to generate an extended primer polynucleotide; performing a single imaging event and collecting the emission signal; wherein...
[0067] The three combinations of nucleotides containing detection tag markers differed in the three rounds of sequencing;
[0068] The types of incorporated nucleotide monomers are distinguished based on the combination of emission signals obtained from three rounds of sequencing, thereby obtaining the initial sequence of the template polynucleotide. The aforementioned three rounds of sequencing include combinations of three nucleotides for detecting the tag, selected from:
[0069] 1) A, C, T / U deoxyribonucleotides are used in one round of sequencing; or,
[0070] 2) A, C, and G deoxyribonucleotides are used in one round of sequencing; or,
[0071] 3) A, G, T / U deoxyribonucleotides are used in one round of sequencing; or,
[0072] 4) C, G, T / U deoxyribonucleotides are used in one round of sequencing.
[0073] For ease of understanding, let's assume that the three nucleotides containing the detection tag in the first three rounds of sequencing are selected as 1), 2), and 3) respectively. The emission signal of the three nucleotides containing the detection tag is 1, and the emission signal of the nucleotide without the detection tag is 0. Then, the signal combinations obtained in the three rounds of sequencing are as follows: A(1, 1, 1), C(1, 1, 0), G(0, 1, 1), T(1, 0, 1). The signals of the four substrate nucleotides are all unique. Therefore, the types of these four nucleotides can be accurately distinguished through three rounds of sequencing, thus obtaining the sequence of the template polynucleotide. The advantages of this method are obvious: it achieves monochromatic sequencing by utilizing the different signal combinations between different nucleotides in multiple rounds of sequencing, without relying on special fluorescent labels, and has a wide range of applications.
[0074] In some preferred embodiments, the template polynucleotide immobilized at the flow cell reaction site is subjected to two rounds of sequencing, each round comprising multiple sequencing cycles. Each sequencing cycle includes: introducing an incorporation mixture containing four types of nucleotide monomers, two of which contain nucleotides labeled with a detection tag; incorporating a single nucleotide from the incorporation mixture into a primer polynucleotide that binds to the template polynucleotide to generate an extended primer polynucleotide; performing a single imaging event and collecting the emission signal; wherein...
[0075] The two combinations of nucleotides containing the detection tag were different in the two rounds of sequencing; one nucleotide was the same, and the other nucleotide was different.
[0076] The type of incorporated nucleotide monomer is distinguished by the combination of emission signals obtained from two rounds of sequencing, thereby obtaining the initial sequence of the template polynucleotide. The aforementioned combination of two nucleotides containing the detection tag is selected from:
[0077] 1) Use A and T / U deoxyribonucleotides in one round of sequencing, or use C and G deoxyribonucleotides; or
[0078] 2) Use A and G deoxyribonucleic acid in one round of sequencing, or use C and T / U deoxyribonucleic acid; or
[0079] 3) Use A and C deoxyribonucleic acid (DNA) or G and T / U DNA in one round of sequencing. Figure 1To illustrate the method of this embodiment in detail, in the first round of sequencing, the two nucleotides containing the detection tag are A and C. For ease of explanation, their relative signal intensity is set to 1. The relative signal intensity of G and T, which do not contain the detection tag, is 0. The combination of the relative signal intensities of the four nucleotides is A,C,G,T (1,1,0,0). In the second round of sequencing, the two nucleotides containing the detection tag are A and T, and their relative signal intensity is 1. The combination of the relative signal intensities of the four nucleotides is A,C,G,T (1,0,0,1). After the two rounds of sequencing, the type of incorporated nucleotide monomer can be distinguished based on the combination of the relative signal intensities obtained from these two rounds of sequencing. For example, the combination corresponding to A is (1,1), the combination corresponding to C is (1,0), the combination corresponding to G is (0,0), and the combination corresponding to T is (0,1). The combination of the signals of these four bases in the two rounds is unique. Therefore, the extended base type can be distinguished based on this, and the sequence of the polynucleotide template to be tested can be deduced according to the base complementary pairing principle. This sequencing method does not rely on specific detection tags (e.g., light-switching tags are specific detection tags), but achieves monochromatic sequencing through different signal combinations between different nucleotides in multiple sequencing rounds, thus broadening its applicability. Furthermore, this implementation only requires two sequencing rounds to obtain the sequence information of the nucleic acid to be tested, significantly improving sequencing speed and efficiency compared to the previous implementation.
[0080] In a preferred embodiment, an additional sequencing round is performed based on the aforementioned two sequencing rounds. This additional round uses a different combination of two nucleotide substrates containing the detection tag than those used in the previous two sequencing rounds, to obtain at least one additional sequence. This additional sequence is then compared with the initial sequence to reduce or eliminate one or more errors contained in the initial sequence. For example, in the aforementioned embodiment, the combinations of two nucleotide substrates containing the detection tag in the first two sequencing rounds were A,C and A,T, respectively. The combination of two nucleotide substrates containing the detection tag used in this sequencing round could be A,G or C,T. Two sequencing rounds already yield high-accuracy sequencing data; performing an additional round on top of this provides redundant information. Error correction codes based on information theory can correct errors from the two sequencing rounds, further improving sequencing accuracy.
[0081] In specific embodiments, the "nucleotide monomer containing the nucleotide of the detection tag" mentioned in this invention can mean that some of the nucleotides contain the detection tag and others do not. In this case, the relative signal intensity of the nucleotide monomer is greater than 0 and less than 1. Alternatively, all nucleotides can contain the detection tag, in which case the relative signal intensity of the nucleotide monomer is 1. Preferably, the relative signal intensity of the nucleotide monomer containing the detection tag is 0.5-1. This is easily understood because the sequencing method of this invention relies on the signal combination of multiple sequencing reactions. For each sequencing reaction, it must provide a signal combination with the highest possible resolution. Taking the combination of two nucleotides containing the detection tag, A and C, as an example, if the relative signal intensity of A and C is greater than 0 but close to 0, the distinction between A and C and G and T is low. Especially as the sequencing reaction progresses, A and C and G and T become increasingly difficult to distinguish. Therefore, the distinction between A and C and G and T must be as high as possible. For example, the relative signal intensity of A and C is 0.5-1, preferably 0.8-1, and more preferably 0.95-1.
[0082] In a more preferred embodiment, the relative signal intensities of the two nucleotide monomers containing the detection tag are both close to 1, ideally both being 1, meaning that both nucleotide monomers containing the detection tag contain the detection tag. In this case, the signal differentiation between the nucleotide monomers containing the detection tag and those without is maximized, ensuring a sufficiently high signal-to-noise ratio and allowing the monochrome sequencing method to maintain an acceptable or even low base recognition error rate in long-read applications.
[0083] In a preferred embodiment of the present invention, the aforementioned nucleotide containing the detection tag has the structure of formula (I):
[0084]
[0085] Wherein, Y is O or S, B is a heterocyclic base, and n is an integer from 0 to 6; R is selected from azidomethyl, amino, allyl, substituted dithioalkyl, substituted methoxymethyl, o-nitrobenzyl, coumarin, phosphate nitrile ethyl ester, trimethylsilyl, tetrahydropyranyl, azido, alkyl hydroxyamino, thiophosphate, malonyl, benzyl, acetal, thiocarbamate, and vinyl; Fluorogenic Dye is selected from anthracene, phenoxazine, acridine, and coumarin with fluorescence switching properties, wherein the anthracene with fluorescence switching properties includes oxanthracene-fluorescein, carbamate-Beijing orange, silanthracene, germananthracene, phosphoroxanthracene, and thioanthracene. Monochrome sequencing reactions using the 5′-terminal phosphate-labeled reversible termination nucleotide disclosed in this embodiment benefit from the following: because the fluorescent group is attached to the terminal phosphate and can be released through enzymatic cleavage, there is no residual group, avoiding molecular scarring. The synthesized polynucleotide chain can maintain its native conformation, which is beneficial for increasing sequencing read length. Furthermore, the fluorescent group being attached to the terminal phosphate makes substrate synthesis cheaper.
[0086] For the monochrome sequencing reaction disclosed in this application, nucleotides exhibiting a "dark state" (i.e., nucleotides without fluorescent tags) are necessary. However, for the aforementioned sequencing using reversible termination nucleotides labeled with 5'-terminal fluorescent phosphate, due to the need for fluorescent activation reagents such as alkaline phosphatase, non-fluorescently labeled nucleotide monomers (e.g., cold nucleotides commonly used in the art, with the structure shown in formula (III), where the terminal phosphate is in an open state and Cap1 is a reversible termination group) are hydrolyzed by alkaline phosphatase and cannot be incorporated into the new synthetic chain by polymerase. For nucleotides containing detection tags (e.g., A, C), this results in all nucleotide molecules incorporated into the new synthetic chain emitting fluorescence, with different nucleotide types having the same relative signal intensity (i.e., 1). For nucleotides without detection tags (e.g., G, T), they are directly hydrolyzed by alkaline phosphatase and cannot be incorporated into the new synthetic chain at all. This results in only some types of nucleotides (A, C) containing detection tags being incorporated, while the remaining nucleotides (G, T) cannot be incorporated, thus making monochrome sequencing impossible.
[0087]
[0088] To overcome this difficulty, the inventors of this application designed a uniquely structured, tagless nucleotide (referred to as dark nucleotide), enabling monochrome sequencing based on a reversible termination nucleotide with a 5'-terminal fluorescent phosphate label. In a specific embodiment of this invention, the structure of the aforementioned dark nucleotide is as shown in formula (II):
[0089]
[0090] Where Y is O or S; X is O or a group that cannot be degraded by phosphatase, where GD is OH or any non-fluorescent organic group; n is an integer from 0 to 6; B is a heterocyclic base; and Cap1 is a reversible termination group.
[0091] Preferably, X is O, and GD is selected from optionally substituted C1-10 alkyl, C1-10 alkyloxy, amino, mono- or disubstituted amino, C6-14 aryl, C6-14 aryloxy, C6-14 arylamino, or C3-12 cycloalkyl, C3-12 cycloalkyloxy, C3-12 cycloalkylamino; C2-13 heteroaryl or C2-13 heterocyclic or combinations thereof; preferably, GD is selected from C1-4 alkyl, C1-4 alkoxy, C1-4 alkylamino or di(C1-4 alkyl)amino, or amino, or phenyl, phenoxy, aniline or combinations thereof;
[0092] Preferably, X is CH2, CF2, or NH, and GD is selected from OH, C1-10 alkyl, C1-10 alkyloxy, mono- or disubstituted amino, C6-14 aryl, C6-14 aryloxy, C6-14 arylamino, or C3-12 cycloalkyl, C3-12 cycloalkyloxy, C3-12 cycloalkylamino; C2-13 heteroaryl or C2-13 heterocyclic or combinations thereof; preferably, GD is OH.
[0093] Furthermore, the sequencing reaction of the template polynucleotide can be performed on an array, such as a chip. The array may include multiple reaction volumes, for example, created by multiple reaction chambers disposed on the array. The template polynucleotide sequence or fragment thereof may be immobilized in the reaction volume, such as by adsorption or specific binding to a trap molecule on a solid support in each reaction volume. After the reaction solution is provided in the reaction mixture and delivered to each reaction volume, each reaction volume may be sealed and / or separated from other reaction volumes on the array. Signals such as fluorescence information can then be detected and / or recorded by each reaction volume. The aforementioned sealing refers to the need to confine the fluorescent group of the 5' phosphate-labeled reversible termination nucleotide in the sequencing reaction to the microreaction chamber to prevent or reduce fluorophore diffusion and solution evaporation, etc., since the fluorescent group can be released through enzymatic cleavage during the sequencing reaction. For example, the microreactor can be sealed with a fluid that is immiscible with water. Such fluids can be oils or gases immiscible with water. Examples of such oils include mineral oils, silicone oils, fluorinated oils (e.g., perfluorocarbons and HFE-7500, 2-trifluoromethyl-3-ethoxydodecylfluorohexane, or Fluorinert), or hydrocarbon oils (e.g., isoparaffins); examples of such gases include air, nitrogen, oxygen, saturated water vapor, or rare gases. In one specific embodiment, the microreactor can be sealed with fluorinated oil. First, the desired sequencing aqueous phase liquid is introduced into the flow cell, and then fluorinated oil is introduced into the flow cell to seal the sequencing aqueous phase liquid within the microreactor and cover it, thereby preventing the diffusion or evaporation of components within the microreactor.
[0094] In a specific embodiment of the present invention, before performing a new round of sequencing reaction, the double strands from the previous sequencing are unwound into single strands, the extended strands of the primers from the previous round are washed away, and then hybridized with the new sequencing primers for the next round.
[0095] In a specific embodiment of the present invention, each sequencing round contains at least 50 cycles, preferably at least 75, or at least 100, or at least 150, or at least 200, or at least 250, or at least 300, or at least 350, or at least 400, or at least 450, or at least 500, or at least 600, or at least 700, or at least 800, or at least 900, or at least 1000 cycles.
[0096] Example 1
[0097] This embodiment describes the synthesis process of a 5'-terminal labeled 3'-azidomethyl-2'-deoxynucleoside derivative that participates as a polymerase substrate in the extension reaction during sequencing. It should be understood that this embodiment only illustrates the general synthetic pathway for the relevant derivatives and does not represent a limitation to the methods shown.
[0098] Part 1: Synthesis of 5'-terminal phosphate-labeled 3'-azidomethyl-2'-deoxynucleoside tetraphosphate. The synthesis of the labeled monophosphate involves first phosphorylating the hydroxyl groups of the molecule to be labeled, as described below:
[0099]
[0100]
[0101] In the formula,
[0102] R4 is: H, Me, MeO, COOH, SO3H, F, Cl;
[0103] R5 is: H, Me, MeO, COOH, SO3H, F, Cl;
[0104] R6 is: Me, Et, MeO, EtO.
[0105] Add the labeled molecule (Ar-OH, 5 mmole) to 200 mL of anhydrous acetonitrile. Place the solid suspension under argon protection in an ice-water bath for 5 min. Add phosphorus oxychloride (5 eq) to the solution with rapid stirring and stir rapidly for 5 min. Continue to add tributylamine solution (10 eq) to the solution and keep the reaction under ice-water bath with stirring for 4 hours. TLC was used to monitor the disappearance of the reactants. After the reactants had largely disappeared, 100 mL of 100 mM TEAA buffer solution (pH = 8.0) was added to the solution. The mixture was stirred in an ice-water bath for 1 hour. The resulting mixture was concentrated to approximately 1 / 4 of its total volume using a rotary evaporator. The solution was then filtered through a 0.22 μm aqueous filter membrane and purified by medium-pressure preparative chromatography and a C18 column (0-50% B gradient, where A = 50 mM TEAA pH = 6.9, B = acetonitrile). The target fraction was collected and concentrated to near dryness using a rotary evaporator. A small amount of purified water was added, and the pH of the solution was adjusted to slightly alkaline (pH = 8) using tetrabutylammonium hydroxide (40% w / v) solution. After freeze-drying, the corresponding ammonium monophosphate solid product was obtained. Some representative molecular information is as follows:
[0106]
[0107] 2) Synthesis of 3'-azidomethyl-2'-deoxynucleoside triphosphate ammonium salt: Commercially available 3'-azidomethyl-2'-deoxynucleoside monophosphate was activated and then reacted with pyrophosphate to obtain the corresponding 3'-azidomethyl-2'-deoxynucleoside triphosphate.
[0108]
[0109] Preparation of 3'-azidomethyl-2'-deoxynucleoside monophosphate ammonium salt: Commercially available 3'-azidomethyl-2'-deoxynucleoside monophosphate (5 mmol / L) (if in sodium phosphate form, the sodium monophosphate needs to be ion-exchanged to phosphate form before this step) was mixed with 200 mL of purified water. The pH of the solution was adjusted to slightly alkaline (pH = 8) with tetrabutylammonium hydroxide (40% w / v) solution, at which point the solution changed from turbid to clear. The adjusted solution was then transferred to a 1 L round-bottom flask and freeze-dried to obtain a white solid of 3'-azidomethyl-2'-deoxynucleoside monophosphate tetrabutylammonium salt. The ammonium salt described in this patent is not limited to tetrabutylammonium salt; tributylammonium salt and triethylammonium salt are also suitable options.
[0110] Preparation of 3'-azidomethyl-2'-deoxynucleoside triphosphate: 5 mmol of solid 3'-azidomethyl-2'-deoxynucleoside monophosphate ammonium salt was dissolved in 100 mL of dry acetonitrile. The solvent was removed by rotary evaporation and the solution was then dried under vacuum. This process was repeated twice with the same volume of acetonitrile added and rotary evaporated twice, followed by drying under vacuum for 2 h. 250 mL of dry anhydrous acetonitrile was added to the flask, and after cooling in an ice-water bath, 2 eq of benzylsulfonic acid imidazole salt activator and 3 eq of diisopropylethylamine solution were added. The mixture was stirred in an ice-water bath for 30 min. Ammonium pyrophosphate salt (5 eq) and anhydrous dry magnesium chloride (5 eq) were added to the activated nucleoside monophosphate solution. The mixture was allowed to rise freely to room temperature with stirring for 24 h. The reaction solution was evaporated using a rotary evaporator to remove most of the acetonitrile. 100 mL of 100 mM TEAA buffer (pH = 8.0) was added to the remaining solution, and the mixture was stirred in an ice-water bath for 1 hour. The resulting mixture was extracted twice with dichloromethane using a separatory funnel (100 mL * 2). The organic phase was removed by separation, and the resulting aqueous phase was filtered through a 0.22 μm aqueous filter membrane. Purification was performed by preparative chromatography using a C18 column (0-30% B for 20 min, where A = 50 mM TEAA pH = 8.0, B = acetonitrile). The target fraction was collected and concentrated to near dryness using a rotary evaporator. The fraction was then evaporated three times (50 mL acetonitrile / evaporation) using dry acetonitrile to effectively remove residual water. The resulting triphosphate product was dissolved in 200 mL of dry acetonitrile, and its concentration was determined by HPLC before being used directly in the next step. (Yield 45%)
[0111] Add CDI activator (5 eq) and diisopropylethylamine (3 eq) to the dissolved 3'-azidomethyl-2'-deoxynucleoside triphosphate acetonitrile solution, and stir at room temperature for 3 h to generate an activated trimetaphosphate nucleoside derivative intermediate. Add methanol (5 eq) to this reaction solution using a microsyringe, stir for 30 min, and then cool in an ice-water bath. Add the previously synthesized labeled monophosphate solid (3 eq) and anhydrous magnesium bromide (5 eq, added at low temperature, exothermic). Continue to cool to room temperature for 40 h. Add 100 mL of 100 mM TEAA buffer solution (pH = 8.0) to this reaction solution, and stir in an ice-water bath for 1 h. Extract the resulting mixed solution twice with dichloromethane using a separatory funnel (100 mL * 2). The organic phase was removed by separation, and the obtained aqueous phase was filtered through a 0.22 μm aqueous phase filter membrane. It was then purified by high-pressure preparative chromatography and C18 column chromatography (0-20% B for 20 min, 20-70% B for 15 min, where A is 50 mM TEAA solution, pH=8.0, and B is acetonitrile). The target fraction with a purity ≥99% was analyzed and collected (the impure part was collected, concentrated, and purified again). The fractions that met the purity requirements were collected, concentrated by rotary evaporator, redissolved in 100 mL of purified water, and freeze-dried to obtain the target product.
[0112] Part 2: Synthesis of 5'-terminal phosphate-labeled 3'-azidomethyl-2'-deoxynucleoside pentaphosphate
[0113] 1) Preparation of ammonium trimetaphosphate: 20g of commercial sodium trimetaphosphate was dissolved in water and exchanged with Dowex 50W X8(H) ion exchange resin to obtain a trimetaphosphate solution. The pH of the solution was adjusted to slightly alkaline (pH=8.0) with tetrabutylammonium hydroxide (40% W / V) solution. After adjustment, the solution was transferred to a 1L round bottom flask and freeze-dried to obtain a white solid of tetrabutylammonium trimetaphosphate.
[0114] 2) Preparation of the labeled ammonium monophosphate: The labeled monophosphate obtained according to the method described in Part 1 of Example 1 was added to 200 mL of anhydrous acetonitrile. The solid suspension was placed under argon protection and kept in an ice-water bath for 5 min. Phosphorus oxychloride (5 eq) was added to the solution under rapid stirring and stirred rapidly for 5 min. Tributylamine solution (10 eq) was added to the solution and the reaction was carried out under ice-water bath stirring for 4 hours. TLC was used to monitor the disappearance of the reaction raw materials. After the raw materials had basically disappeared, 100 mL of 100 mM TEAA buffer solution (pH = 8.0) was added to the solution. The mixture was stirred and reacted for 1 hour under an ice-water bath. The resulting mixed solution was concentrated to about 1 / 4 of the total volume using a rotary evaporator. The solution was filtered through a 0.22 μm aqueous filter membrane and purified by medium-pressure preparative chromatography and a C18 column (0-50% B gradient, where A = 50 mM TEAA pH = 6.9, B = acetonitrile). The target fraction was collected and concentrated to near dryness using a rotary evaporator. A small amount of purified water was added, and the pH of the solution was adjusted to slightly alkaline (pH = 8.0) using tetrabutylammonium hydroxide (40% W / V) solution. After freeze-drying, the corresponding ammonium monophosphate solid product was obtained.
[0115] 3) Synthesis of 5'-terminal phosphate-labeled 3'-azidomethyl-2'-deoxynucleoside pentaphosphate:
[0116] 2 mmol of solid ammonium trimetaphosphate was placed in a 500 mL round-bottom flask, and 50 mL of anhydrous dry acetonitrile was added. The solvent was removed by rotary evaporation and the mixture was then dried under vacuum. The same volume of acetonitrile was added and rotary evaporated twice, followed by drying under vacuum for 2 h. 250 mL of dry anhydrous acetonitrile was added to the flask, and after cooling in an ice-water bath, 1.5 eq of benzylsulfonic acid imidazole salt activator and 3 eq of diisopropylethylamine solution were added. The mixture was stirred in an ice-water bath for 30 min. 1 eq of dried labeled ammonium monophosphate was added to the activated trimetaphosphate solution, and the mixture was kept in an ice-water bath for 3 h. 1 eq of dried 3'-azidomethyl-2'-deoxynucleoside ammonium monophosphate and 5 eq of anhydrous dry magnesium chloride were then added to the solution. The mixture was allowed to rise freely to room temperature for 24 h with stirring. The reaction solution was removed by rotary evaporation to remove most of the acetonitrile. 100 mL of 100 mM TEAA buffer solution (pH = 8.0) was added to the remaining solution, and the mixture was stirred in an ice-water bath for 1 hour. The resulting mixture was extracted twice with dichloromethane using a separatory funnel (100 mL * 2). The organic phase was removed by separation, and the resulting aqueous phase was filtered through a 0.22 μm aqueous filter membrane. Purification was performed by preparative chromatography and a C18 column (0-30% B for 20 min, where A = 50 mM TEAA pH = 8.0, B = acetonitrile). The fractions with ≥99% purity were collected and analyzed (impure fractions were collected, concentrated, and purified again). Fractions meeting the purity requirements were concentrated by rotary evaporation, redissolved in 100 mL of purified water, and freeze-dried to obtain the target product.
[0117] Part 3: Synthesis of 5'-terminal phosphorylated fluorescently labeled 3'-azidomethyl-2'-deoxyribonucleophosphate
[0118] 1) Preparation of ammonium trimetaphosphate: 20g of commercial sodium trimetaphosphate was dissolved in water and exchanged with Dowex 50W X8(H) ion exchange resin to obtain a trimetaphosphate solution. The pH of the solution was adjusted to slightly alkaline (pH=8.0) with tetrabutylammonium hydroxide (40% W / V) solution. After adjustment, the solution was transferred to a 1L round bottom flask and freeze-dried to obtain a white solid of tetrabutylammonium trimetaphosphate.
[0119] 2) Preparation of 3'-azidomethyl-2'-deoxynucleoside triphosphate ammonium salt: The 3'-azidomethyl-2'-deoxynucleoside triphosphate ammonium salt obtained in accordance with the method described in Part 1 of Example 1 was prepared.
[0120] 3) Synthesis of 5'-terminal phosphate-labeled 3'-azidomethyl-2'-deoxynucleoside hexaphosphate: 2 mmol of solid ammonium trimetaphosphate was placed in a 500 mL round-bottom flask, and 50 mL of anhydrous dry acetonitrile was added. The solvent was removed by rotary evaporation and the mixture was then dried under vacuum. The same volume of acetonitrile was added and rotary evaporated twice, followed by drying under vacuum for 2 h. 250 mL of dry anhydrous acetonitrile was added to the flask, and after cooling in an ice-water bath, 1.5 eq of benzylsulfonic acid imidazole salt activator and 3 eq of diisopropylethylamine solution were added. The mixture was stirred in an ice-water bath for 30 min. 1 eq of dried ammonium trimetaphosphate was added to the activated trimetaphosphate solution, and the reaction was maintained in an ice-water bath for 3 h. Then, 1 eq of dried 3'-azidomethyl-2'-deoxynucleoside ammonium triphosphate and 5 eq of anhydrous dry magnesium chloride were added to the solution. The mixture was allowed to rise freely to room temperature for 24 h with stirring. The reaction solution was removed by rotary evaporation to remove most of the acetonitrile. 100 mL of 100 mM TEAA buffer solution (pH = 8.0) was added to the remaining solution, and the mixture was stirred in an ice-water bath for 1 hour. The resulting mixture was extracted twice with dichloromethane using a separatory funnel (100 mL * 2). The organic phase was removed by separation, and the resulting aqueous phase was filtered through a 0.22 μm aqueous filter membrane. Purification was performed by preparative high-pressure chromatography and a C18 column (0-30% B for 20 min, where A = 50 mM TEAA, pH = 8.0, B = acetonitrile). The fractions with ≥99% purity were collected and analyzed (impure fractions were collected, concentrated, and purified again). Fractions meeting the purity requirements were concentrated by rotary evaporation, redissolved in 100 mL of purified water, and freeze-dried to obtain the target product solid.
[0121] Part 4: Synthesis of 5'-nonphosphatase hydrolysed 3'-azidomethyl-2'-deoxynucleotide triphosphate
[0122] 1) Preparation of 3'-azidomethyl-2'-deoxynucleoside monophosphate ammonium salt:
[0123] Commercially available 3'-azidomethyl-2'-deoxynucleoside monophosphate (5 mmole) (if in Na salt form, the sodium monophosphate needs to be ion-exchanged to H form before this step) was mixed with 200 mL of purified water. The pH of the solution was adjusted to slightly alkaline (pH = 8.0) with tetrabutylammonium hydroxide (40% W / V) solution, at which point the solution changed from turbid to clear. The adjusted solution was then transferred to a 1 L round-bottom flask and freeze-dried to obtain a white solid of 3'-azidomethyl-2'-deoxynucleoside monophosphate tetrabutylammonium salt.
[0124] 2) Preparation of methylene diphosphate ammonium salt: (Difluoromethylene diphosphate or amino diphosphate are prepared in the same way) Mix commercial methylene diphosphate (2 mmol e) with 100 mL of purified water, and adjust the pH of the solution to be slightly alkaline (pH = 8) with tetrabutylammonium hydroxide (40% W / V) solution. At this time, the solution changes from turbid to clear. After adjustment, the solution is transferred to a 500 mL round bottom flask and freeze-dried to obtain solid tetrabutylammonium methylene diphosphate.
[0125] 3) Preparation of 3'-azidomethyl-2'-deoxynucleoside-substituted triphosphate derivatives
[0126] 3'-Azidomethyl-2'-deoxynucleoside monophosphate (2 mmol e) solid was dissolved in 100 mL of dry acetonitrile. The solvent was removed by rotary evaporation and the solution was then dried under vacuum. The same volume of acetonitrile was added twice, and the solution was rotary evaporated twice, followed by drying under vacuum for 2 h. 100 mL of dry anhydrous acetonitrile was added to this flask, and after cooling in an ice-water bath, 2 eq of benzylsulfonic acid imidazole salt activator and 3 eq of diisopropylethylamine solution were added. The mixture was stirred in an ice-water bath for 30 min. Methylene diphosphate ammonium salt (5 eq) and anhydrous dry magnesium chloride (5 eq) were added to this activated nucleoside monophosphate solution. The mixture was allowed to rise freely to room temperature for 24 h with stirring. Most of the acetonitrile was removed by rotary evaporation. 100 mL of 100 mM TEAA buffer solution (pH = 8.0) was added to the remaining solution, and the mixture was stirred in an ice-water bath for 1 h. The resulting mixture was extracted once with 100 mL of dichloromethane using a separatory funnel. The organic phase was removed by separation, and the resulting aqueous phase was filtered through a 0.22 μm aqueous filter membrane. Purification was then performed by preparative high-pressure chromatography and a C18 column (0-20% B for 30 min, where A = 50 mM TEAA, pH = 8.0, B = acetonitrile). Fractions with a purity ≥98% were collected and analyzed (impure fractions were collected, concentrated, and purified again). Fractions meeting the purity requirements were concentrated by rotary evaporation, redissolved in 50 mL of purified water, and freeze-dried to obtain the target product with a yield of 15%. The synthesis of nucleoside triphosphate derivatives corresponding to difluoromethylene diphosphate or amino diphosphate is the same.
[0127] The representative compound structures and mass spectrometry detection data are as follows:
[0128]
[0129] Part 5: Synthesis of 5'-nonphosphatase hydrolysed 3'-azidomethyl-2'-deoxyribonucleoside tetraphosphate
[0130] 1) Preparation of ammonium trimetaphosphate: Dissolve 20g of commercial sodium trimetaphosphate in water, and then... (The sentence is incomplete and requires more context to translate accurately.)
[0131] (H) Ion exchange resin was used to exchange the solution with trimetaphosphate. The pH of the solution was adjusted to slightly alkaline (pH=8.0) with tetrabutylammonium hydroxide (40% W / V) solution. After adjustment, the solution was transferred to a 1L round bottom flask and freeze-dried to obtain a white solid of tetrabutylammonium trimetaphosphate.
[0132] 2) Synthesis of 3'-azidomethyl-2'-deoxynucleotide tetraphosphate labeled with 5'-terminal phosphate:
[0133] 2 mmol of solid ammonium trimetaphosphate was placed in a 500 mL round-bottom flask, and 50 mL of anhydrous dry acetonitrile was added. The solvent was removed using a rotary evaporator, and the mixture was then dried under vacuum. The same volume of acetonitrile was added and the mixture was rotary evaporated twice, followed by drying under vacuum for 2 h. 250 mL of dry anhydrous acetonitrile was added to the flask, and after cooling in an ice-water bath, 2 eq of benzylsulfonic acid imidazole salt activator and 6 eq of diisopropylethylamine solution were added. The mixture was stirred in an ice-water bath for 30 min. 2 mmol of solid ammonium 3'-azidomethyl-2'-deoxynucleoside monophosphate and 5 eq of anhydrous dry magnesium chloride were added to the activated trimetaphosphate solution. The mixture was kept in an ice-water bath and reacted for 2 h. 5 eq of dimethylamine solution was then added to the solution. The mixture was allowed to heat freely to room temperature for 24 h with stirring. The reaction solution was removed by rotary evaporation to remove most of the acetonitrile. 100 mL of 100 mM TEAA buffer solution (pH = 8.0) was added to the remaining solution, and the mixture was stirred in an ice-water bath for 1 hour. The resulting mixture was extracted twice with dichloromethane using a separatory funnel (100 mL * 2). The organic phase was removed by separation, and the resulting aqueous phase was filtered through a 0.22 μm aqueous filter membrane. Purification was performed by preparative chromatography and a C18 column (0-30% B for 20 min, where A = 50 mM TEAA pH = 8.0, B = acetonitrile). The fractions ≥99% of the target product were collected and analyzed (impure fractions were collected, concentrated, and purified again). Fractions meeting the purity requirements were concentrated by rotary evaporation, redissolved in 100 mL of purified water, and freeze-dried to obtain the target product.
[0134]
[0135] Part VI: Synthesis of 5'-Nonphosphatase Hydrolysis of 3'-Azidemethyl-2'-Deoxynucleotide Pentaphosphate
[0136] Preparation of 3'-azidomethyl-2'-deoxynucleoside triphosphate ammonium salt: The 3'-azidomethyl-2'-deoxynucleoside triphosphate ammonium salt obtained in accordance with the method described in Part 1 of Example 1 was prepared.
[0137] 1 mmol of 3'-azidomethyl-2'-deoxynucleoside triphosphate ammonium salt was placed in a 100 mL round-bottom flask, and 30 mL of dry acetonitrile was added. The solvent was removed by rotary evaporation and the mixture was then dried under vacuum. This process was repeated twice with the same volume of acetonitrile added and rotary evaporated twice, followed by 2 hours of drying under vacuum. 50 mL of dry acetonitrile was added to the round-bottom flask, a magnetic stir bar was added, and the mixture was thoroughly dissolved. Under stirring, CDI activator (5 eq) and diisopropylethylamine (3 eq) were added to the solution. The mixture was stirred at room temperature for 3 hours to generate an activated triphosphate nucleoside derivative intermediate. Methanol (5 eq) was added to this reaction solution using a microsyringe, and the mixture was stirred for 30 minutes. The solution was then cooled in an ice-water bath. The diphosphate derivative ammonium salt (5 eq) and anhydrous dry magnesium chloride (5 eq) were added to this activated nucleoside monophosphate solution. The mixture was allowed to rise freely to room temperature for 24 hours with stirring. The reaction solution was removed by rotary evaporation to remove most of the acetonitrile. 100 mL of 100 mM TEAA buffer solution (pH = 8.0) was added to the remaining solution, and the mixture was stirred in an ice-water bath for 1 hour. The resulting mixture was extracted once with 100 mL of dichloromethane using a separatory funnel. The organic phase was removed by separation, and the resulting aqueous phase was filtered through a 0.22 μm aqueous filter membrane. The aqueous phase was then purified by preparative high-pressure chromatography and a C18 column (0-20% B for 30 min, where A = 50 mM TEAA pH = 8.0, B = acetonitrile). Fractions with a purity ≥98% were collected and analyzed (impure fractions were collected, concentrated, and purified again). Fractions meeting the purity requirements were concentrated by rotary evaporation, redissolved in 50 mL of purified water, and freeze-dried to obtain the target product.
[0138] Example 2
[0139] Prepare the following solution:
[0140] First sequencing reaction solution: 20mM Tris-HCl pH=8.3; 10mM (NH4)2SO4; 50mM KCl; 1mM MgCl2; 0.1% Tween; 9°N polymerase; AF532 fluorescently labeled dATP and AF532 fluorescently labeled dCTP (the structural formulas of the two are as follows, The simplified base representation is B), and the unlabeled dGTP and dTTP (their structural formulas are shown below) The simplified representation of the base is B).
[0141] Second sequencing reaction solution: 20mM Tris-HCl pH=8.3; 10mM (NH4)2SO4; 50mM KCl; 1mM MgCl2; 0.1% Tween; 9°N polymerase; AF532 fluorescently labeled dATP and AF532 fluorescently labeled dGTP (the structural formulas of the two are as follows, The simplified base representation is B), and the unlabeled dCTP and dTTP (their structural formulas are shown below) The simplified representation of the base is B).
[0142] Deblocking reaction solution: from MGISEQ-2000RS high-throughput sequencing kit (PE150), catalog number: 1000012537.
[0143] Cleaning reagent: from MGISEQ-2000RS high-throughput sequencing kit (PE150), catalog number: 1000012537.
[0144] After constructing a library from the genomic DNA of lambda phage, immobilize it onto a sequencing chip, place the sequencing chip in a sequencer, and construct nucleic acid clusters through amplification; proceed as follows:
[0145] 1) Hybridize the sequencing primers onto the prepared DNA array;
[0146] Perform the first round of sequencing:
[0147] 2) Introduce the first sequencing reaction solution into the chip;
[0148] 3) Heat the chip to 55°C to carry out nucleotide polymerization, so that nucleotides are incorporated into the 3′ end of the growing nucleic acid chain and the growth is terminated; extend for 60 seconds;
[0149] 4) Clean to remove the reaction solution;
[0150] 5) Take a picture, detect the fluorescence signal using a 532nm light source as the excitation wavelength, take a picture, and store the image;
[0151] 6) Add 200 μL of blocking reaction solution to the chip, set the chip temperature to 55 °C, and react for 120 s to remove the blocking group azidomethyl;
[0152] 7) Cleaning: Introduce 400 μL of cleaning reagent into the chip and clean it 3 times;
[0153] 8) Repeat steps 2)-7) to perform the next cycle sequencing, for a total of 150 cycles.
[0154] 9) Introduce formamide to unwind the double-stranded nucleic acid molecules obtained from the previous sequencing reaction, and wash to remove the supernatant.
[0155] After one round of primer extension, the sequencing primers are re-hybridized for a second round of sequencing.
[0156] 10) Introduce the second sequencing reaction solution into the chip;
[0157] 11) Heat the chip to 55°C and perform nucleotide polymerization to incorporate nucleotides into the 3′ end of the growing nucleic acid chain and terminate further growth; extend for 60 seconds;
[0158] 12) Clean and remove the reaction solution;
[0159] 13) Take a picture, detect the fluorescence signal using a 532nm light source as the excitation wavelength, take a picture, and store the image;
[0160] 14) Add 200 μL of blocking reaction solution to the chip, set the chip temperature to 55 °C, and react for 120 s to remove the blocking group azidomethyl;
[0161] 15) Cleaning: Introduce 400 μL of cleaning reagent into the chip and clean 3 times;
[0162] 16) Repeat steps 10)-15) to perform the next cycle sequencing, for a total of 150 cycles.
[0163] 17) Data processing and analysis of sequencing results. In the first round of sequencing, each sequencing cycle yields degenerate base information based on the presence or absence of a fluorescence signal. If fluorescence is present, the base is M (i.e., A or C); if no fluorescence is present, the base is K (i.e., G or T). In the second round of sequencing, each sequencing cycle yields degenerate base information based on the presence or absence of a fluorescence signal. If fluorescence is present, the base is R (i.e., A or G); if no fluorescence is present, the base is Y (i.e., C or T).
[0164] T); By combining the results of two rounds of sequencing, the specific base can be determined. For example, if it is both M and R, then it is...
[0165] A (It should be noted that when determining the specific base type, the specific base can be determined directly by taking the intersection of the results of the two rounds of sequencing, without first obtaining degenerate base information); using the above automated sequencing analysis program, the sequencing results are as follows Figure 2 As shown, two rounds of sequencing reactions, each 150bp, were performed, and the average error rate was calculated to be approximately 0.36%, with 86% of the samples having a Q30 or higher.
[0166] Comparative Example
[0167] Library construction was performed on the genomic DNA of lambda phage, and sequencing was carried out using the MGISEQ-2000 sequencer and the MGISEQ-2000RS high-throughput sequencing kit (PE150) manufactured by BGI Genomics. The sequencer is a dual-channel sequencer, and imaging was performed at excitation wavelengths of 532nm and 660nm, respectively. The sequencing mode was selected as PE150. The sequencing results are as follows: the average error rate was 0.80%, and the proportion of Q30 and above was 88.46%. It can be seen that the single-color sequencing method of Example 2 of this invention can obtain sequencing results that are close to those of dual-color sequencing, and the single-color sequencing method of Example 2 has a lower sequencing error rate.
[0168] Example 3
[0169] Prepare the following solution:
[0170] The third sequencing reaction solution consisted of: 20 mM Tris-HCl pH 8.3; 10 mM (NH4)2SO4; 50 mM KCl; 1 mM MgCl2; 0.1% Tween; 9°N polymerase; AF532 fluorescently labeled dATP and AF532 fluorescently labeled dTTP (the structural formulas of which are shown below). The simplified base representation is B), and the unlabeled dCTP and dGTP (their structural formulas are shown below) The simplified representation of the base is B).
[0171] The remaining solutions are the same as in Example 2.
[0172] Based on Example 2, an additional sequencing reaction is added, wherein steps 1)-16) are the same as in Example 2, and will not be repeated here:
[0173] 17) Introduce formamide to unwind the double-stranded nucleic acid molecules obtained from the previous sequencing reaction, wash away the primer extension strands from the previous round, and then re-hybridize the sequencing primers for the third round of sequencing.
[0174] 18) Introduce the third sequencing reaction solution into the chip;
[0175] 19) Heat the chip to 55°C and perform nucleotide polymerization to incorporate nucleotides into the 3′ end of the growing nucleic acid chain and terminate further growth; extend for 60 seconds;
[0176] 20) Clean and remove the reaction solution;
[0177] 21) Take a picture, detect the fluorescence signal using a 532nm light source as the excitation wavelength, take a picture, and store the image;
[0178] 22) Add 200 μL of blocking reaction solution to the chip, set the chip temperature to 55 °C, and react for 120 s to remove the blocking group azidomethyl;
[0179] 23) Cleaning: Introduce 400 μL of cleaning reagent into the chip and clean 3 times;
[0180] 24) Repeat steps 18)-23) to perform the next cycle sequencing, for a total of 150 cycles.
[0181] 25) Data processing and sequencing result analysis. Each of the three sequencing reactions yields corresponding degenerate base information, resulting in a total of three sets of degenerate base information. Because the combined determination of bases across the three rounds has a higher tolerance for errors in single-round signals, the final error rate is lower. Specific results are as follows: Figure 3 As shown, the blue line represents the sequencing error rate per bp of the target nucleic acid after only two sequencing rounds, and the orange line represents the sequencing error rate per bp of the target nucleic acid after three sequencing rounds.
[0182] The corresponding sequencing error rates show that performing three rounds of sequencing significantly reduces the error rate. According to specific statistical results, the average error rate for three rounds of sequencing is 0.1%, with 95.0% achieving Q30 or higher, significantly better than the results of only two rounds of sequencing (average error rate approximately 0.36%).
[0183] (86% of those with Q30 or above).
[0184] Example 4
[0185] Prepare the following solution:
[0186] First sequencing reaction solution: 20mM Tris-HCl pH=8.3; 10mM (NH4)2SO4; 50mM KCl; 1mM MgCl2; 0.1% Tween; 9°N polymerase; AF532 fluorescently labeled dATP and unlabeled dATP, the ratio of these two nucleotides was adjusted so that the relative signal intensity of dATP after sequencing was 0.5; AF532 fluorescently labeled dCTP and unlabeled dCTP, the ratio of these two nucleotides was adjusted so that the relative signal intensity of dCTP after sequencing was 0.5; unlabeled dGTP and dTTP, the structure is the same as in Example 2.
[0187] Second sequencing reaction solution: 20mM Tris-HCl pH=8.3; 10mM (NH4)2SO4; 50mM KCl; 1mM MgCl2; 0.1% Tween; 9°N polymerase; AF532 fluorescently labeled dATP and unlabeled dATP, the ratio of these two nucleotides was adjusted so that the relative signal intensity of dATP after sequencing was 0.5; AF532 fluorescently labeled dGTP and unlabeled dGTP, the ratio of these two nucleotides was adjusted so that the relative signal intensity of dGTP after sequencing was 0.5; unlabeled dCTP and dTTP, the structure is the same as in Example 2.
[0188] The remaining solutions and sequencing steps are the same as in Example 2.
[0189] The data from this set of experiments were processed and compared with the sequencing data from Example 2. The sequencing results are as follows: Figure 4 As shown, the blue line represents the sequencing error rate per bp in Example 2, and the orange line represents the sequencing error rate per bp in Example 4. It can be seen that in the early stage of sequencing (about the first 90 bp), the error rates of the two are basically the same. However, as the sequencing reaction progresses to the later stage, the error rate of the group with relative signals of (1, 1, 0, 0) (i.e., Example 2) is significantly lower than the error rate of the group with relative signals of (0.5, 0.5, 0, 0) (i.e., Example 4). This indicates that under the same experimental conditions, the greater the difference in relative signals between the two nucleotides that emit fluorescent signals and the two nucleotides that do not emit fluorescent signals (the difference between 0 and 1), the higher the signal discrimination and the lower the sequencing error rate.
[0190] Example 5
[0191] Prepare the following solution:
[0192] First sequencing reaction solution: 20mM Tris-HCl pH=8.3; 10mM (NH4)2SO4; 50mM KCl; 1mM MnCl2; 100mM glycine; 0.1% Tween; 9°N polymerase; CIP (alkaline phosphatase, bovine intestine); PO fluorescently labeled dA4Ps; PO fluorescently labeled dC4Ps (the structural formulas of both are as follows: The base is simplified to B), and the unlabeled dG4Ps and dT4Ps (their structural formulas are as follows) The simplified representation of the base is B).
[0193] Second sequencing reaction solution: 20mM Tris-HCl pH=8.3; 10mM (NH4)2SO4; 50mM KCl; 1mM MnCl2; 100mM glycine; 0.1% Tween; 9°N polymerase; CIP (alkaline phosphatase, bovine intestine); PO fluorescently labeled dA4Ps; PO fluorescently labeled dG4Ps (the structural formulas of both are as follows: The simplified base representation is B), and the unlabeled dC4Ps and dT4Ps (their structural formulas are as follows: The simplified representation of the bases is as follows: B) Deblocking reaction solution: 20mM tris(3-hydroxypropyl)phosphine (THPP), 0.5M NaCl, 50mM Tris-HCl, pH=8.6, 0.05% Tween-20;
[0194] Cleaning solution: 20 mM Tris-HCl pH=8.3; 10 mM (NH4)2SO4; 50 mM KCl; 0.1 mM EDTA and 0.1% Tween 20.
[0195] Oil sealant: Novec TM 7500 electronic fluorinated liquid (purchased from 3M) TM ).
[0196] After constructing a library from the lambda phage genomic DNA, it was immobilized onto a sequencing chip, which was then placed in a sequencer for amplification to construct nucleic acid molecular clusters; the following steps were followed:
[0197] 1) Hybridize the sequencing primers onto the prepared DNA array;
[0198] Perform the first round of sequencing:
[0199] 2) Control the chip temperature to 4℃ and introduce the first sequencing reaction solution into the chip;
[0200] 3) Introduce 100-200 μL of sealing fluid into the chip;
[0201] 4) Heat the chip to 55°C to perform a nucleotide polymerization reaction, allowing nucleotides to be incorporated into the growing nucleic acid chains.
[0202] 3′ end and stop further growth; elongation 60s;
[0203] 5) Take a picture, detect the fluorescence signal using a 540nm light source as the excitation wavelength, take a picture, and store the image;
[0204] 6) First, clean off the sealing fluid, then add 200μL of desealing reaction solution into the chip, and set the chip temperature to [temperature value missing].
[0205] The reaction was carried out at 55°C for 120 seconds to remove the blocking group azidomethyl.
[0206] 7) Cleaning: Pour 400uL of cleaning solution into the chip and clean it 3 times;
[0207] 8) Repeat steps 2)-7) to perform the next cycle sequencing, for a total of 150 cycles.
[0208] 9) Introduce formamide to unwind the double-stranded nucleic acid molecules obtained from the previous sequencing reaction, and wash away the residue from the previous reaction.
[0209] The primer extension strands are then re-hybridized with sequencing primers for a second round of sequencing.
[0210] 10) Control the chip temperature to 4°C and introduce the second sequencing reaction solution into the chip;
[0211] 11) Introduce 100-200 μL of sealing fluid into the chip;
[0212] 12) Heat the chip to 55°C to perform nucleotide polymerization, causing nucleotides to incorporate into the growing nucleic acid chains.
[0213] 3′ end and stop further growth; elongation 60s;
[0214] 13) Take a picture, detect the fluorescence signal using a 540nm light source as the excitation wavelength, take a picture, and store the image;
[0215] 14) First, clean off the sealing fluid, then add 200μL of desealing reaction solution into the chip, and set the chip temperature to [temperature value missing].
[0216] The reaction was carried out at 55°C for 120 seconds to remove the blocking group azidomethyl.
[0217] 15) Cleaning: Introduce 400uL of cleaning solution into the chip and clean it 3 times.
[0218] 16) Repeat steps 10)-15) for the next sequencing cycle, for a total of 150 sequencing cycles. 17) Process the data from this experiment and analyze the sequencing results as follows: Figure 5 As shown, the orange lines represent
[0219] Table 2 shows the sequencing error rate per bp in Example 2. The blue line represents the sequencing error rate per bp in this example.
[0220] The corresponding sequencing error rates show a significant difference between the two. Sequencing using the 5'-terminal phosphorylated fluorescently labeled substrate nucleotides in this embodiment can achieve a significantly lower error rate, meaning that an extremely high sequencing accuracy can be obtained.
[0221] Example 6
[0222] Since the error rate in Example 5 was still extremely low at 150bp, it is speculated that this sequencing method can be used to achieve long-read sequencing. Based on Example 5, a sequencing reaction of 300 cycles was performed (all other steps were the same and will not be repeated here). The sequencing results are as follows: Figure 6 As shown in the figure, the error rate is less than 0.7% when sequencing reaches 300bp, which meets the sequencing requirements. This result is even better than the sequencing results of only 150bp using conventional sequencing substrates (e.g., the substrate used in Example 2) (endpoint error rate close to 0.8%).
[0223] The detailed description and embodiments above are provided only to provide a clear understanding of the invention. They should not be construed as unnecessary limitations. The invention is not limited to the exact details shown and described, as variations that will be apparent to those skilled in the art will be included within the scope of the invention as defined by the claims.
Claims
1. A method for determining the sequence of a template polynucleotide, characterized in that, include: At least two rounds of sequencing are performed on the template polynucleotide immobilized at the flow cell reaction site. Each sequencing round comprises multiple sequencing cycles. Each sequencing cycle includes: introducing an incorporation mixture containing four types of nucleotide monomers, wherein at least one type and up to three types of nucleotide monomers contain nucleotides labeled with a detection tag, and the incorporation mixture also includes nucleotides without a detection tag; incorporating a single nucleotide from the incorporation mixture into a primer polynucleotide that binds to the template polynucleotide to generate an extended primer polynucleotide; performing a single imaging event and collecting the emission signal; wherein... At least one and at most three combinations of nucleotides containing the detection tag must be different in at least two rounds of sequencing; The type of incorporated nucleotide monomer is distinguished by the combination of emission signals obtained from at least two rounds of sequencing, thereby obtaining the initial sequence of the template polynucleotide.
2. The method according to claim 1, characterized in that, Two rounds of sequencing were performed on the template polynucleotides immobilized at the flow cell reaction site. Two types of nucleotide monomers contained nucleotides labeled with the detection tag. The combinations of the two types of nucleotides containing the detection tag were different in the two rounds of sequencing. One type of nucleotide was the same, and the other type of nucleotide was different.
3. The method according to claim 2, characterized in that, Also includes: An additional sequencing round is performed, in which the combination of two nucleotide substrates containing the detection tag used in this round of sequencing differs from the combination of two nucleotides containing the detection tag used in the previous two rounds of sequencing, to obtain at least one additional sequence, and the additional sequence is compared with the initial sequence to reduce or eliminate one or more errors contained in the initial sequence.
4. The method according to claim 2, wherein the relative signal intensity of the nucleotide monomer containing the detection tag is 0.5-1, preferably 0.8-1, and more preferably 0.95-1.
5. The method according to claim 4, characterized in that, The relative signal intensities of the two nucleotide monomers containing detection tags are both close to 1.
6. The method according to claim 1 or 2, characterized in that, The one containing the detection tag marking The structure of a nucleotide is shown in formula (Ⅰ): Wherein, Y is O or S, B is a heterocyclic base, and n is an integer from 0 to 6; R is selected from azidomethyl, amino, allyl, substituted dithioalkyl, substituted methoxymethyl, o-nitrobenzyl, coumarin, phosphate nitrile ethyl ester, trimethylsilyl, tetrahydropyranyl, azido, alkyl hydroxyamino, thiophosphate, malonyl, benzyl, acetal, thiocarbamate, and vinyl; Fluorogenic Dye is selected from anthracene, phenoxazine, acridine, and coumarin with fluorescence switching properties, wherein the anthracene with fluorescence switching properties includes one or more of oxanthracene-fluorescein, carbamate-Beijing orange, silanthracene, germananthracene, phosphoroxanthracene, and thioanthracene.
7. The method according to claim 6, characterized in that, The reaction sites comprise multiple reaction volumes created by multiple reaction chambers disposed on an array, in which the template polynucleotide is immobilized; after the incorporation mixture is delivered to each reaction volume, each reaction volume may be closed and / or separated from other reaction volumes on the array; The emission signal from the detection tag can then be detected and / or recorded for each reaction volume.
8. The method according to claim 6, characterized in that, The structure of the unlabeled nucleotide is shown in formula (II): Where Y is O or S; X is O or a group that cannot be degraded by phosphatase, where GD is OH or any non-fluorescent organic group; n is an integer from 0 to 6; B is a heterocyclic base; and Cap1 is a reversible termination group. Preferably, X is O, and GD is selected from optionally substituted C1-10 alkyl, C1-10 alkyloxy, amino, mono- or disubstituted amino, C6-14 aryl, C6-14 aryloxy, C6-14 arylamino, or C3-12 cycloalkyl, C3-12 cycloalkyloxy, C3-12 cycloalkylamino; C2-13 heteroaryl or C2-13 heterocyclic or combinations thereof; preferably, GD is selected from C1-4 alkyl, C1-4 alkoxy, C1-4 alkylamino or di(C1-4 alkyl)amino, or amino, or phenyl, phenoxy, aniline or combinations thereof; Preferably, X is CH2, CF2, or NH, and GD is selected from OH, C1-10 alkyl, C1-10 alkyloxy, mono- or disubstituted amino, C6-14 aryl, C6-14 aryloxy, C6-14 arylamino, or C3-12 cycloalkyl, C3-12 cycloalkyloxy, C3-12 cycloalkylamino; C2-13 heteroaryl or C2-13 heterocyclic or combinations thereof; preferably, GD is OH.
9. The method according to claim 2, characterized in that, The combination of the two types of nucleotide monomers containing the detection tag is selected from: 1) Use A and T / U deoxyribonucleotides in one round of sequencing, or use C and G deoxyribonucleotides; or 2) Use A and G deoxyribonucleic acid in one round of sequencing, or use C and T / U deoxyribonucleic acid; or 3) Use A and C deoxyribonucleic acid in one round of sequencing, or use G and T / U deoxyribonucleic acid.
10. The method according to any one of claims 1-9, further comprising, wherein the 3' end of the four types of nucleotide monomers contains a reversible termination group, and the reversible termination group at the 3' end of the incorporated nucleotide is removed before the next sequencing cycle, such that the nucleotide becomes an extendable state.
Citation Information
Patent Citations
Poly phosphoric acid end fluorescent labeled nucleotide and application thereof
CN104910229A
Reversibly blocked nucleoside analogues and their uses
CN109790196B
Methods and Compositions for Nucleic Acid Sequencing Using Photo-Switchable Tags
CN112639128B
Modified nucleosides or nucleotides
WO2014139596A1