Modified nucleotide for improving strand bias, and preparation method therefor and use thereof

By introducing modified nucleotides with specific fluorescent dye-labeled modified nucleotides on the nucleotides, the problem of inconsistency in variant detection caused by chain bias in high-throughput sequencing is solved, and the sequencing accuracy is improved, especially the chain bias of T to G and A to C, which is suitable for the detection of single nucleotide polymorphisms.

WO2025151976A1PCT designated stage expired Publication Date: 2025-07-24MGI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/072251
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

In existing high-throughput sequencing technologies, the chain bias problem leads to inconsistent variation detection of sense strands and antisense strands, making it difficult to accurately detect single nucleotide polymorphisms, and lacks effective solutions.

Method used

A structurally novel modified nucleotide was designed to improve strand bias problems and improve sequencing accuracy by introducing specific fluorescent dye markers on the nucleotides.

Benefits of technology

It effectively solves the chain bias problem between T to G and A to C, improves the accuracy of nucleic acid sequencing, and is suitable for the detection of single nucleotide polymorphisms in high-throughput sequencing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024072251_24072025_PF_FP_ABST
    Figure CN2024072251_24072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the field of sequencing. In particular, the present invention relates to a modified nucleotide as represented by formula (I-1), and a preparation method therefor and the use thereof for sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

A modified nucleotide for improving chain deviation, preparation method and application thereof Technical Field

[0001] The present invention relates to the field of sequencing, and in particular to a modified nucleotide, a method for preparing the same, and use thereof in sequencing. Background Art

[0002] Although the data generated by high-throughput sequencing (NGS) is richer and contains more information than other traditional methods, the subsequent analysis of NGS data has also encountered many new difficulties. The main challenge for sequencing data is how to accurately detect SNP variations, that is, DNA sequence polymorphisms caused by variations in single nucleotides. In the process of applying high-throughput sequencing technology to SNP detection, people found that when reads were spliced ​​and aligned to the genome sequence reference, sometimes the positive chain reads and the negative chain reads showed significantly different types of variations. One may show a homozygous mutation and the other a heterozygous mutation. This inconsistent phenomenon is called bias. After research, people found that bias is specifically divided into orientation bias and strand bias, both of which are caused by the limitations of existing sequencing technology.

[0003] Technicians can easily detect the presence of identity bias and try to eliminate it whenever possible. For example, during the PCR step of library preparation, the polymerase can incorporate incorrect bases during extension. Theoretically, the probability of such amplification errors occurring is equal for both positive and negative-sense genomic fragments, meaning that approximately 50% of all sequencing error-prone reads are from positive-sense reads and approximately 50% are from negative-sense reads. The solution is also relatively simple, such as selecting a high-fidelity polymerase.

[0004] As for the issue of strand bias, it wasn't noticed until 2012 (for details, see the academic paper "The effect of strand bias in Illumina short-read sequencing data," doi:10.1186 / 1471-2164-13-666). Strand bias manifests itself as inconsistent base site preferences on either the sense or antisense strand after reads are spliced ​​and aligned to a genomic sequence reference, meaning the ratio between the two deviates significantly from 50%:50. In a sense, strand bias isn't used to characterize the error rate, but rather to assess the uniformity of error events across the entire system. However, to this day, there's still no consensus on the specific mechanisms that produce strand bias, and it's difficult to come up with a solution.

[0005] Summary of the Invention

[0006] The inventors speculate that one possible cause of strand deviation is that the N3-linker on the commonly used dNTP substrates strongly binds to the sequencing enzyme. While this improves incorporation efficiency, it can reduce the enzyme's elution efficiency after base incorporation. Under certain sequence conditions, the enzyme's binding to the DNA template strand is further enhanced, preventing efficient enzyme elution. The negatively charged AF532 dye, which marks the A base, makes incomplete elution susceptible to binding to the enzyme's positively charged finger region, resulting in poor dye excitation and, for example, strand deviation from T to G and A to C (T->G, A->C).

[0007] In order to solve the above problems, the present invention provides a modified nucleotide with a novel structure, which can be used as a dNTP for sequencing to solve the chain deviation problem of T to G, A to C (T->G, A->C).

[0008] Modified nucleotides

[0009] The present application provides a modified nucleotide, a salt thereof, or an ester thereof, wherein the modified nucleotide has a structure shown in the general formula I-1:

[0010] Wherein, D is a nucleotide;

[0011] Dye is a fluorescent dye;

[0012] m1 is selected from 1, 2, 3, 4, 5; r1 is selected from 1, 2, 3, 4, 5; n1 is selected from 1, 2, 3, 4, 5, 6, 7.

[0013] In some embodiments, m1 is 1, r1 is 1, and n1 is 7.

[0014] In some embodiments, D is dNTP, ie, deoxyribonucleoside triphosphate, which can be selected from dATP, dGTP, dTTP, and dCTP.

[0015] In some embodiments, D is rNTP, ie, ribonucleoside triphosphate, which can be selected from ATP, GTP, CTP, and UTP.

[0016] In some embodiments, D is modified with a reversible blocking group, for example, the 3'-O of the deoxyribose is modified with an azidomethylene (-CH2-N3) or allyl group.

[0017] In some embodiments, the modified nucleotide is selected from the following structures:

[0018] In some embodiments, the fluorescent dyes are each independently selected from cyanine dyes, fluorescein dyes, rhodamine dyes, and AF series dyes. In some embodiments, the fluorescent dyes are each independently selected from AF532 or Cy5.

[0019] The dye molecule can be attached to any position on the nucleotide base via a linker, provided that Watson-Crick base pairing can still occur. Specific nucleobase labeling sites include the C5 position of pyrimidine bases or the C7 position of 7-deazapurine bases. The modified nucleotides of the present invention include, but are not limited to:

[0020] As used herein, the term "salt" refers to (i) a salt formed by an acidic functional group (e.g., -COOH) present in the compounds provided herein with a suitable inorganic or organic cation (base), and includes, but is not limited to, alkali metal salts, such as sodium salts, potassium salts, lithium salts, etc.; alkaline earth metal salts, such as calcium salts, magnesium salts, etc.; other metal salts, such as aluminum salts, iron salts, zinc salts, copper salts, nickel salts, cobalt salts, etc.; inorganic base salts, such as ammonium salts; organic base salts, such as tert-octylamine salts, dibenzylamine salts, morpholine salts, glucosamine salts, phenylglycine alkyl ester salts, ethylenediamine salts, N-methylglucosamine salts, guanidine salts, diethylamine salts, triethylamine salts, dicyclohexylamine salts, N,N'-dibenzylethylenediamine salts, chloroprocaine salts, procaine salts, diethanolamine salts, N-benzyl-phenethylamine salts, piperazine salts, tetramethylamine salts, tris(hydroxymethyl)aminomethane salts. and (ii) salts formed by basic functional groups (e.g., -NH2) present in the compounds provided by the present invention and appropriate inorganic or organic anions (acids), including but not limited to hydrohalides, such as hydrofluorides, hydrochlorides, hydrobromides, hydroiodides, etc.; inorganic acid salts, such as nitrates, perchlorates, sulfates, phosphates, etc.; lower alkanesulfonates, such as methanesulfonates, trifluoromethanesulfonates, ethanesulfonates, etc.; arylsulfonates, such as benzenesulfonates, p-toluenesulfonates, etc.; organic acid salts, such as acetates, malates, fumarates, succinates, citrates, tartrates, oxalates, maleates, etc.; amino acid salts, such as glycine, trimethylglycine, arginine, ornithine, glutamate, aspartate, etc.

[0021] As used herein, term " ester " refers to the ester that the-COOH existing in the compound provided by the present invention forms with suitable alcohol, or the ester that the-OH existing in the compound provided by the present invention forms with suitable acid (for example, carboxylic acid or oxygen-containing inorganic acid). Suitable ester groups include but are not limited to formates, acetates, propionates, butyrates, acrylates, ethyl succinates, stearic acid esters or palmitates. In the presence of acid or alkali, esters can undergo hydrolysis reactions to generate corresponding acids or alcohols.

[0022] Sequencing methods

[0023] The modified nucleotides of the present invention can be used in any analytical method requiring detection of fluorescent markers attached to nucleotides or nucleosides, whether by themselves or incorporated into or associated with larger molecular structures or conjugates. Certain embodiments of the present application relate to methods of sequencing, comprising: (a) incorporating at least one modified nucleotide as described herein into a polynucleotide; and (b) detecting the modified nucleotide incorporated into the polynucleotide by detecting a fluorescent signal from a fluorescent dye attached to the modified nucleotide.

[0024] In certain embodiments, at least one modified nucleotide is incorporated into a polynucleotide by the action of a polymerase during the synthesis step. However, other methods for incorporating modified nucleotides into polynucleotides, such as chemical oligonucleotide synthesis or the connection of labeled oligonucleotides to unlabeled oligonucleotides, are not excluded. Therefore, the term "incorporating" a nucleotide into a polynucleotide encompasses polynucleotide synthesis by chemical methods as well as enzymatic methods.

[0025] In specific non-limiting embodiments, nucleic acids labeled with modified nucleotides according to the present invention can be used in methods for nucleic acid sequencing, resequencing, whole genome sequencing, single nucleotide polymorphism scoring, any other application involving the detection of modified nucleotides or nucleosides when incorporated into a polynucleotide, or any other application requiring the use of polynucleotides labeled with modified nucleotides comprising the present invention.

[0026] In a specific embodiment, the application provides the purposes of the modified nucleotide comprising the dye compound of the present invention in polynucleotide " sequencing by synthesis " reaction. Sequencing by synthesis is generally related to using polymerase or ligase in 5 ' to 3 ' direction by one or more nucleotides or oligonucleotides are added to the growing polynucleotide chain in succession, so as to form an extended polynucleotide chain complementary to the template nucleic acid to be sequenced. The identity (identity) of the bases present in one or more of the nucleotides added is determined in a detection step or " imaging " step. The identity of the bases added can be determined after each nucleotide is incorporated into the step. Then, the sequence of the template can be inferred using conventional Watson-Crick base pairing rules. Using a modified nucleotide labeled with content of the present disclosure for determining the identity of a single base can be useful, for example, in the scoring of a single nucleotide polymorphism, and such a single base extension reaction is within the scope of the application.

[0027] In an embodiment, the sequence of the template polynucleotide is determined by detecting the incorporation of one or more nucleotides into a nascent chain complementary to the template polynucleotide to be sequenced via detection of a fluorescent marker attached to the incorporated nucleotide. The nucleic acid template to be sequenced can be DNA or RNA, or even a hybrid molecule comprising deoxynucleotides and ribonucleotides. The nucleic acid template can comprise naturally occurring nucleotides and / or non-naturally occurring nucleotides and natural or non-natural backbone linkages, provided that these do not prevent replication of the template in the sequencing reaction.

[0028] While one application of the modified nucleotides of the present invention is in sequencing by synthesis reactions, the utility of such labeled nucleotides is not limited to such methods. In fact, the nucleotides can be advantageously used in any sequencing method requiring detection of fluorescent labels attached to nucleotides incorporated into a polynucleotide.

[0029] In certain embodiments, the present invention provides a method for determining the sequence of a target single-stranded polynucleotide comprising the steps of:

[0030] (a) providing a duplex, nucleotides, a polymerase, and an excision reagent; the duplex comprises a growing nucleic acid chain and a nucleic acid molecule to be sequenced;

[0031] (b) performing a reaction cycle comprising the following steps (i), (ii) and (iii):

[0032] Step (i): using a polymerase to incorporate nucleotides into the growing nucleic acid chain to form a nucleic acid intermediate comprising a blocking group and a detectable label;

[0033] Step (ii): detecting the detectable label on the nucleic acid intermediate;

[0034] Step (iii): using a cleavage reagent to remove the blocking group on the nucleic acid intermediate.

[0035] In certain embodiments, the reaction cycle further comprises step (iv): removing the detectable label on the nucleic acid intermediate using a cleavage reagent.

[0036] In the present invention, nucleic acids may include nucleotides or nucleotide analogs. Nucleotides generally contain a sugar, a nucleobase, and at least one phosphate group. Nucleotides include deoxyribonucleotides, modified deoxyribonucleotides, ribonucleotides, modified ribonucleotides, peptide nucleotides, modified peptide nucleotides, modified phosphate sugar backbone nucleotides, and mixtures thereof. Examples of nucleotides include, for example, adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), cytidylic acid (CMP), cytidylic acid diphosphate (CDP), cytidylic acid triphosphate (CTP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), deoxyadenosine monophosphate (dAMP), Deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxycytidine diphosphate (dCDP), deoxycytidine triphosphate (dCTP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP) and deoxyuridine triphosphate (dUTP). Nucleotide analogs comprising modified nucleobases can also be used in the methods described herein. Exemplary modified nucleobases that can be included in polynucleotides, whether having a natural backbone or an analogous structure, include, for example, inosine, xanthine, hypoxanthine, isocytosine, isoguanine, 2-aminopurine, 5-methylcytosine, 5-hydroxymethylcytosine, 2-aminoadenine, 6-methyladenine, 6-methylguanine, 2-propylguanine, 2-propyladenine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 15-halouracil, 15-halocytosine, 5-propynyluracil, 5-propynyl Cytosine, 6-azouracil, 6-azocytosine, 6-azothymidine, 5-uracil, 4-thiouracil, 8-haloadenine or guanine, 8-aminoadenine or guanine, 8-thioadenine or guanine, 8-sulfanyladenine or guanine, 8-hydroxyadenine or guanine, 5-halosubstituted uracil or cytosine, 7-methylguanine, 7-methyladenine, 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, etc. As is known in the art, certain nucleotide analogs cannot be incorporated into polynucleotides, for example, nucleotide analogs such as adenosine 5'-phosphosulfate.

[0037] In some embodiments, the nucleotides used in the sequencing methods of the present invention include dATP, dTTP, dCTP, dGTP, wherein one, two, three, or four of the nucleotides are modified nucleotides of the present invention. In some embodiments, one, two, three, or four of the dATP, dTTP, dCTP, dGTP used in the sequencing methods include modified nucleotides of the present invention.

[0038] In some embodiments, the sequencing method of the present invention uses two-color sequencing technology, that is, four bases are mixed and labeled with two fluorescent signals, and then different light signal combinations are used to collect and convert them into gene sequences with a high-resolution camera.

[0039] In some embodiments, the method comprises using a first, second, third, and fourth nucleotide that are different from each other; wherein,

[0040] The first nucleotide carries a first fluorescent label detectable at a first emission wavelength;

[0041] the second nucleotide carries a second fluorescent label detectable at a second emission wavelength;

[0042] Some of the third nucleotides carry a fluorescent label detectable at a first emission wavelength but with a different brightness than the first fluorescent label, and some of the third nucleotides carry a fluorescent label detectable at a second emission wavelength but with a different brightness than the second fluorescent label; and

[0043] The fourth nucleotide does not contain a fluorescent label.

[0044] In dual-color sequencing, two different fluorescent dyes can be used to label the same base, for example, 50% of base A is labeled with the dye AF532 and the other 50% with the dye Cy5. Therefore, in such an embodiment, the same type of nucleotide (e.g., dATP) labeled with different dyes can be used simultaneously. In some embodiments, the fourth nucleotide without a fluorescent label is dGTP.

[0045] In other embodiments, the modified nucleotides of the present invention can also be used in four-color sequencing technology, that is, using four fluorescent dyes to respectively label the four bases A, T / U, C and G to achieve the identification and differentiation of these four bases.

[0046] In some embodiments, the method comprises using a first, second, third, and fourth nucleotide that are different from each other; wherein,

[0047] The first nucleotide carries a first fluorescent label detectable at a first emission wavelength;

[0048] the second nucleotide carries a second fluorescent label detectable at a first emission wavelength;

[0049] The third nucleotide carries a third fluorescent label detectable at a second emission wavelength; and

[0050] The fourth nucleotide carries a fourth fluorescent label detectable at a second emission wavelength.

[0051] In the method of the present invention, the nucleic acid molecule to be sequenced is not limited by its length. In certain preferred embodiments, the length of the nucleic acid molecule to be sequenced can be at least 10bp, at least 20bp, at least 30bp, at least 40bp, at least 50bp, at least 100bp, at least 200bp, at least 300bp, at least 400bp, at least 500bp, at least 1000bp, or at least 2000bp. In certain preferred embodiments, the length of the nucleic acid molecule to be sequenced can be 10-20bp, 20-30bp, 30-40bp, 40-50bp, 50-100bp, 100-200bp, 200-300bp, 300-400bp, 400-500bp, 500-1000bp, 1000-2000bp, or more than 2000bp. In certain preferred embodiments, the nucleic acid molecule to be sequenced can have a length of 10-1000bp, to facilitate high-throughput sequencing.

[0052] In certain preferred embodiments, the nucleic acid molecules may be pretreated before being fixed to a support. Such pretreatments include, but are not limited to, fragmentation of the nucleic acid molecules, end-padding, addition of adapters, addition of tags, amplification of the nucleic acid molecules, separation and purification of the nucleic acid molecules, and any combination thereof.

[0053] In certain embodiments, the solid support surface may have reactive functional groups that react with complementary functional groups on the polynucleotide molecules to form covalent bonds, for example, using the same techniques used to attach cDNA to microarrays, for example, see Smirnov et al. (2004), Genes, Chromosomes & Cancer, 40: 72-77 and Beaucage (2001), Current Medicinal Chemistry, 8: 1213-1244, both of which are incorporated herein by reference. DNB can also effectively attach to hydrophobic surfaces, such as clean glass surfaces with low concentrations of various reactive functional groups (e.g., -OH groups). Attachment via covalent bonds formed between polynucleotide molecules and reactive functional groups on the surface is also referred to herein as "chemical attachment."

[0054] In other embodiments, the polynucleotide molecules can be adsorbed onto a surface. In this embodiment, the polynucleotide is immobilized by non-specific interactions with the surface, or by non-covalent interactions such as hydrogen bonds, van der Waals forces, etc.

[0055] In other embodiments, the nucleic acid library can be double-stranded nucleic acid fragments, which are immobilized on the surface of a solid support by ligation reaction with oligonucleic acids immobilized on the surface of a solid support, and then subjected to rolling circle amplification reaction to prepare a sequencing library.

[0056] Rolling Circle Amplification (RCA) is a constant-temperature nucleic acid amplification technology based on the rolling circle replication of circular pathogenic microbial DNA molecules in nature. The RCA technique allows circular single-stranded DNA to form multiple copies linked end-to-end, which then fold freely into nanosphere structures, known as DNA nanoballs (DNBs), in solution. DNA nanoballs can be immobilized on arrayed silicon chips using DNB loading technology and sequenced. Therefore, in certain embodiments, the nucleic acid molecules to be sequenced can be DNA nanoballs.

[0057] Reagent test kit

[0058] In another aspect, the application provides a kit comprising a modified nucleotide labeled with the present invention. In certain embodiments, the kit comprises one or more nucleotides, wherein at least one nucleotide is a modified nucleotide of the present invention. In certain embodiments, the kit may comprise two or more labeled nucleotides. The modified nucleotides or kit of the present invention can be used for sequencing, expression analysis, hybridization analysis, gene analysis, RNA analysis or protein binding assays. This use can be performed on an automatic sequencing instrument. The sequencing instrument can include two lasers operating at different wavelengths.

[0059] In the case that test kit comprises the multiple Nucleotide, especially two Nucleotide and four kinds of Nucleotide that are marked with dye compounds, different Nucleotide can be marked with identical or different dye compounds, or a kind of Nucleotide can not mark dye compounds.In the case that different Nucleotide marks have identical or different dye compounds, the feature of test kit is that the Nucleotide of described dye compound mark can be distinguished by fluorescence spectrum and algorithm.When two kinds of Nucleotide that are marked with fluorescent dye compounds are supplied with test kit form, in certain embodiments, spectrally distinguishable fluorescent dye can be excited at identical wavelength (such as for example by identical laser).When four kinds of Nucleotide that are marked with fluorescent dye compounds are supplied with test kit form, in certain embodiments, spectrally distinguishable two kinds can both be excited at a wavelength, and other two kinds of spectrally distinguishable dye can both be excited at another wavelength.

[0060] In certain embodiments, the kit of the present invention may further comprise: a reagent for fixing the nucleic acid molecule to be sequenced to a support (e.g., by covalent or non-covalent attachment); a primer for initiating nucleotide polymerization; a polymerase for carrying out nucleotide polymerization; one or more buffer solutions; one or more washing solutions; or any combination thereof.

[0061] In certain embodiments, the test kit of the present invention may also include reagents and / or devices for extracting nucleic acid molecules from a sample. Methods for extracting nucleic acid molecules from a sample are well known in the art. Therefore, various reagents and / or devices for extracting nucleic acid molecules may be configured as needed in the test kit of the present invention, such as reagents for crushing cells, reagents for precipitating DNA, reagents for washing DNA, reagents for dissolving DNA, reagents for precipitating RNA, reagents for washing RNA, reagents for dissolving RNA, reagents for removing protein, reagents for removing DNA (for example, when the target nucleic acid molecule is RNA), reagents for removing RNA (for example, when the target nucleic acid molecule is DNA), and any combination thereof.

[0062] In certain embodiments, test kit of the present invention also comprises, and is used for the reagent of pre-treatment nucleic acid molecule.In test kit of the present invention, the reagent for pre-treatment nucleic acid molecule is not subject to additional restriction, and can be selected according to actual needs.The described reagent for pre-treatment nucleic acid molecule comprises for example, and is used for the reagent (for example DNA enzyme I) of nucleic acid molecule fragmentation, and is used for the reagent (for example DNA polymerase, for example T4DNA polymerase, Pfu DNA polymerase, Klenow DNA polymerase) of polishing nucleic acid molecule end, joint molecule, label molecule, and is used for the reagent (for example ligase, for example T4DNA ligase) that joint molecule is connected with target nucleic acid molecule, and is used for the reagent (for example, losing 3'-5' exonuclease activity but showing 5'-3' exonuclease activity) of repairing nucleic acid end, and is used for the reagent (for example, DNA polymerase, primer, dNTP) of amplifying nucleic acid molecule, and is used for the reagent (for example chromatography column) of separation and purification nucleic acid molecule, and its any combination.

[0063] In certain embodiments, the kit of the present invention further comprises a support for fixing the nucleic acid molecules to be sequenced. Typically, the support for fixing the nucleic acid molecules to be sequenced is in a solid phase for ease of operation. Therefore, in this disclosure, "support" is sometimes also referred to as "solid support" or "solid phase support". However, it should be understood that the "support" mentioned herein is not limited to solid, and it can also be a semi-solid (e.g., gel).

[0064] As used herein, the terms "loaded," "fixed," and "attached" when used in reference to nucleic acids, mean attached directly or indirectly to a solid support via covalent or non-covalent bonds. In certain embodiments of the present disclosure, the methods of the present invention include immobilizing nucleic acids on a solid support via covalent attachment. Generally, however, it is only required that the nucleic acids remain fixed or attached to the solid support under conditions where the solid support is desired to be used (e.g., in applications where nucleic acid amplification and / or sequencing is required). In certain embodiments, immobilizing nucleic acids on a solid support can include immobilizing an oligonucleotide to be used as a capture primer or an amplification primer on a solid support such that the 3' end is available for enzymatic extension and at least a portion of the primer sequence is capable of hybridizing to a complementary nucleic acid sequence; the nucleic acid to be immobilized is then hybridized to the oligonucleotide, in which case the immobilized oligonucleotide or polynucleotide can be in a 3'-5' direction. In certain embodiments, immobilizing nucleic acids on a solid support can include binding a nucleic acid binding protein to the solid support by amino modification, and capturing nucleic acid molecules by the nucleic acid binding protein. Alternatively, loading can occur by other means besides base pair hybridization, such as covalent attachment as described above. Non-limiting examples of nucleic acid attachment to a solid support include nucleic acid hybridization, biotin-streptavidin binding, sulfhydryl binding, photoactivated binding, covalent binding, antibody-antigen, physical confinement via a hydrogel or other porous polymer, and the like. Various exemplary methods for immobilizing nucleic acids on solid supports can be found in, for example, G. Steinberg-Tatman et al., Bioconjugate Chemistry 2006, 17, 841-848; Xu X. et al. Journal of the American Chemical Society 128 (2006) 9286-9287; U.S. patent applications US 5639603, US 5641658, US2010248 991; international patent applications WO 2001062982, WO 2001012862, WO 2007111937, WO0006770, all of which are incorporated herein by reference in their entirety for all purposes, in particular for all teachings relating to the preparation of solid supports having nucleic acids immobilized thereon.

[0065] In the present invention, the support can be made of various suitable materials. Such materials include, for example, inorganic substances, natural polymers, synthetic polymers, and any combination thereof. Specific examples include, but are not limited to, cellulose, cellulose derivatives (e.g., nitrocellulose), acrylic resins, glass, silica gel, silicon dioxide, polystyrene, gelatin, polyvinyl pyrrolidone, copolymers of vinyl and acrylamide, cross-linked polystyrenes such as divinylbenzene (see, for example, Merrifield Biochemistry 1964, 3, 1385-1390), polyacrylamide, latex, dextran, rubber, silicon, plastics, natural sponges, metal plastics, cross-linked dextran (e.g., Sephadex TM ), agarose gel (Sepharose TM ), and other supports known to those skilled in the art.

[0066] In certain preferred embodiments, the support used to immobilize the nucleic acid molecules to be sequenced can be a solid support comprising an inert substrate or matrix (e.g., a glass slide, polymer beads, etc.) that has been functionalized, for example, by applying an intermediate material containing reactive groups that allow for covalent attachment of biomolecules such as polynucleotides. Examples of such supports include, but are not limited to, polyacrylamide hydrogels supported on an inert substrate such as glass, particularly the polyacrylamide hydrogels described in WO 2005 / 065814 and US 2008 / 0280773, the contents of which are incorporated herein by reference in their entirety. In such embodiments, the biomolecules (e.g., polynucleotides) can be directly covalently attached to the intermediate material (e.g., hydrogel), while the intermediate material itself can be non-covalently attached to the substrate or matrix (e.g., a glass substrate). In certain preferred embodiments, the support is a glass or silicon wafer whose surface is modified with a layer of avidin, amino, acrylamide silane, or aldehyde chemical groups.

[0067] In the present invention, the support or solid support is not limited to its size, shape and configuration. In some embodiments, the support or solid support is a planar structure, such as a slide, chip, microchip and / or array. The surface of such a support can be in the form of a planar layer.

[0068] In certain preferred embodiments, the support for immobilizing the nucleic acid molecules to be sequenced is an array of beads or wells (also referred to as a chip). The array can be prepared using any of the materials outlined herein for preparing solid supports, and preferably, the surface of the beads or wells on the array is functionalized to facilitate the immobilization of nucleic acid molecules. The number of beads or wells on the array is not limited. For example, each array may contain 10-10 2 , 10 2-10 3 , 10 3 -10 4 , 10 4 -10 5 , 10 5 -10 6 , 10 6 -10 7 , 10 7 -10 8 , 10 8 -10 9 , 10 10 -10 11 , 10 11 -10 12 In certain exemplary embodiments, the surface of each bead or hole can be fixed with one or more nucleic acid molecules. Accordingly, each array can be fixed with 10-10 2 , 10 2 -10 3 , 10 3 -10 4 , 10 4 -10 5 , 10 5 -10 6 , 10 6 -10 7 , 10 7 -10 8 , 10 8 -10 9 , 10 10 -10 11 , 10 11 -10 12 or more nucleic acid molecules. Therefore, such arrays can be particularly advantageously used for high-throughput sequencing of nucleic acid molecules.

[0069] In certain preferred embodiments, the kit of the present invention further comprises a reagent for fixing the nucleic acid molecule to be sequenced to the support (e.g., by covalent or non-covalent attachment). Such reagents include, for example, reagents for activating or modifying nucleic acid molecules (e.g., their 5' ends), such as phosphoric acid, thiol, amine, carboxylic acid, or aldehyde; reagents for activating or modifying the surface of the support, such as amino-alkoxysilanes (e.g., aminopropyltrimethoxysilane, aminopropyltriethoxysilane, 4-aminobutyltriethoxysilane, etc.); cross-linking agents, such as succinic anhydride, phenyl diisocyanate (Guo et al., 1994), maleic anhydride (Yang et al., 1998), 1-ethyl-3-(3-dimethoxysilane)-1-ol; methylaminopropyl)-carbodiimide hydrochloride (EDC), m-maleimidobenzoic acid-N-hydroxysuccinimide ester (MBS), N-succinimidyl[4-iodoacetyl]aminobenzoic acid (SIAB), 4-(N-maleimidomethyl)cyclohexane-1-carboxylic acid succinimide (SMCC), N-γ-maleimidobutyryloxy-succinimide ester (GMBS), 4-(p-maleimidophenyl)butyric acid succinimide (SMPB); and any combination thereof.

[0070] In certain preferred embodiments, the kit of the present invention also includes a primer for initiating nucleotide polymerization. In the present invention, the primer is not subject to additional restrictions, as long as it can specifically anneal to a region of the target nucleic acid molecule. In some exemplary embodiments, the length of the primer can be 5-50bp, such as 5-10, 10-15, 15-20, 20-25, 25-30, 30-35, 35-40, 40-45, 45-50bp. In some exemplary embodiments, the primer can include naturally occurring or non-naturally occurring nucleotides. In some exemplary embodiments, the primer includes naturally occurring nucleotides or is composed of naturally occurring nucleotides. In some exemplary embodiments, the primer includes modified nucleotides, such as locked nucleic acid (LNA). In certain preferred embodiments, the primer includes a universal primer sequence.

[0071] In some preferred embodiments, test kit of the present invention also comprises the polymerase for carrying out nucleotide polymerization reaction.In the present invention, various suitable polymerases can be used to carry out polyreaction.In some exemplary embodiments, described polymerase can be the new DNA chain (such as DNA polymerase) synthesized with DNA as a template.In some exemplary embodiments, described polymerase can be the new DNA chain (such as reverse transcriptase) synthesized with RNA as a template.In some exemplary embodiments, described polymerase can be the new RNA chain (such as RNA polymerase) synthesized with DNA or RNA as a template.Therefore, in some preferred embodiments, described polymerase is selected from DNA polymerase, RNA polymerase, and reverse transcriptase.

[0072] In certain preferred embodiments, the kit of the present invention further comprises one or more excision reagents. In certain embodiments, the excision reagents are selected from one or more of the following reagents: endonuclease IV, alkaline phosphatase, an organic phosphine (e.g., tris(3-hydroxypropyl)phosphine (THPP), tris(2-carboxyethyl)phosphine hydrochloride (TCEP)), or a complex of PdCl2 and sulfonated triphenylphosphine.

[0073] In certain preferred embodiments, the kit of the present invention further comprises one or more buffer solutions. Such buffer solutions include, but are not limited to, buffer solutions for DNA enzyme I, buffer solutions for DNA polymerase, buffer solutions for ligase, buffer solutions for eluting nucleic acid molecules, buffer solutions for dissolving nucleic acid molecules, buffer solutions for carrying out nucleotide polymerization reactions (e.g., PCR), and buffer solutions for carrying out ligation reactions. The kit of the present invention may comprise any one or more of the above-mentioned buffer solutions.

[0074] In certain preferred embodiments, the kit of the present invention further comprises one or more washing solutions. Examples of such washing solutions include, but are not limited to, phosphate buffer, citrate buffer, Tris-HCl buffer, acetate buffer, carbonate buffer, and the like. The kit of the present invention may comprise any one or more of the above-mentioned washing solutions.

[0075] In certain preferred embodiments, the kit of the present invention comprises a sequencing reagent. In certain preferred embodiments, the sequencing reagent comprises at least one of a dNTPs mixture, a nucleic acid polymerase mixture, and an eluent.

[0076] In certain preferred embodiments, the dNTPs mixture contains at least one modified nucleotide of the present invention, or a salt or ester thereof.

[0077] In certain preferred embodiments, the eluent contains an excision reagent; optionally, the excision reagent is selected from one or more of the following reagents: endonuclease IV, alkaline phosphatase, an organic phosphine (e.g., tris(3-hydroxypropyl)phosphine (THPP), tris(2-carboxyethyl)phosphine hydrochloride (TCEP)) or a complex of PdCl2 and sulfonated triphenylphosphine.

[0078] In another aspect, the present application provides use of the modified nucleotide, salt or ester thereof, or kit of the present invention for determining the sequence of a target polynucleotide.

[0079] Method for preparing nucleotides

[0080] The present application also provides a method for preparing the modified nucleotide of the present invention, the method comprising:

[0081] 1) Synthetic linker,

[0082] 2) coupling the linker to a fluorescent dye derivative (e.g., an active ester thereof),

[0083] 3) coupling the product of step 2) with a nucleotide derivative modified with an ethynylamino group.

[0084] The modified nucleotides shown in Formula I-1 can be prepared according to the following synthetic route:

[0085] Wherein, D is a nucleotide,

[0086] Dye is a fluorescent dye.

[0087] In some embodiments, D is dNTP, ie, deoxyribonucleoside triphosphate, which can be selected from dATP, dGTP, dTTP, and dCTP.

[0088] In some embodiments, D is modified with a reversible blocking group, for example, the 3'-O of deoxyribose is modified with -CH2-N3. Beneficial effects

[0089] The present invention solves the chain deviation problem by designing nucleotides modified with fluorescent groups with novel structures, improves the accuracy of nucleic acid sequencing, and has broad application prospects.

[0090] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples. However, it will be understood by those skilled in the art that the following drawings and examples are intended only to illustrate the present invention and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art based on the following detailed description of the accompanying drawings and preferred embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] Figure 1 shows the hot dATP SSEB Linker-PEG8-AF532 (Compound 8) 31 P NMR spectrum.

[0092] Figure 2 shows the hot dATP SSEB Linker-PEG8-Cy5 (Compound 10) 31 P NMR spectrum.

[0093] Figure 3 shows the hot dCTP SSEB Linker-PEG8-Cy5 (Compound 11) 31 P NMR spectrum.

[0094] FIG4 shows the effect of dTTP base pairs on strand deviation using different structures in Example 6.

[0095] FIG5 shows the results of the test using different base combinations with the new SSEB-PEG8 structure linker in Example 7. DETAILED DESCRIPTION

[0096] The embodiments of the present invention will be described in detail below with reference to the examples, but it will be understood by those skilled in the art that the following examples are merely illustrative of the present invention and should not be construed as limiting the scope of the invention. Where specific conditions are not specified in the examples, the methods were performed according to conventional conditions or the conditions recommended by the manufacturer. Where the manufacturers of the reagents or instruments are not specified, they are all conventional products that can be obtained commercially.

[0097] 1. Compound Synthesis

[0098] Reagents and instruments

[0099] 1.1 Instruments:

[0100] DIONEX UltiMate 3000 liquid chromatography-mass spectrometry (degasser: SRD-3400, binary pump: HPG-3400RS, autosampler: WPS-3000TRS, column oven: TCC-3000SD, detector: DAD-3000, mass spectrometer: ISQ EM-Mass spectrometer, Thermo Fisher Scientific), DIONEX UltiMate 3000 preparative liquid chromatography (binary pump: HPG-3200BX, autosampler: WPS-3000TSL, detector: DAD-3000, fraction collector: Fraction Collector F, Thermo Fisher Scientific), KQ-300E ultrasonic cleaner (Kunshan Ultrasonic Instrument Co., Ltd.), electronic balance (model: BCA324I-10CN, Sartorius Scientific Instrument (Beijing) Co., Ltd.), biotage automatic column machine (Biotech Trading (Shanghai) Co., Ltd.), vacuum freeze dryer (Boyikang (Beijing) Instrument Co., Ltd.), and pipettes (eppendorf, 200 μL; 1000 μL; 5000 μL).

[0101] 1.2 Reagents:

[0102] Acetonitrile (Batch No. JA113030, chromatographic grade, Sigma-Aldrich reagent), ultrapure water, N'N-diisopropylethylamine (DIPEA) (Batch No. C12677270, 99%, Shanghai MacLean Biochemical Technology Co., Ltd.), anhydrous N'N-dimethylformamide (DMF) (Batch No. C12504625, 99.8%, Shanghai MacLean Biochemical Technology Co., Ltd.), N,N'-disuccinimidyl carbonate (DSC) (Batch No. BCCB6748, 95%, Sigma-Aldrich reagent), 4-dimethylaminopyridine (DMAP) (Batch No. C12509524, 99%, Shanghai MacLean Biochemical Technology Co., Ltd.), trifluoroacetic acid (TFA) (99%, Shanghai MacLean Biochemical Technology Co., Ltd.), triethylamine (Et3N) (Batch No. STBK3222, 99.8%, Shanghai MacLean Biochemical Technology Co., Ltd.)5%, Sigma-Aldrich reagent), ammonia (NH3·H2O) (Batch No.: C14619978, 25-28%, Shanghai MacLean Biochemical Technology Co., Ltd.), Fmoc-octaethylene glycol propionic acid (Batch No.: YP2212031423001, Xiamen Sinobond Biotechnology Co., Ltd.), AF532 active ester (Batch No.: KS23D323-202304, Beijing Okainas Technology Co., Ltd.), Cy5 active ester (97%, Beijing Okainas Technology Co., Ltd.), SS-Linker (Batch No.: WHR059-033-B3, Wuhan BGI Life Sciences Institute), dTTP dry powder (Batch No.: HTA022206002, Wuhan BGI Life Sciences Institute), dATP dry powder (Batch No.: HAA032206001, Wuhan BGI Life Sciences Institute), dCTP dry powder (Batch No.: HCA032206001, Wuhan BGI Life Sciences Institute), ethyl 3-hydroxybenzoate (Batch No.: C11020977, 99%, Shanghai MacLean Biochemical Technology Co., Ltd.), 2-bromomethyl-1,3-dioxolane (Batch No.: C11148628, 97%, Shanghai MacLean Biochemical Technology Co., Ltd.), potassium carbonate (Batch No.: C14244715, 99%, Shanghai MacLean Biochemical Technology Co., Ltd.), sodium iodide (Batch No.: C12666357, 99%, Shanghai MacLean Biochemical Technology Co., Ltd.), trimethylazide silane (Batch No.: C12460006, 93%, Shanghai MacLean Biochemical Technology Co., Ltd.), tin tetrachloride (Batch No.: C12848556, AR, Shanghai MacLean Biochemical Technology Co., Ltd.), lithium hydroxide (Batch No.: 20210125, 98%, Beijing Wokai Biotechnology Co., Ltd.), N-Boc-ethylenediamine (Batch No.: C11090125, 98%, Shanghai MacLean Biochemical Technology Co., Ltd.), O-(7-azabenzotriazol-1-yl)-N,N,N',N'-tetramethyluronium hexafluorophosphate (HATU) (Batch No.: A2112243, 99%, Aladdin Biochemical Technology Co., Ltd.), nitroxide piperidinol (TEMPO) (Batch No.: C131297 56, 98%, Shanghai Maclean Biochemical Technology Co., Ltd.), sodium dihydrogen phosphate (Batch No.: 72817076, AR, Shanghai Mairui Biochemical Technology Co., Ltd.), disodium hydrogen phosphate (Batch No.: 71817141, AR, Shanghai Mairui Biochemical Technology Co., Ltd.), sodium chlorite (Batch No.: C14281489, 80%, Shanghai Mairui Biochemical Technology Co., Ltd.), 2-(aminomethyl)-1,3-dioxolane (Batch No.: C13129734, 98%, Shanghai Maclean Biochemical Technology Co., Ltd.), 9-fluorenylmethyl chloroformate (Fmoc-Cl) (Batch No.: 77907028, 98%, Shanghai Mairui Biochemical Technology Co., Ltd.), pyridine (Batch No.: C14354827, 99.5%, Shanghai Maclean Biochemical Technology Co., Ltd.), tetrabutylammonium fluoride (TBAF) (Batch No.: C14350370, 1.0 M, Shanghai Maclean Biochemical Technology Co., Ltd.), 2-[2-(2-aminoethoxy)ethoxy]acetic acid (Batch No.: 9OAHRWNU, 98%, Anhui Zesheng Technology Co., Ltd.), di-tert-butyl dicarbonate (Batch No.: C13351048, 98%, Shanghai Maclean Biochemical Technology Co., Ltd.), sodium hydroxide (Batch No.: 20230207, AR, Sinopharm Chemical Reagent Co., Ltd.), tert-butyl bromoacetate (Batch No.: C12956303, 98%, Shanghai Maclean Biochemical Technology Co., Ltd.).

[0103] Example 1 Synthesis of hot dTTP SSEB Linker-PEG8-AF532 (Compound 7)

[0104] Synthesis of compound 2

[0105] Compound 1 (500 mg, 0.75 mmol) was weighed into a 20 mL brown reaction bottle and dissolved in anhydrous DMF (4 mL). DSC (386 mg, 1.5 mmol) and DMAP (18 mg, 0.15 mmol) were added and stirred at room temperature for 5 h. The reaction solution was filtered and the filtrate was purified by reverse-phase biotage column chromatography (welflash C18-I, regular C18 20-40 μm, 120 g, 0.1% FA / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 2. LCMS: calculated for C 38 H 52 N2O 14 [M+H] + :761.35.Found,m / z,[M+NH4] + :778.62.

[0106] Synthesis of compound 4

[0107] Compound 3 (200 mg, 0.34 mmol) was weighed into a 50 mL round-bottom flask and dissolved in methanol (4 mL). Ammonia (2 mL) was added and stirred at room temperature for 15 h. The reaction solution was concentrated under reduced pressure, dissolved in methanol (4 mL), filtered, and the filtrate was purified by reverse-phase biotage column chromatography (welflash C18-I, regular C18 20-40 μm, 120 g, 0.1% FA / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 4. LCMS: calculated for C 20 H 31N2O8S2[M+H] + :491.15.Found,m / z,[M+H] + :491.20.

[0108] Synthesis of compound 5

[0109] In a 20 mL brown reaction flask, compound 4 (150 mg, 0.3 mmol) and compound 2 (194 mg, 0.25 mmol) were weighed and dissolved in anhydrous DMF (3 mL). TEA (176 μL, 1.27 mmol) was added and stirred for 15 h at room temperature. The reaction solution was filtered and the filtrate was purified by reverse-phase biotage column chromatography (welflash C18-I, regular C18 20-40 μm, 120 g, 0.1% TFA / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 5. LCMS: calculated for C 39 H 67 N3O 17 S2[M+H]+:914.40.Found,m / z,[M+H]+:914.58.

[0110] Synthesis of compound 6

[0111] Compound 5 (250 mg, 0.27 mmol) and AF532 active ester (198 mg, 0.27 mmol) were weighed into a 20 mL brown reaction flask and dissolved in anhydrous DMF (4 mL). DIPEA (226 μL, 1.37 mmol) was added and stirred for 5 h at room temperature. The reaction solution was filtered and the filtrate was purified by reverse-phase biotage column chromatography (welflash C18-I, regular C18 20-40 μm, 120 g, 0.1% TFA / acetonitrile, 5-95% acetonitrile content) and then by preparative HPLC (COSMOSIL, 5C18-MS-II, 20 ID × 250 mm, Code: 18420-21, 0.1% TFA / acetonitrile, 5-95% acetonitrile content). The solution was concentrated under reduced pressure and freeze-dried to obtain compound 6. LCMS: calculated for C 69 H 95 N5O 25 S4[M+H]+:1522.53.Found,m / z,[M+H]+:1522.67.

[0112] Synthesis of compound 7

[0113] Compound 6 (30 mg, 0.02 mmol) was weighed into a 20 mL brown reaction flask and dissolved in anhydrous DMF (2 mL). DSC (10 mg, 0.04 mmol) and DMAP (0.5 mg, 0.004 mmol) were added and stirred at room temperature for 6 h. dTTP (11 mg, 0.02 mmol) and DIPEA (16 μL, 0.1 mmol) were then added and stirred at room temperature for 15 h. The reaction mixture was filtered and the filtrate was purified by preparative HPLC (COSMOSIL, 5C18-MS-Ⅱ, 20 ID × 250 mm, Code: 18420-21, 0.1 M TEAB / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 7.

[0114] 1 H NMR(600MHz,D2O)δ8.04(d,J=8.7Hz,2H),7.97(d,J=3.1Hz,1H),7.80(dd,J=18.6,2.1Hz,1H),7.7 2(t,J=8.2Hz,1H),7.50(d,J=7.9Hz,2H),7.42(dd,J=13.8,9.4Hz,1H),6.87(s,2H),6.13(t,J=7. 0Hz,1H),5.08(q,J=6.8Hz,1H),4.58(d,J=3.0Hz,1H),4.52–4.41(m,2H),4.34(s,1H),4.18(d,J= 5.0Hz,4H),4.14(d,J=3.0Hz,2H),3.94(s,2H),3.90(q,J=6.5Hz,2H),3.80(t,J=5.2Hz,2H),3.78 –3.76(m,2H),3.75–3.73(m,4H),3.72–3.70(m,2H),3.69–3.66(m,7H),3.64(t,J=5.2Hz,4H),3.6 1(d,J=1.9Hz,3H),3.58(d,J=3.3Hz,4H),3.55(d,J=2.3Hz,14H),3.39(t,J=5.3Hz,2H),3.06(q,J =7.3Hz,1H),2.56–2.51(m,1H),2.45(t,J=5.8Hz,3H),2.31(dq,J=13.7,7.4Hz,1H),2.04(d,J=7. 3Hz,3H),1.58(dd,J=7.0,3.3Hz,4H),1.22(d,J=6.7Hz,6H),1.18(s,6H),1.06(s,6H).LCMS:calcd for C 82 H 112 N11 O 38 P3S4[(M-2) / 2] - :1038.77.Found,m / z,[(M-2) / 2] - :1038.94.

[0115] Example 2 Synthesis of hot dATP SSEB Linker-PEG8-AF532 (Compound 8):

[0116] Synthesis of compound 8

[0117] Compound 6 (60 mg, 0.04 mmol) was weighed into a 20 mL brown reaction flask and dissolved in anhydrous DMF (2 mL). DSC (20 mg, 0.08 mmol) and DMAP (1 mg, 0.008 mmol) were added and stirred at room temperature for 6 h. dATP (24 mg, 0.04 mmol) and DIPEA (32 μL, 0.2 mmol) were then added and stirred at room temperature for 15 h. The reaction mixture was filtered and the filtrate was purified by preparative HPLC (COSMOSIL, 5C18-MS-II, 20 ID × 250 mm, Code: 18420-21, 0.1 M TEAB / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 8.

[0118] 1H NMR(400MHz,D2O)δ8.02(d,J=7.7Hz,2H),7.98(s,1H),7.66(dd,J=8.5,4.7Hz,2H),7.55(s,1H),7.36–7.29(m,2H),7.25(d,J=8.6Hz,1H),6. 72(dd,J=9.7,2.9Hz,2H),6.45(t,J=7.1Hz,1H),5.05–4.92(m,2H),4.86(d,J=9.2Hz,1H),4.65(s,1H),4.48(s,2H),4.37(s,1H),4.27(s,2H) ,4.16(d,J=12.9Hz,6H),3.95(s,2H),3.78(t,J=5.3Hz,3H),3.75–3.65(m,15H),3.62(d,J=4.5Hz,8H),3.57(d,J=8.7Hz,15H),3.38(t,J=5. 3Hz,2H),3.11–3.00(m,1H),2.46(t,J=6.2Hz,4H),2.05–1.96(m,3H),1.47(t,J=5.4Hz,3H),1.20(d,J=6.2Hz,6H),1.11(s,6H),0.99(s,6H). 31 P NMR(162MHz,D2O)δ-10.91(d,J=19.9Hz),-11.55(d,J=19.8Hz),-23.32(t,J=20.0Hz).LCMS:calcd for C 84 H 114 N 13 O 36 P3S4[(M-2) / 2] - :1050.28.Found,m / z,[(M-2) / 2] - :1050.58. Figure 1 shows the 31 P NMR spectrum.

[0119] Example 3 Synthesis of hot dATP SSEB Linker-PEG8-Cy5 (Compound 10):

[0120] Synthesis of compound 9

[0121] In a 20 mL brown reaction flask, compound 5 (100 mg, 0.1 mmol) and Cy5 active ester (82 mg, 0.1 mmol) were weighed and dissolved in anhydrous DMF (4 mL). DIPEA (90 μL, 0.5 mmol) was added and stirred for 4 h at room temperature. The reaction solution was filtered and the filtrate was purified by reverse-phase biotage column chromatography (welflash C18-I, regular C18 20-40 μm, 120 g, 0.1% TFA / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 9. LCMS: calculated for C 72 H 105 N5O 24 S4[M+H]+:1552.61.Found,m / z,[M+H]+:1552.86.

[0122] Synthesis of compound 10

[0123] Compound 9 (20 mg, 0.013 mmol) was weighed into a 20 mL brown reaction flask and dissolved in anhydrous DMF (2 mL). DSC (6.6 mg, 0.026 mmol) and DMAP (0.3 mg, 0.002 mmol) were added and stirred at room temperature for 5 h. dATP (8 mg, 0.013 mmol) and DIPEA (11 μL, 0.065 mmol) were then added and stirred at room temperature for 15 h. The reaction mixture was filtered and the filtrate was purified by preparative HPLC (COSMOSIL, 5C18-MS-II, 20 ID × 250 mm, Code: 18420-21, 0.1 M TEAB / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 10.

[0124] 1H NMR(400MHz,D2O)δ8.04–7.90(m,3H),7.90–7.81(m,4H),7.78(dd,J=8.7,4.2 Hz,1H),7.65(d,J=9.9Hz,2H),7.40–7.26(m,3H),6.50–6.38(m,2H),6.18(dd,J=13.7,6.6Hz,2H),5.12–5.05(m,1H),4.94(dd,J=9.1,4.5Hz,1 H),4.85(d,J=9.4Hz,1H),4.69(s,1H),4.53(s,1H),4.37(s,1H),4.24–4.15(m,5H),4.11(d,J=3.6Hz,3H),4.02(dd,J=16.4,5.9Hz,6H),3.70( d,J=14.1Hz,6H),3.63(d,J=2.2Hz,3H),3.60–3.54(m,28H),3.51(t,J= 5.2Hz,3H),3.39(s,2H),3.33–3.24(m,3H),2.62(d,J=8.1Hz,1H),2.50 (d,J=7.0Hz,1H),2.45(t,J=6.0Hz,2H),2.22(t,J=7.2Hz,2H),2.08(d,J=10.4Hz,3H),1.76(s,2H),1.66–1.52(m,17H),1.32(d,J=7.2Hz,4H). 31 P NMR(162MHz,D2O)δ-6.34(d,J=20.9Hz),-11.45(d,J=19.1Hz),-22.53(t,J=19.4Hz).LCMS:calcd for C 87 H 124 N 13 O 35 P3S4[(M-2) / 2] - :1064.82.Found,m / z,[(M-2) / 2] - :1065.22. Figure 2 shows the 31 P NMR spectrum.

[0125] Example 4 Synthesis of hot dCTP SSEB Linker-PEG8-Cy5 (Compound 11):

[0126] Synthesis of compound 11

[0127] Compound 9 (30 mg, 0.02 mmol) was weighed into a 20 mL brown reaction flask and dissolved in anhydrous DMF (3 mL). DSC (15 mg, 0.06 mmol) and DMAP (0.5 mg, 0.004 mmol) were added and stirred at room temperature for 6 h. dCTP (11 mg, 0.02 mmol) and DIPEA (16 μL, 0.1 mmol) were then added and stirred at room temperature for 15 h. The reaction mixture was filtered and the filtrate was purified by preparative HPLC (COSMOSIL, 5C18-MS-II, 20 ID × 250 mm, Code: 18420-21, 0.1 M TEAB / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 11.

[0128] 1 H NMR(400MHz,D2O)δ8.08–7.97(m,3H),7.89–7.75(m,6H),7.50(t,J=7.6Hz,1H),7.34( t,J=9.0Hz,2H),6.53(t,J=12.4Hz,1H),6.23(d,J=13.7Hz,2H),6.12–6.03(m,1H),5. 13(dd,J=6.9,3.8Hz,1H),4.86(dd,J=9.1,5.6Hz,2H),4.78–4.72(m,1H),4.55(d,J=1 5.6Hz,2H),4.35(s,1H),4.23(dd,J=9.0,4.4Hz,3H),4.20–4.11(m,6H),4.07(d,J=8. 8Hz, 4H), 3.99 (s, 2H), 3.71 (dd, J=12.4, 6.4Hz, 7H), 3.64 (t, J=5.3Hz, 3H), 3.60–3.56 (m,26H),3.51(t,J=5.3Hz,2H),3.40(t,J=5.3Hz,2H),3.32–3.25(m,3H),2.54(d,J=1 4.8Hz,1H),2.47(t,J=6.1Hz,2H),2.23(t,J=7.1Hz,2H),2.10(d,J=8.6Hz,3H),1.79( d,J=8.2Hz,2H),1.66(d,J=10.3Hz,12H),1.60(d,J=6.9Hz,3H),1.34(q,J=7.4Hz,6H). 31 P NMR(162MHz,D2O)δ-6.34(d,J=20.9Hz),-11.61(d,J=21.5Hz),-22.42(t,J=21.1Hz).LCMS:calcd for C 85H 123 N 12 O 36 P3S4[(M-2) / 2]-:1053.32.Found,m / z,[(M-2) / 2]-:1053.67. Figure 3 shows the 31 P NMR spectrum.

[0129] Example 5 Synthesis of hot dTTP-short-SSEB Linker 1-AF532 (Compound 45, Control):

[0130] Synthesis of compound 38

[0131] Compound 37 (5.00 g, 30.6 mmol) was weighed into a 250 mL round-bottom flask and dissolved in THF (50.0 mL) and H₂O (20.0 mL). Boc₂O (10.0 g, 46.0 mmol) and NaOH (400 mg, 10.0 mmol) were added at 0°C. The mixture was stirred for 12 h at room temperature. The reaction mixture was adjusted to pH 5.0 with aqueous hydrochloric acid (0.50 M) and extracted with ethyl acetate (50.0 mL x 3). The organic phases were combined, washed with saturated sodium chloride solution (100 mL), dried over anhydrous sodium sulfate, filtered, and concentrated under reduced pressure to obtain the crude product. The product was purified by biotage column chromatography (silica gel, eluent: dichloromethane / methanol = 20 / 10 to 5 / 1) to obtain compound 38.

[0132] Synthesis of compound 40

[0133] Take a 100mL round-bottom flask, weigh compound 38 (1.08g, 4.11mmol), add DMF (20.0mL) to dissolve, add HATU (1.48g, 4.11mmol) and DIPEA (1.59g, 12.3mmol), stir at room temperature for 8h, then add compound 39 (0.50g, 2.05mmol), stir at room temperature for 8h. Add ethyl acetate (150mL) to the reaction solution, wash with saturated sodium chloride solution (200mL*), dry over anhydrous sodium sulfate, filter, and concentrate under reduced pressure to obtain a crude product. Purify by reverse phase biotage automatic column machine (welflash C18-I, regular C18 20-40μm, 330g, 0.1% FA / acetonitrile, acetonitrile content 5-95%), concentrate the prepared solution under reduced pressure, and freeze-dry to obtain compound 40. LCMS: calculated for C 21 H 32 N2O7S2[M+H] + :489.17.Found,m / z,[M+H] +:489.20.

[0134] Synthesis of compound 42

[0135] Take a 100mL round-bottom flask, weigh compound 40 (500mg, 1.02mmol), add DMF (10.0mL) to dissolve, add potassium carbonate (170mg, 1.23mmol), and stir under electromagnetic stirring for 0.5h at room temperature. Then add compound 41 (240mg, 1.23mmol) and stir under electromagnetic stirring for 2h. The reaction system is filtered, and the filtrate is purified by reverse phase biotage automatic column machine (welflash C18-I, regular C18 20-40μm, 330g, 0.1% FA / acetonitrile, acetonitrile content 5-95%). The prepared solution is concentrated under reduced pressure and freeze-dried to obtain compound 42. LCMS: calculated for C 27 H 42 N2O9S2[M+H] + :603.23.Found,m / z,[M+H] + :603.25.

[0136] Synthesis of compound 43

[0137] Compound 42 (300 mg, 0.50 mmol) was weighed into a 100 mL round-bottom flask and dissolved in dichloromethane (8.00 mL). Trifluoroacetic acid (2.00 mL) was added and stirred at room temperature for 2 h. The reaction solution was concentrated under reduced pressure, dissolved in acetonitrile (5 mL), filtered, and the filtrate was purified by reverse-phase biotage column chromatography (welflash C18-I, regular C18 20-40 μm, 80 g, 0.1% TFA / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 43. LCMS: calculated for C 22 H 34 N2O7S2[M+H] + :503.18.Found,m / z,[M+H] + :503.20.

[0138] Synthesis of compound 44

[0139] Compound 43 (60.0 mg, 134 μmol) and AF532 active ester (107 mg, 148 μmol) were weighed into a 20 mL brown reaction bottle and dissolved in anhydrous DMF (3 mL). DIPEA (104 mg, 806 μmol) was added and stirred at room temperature for 12 h. The reaction mixture was filtered and the filtrate was purified by preparative HPLC (COSMOSIL, 5C18-MS-Ⅱ, 20 ID × 250 mm, Code: 18420-21, 0.1% FA / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 44. LCMS: calculated for C 48 H 54 N4O 15 S4[M+H]+:1055.25.Found,m / z,[M+H]+:1055.35. Synthesis of compound 45

[0140] Compound 44 (20.0 mg, 19 μmol) was weighed into a 20 mL brown reaction bottle and dissolved in anhydrous DMF (2 mL). DSC (9.7 mg, 38 μmol) and DMAP (0.5 mg, 4 μmol) were added and stirred at room temperature for 2 h. dTTP (12.0 mg, 21 μmol) and DIPEA (19.6 mg, 152 μmol) were then added and stirred at room temperature for 15 h. The reaction mixture was filtered and the filtrate was purified by preparative HPLC (COSMOSIL, 5C18-MS-II, 20 ID × 250 mm, Code: 18420-21, 0.1 M TEAB / acetonitrile, acetonitrile content 5-95%). The prepared solution was concentrated under reduced pressure and freeze-dried to obtain compound 45. LCMS: calculated for C 61 H 71 N 10 O 28 P3S4[M-1] - :1611.25.Found,m / z,[M-1] - :1611.30.

[0141] 2. Gene Sequencing Experiment

[0142] 1. Experimental equipment: MGISEQ-200RS sequencer, MGISEQ-200RS sequencing slide, the excitation wavelengths of the instrument are 532 and 650 nm respectively.

[0143] 2. Reagents and raw materials used in the experiment

[0144] The new structure of reversibly blocked modified nucleotides, ultrapure water, and E. coli single-stranded circular DNA as template (Standard Library Reagent V3.0) were used. The primer sequence is CAACTCCTTGGCTCACAGAACATGGCTACGATCCGACTT (SEQ ID NO: 1). DNA polymerase and DNA nanospheres were also from BGI.

[0145] The following experiments all used E. coli single-stranded circular DNA as a template and used the MGISEQ-200RS High-Throughput Sequencing Kit to prepare DNA nanospheres and load them onto the chip for subsequent sequencing.

[0146] Control group: SE50 sequencing was performed on the MGISEQ-200RS sequencing platform using the MGISEQ-200RS high-throughput sequencing kit according to the experimental protocol. Briefly, a mixture of fluorescently modified and reversibly blocked nucleotides was sequentially polymerized on the MGISEQ-200RS sequencing platform. Free nucleotides were then eluted using an elution reagent, signals were acquired in a photosensitive solution, protecting groups were removed using an excision reagent, and washing was performed using an elution reagent. Sequence files for each experiment were then obtained and subjected to genomic alignment to obtain error rates and related files. Tools were then used to count the types and corresponding numbers of strand deviations.

[0147] The structure of the nucleotides included in the MGISEQ-200RS High-Throughput Sequencing Kit is as follows:

[0148] Experimental group: Use the MGISEQ-200RS high-throughput sequencing kit, then remove the nucleotide mixture in well #1 of the kit and replace it with the new structure nucleotide mixture of the present invention (including a certain base or several bases). The fluorescently modified and reversibly blocked nucleotide mixtures are polymerized sequentially on the MGISEQ-200RS sequencing platform, and the free nucleotides are eluted with an elution reagent. This allows polymerization and leveling to be performed while signal acquisition is performed. At this time, the polymerized nucleotide mixture is a reversibly blocked modified nucleotide with a new structure. Then, the FQ sequence file of each experiment is obtained, and the genome alignment is performed on the file to obtain the error rate and related files. The type and corresponding number of chain deviations are counted using tools.

[0149] 3. Experimental Methods and Procedures

[0150] (a) Reagent tank configuration:

[0151] Configure the reagent tank according to the MGISEQ-200RS reagent tank instructions. When testing the experimental group, replace the nucleotides with the new structure and reversible blocking modification in sequence.

[0152] (b) DNA sequencing method

[0153] Step 1: Load DNA nanospheres onto the prepared chip

[0154] Step 2: Pump the prepared dNTP molecule mixture into the chip and use DNA polymerase to add dNTP molecules to the parent chain of DNA

[0155] Step 3: Scan the reagents and take pictures to determine the type of base

[0156] Step 4: Use the removal reagent THPP to remove the reversible blocking group

[0157] Step 5: Repeat steps 2-4 to perform a second cycle of sequencing.

[0158] Repeat the above cycle until the sequencing is completed.

[0159] Example 6 Sequencing experiments using dTTP with different linkers

[0160] Grouped according to Table 1, dTTP sequencing using different linkers.

[0161] Table 1

[0162] In Figure 4, different colors represent the probability of strand deviation, with increasing severity from red to blue. As can be seen from Figure 4, using SSEB linker nucleotides has a significant effect on reducing T>G strand deviation.

[0163] Example 7 Sequencing experiments using different base combinations with a linker structure of SSEB-PEG8

[0164] According to the grouping in Table 2, sequencing was performed using different base combinations with a linker structure of SSEB-PEG8.

[0165] Table 2

[0166] Table 3 and Figure 5 show the results of testing different base combinations using a new SSEB-PEG8 linker structure. For the chain deviation results of replacing only one T base, replacing two bases (T and A) at the same time, and replacing three bases (T, A, and C) at the same time, the T>G ratio decreased significantly.

Claims

1. A modified nucleotide, its salt or its ester, wherein the modified nucleotide has a structure represented by General Formula I-1 or I-2: Among them, D is a nucleotide; dye is a fluorescent dye; m1 is selected from 1, 2, 3, 4, 5; r1 is selected from 1, 2, 3, 4, 5; n1 is selected from 1, 2, 3, 4, 5, 6, 7.

2. The modified nucleotide, its salt or its ester of claim 1, wherein, D is a dNTP or rNTP, such as dATP, dGTP, dTTP, dCTP, ATP, GTP, CTP or UTP.

3. The modified nucleotide, salt or ester thereof according to claim 1 or 2, wherein, D is modified with a reversible blocking group, such as azidomethyl or allyl at the 3'-O of deoxyribose.

4. A modified nucleotide, a salt thereof or an ester thereof according to any one of claims 1 to 3, wherein the modified nucleotide is selected from the following structures:

5. The modified nucleotide, its salt or its ester according to any one of claims 1-4, wherein the fluorescent dyes are independently selected from cyanine dyes, fluorescein dyes, rhodamine dyes, and AF series dyes; Preferably, the fluorescent dyes are independently selected from AF532 or Cy5.

6. The modified nucleotide, its salt or its ester according to any one of claims 1-5, including but not limited to:

7. A method for preparing the modified nucleotide according to any one of claims 1-6, the method comprising: 1) Synthesizing a linker, 2) Coupling the linker with a fluorescent dye derivative (such as its active ester), 3) Coupling the product of step 2) with a nucleotide derivative bearing an ethynylamino modification.

8. The method of claim 7, wherein, The modified nucleotide is prepared according to the following synthetic route: Wherein, D is a nucleotide, dye is a fluorescent dye.

9. A sequencing method, the sequencing method comprising: (a) Incorporating at least one of the modified nucleotides, its salt or its ester according to any one of claims 1-6 into a polynucleotide; and (b) Detecting the modified nucleotide, its salt or its ester incorporated into the polynucleotide by detecting the fluorescence signal from the fluorescent dye attached to the modified nucleotide, its salt or its ester.

10. The sequencing method of claim 9, the method comprising the following steps: (a) Providing a duplex, nucleotides, a polymerase and an excision reagent; the duplex comprises a growing nucleic acid strand and a nucleic acid molecule to be sequenced; (b) Performing a reaction cycle comprising the following steps (i), (ii) and (iii): Step (i): Using a polymerase, incorporating nucleotides into the growing nucleic acid strand to form a nucleic acid intermediate comprising a blocking group and a detectable label; Step (ii): Detecting the detectable label on the nucleic acid intermediate; Step (iii): Removing the blocking group on the nucleic acid intermediate using an excision reagent; Optionally, the reaction cycle further comprises step (iv): Removing the detectable label on the nucleic acid intermediate using an excision reagent.

11. The sequencing method of claim 9 or 10, the method using a two-color sequencing technique, the method comprising using first, second, third and fourth nucleotides that are different from each other; wherein, The first nucleotide bears a first fluorescent label detectable at a first emission wavelength; The second nucleotide bears a second fluorescent label detectable at a second emission wavelength; Part of the third nucleotides bear a fluorescent label detectable at the first emission wavelength but having a different brightness from the first fluorescent label, and part of the third nucleotides bear a fluorescent label detectable at the second emission wavelength but having a different brightness from the second fluorescent label; and The fourth nucleotide does not contain a fluorescent label.

12. The sequencing method of claim 9 or 10, the method using a four-color sequencing technique, the method comprising using first, second, third and fourth nucleotides that are different from each other; wherein, The first nucleotide bears a first fluorescent label detectable at a first emission wavelength; The second nucleotide bears a second fluorescent label detectable at the first emission wavelength; The third nucleotide bears a third fluorescent label detectable at a second emission wavelength; and The fourth nucleotide bears a fourth fluorescent label detectable at the second emission wavelength.

13. The sequencing method of claim 9, wherein the nucleic acid molecule to be sequenced is a DNA nanosphere.

14. A kit comprising one or more nucleotides, wherein at least one nucleotide is a modified nucleotide, its salt or its ester according to any one of claims 1-6; Preferably, the kit comprises two or more labeled nucleotides.

15. The kit of claim 14, wherein the kit comprises sequencing reagents; Preferably, the sequencing reagents comprise at least one of: a dNTPs mixture, a nucleic acid polymerase mixture, and an eluent.

16. The kit of claim 15, wherein the dNTPs mixture contains at least one modified nucleotide, its salt or its ester according to any one of claims 1-6; Preferably, the eluent contains an excision reagent; Optionally, the excision reagent is selected from one or more of the following reagents: endonuclease IV, alkaline phosphatase, organophosphides (such as tris(3-hydroxypropyl)phosphine (THPP), tris(2-carboxyethyl)phosphine hydrochloride (TCEP)), or a PdCl2 complex with sulfonated triphenylphosphine.

17. Use of a modified nucleotide, its salt or its ester according to any one of claims 1-6 or a kit according to claims 14-16 in sequencing.

Citation Information

Patent Citations

  • Modified nucleoside or nucleotide

    WO2022083686A1

  • Nucleotide analogue for sequencing

    WO2022206922A1