Nucleic acid synthesis and sequencing using tethered nucleoside triphosphates

By using polymerase and nucleoside triphosphate conjugates to cleave bonds, a two-step cyclic synthesis and sequencing of nucleic acids was achieved, solving the problems of high error rate and complex sequence synthesis in existing technologies, and improving the accuracy and yield of nucleic acid synthesis.

CN114874337BActive Publication Date: 2026-05-29RGT UNIV OF CALIFORNIA

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RGT UNIV OF CALIFORNIA
Filing Date
2017-06-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing DNA synthesis technologies suffer from high error rates and difficulty in synthesizing complex sequences, especially in oligonucleotide assembly, resulting in low yields and making them unsuitable for research and engineering applications.

Method used

A two-step cyclic synthesis and sequencing of nucleic acids is achieved by using a conjugate containing polymerase and nucleoside triphosphates linked by cleavable bonds. Cleavable adapters are used to cleave nucleotides after incorporation to achieve further extension and sequence determination of nucleic acids.

Benefits of technology

It improves the accuracy and yield of nucleic acid synthesis, enables the synthesis of complex sequences, is suitable for DNA and RNA synthesis and sequencing, reduces error rates, and improves product purity and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114874337B_ABST
    Figure CN114874337B_ABST
Patent Text Reader

Abstract

Among other things, the present invention provides a conjugate comprising a polymerase and a nucleoside triphosphate, wherein the polymerase and the nucleoside triphosphate are covalently linked via a linker comprising a cleavable bond. Also provided is a set of such conjugates, wherein the conjugates correspond to G, A, T (or U), and C. Also provided is a method for synthesizing a nucleic acid of a defined sequence. The conjugates can also be used in sequencing applications.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of PCT international application PCT / US2017 / 039120 filed on June 23, 2017, which entered the Chinese national phase on February 25, 2019, with application number 201780052080.7 and invention title "Nucleic Acid Synthesis and Sequencing Using Tethered Nucleoside Triphosphates".

[0002] Cross-references

[0003] This application claims the benefit of U.S. Patent Application No. 62 / 354,635, filed June 24, 2016, which is incorporated herein by reference. background

[0004] Most current next-generation sequencing (NGS) is based on sequencing by synthesis (SBS), in which the sequence of the initiated template molecule is determined by a signal generated by the stepwise incorporation of complementary nucleotides by polymerase (Goodwin et al., Nat RevGenet. 2016 May 17; 17(6):333-51). Currently, the most popular SBS method uses fluorescent reversible terminator nucleotides (RTdNTPs)—chemically modified nucleotides that, once incorporated into primers, block extension by polymerase. After incorporation, the free RTdNTPs are removed, and the identity of the added bases is determined by the fluorescent signal from the incorporated nucleotides. Next, the fluorescent reporter and terminator groups are removed from the incorporated nucleotides, rendering the primers non-fluorescent and ready for subsequent extension by polymerase. By repeating this cycle of template-dependent extension, detection, and deprotection, the sequence of the template molecule is inferred from the sequence of the fluorescent signal.

[0005] Modern DNA synthesis began with the chemical synthesis of dozens to hundreds of 50-200 nt oligonucleotides using the phosphoramidite method (Beaucage and Caruthers, Tetrahedron Letters 22.20 (1981):1859-1862). These oligonucleotides were assembled into kilobase-sized products, which were then isolated, sequence-verified, and amplified for subsequent reconstruction into full-length target sequences if necessary (Kosuri and Church, Nature methods 11.5 (2014):499-507). Despite decades of incremental improvements, each chemical step in oligonucleotide synthesis produces 0.5-1.0% unreacted (or side-reaction) product, and these small losses compound exponentially, severely compromising the yield of full-length oligonucleotides. Because many oligonucleotides are assembled into products each kilobase-sized, even the presence of a small fraction of erroneous oligonucleotides in the assembly reaction will result in most products containing at least one error. Current gene synthesis technologies employ various "error correction" strategies to enrich error-free oligonucleotides or assembly products. However, erroneous oligonucleotides have been reported to be "the most critical factor in current DNA synthesis protocols" (Czar et al., Trends in biotechnology 27.2(2009):63-72). Furthermore, many biologically relevant sequences, such as those with repetitive or structurally forming regions and / or high or low G / C content, are difficult to construct via oligonucleotide assembly, if not impossible, hindering their use in research and engineering.

[0006] To date, there is no practical method for de novo DNA synthesis using polymerases to extend nucleic acids in a cyclic manner similar to SBS.

[0007] The key advancement leading to the NGS revolution can be attributed to the development of reversible terminator deoxynucleoside triphosphates (RTdNTPs), which can be incorporated into DNA via polymerase and reversibly terminate further addition of dNTPs. This improved system allows for single nucleotide elongation of grown nucleic acids to favor SBS (Self-Brain Synthesis), enabling practical enzymatic de novo DNA synthesis.

[0008] Overview

[0009] Among other aspects, the present invention provides conjugates comprising a polymerase and a nucleoside triphosphate, wherein the polymerase and the nucleoside triphosphate are covalently linked by a linker comprising a cleavable bond. A set of such conjugates corresponding to G, A, T (or U), and C is also provided. A method for synthesizing nucleic acids with defined sequences is also provided. The conjugates can also be used for sequencing applications. Brief description of the attached diagram

[0010] Those skilled in the art will understand that the accompanying drawings described below are for illustrative purposes only. The drawings are not intended to limit the scope of the teachings of the invention in any way.

[0011] Figure 1A A two-step cyclic nucleic acid synthesis protocol using a polymerase-nucleotide conjugate. In the first step, the conjugate uses its linked dNTP portion to extend a DNA molecule; in the second step, the bond between the polymerase and the extended DNA molecule is cleaved, deprotecting the DNA molecule for subsequent extension.

[0012] Figure 1B A two-step cyclic nucleic acid synthesis protocol using a TdT-dNTP conjugate comprising a TdT molecule specifically labeled with a dNTP site via a cleavable linker.

[0013] Figure 2 The co-crystal structure of .TdT (PDB ID: 4I27) with oligonucleotides and dNTPs is shown in the annotations, which indicate the possible linker sites for dNTPs to polymerases.

[0014] Figures 3A-3E The protocol for tethering dUTP to TdT and the chemical details of using conjugates to extend nucleic acids.

[0015] Figure 3A Starting material for preparing thiol-reactive linker-nucleotide OPSS-PEG4-amino-allyl-dUTP.

[0016] Figure 3B The structure of the thiol reactive linker-nucleotide OPSS-PEG4-amino-allyl-dUTP.

[0017] Figure 3C Polymerase-nucleotide conjugate prepared by labeling TdT with OPSS-PEG4-amino-allyl-dUTP.

[0018] Figure 3D DNA molecules are extended via polymerase-nucleotide conjugates.

[0019] Figure 3E The cleavage of the bond between the extended DNA molecule and TdT.

[0020] Figures 4A-4E The scheme for tethering dCTP to TdT based on photolytically cleavable linkers and the chemical details of using conjugates to extend nucleic acids.

[0021] Figure 4A Starting material for the preparation of thiol-reactive linker-nucleotide BP-23354-propyneamino-dCTP.

[0022] Figure 4B The structure of the thiol-reactive, photolyzable linker-nucleotide BP-23354-propyne-amino-dCTP.

[0023] Figure 4C Polymerase-nucleotide conjugate prepared by labeling TdT with BP-23354-propynylamino-dCTP.

[0024] Figure 4D DNA molecules are extended via polymerase-nucleotide conjugates.

[0025] Figure 4E The cleavage of the bond between the extended DNA molecule and TdT.

[0026] Figures 5A-5E A protocol for real-time error correction in DNA synthesis using fluorescent polymerase-nucleotide conjugates.

[0027] Figure 5A The reaction chamber contains a single molecule loaded with primers.

[0028] Figure 5B Primers extend through binding agents.

[0029] Figure 5C The extension reaction is confirmed by detecting the reporter portion of the conjugate. If the reporter is not detected, the extension reaction is repeated.

[0030] Figure 5D The primers are deprotected by cleavage linkers, releasing polymerase and reporter.

[0031] Figure 5E The deprotection reaction is confirmed by detecting the absence of the reporter substance. If the reporter substance is not detected, the deprotection reaction is repeated.

[0032] Figure 6 A schematic diagram of an integrated microfluidic device for DNA synthesis using fluorescent polymerase-nucleotide conjugates detectable by TIRF microscopy.

[0033] Figures 7A-7E A DNA sequencing protocol using fluorescent polymerase-nucleotide conjugates.

[0034] Figure 7A The reaction chamber is loaded with primer-template duplexes.

[0035] Figure 7B The primer-template duplex was exposed to a mixture of four nucleotides labeled with different fluorophores and extended by a conjugate complementary to the first template base to be sequenced.

[0036] Figure 7C Template bases are determined by detecting the report portion of the conjugate.

[0037] Figure 7D The primers are deprotected by cleavage linkers, releasing polymerase and reporter.

[0038] Figure 7E The absence of the detected component in the report can be used to optionally confirm the deprotection reaction.

[0039] Figures 8A-8B Primer extension was confirmed on SDS-PAGE using polymerase-nucleotide conjugates with varying numbers of tethered nucleotides.

[0040] Figure 8A Primers were extended using wild-type TdT (with up to five tethered nucleotides) and TdT mutants with only one tethered nucleotide.

[0041] Figure 8B Primers were extended using conjugates of TdT mutants with various nucleotide linkage sites for tethering.

[0042] Figures 9A-9B DNA synthesis cycle was demonstrated by SDS-PAGE and capillary electrophoresis (CE).

[0043] Figure 9A SDS-PAGE analysis of protein-DNA complex formation and dissociation during polymerase-nucleotide conjugate primer extension and cleavage linker cleavage.

[0044] Figure 9B From Figure 9A Capillary electrophoresis diagram of the reaction products.

[0045] Figure 10 25 nM DNA primers were extended using 16 μM TdT-dATP, -dCTP, -dGTP and -dTTP conjugates, followed by capillary electrophoresis of the reaction time path of photolysis.

[0046] Figure 11A-11B Evidence of the synthesis of .4-mer (5'-CTAG-3').

[0047] Figure 11A The process of synthesizing and verifying the sequence of the extended product. Step 2: SEQ ID NO:15, Step 3: from top to bottom: SEQ ID NO:16, SEQ ID NO:23.

[0048] Figure 11B Sequencing electrophoresis image of one of the clones (SEQ ID NO:17).

[0049] Figure 12 CE analysis demonstrated that free nucleotides are incorporated into the polymerase primers via their tethered nucleotides.

[0050] Figures 13A-13C An experimental setup to demonstrate that scar DNA can serve as a template for precise complementary DNA synthesis.

[0051] Figure 13A A scheme for synthesizing polynucleotides consisting of nucleotides modified with 3-acetaminopropynyl (“scar”).

[0052] Figure 13B Capillary electrophoretic analysis of synthesized “scar” polynucleotides.

[0053] Figure 13C qPCR amplification of “scar” polynucleotides.

[0054] Figures 14A-14B Evidence of the synthesis of .10-mer (5'-CTACTGACTG-3') (SEQ ID NO:18).

[0055] Figure 14A The process of synthesizing and verifying the sequence of the extended product. Step 1: SEQ ID NO:19, Step 2: SEQ ID NO:20, Step 3: from top to bottom: SEQ ID NO:21, SEQ ID NO:24.

[0056] Figure 14B Analysis of sequencing electrophoresis diagram and synthesis steps of clone one (SEQ ID NO:22).

[0057] Detailed Explanation

[0058] This invention provides a conjugate comprising a polymerase and a nucleoside triphosphate, wherein the polymerase and the nucleoside triphosphate are linked by a linker comprising a cleavable bond. Examples of such conjugates are shown in [the following text is missing from the original]. Figure 3C and Figure 4C In the middle. The polymerase moiety of the conjugate can use its linked nucleoside triphosphate to extend the nucleic acid (i.e., the polymerase can catalyze the attachment of the nucleotide to it to the nucleic acid) and maintain the attachment to the extended nucleic acid through the linker until the linker is cleaved.

[0059] In some embodiments, once the polymerase of the conjugate has incorporated its tethered nucleotide into the nucleic acid, further extension of the nucleic acid by other polymerase-nucleotide conjugates is prevented by an effect referred to herein as “shielding,” where “shielding” refers to the following phenomena: 1) the conjugate polymerase molecule prevents other conjugate molecules from approaching the 3'OH of the extended DNA molecule, and 2) the nucleotide triphosphate molecule tethered to the other conjugate molecule prevents the polymerase from approaching the catalytic site of the polymerase already attached to the end of the extended nucleic acid. In some embodiments, further extension of the nucleic acid can be terminated without the need for additional blocking groups on the tethered nucleotide triphosphate. Termination of extension caused by the shielding effect can be reversed by cleavage of the linker, which releases the tethered polymerase and thereby exposes the 3' end of the extended nucleic acid for subsequent extension by another conjugate.

[0060] In any embodiment, the conjugate may contain an additional portion that, once the tethered nucleotide has been incorporated, facilitates the termination of nucleic acid elongation. For example, well-known and numerously reviewed publications, including Chen, Fei, et al. (Genomics, Proteomics & Bioinformatics 11.1 (2013): 34-40), describe reversible terminator nucleotides that can be incubated in a solution containing polymerase and nucleic acid and, once incorporated into the nucleic acid molecule, inhibit further elongation in the reaction. When conjugates containing polymerase and RTdNTPs are used to elongate nucleic acids, in addition to linker cleavage, deprotection of the RTdNTPs may be required to allow the elongated nucleic acid to undergo further nucleotide addition.

[0061] In some embodiments, the conjugate may be fluorescent, which could be useful in sequencing applications. In some embodiments, the nucleoside triphosphate may be linked to a cysteine ​​residue in the polymerase. However, other chemicals can be used to link the protein and the nucleoside triphosphate, so in some cases, the nucleoside triphosphate may be linked to a non-cysteine ​​residue in the polymerase.

[0062] The cleavable linker should be able to be selectively cleaved by a stimuli (e.g., light, changes in its environment, or exposure to chemicals or enzymes) without disrupting other bonds in the nucleic acid. In some embodiments, the cleavable bond may be a disulfide bond, which can be readily broken using a reducing agent (e.g., β-mercaptoethanol, etc.). Possibly suitable cleavable bonds may include, but are not limited to, the following: base-cleavable sites, such as esters, particularly succinates (which can be cleaved, for example, by ammonia or trimethylamine), quaternary ammonium salts (which can be cleaved, for example, by diisopropylamine), and carbamates (which can be cleaved by aqueous sodium hydroxide solution); acid-cleavable sites, such as benzyl alcohol derivatives (which can be cleaved using trifluoroacetic acid), teicoplanin (which can be cleaved by trifluoroacetic acid, followed by base cleavage), acetals and thioacetals (which can also be cleaved by trifluoroacetic acid), thioethers (which can be cleaved, for example, by HF or cresol), and sulfonyl groups. (Cleavable by trifluoromethanesulfonic acid, trifluoroacetic acid, thioanisole, etc.); nucleophilic cleavable sites, such as phthalamides (cleavable by substituted hydrazides), esters (cleavable by, for example, aluminum trichloride); and Weinreb amides (cleavable by lithium aluminum hydride); and other types of chemically cleavable sites, including thiophosphates (cleavable by silver or mercury ions), diisopropyldialkoxysilyl (cleavable by fluoride ions), diols (cleavable by sodium periodate), and azobenzene (cleavable by sodium dithionite). Other cleavable bonds will be apparent to those skilled in the art or described in relevant literature and textbooks (e.g., Brown (1997) Contemporary Organic Synthesis 4(3); 216-237).

[0063] In certain embodiments, photodegradable (“PC”) connectors (e.g., UV-degradable connectors) may be employed. Suitable photodegradable connectors available for use may include o-nitrobenzyl connectors, benzoylmethyl connectors, alkoxybenzoin connectors, chromium aromatic complex connectors, NpSSMpact connectors, and neopentanoyl glycol connectors, as described by Guillier et al. (Chem Rev. 2000 Jun 14; 100(6):2091-158). Exemplary linker groups that can be used in the subject method can be described in Guillier et al., ibid. and Olejnik et al. (Methods in Enzymology 1998 291:135-154), and further described in USPN 6,027,890; Olejnik et al. (Proc. Natl. Acad Sci. 92:7590-94); Ogata et al. (Anal. Chem. 2002 74:4702-4708); Bai et al. (Nucl. Acids Res. 2004 32:535-541); Zhao et al. (Anal. Chem. 2002 74:4259-4268); and Sanford et al. (Chem Mater. 1998 10:1510-20), and are available from Ambergen (Boston, MA; NHS-PC-LC-Biotin), Link Technologies (Bellshill, Scotland), Fisher Scientific (Pittsburgh, PA), and Calbiochem-Novabiochem Corp. (La Jolla, CA).

[0064] In other embodiments, the bonds can be cleaved by enzymes. For example, amide bonds can be cleaved by proteases, ester bonds by esterases, and glycosidic bonds by glycosidases. In some embodiments, the cleavage agent can also break the bonds in the linked polymerase; for example, the protease can also digest the polymerase.

[0065] In the conjugate, the linker is considered to be at least a C-linker that connects a base, sugar, or α-phosphate of a nucleotide to the C-linker in the polymerase backbone. α Atoms of atoms. In some embodiments, the polymerase and nucleotide are covalently linked, and the linking atom of the nucleotide is a C atom in the polymerase backbone to which it is linked. α The distance between atoms can be For example, 15-40 or Within a certain range, but this distance can vary depending on the location of the nucleoside triphosphate tether. In some embodiments, the linker can be a PEG or peptide linker, but there is considerable flexibility in the type of linker used. In some embodiments, the linker should be attached to the base of the nucleotide at an atom that does not participate in base pairing. In such embodiments, the linker is considered to be at least the C in the polymerase backbone. α An atom attached to any atom in a monocyclic or polycyclic system (e.g., pyrimidine, purine, 7-denitropurine, or 8-aza-7-denitropurine) bonded to the 1' position of a sugar. For example, in Figure 4D In the depicted complex, the linker is attached to the carbon atom at position 5 of the cytosine nucleobase and to the C atom at position 5 of the cysteine ​​residue of the polymerase. α Atom. In other embodiments, the linker should be attached to a base of the nucleotide at the atom involved in base pairing. In other embodiments, the linker should be attached to a sugar or to an α-phosphate of the nucleotide.

[0066] In all embodiments, the linker used should be long enough to allow the nucleoside triphosphate to approach the active site of its tethered polymerase. As will be described in more detail below, the polymerase of the conjugate is capable of catalyzing the addition of its linked nucleotide to the 3' end of the nucleic acid.

[0067] The nucleic acid can be at least 3 nucleotides, at least 10 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 500 nucleotides, at least 1,000 nucleotides, or at least 5,000 nucleotides in length, and can be entirely single-stranded or at least partially double-stranded, for example, hybridizing with another molecule (i.e., a portion of a double-stranded structure) or with itself (e.g., in a hairpin configuration). In any embodiment, the nucleic acid can be an oligonucleotide, which can be at least 3 nucleotides long, for example, at least 10 nucleotides, at least 50 nucleotides, at least 100 nucleotides long, at least 500 nucleotides long to at most 1,000 nucleotides or more, and can be entirely single-stranded or at least partially double-stranded, for example, hybridizing with another molecule (i.e., a portion of a double-stranded structure) or with itself (e.g., in a hairpin configuration). In some embodiments, the oligonucleotide can hybridize with a template nucleic acid. In these embodiments, the template nucleic acid can be at least 20 nucleotides long, for example, at least 80 nucleotides long, at least 150 nucleotides long, at least 300 nucleotides long, at least 500 nucleotides long, at least 2000 nucleotides long, at least 4000 nucleotides long, or at least 10,000 nucleotides long. In some cases, the nucleic acid can be a portion of a natural DNA substrate; for example, it can be a strand of a plasmid. If the nucleic acid is double-stranded, it can have 3' overhangs.

[0068] The present invention also provides a set of conjugates summarized above, wherein the conjugates correspond to (i.e. have the same base pairing ability) G, A, T (or U) and C (i.e. deoxyadenosine triphosphate (dATP), deoxyguanosine triphosphate (dGTP), deoxycytidine triphosphate (dCTP), deoxythymidine triphosphate (dTTP)).

[0069] In some embodiments, these conjugates are in different containers. In other embodiments, the conjugates may be in the same container, particularly if they will be used for sequencing. The nucleotides used herein may contain adenine, cytosine, guanine, and thymine bases, and / or bases that pair with complementary nucleotide bases and can be used as templates by DNA or RNA polymerases, such as 7-deazo-7-propyneamino-adenine, 5-propyneamino-cytosine, 7-deazo-7-propyneamino-guanosine, 5-propyneamino-uridine, 7-deazo-7-hydroxymethyl-adenine, 5-hydroxymethyl-cytosine, 7-deazo-7-hydroxymethyl-guanosine, 5-hydroxymethyl-uridine, 7-deazo-adenine, 7-deazo-guanine, adenine, guanine, Cytosine, thymine, uracil, 2-deazo-2-thio-guanosine, 2-thio-7-deazo-guanosine, 2-thio-adenine, 2-thio-7-deazo-adenine, isoguanine, 7-deazo-guanine, 5,6-dihydrouridine, 5,6-dihydrothymine, xanthine, 7-deazo-xanthine, hypoxanthine, 7-deazo-xanthine, 2,6-diamino-7-deazopurine, 5-methyl-cytosine, 5-propynyl-uridine, 5-propynyl-cytidine, 2-thio-thymine or 2-thio-uridine are examples of such bases, but others are also known. An exemplary set of conjugates for synthesizing and / or sequencing DNA molecules may include a DNA polymerase linked to a deoxyribonucleotide triphosphate selected from: deoxyribonucleotide adenosine triphosphate (dATP), deoxyribonucleotide guanosine triphosphate (dGTP), deoxyribonucleotide cytidine triphosphate (dCTP), deoxyribonucleotide thymidine triphosphate (dTTP), and / or other deoxyribonucleotides that base-pair in the same manner as those deoxyribonucleotides. An exemplary set of conjugates for synthesizing RNA molecules may include an RNA polymerase linked to a ribonucleotide triphosphate selected from: adenosine triphosphate (ATP), guanosine triphosphate (GTP), cytidine triphosphate (dCTP), and uridine triphosphate (UTP), and / or other ribonucleotides that base-pair in the same manner as those ribonucleotide triphosphates.

[0070] The aforementioned conjugates can be used in nucleic acid synthesis methods. In some embodiments, the method may include incubating the nucleic acid with the first conjugate under conditions in which a polymerase catalyzes the covalent addition of a nucleotide of the first conjugate to the 3' hydroxyl group of the nucleic acid to prepare an extension product. This reaction may be carried out using nucleic acids ligated to a solid support or in solution, i.e., not tethered to a solid support. After extending the nucleic acid with a first desired nucleotide, the method may include a deprotection step in which the cleavable bonds of the linker are cleaved, thereby releasing the polymerase from the extension product. If the cleavable bonds are disulfide bonds, this may be accomplished by exposing the reaction product to reducing conditions. However, this step may use other chemicals and reagents. In some embodiments, the nucleoside triphosphate may be an RTdNTP, and the deprotection step of the method further includes removing the blocking group (i.e., removing the terminator group) from the added nucleotide to prepare a deprotected extension product. Deprotection enables subsequent extension of the nucleic acid, thereby allowing these steps to be repeated cyclically to prepare an extension product with a defined sequence. Specifically, in some embodiments, the method may further include, after deprotection, incubating the deprotected extension product with the second conjugate under conditions in which a polymerase catalyzes the covalent addition of a nucleotide of the second conjugate to the 3' end of the extension product.

[0071] In some embodiments, the method may include (a) incubating the nucleic acid with the first conjugate under conditions where a polymerase catalyzes the covalent addition of a nucleotide (i.e., a single nucleotide) of the first conjugate to the 3' hydroxyl group of the nucleic acid to prepare an extension product; (b) cleaving the cleavable bonds of the linker, thereby releasing the polymerase from the extension product and deprotecting the extension product; (c) incubating the deprotected extension product with the second conjugate of claim 1 under conditions where a polymerase catalyzes the covalent addition of a nucleotide of the second conjugate to the 3' end of the extension product to prepare a second extension product; and (d) repeating steps (b)-(c) multiple times (e.g., 2 to 100 or more times) with the second extension product to prepare an extended oligonucleotide of a defined sequence. Steps (b)-(c) may be repeated multiple times as needed until an extension product of a defined sequence and length is synthesized. The length of the final product may be 2-100 bases, but theoretically, the method can be used to prepare products of any length, including products greater than 200 bases or greater than 500 bases.

[0072] In some implementations, cleavage of the linker may leave a “scar” (i.e., a portion of the linker) on each or some of the added nucleotides. In other implementations, cleavage of the linker does not produce a scar.

[0073] In some embodiments, the scar can be further derivatized after each deprotection step (e.g., by alkylation of thiol-containing scars using iodoacetamide). In other embodiments, all scars in the final product can be derivatized simultaneously (e.g., by acetylation of propargylamino scars using NHS acetate).

[0074] In some implementations, the product can be amplified, for example by PCR or some other method, to produce a scarless product (as shown in Example 4).

[0075] A sequencing method is also provided. These methods may include incubating a duplex containing primers and a template with a composition containing a set of binders, wherein the binders correspond to G, A, T, and C and are distinguishably labeled, for example, with fluorescent labels; detecting which nucleotide has been added to the primers by detecting a label attached to a polymerase that has added the nucleotides to the primers; deprotecting the extension product by cleaving the linker; and repeating the incubation, detection, and deprotection steps to obtain at least a portion of the sequence of the template.

[0076] A set of reagents for preparing the above-described conjugates is also provided. In some embodiments, the set of reagents may include: a polymerase modified to contain a single cysteine ​​residue on its surface; and a set of nucleoside triphosphates, wherein each of the nucleoside triphosphates is linked to a thiol reactive group. In some embodiments, the nucleoside triphosphates correspond to G, A, T, and C. As mentioned above, the nucleoside triphosphates may be reversible terminators. In this set of reagents, the nucleoside triphosphates may contain a length of... For example or Connectors within the specified range.

[0077] In any implementation, the polymerase may be a template-independent polymerase, namely a terminal deoxynucleotidyl transferase or a DNA extranucleotidyl transferase, these terms are used interchangeably to refer to an enzyme with activity 2.7.7.31 using the IUBMB nomenclature. Descriptions of these enzymes can be found in Bollum, FJ Deoxynucleotide-polymerizing enzymes of calf thymus gland. V. Homogeneous terminal deoxynucleotidyltransferase. J. Biol. Chem. 246 (1971) 909-916; Gottesman, ME and Canellakis, ES The terminal nucleotidyltransferases of calf thymus nuclei. J. Biol. Chem. 241 (1966) 4339-4352; and Krakow, JS, Coutsogeorgopoulos, C. and Canellakis, ES Studies on the incorporation of deoxyribonucleic acid. Biochim. Biophys. Acta 55 (1962) 639-650.

[0078] Terminal transferase implementation schemes can be used for DNA synthesis.

[0079] In any implementation, the polymerase may be a template-dependent polymerase, i.e., a DNA-directed DNA polymerase (these terms are used interchangeably to refer to an enzyme with activity 2.7.7.7 using the IUBMB nomenclature), or a DNA-directed RNA polymerase. Descriptions of these enzymes can be found in Richardson, A. Enzymatic synthesis of deoxyribonucleic acid. deoxythymidylate. J. Biol. Chem. 235 (1960) 3242-3249; and Zimmerman, BKP Purification and properties of deoxyribonucleic acid polymerase from Micrococcus lysodeikticus. J. Biol. Chem. 241 (1966) 2035-2041.

[0080] In any of the above embodiments, the nucleoside triphosphate can be deoxyribonucleoside triphosphate or ribonucleoside triphosphate. In some embodiments, the conjugate may contain an RNA polymerase linked to a ribonucleoside triphosphate. In these embodiments, the nucleotide added to the nucleic acid may be a ribonucleotide. In other embodiments, the conjugate contains a DNA polymerase linked to a deoxyribonucleoside triphosphate. In these embodiments, the nucleotide added to the nucleic acid may be a deoxyribonucleotide.

[0081] In any implementation, the polymerase used may have an amino acid sequence that is at least 80% identical to, for example, at least 90% or at least 95% identical to, that of the wild-type polymerase.

[0082] In some embodiments, the yield of each nucleotide addition step can be at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99%, such as 91% or 99.5%. The yield of each step in any embodiment of the method can be improved by optimizing the conditions. As will be appreciated, nucleic acids prepared by the proposed method can be purified, for example, by liquid chromatography before use.

[0083] In any embodiment, the conjugate may additionally comprise an additional polypeptide domain fused to the polymerase. For example, a maltose-binding protein may be fused to the N-terminus of a terminal deoxynucleotidyl transferase to enhance its solubility and / or achieve amylose affinity purification. In any embodiment, the nucleoside triphosphate may be a reversible terminator.

[0084] All patents and publications mentioned herein, including all sequences disclosed therein, are expressly incorporated herein by reference.

[0085] Further details of the reagents and methods described above can be found below. Some of the descriptions involve TdT. The principles described can also be applied to other template-independent and template-dependent polymerases.

[0086] Thrusted nucleotides can achieve high effective concentrations, enabling rapid incorporation kinetics.

[0087] Depending on the length and geometry of the linker and its attachment site on the protein, the tethered nucleotide will have a specific occupancy rate at the polymerase's active site. This rate can be expressed as the effective concentration (which will give the concentration of free nucleotides at a corresponding occupancy rate). By changing the linker properties and attachment sites, it is possible to control the effective concentration of nucleotides, achieving high effective concentrations and thus rapid incorporation. For example, very rough calculations show that... The effective concentration of dNTPs in the connector system will be ~50 mM (with (A molecule in the volume of a sphere of radius). In this example, the local concentration of dNTPs can be increased by shortening the joint, or decreased by extending the joint.

[0088] Linkage site of the adapter on the polymerase

[0089] In some implementations, the linker specifically attaches to the amino acid of the polymerase (see...). Figure 2(See schematic diagram). In these cases, it is preferable to connect the linker to an amino acid at a position where it can be mutated without loss of polymerase activity, for example, positions 180, 188, 253, or 302 of mouse TdT (as numbered in the crystal structure PDB ID: 4I27). It is preferable not to connect the linker to an amino acid involved in the catalytic activity of the polymerase to avoid interference with catalysis. The residues involved in catalysis and the methods for determining whether a residue is involved in catalysis (e.g., by site-specific mutagenesis) are known to those skilled in the art and have been reviewed in the literature (e.g., Joyce et al. (Journal of Bacteriology 177.22(1995):6321.) and Jara and Martinez (The Journal of Physical Chemistry B 120.27(2016):6504-6514.)).

[0090] Length of connector

[0091] In any implementation, the linker length can be longer than the distance between the linker's attachment site on the polymerase and the attachment site on which the nucleoside triphosphate binds to the catalytic site. In some cases, spatial constraints, such as due to the polymerase or because the linker can restrict the mobility of the tethered nucleoside triphosphate, necessitate an increased linker length to allow the tethered nucleoside triphosphate to approach the catalytic site of the polymerase that produces the conformation. For example, the linker length can exceed the distance between its two attachment sites. or or Or longer.

[0092] Site-specific ligation strategy between the linker and polymerase.

[0093] In some embodiments, thiol-specific linking chemicals are used to specifically link the tethered nucleoside triphosphate to a cysteine ​​residue of the polymerase. Possible thiol-specific linking chemicals include, but are not limited to, o-pyridyl disulfide (OPSS) (exemplified in Figure 3 and illustrated in Example 1), maleimide functional groups (exemplified in Figure 4 and illustrated in Example 2), 3-arylpropynitrile functional groups, allylamine functional groups, haloacetyl functional groups such as iodoacetyl or bromoacetyl, alkyl halides, or perfluoroaryl groups (Zhang, Chi, et al., Nature Chemistry 8, (2015) 120–128). Other linking chemicals for specifically labeling cysteine ​​residues will be apparent to those skilled in the art or described in relevant literature and textbooks (e.g., Kim, Younggyu, et al., Bioconjugate Chemistry 19.3 (2008): 786–791).

[0094] In other embodiments, the linker can be attached to a lysine residue via an amine-reactive functional group (e.g., NHS ester, sulfonyl-NHS ester, tetrafluoro or pentafluorophenyl ester, isothiocyanate, sulfonyl chloride, etc.).

[0095] In other embodiments, the linker can be linked to the polymerase by connecting to a non-natural amino acid inserted in the form of a gene, such as p-propoxyphenylalanine or p-azidophenylalanine which can undergo azido-alkynyl-Huisgen cycloaddition, but there are many suitable non-natural amino acids suitable for site-specific labeling and can be found in the literature (e.g., as described in Lang and Chin., Chemical reviews 114.9(2014):4764-4806).

[0096] In other embodiments, the linker can be specifically attached to the N-terminus of the polymerase. In some embodiments, the polymerase is mutated to have an N-terminal serine or threonine residue, which can be specifically oxidized to produce an N-terminal aldehyde for subsequent coupling to, for example, an acylhydrazine. In other embodiments, the polymerase is mutated to have an N-terminal cysteine ​​residue, which can be specifically labeled with an aldehyde to form a thiazolidinyl ester. In other embodiments, the N-terminal cysteine ​​residue can be labeled with a peptide linker via a natural chemical linker.

[0097] In other embodiments, a peptide tag sequence may be inserted into the polymerase, the peptide tag sequence being specifically labeled by the enzyme with synthetic groups, for example using biotin ligase, transglutaminase, lipoic acid ligase, bacterial sorting enzyme, and phosphoproteopanthenylaminotransferase (e.g., as described in references 74-78 of Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).

[0098] In other embodiments, the adapter is attached to a labeling domain fused to the polymerase. For example, adapters with corresponding reactive moieties can be used to covalently label SNAP tags, CLIP tags, HaloTags, and acyl carrier protein domains (e.g., as described in references 79-82 of Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876–884).

[0099] In other embodiments, the linker is attached to an aldehyde specifically generated within the polymerase, as described by Carrico et al. (Nat. Chem. Biol. 3, (2007) 321–322). For example, after inserting an amino acid sequence recognized by the enzyme formylglycine synthase (FGE) into the polymerase, it can be exposed to FGE, which will specifically convert cysteine ​​residues in the recognized sequence to formylglycine (i.e., generate an aldehyde). The aldehyde can then be specifically labeled with, for example, an acylhydrazine or aminooxy group of the linker.

[0100] In some implementations, the linker can be attached to the polymerase via non-covalent binding of the linker portion to a portion fused to the polymerase. Examples of such ligation strategies include fusing the polymerase to streptavidin, which can bind the biotin portion of the linker, or fusing the polymerase to anti-digitoxin, which can bind the digitoxin portion of the linker.

[0101] In some implementations, site-specific labeling can result in reversible linker ligation to polymerase (e.g., the o-pyridyl disulfide (OPSS) group that forms a disulfide bond with cysteine, which can be cleaved using a reducing agent, such as TCEP), while other linker chemicals will produce permanent ligation.

[0102] In any implementation, the polymerase can be mutated to ensure that the tethered nucleotides are specifically linked to a particular site on the polymerase, as will be apparent to those skilled in the art. For example, for thiol-specific linking chemicals such as maleimide or o-pyridyl disulfide, accessible cysteine ​​residues in the wild-type polymerase can be mutated to non-cysteine ​​residues to prevent labeling at those sites. In this context of "non-reactive cysteines," cysteine ​​residues can be introduced by mutation at the desired linking site. These mutations preferentially do not interfere with the activity of the polymerase.

[0103] Other strategies for site-specific ligation of synthetic groups to proteins are readily apparent to those skilled in the art and have been reviewed in the literature (e.g., Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876–884).

[0104] The strategy of connecting the adapter to nucleoside triphosphates.

[0105] In some embodiments, the linker is attached at the 5-position of pyrimidine or the 7-position of 7-denitropurine. In other embodiments, the linker may be attached to an exocyclic amine of the nucleobase, for example, by N-alkylating the exocyclic amine of cytosine with a nitrobenzyl moiety discussed below. In other embodiments, the linker may be attached to any other atom in the nucleobase, sugar, or α-phosphate, as will be apparent to those skilled in the art.

[0106] Some polymerases are highly tolerant to modifications of certain parts of nucleotides; for example, some polymerases are well tolerant to modifications at the 5-position of pyrimidines and the 7-position of purines (He and Seela. Nucleic Acids Research 30.24(2002):5485-5496; or Hottin et al., Chemistry. 2017 Feb 10; 23(9):2109-2118). In some embodiments, the linker is attached to these positions.

[0107] In some instances, polymerase-nucleotide conjugates are prepared by first synthesizing an intermediate compound containing a linker and a nucleoside triphosphate (referred to herein as a "linker-nucleotide"), and then linking this intermediate compound to a polymerase. In some instances, nucleosides with substitutions compared to natural nucleosides, such as pyrimidines with 5-hydroxymethyl or 5-propyne amino substituents, or 7-denitropurines with 7-hydroxymethyl or 7-propyne amino substituents, can be useful starting materials for the preparation of linker-nucleotides. An exemplary set of nucleosides with 5- and 7-hydroxymethyl substituents that can be used to prepare linker-nucleotides is shown below:

[0108]

[0109] An exemplary set of nucleosides having 5- and 7-deazo-7-propyne amino substituents that can be used to prepare linker nucleotides are shown below:

[0110]

[0111] These nucleosides are also commercially available as deoxyribonucleoside triphosphates.

[0112] In Example 2, a linker nucleotide containing a 3-(((2-nitrobenzyl)oxy)carbonyl)aminopropynyl group linked to the 5- and 7-propynylamino substituents was prepared by reacting a nucleoside triphosphate containing 5- and 7-propynylamino substituents with a precursor molecule containing nitrobenzyl NHS carbonate (as shown in Figure 4).

[0113] Connector Disintegration Strategy

[0114] As mentioned above, linkers can be attached to various positions on a nucleotide and can be used using a variety of cleavage strategies. These strategies may include, but are not limited to, the following examples:

[0115] In some embodiments, the linker can be cleaved by exposure to a reducing agent such as dithiothreitol (DTT). For example, a linker containing a 4-(dithioalkyl)butyryloxy-methyl group linked to the 5-position of a pyrimidine or the 7-position of a 7-deazopurine can be cleaved by a reducing agent (e.g., DTT) to produce a 4-mercaptobutyryloxymethyl scar on the nucleobase. This scar can undergo intramolecular thiolactoneation to eliminate the 2-oxothiopental ring, leaving a smaller hydroxymethyl scar on the nucleobase. An example of such a linker linked to the 5-position of a cytosine is shown below, but this strategy is applicable to any suitable nucleobase:

[0116]

[0117] In other embodiments, the linker can be cleaved by exposure to light. For example, a linker containing (2-nitrobenzyl)oxymethyl can be cleaved with 365 nm light, leaving a hydroxymethyl scar, as shown below for cytosine, but applicable to any suitable nucleobase:

[0118]

[0119] (Where, for example, R” = H or R” = CH3 or R” = t-Bu.)

[0120] In other embodiments, the linker may comprise a 3-(((2-nitrobenzyl)oxy)carbonyl)aminopropynyl group, which can be photolyzed at 365 nm to release the nucleobase with the propynyl amino scar. This strategy was used in Example 2 and is shown below for cytosine, but is applicable to any suitable nucleobase:

[0121]

[0122] In other embodiments, the linker may contain an acyloxymethyl group, which can be cleaved with a suitable esterase to release a nucleobase with a hydroxymethyl scar, for example, as shown below for cytosine, but applicable to any suitable nucleobase:

[0123]

[0124] In such an implementation, the linker may contain additional atoms adjacent to the ester (included in R' above), which increases the activity of the esterase for the ester bond.

[0125] In other embodiments, the linker may comprise an N-acyl-aminopropynyl group, which can be cleaved by a peptidase to release a nucleobase with a propynyl amino scar, for example, as described below for 5-propynylaminocytosine, but applicable to any suitable nucleobase:

[0126]

[0127] In such embodiments, the linker may contain additional atoms adjacent to the amide (included in R' above), which increase the peptidase's activity against the amide bond. In some embodiments, R' is a peptide or polypeptide.

[0128] Cleavage peptide bonds to separate tethered nucleotides

[0129] In some embodiments, one or more amino acids are inserted into the polymerase, and these amino acids can act as portions of a cleavable linker for the specific ligation of nucleotides. In this case, the linker contains the inserted amino acid, and the cleavable bond is considered to be one or more bonds of the inserted amino acid (multiple inserted amino acids). For example, peptide bonds can be cleaved using a peptidase (these terms are used interchangeably to refer to an enzyme with activity 3.4 using the IUBMB nomenclature), such as proteinase K (EC 3.4.21.64, using the IUBMB nomenclature).

[0130] In some implementations, the protein itself can be cleaved to separate the tethered nucleoside triphosphates from the polymerase. For example, a peptidase cleavage linker can be used to cleave peptide bonds before and / or after the cleavage site.

[0131] In some embodiments, the amino acid position near the linker can be mutated to ensure that the peptide sequence near the linker is a good substrate for the protease, as will be apparent to those skilled in the art. For example, mutations can be introduced into aliphatic amino acids such as leucine or phenylalanine to achieve rapid cleavage with proteinase K.

[0132] RTdNTP-polymerase conjugate

[0133] In some embodiments, the tethered nucleotide analogue is such that, when it is available free in solution, incorporation does not terminate DNA synthesis. However, in other embodiments, the tethering is performed on a nucleotide analogue containing a reversible terminator group, such as an O-azidomethyl or O-NH2 group at the 3' position of a sugar, or a (α-tert-butyl-2-nitrobenzyl)oxymethyl group at the 5' position of a pyrimidine or the 7' position of a 7-dehydropurine (see, for example, Chen et al., Genomics, Proteomics & Bioinformatics 2013 11:34–40 for an overview). In these embodiments, the nucleotide analogue prevents or hinders further elongation once incorporated into the nucleic acid, and thus may contribute to the ability of the conjugate to achieve termination in addition to other termination-contributing properties (e.g., shielding). In cases where RTdNTP-polymerase conjugates do not rely on shielding to achieve termination, for example, when a 3'-modified RTdNTP is tethered to the polymerase, the linker length used can exceed [a certain value]. or

[0134] Shielding effect of polymerase-nucleotide conjugates

[0135] When incubating a conjugate containing a polymerase and a nucleoside triphosphate with nucleic acids, it is preferable to use the tethered nucleotides to extend the nucleic acid (as opposed to using nucleotides of another conjugate molecule). As described above, the polymerase then maintains its connection with the nucleic acid through its tethering with the added nucleotides (e.g., Figure 3D and Figure 4D The polymerase-nucleotide conjugate continues to extend until it is exposed to a stimulus that causes the bond of the added nucleotide to be broken. In this case, further extension of the polymerase-nucleotide conjugate is hindered by “shielding” when: 1) the conjugated polymerase molecule prevents other conjugates from approaching the 3'OH of the extended DNA molecule, and 2) other nucleoside triphosphates in the system prevent other conjugates from approaching the catalytic site of the polymerase that holds the conjugate to the 3' end of the extended nucleic acid. (The degree of shielding can be described as the extent to which these two interactions are hindered.) To allow for subsequent extension, the conjugated nucleotide tethered to the polymerase linker can be cleaved, releasing the polymerase from the nucleic acid and thus re-exposing its 3'OH group for subsequent extension (e.g., as shown in the image). Figure 3E and Figure 4E (As shown).

[0136] The nucleic acid synthesis and sequencing method described herein, which utilizes shielding to achieve termination, includes an extension step in which nucleic acids are preferentially exposed to the conjugate in the absence of free (i.e., untethered) nucleoside triphosphates, because the shielding termination mechanism may not prevent their incorporation into the nucleic acid. As shown in Example 3, exposing primers that have already been extended by the TdT-dCTP conjugate to free dCTP results in several additional extensions.

[0137] In some implementations, termination of further extension can be “complete,” meaning that no further extension can occur during the reaction after the nucleic acid molecule has been extended by the conjugate. In other implementations, termination of further extension can be “incomplete,” meaning that further extension can occur during the reaction, but at a significantly reduced rate compared to the initial extension, such as 100-fold, 1000-fold, 10,000-fold, or more. When the reaction is stopped after an appropriate time, the conjugate achieving incomplete termination can still be used to extend nucleic acids primarily by single nucleotides (e.g., in nucleic acid synthesis and sequencing methods).

[0138] In some implementations, the reagent containing the conjugate may additionally contain a polymerase of nucleoside triphosphates without tethering, but those polymerases should not significantly affect the reaction because there are no free dNTPs in the mixture.

[0139] The reagents for conjugates that terminate using shielding mechanisms preferably contain only polymerase-nucleotide conjugates, wherein all polymerases are folded in their active conformation. In some cases, if the polymerase moiety of the conjugate is unfolded, its tethered nucleoside triphosphates can become more readily accessible to the polymerase moiety of other conjugate molecules. In these cases, the unshielded nucleotides can be more easily incorporated into other conjugate molecules, circumventing the termination mechanism.

[0140] Polymerase-nucleotide conjugates terminated using shielding are preferably labeled with only a single nucleoside triphosphate moiety. In some cases, polymerase-nucleotide conjugates labeled with multiple nucleoside triphosphates that are accessible to the catalytic site can incorporate multiple nucleoside triphosphates into the same nucleic acid (e.g., as shown in the conjugate of wt TdT labeled with up to 5 nucleoside triphosphates in Example 1). Therefore, additional tethered nucleotides during the reaction may lead to the incorporation of other unwanted nucleotides into the nucleic acid. Furthermore, since only one tethered nucleoside triphosphate can occupy its (buried) catalytic site of the polymerase at a time, other tethered nucleoside triphosphates (multiple tethered nucleoside triphosphates) can have increased accessibility to the polymerase moiety of other conjugate molecules, as described below. The strategy of specifically tethering at most one nucleoside triphosphate site to the polymerase has been described above.

[0141] Polymerase-nucleotide conjugates that utilize shielding for termination preferably contain the shortest possible linker, which still allows the nucleoside triphosphate to frequently approach the catalytic site of the polymerase molecule it is linked to in order to achieve rapid nucleotide incorporation into nucleic acids. Such conjugates may also preferably use a linker as close as possible to the catalytic site of the polymerase, allowing for the use of even shorter linkers. The length of the linker will determine the maximum distance that can be reached from the linker point of the tethered nucleoside triphosphate or tethered nucleic acid. A smaller distance can lead to reduced accessibility of the tethered portion to other polymerase-nucleotide molecules, as described below. The linkers used in Examples 1 and 2 are approximately 24 and... Shorter connectors, for example, those with a length of Longer connectors can increase shielding; longer connectors, for example, those longer than [a certain length], can also increase shielding. or The connector can reduce shielding.

[0142] The shielding effect may be affected by a combination of factors, including but not limited to the polymerase structure, the length of the linker, the linker structure, the linker-polymerase connection site, the binding affinity of nucleoside triphosphates to the polymerase catalytic site, the binding affinity of nucleic acids to the polymerase, the preferred conformation of the polymerase, and the preferred conformation of the linker.

[0143] One contributing factor to shielding can be steric effects, which block the 3'OH of a nucleic acid extended by a conjugate from reaching the catalytic site of the polymerase moiety of another conjugate. Steric effects can also impede the reach of tethered nucleoside triphosphates to the catalytic site of another polymerase-nucleotide conjugate molecule due to collisions that will occur between conjugates during this process. If these steric effects completely block the productive interaction between the tethered nucleoside triphosphate (or extended nucleic acid) of one conjugate molecule and another conjugate molecule, it can lead to complete termination; if these steric effects only impede this intermolecular interaction, it can lead to incomplete termination.

[0144] Another contributing factor to shielding stems from the binding affinity of the tethered nucleoside triphosphate (NPPT) to the catalytic site of the polymerase. The NPPT of the conjugate has a high effective concentration at its tethered polymerase catalytic site, thus it remains bound to that site for most of the time. When the NPPT binds to the catalytic site of its tethered polymerase molecule, it cannot be used for incorporation by other polymerase molecules. Therefore, tethering reduces the effective concentration of NPPT available for intermolecular incorporation (i.e., incorporation catalyzed by polymerase molecules of untethered nucleotides). This shielding effect can be enhanced and terminated by reducing the rate of elongation of nucleic acids, which utilizes the NP portion of one conjugate molecule through the polymerase portion of another conjugate molecule.

[0145] Another contributing factor to shielding stems from the binding affinity of the 3' region of the nucleic acid molecule to the catalytic site of the polymerase molecule. After extension via a conjugate, the nucleic acid is tethered to the conjugate via its 3' terminal nucleotide and has a high effective concentration at the tethered polymerase catalytic site, thus maintaining its binding to that site for most of the time. When the nucleic acid is bound to the catalytic site of its tethered polymerase molecule, it cannot be used for extension via other conjugate molecules. This effect can enhance termination by reducing the rate at which nucleic acid already extended by the first conjugate is further extended by other conjugate molecules.

[0146] Add elements with space constraints to increase shielding effectiveness.

[0147] In some embodiments, the polymerase-nucleotide conjugate includes an additional portion that spatially prevents the tethered nucleoside triphosphate (or extended linked nucleic acid) from approaching the catalytic site of another conjugate molecule. Such portions include polypeptide or protein domains that can be inserted into the polymerase ring, as well as polymers that can be site-specifically linked to, for example, inserted non-natural amino acids or specific polypeptide tags.

[0148] The combination of RTdNTP termination mechanism and shielding effect.

[0149] As described above, in some embodiments, the conjugate may comprise a polymerase and a tethered reversible terminator nucleoside triphosphate. Some RTdNTPs (particularly 3'O-unblocked RTdNTPs) achieve incomplete termination when used in free form in solution. Polymerase-nucleotide conjugates of such RTdNTPs, which also employ shielding, can achieve more complete termination than the RTdNTPs used on their own. In some embodiments, RTdNTPs with a (α-tert-butyl-2-nitrobenzyl)oxymethyl group linked to the 5-position of pyrimidine or the 7-position of 7-dehydropurine in the nucleoside triphosphate may be used, for example, as described in Gardner et al. (Nucleic Acids Res. 2012 Aug; 40(15): 7404-1) or Stupi et al. (Angewandte Chemie International Edition 51.7(2012): 1724-1727). In some embodiments, the linker is attached to an atom in the termination portion of the RTdNTP. In other embodiments, the connector is atomically connected to RTdNTPs that are not in the termination portion.

[0150] Effective concentration of nucleoside triphosphates tethered to other polymerase-nucleotide conjugates

[0151] As previously mentioned, tethered nucleoside triphosphates exhibit high effective concentrations at the catalytic sites of the polymerase moieties they are linked to, enabling rapid incorporation. For the catalytic sites of other polymerase moieties, the same nucleoside triphosphates have much lower concentrations, resulting in slower intermolecular nucleotide incorporation rates if intermolecular incorporation is possible. The effective concentration of tethered nucleoside triphosphates with other conjugates is at most the absolute concentration of the conjugate, since each conjugate molecule contains a single nucleotide. The effective concentration is further reduced due to shielding effects that impede the accessibility of these nucleoside triphosphates.

[0152] To prevent extended nucleic acids from translocating at the polymerase catalytic site

[0153] Further termination by the polymerase-nucleotide conjugate can be achieved by selecting a linker and a binding site that prevents the extended (and therefore tethered) nucleic acid from moving its tethered 3' end to a position where its 3'OH can be activated to attack the incoming nucleoside triphosphate. This effect can be achieved if the tethered nucleoside triphosphate can approach the polymerase's nucleoside triphosphate binding site but cannot reach the position corresponding to its 3'OH of the incoming nucleic acid.

[0154] Application of polymerase-nucleotide conjugates in de novo nucleic acid synthesis

[0155] This article describes a method for de novo synthesis of nucleic acids using a conjugate comprising a polymerase and a nucleoside triphosphate. In some embodiments of this method, the conjugate comprises a polymerase terminal deoxynucleotidyl transferase (TdT). In other embodiments, the method may use a conjugate comprising another template-independent or template-dependent polymerase.

[0156] Figure 1A This describes a typical method for stepwise synthesis of a defined sequence using a template-independent polymerase. The nucleic acid (i.e., the “starting molecule”) that serves as the initial substrate for extension is incubated with a first polymerase-nucleotide conjugate. Once the nucleic acid is extended by the tethered nucleotides of the conjugate, no further extension occurs because the conjugate provides a termination mechanism, for example, based on a shielding effect. In the second step of the method, the linker is cleaved to release the polymerase and reverse the termination mechanism, thereby enabling subsequent extension. The extension product is then exposed to a second conjugate, and these two steps are repeated to extend the nucleic acid with the defined sequence. Figure 1B The process of synthesizing a compound containing TdT and a photolytically spliable connector, as practiced in Example 2, is described. As mentioned above, other strategies can be used for connector joining and splitting.

[0157] For DNA synthesis applications, template-independent polymerases, specifically terminal deoxynucleotidyl transferases or DNA extranucleotidyl transferases, can be used. These terms are used interchangeably to refer to enzymes with activity 2.7.7.31. Polymerases capable of extending single-stranded nucleic acids include, but are not limited to, polymerase θ (Kent et al., Elife 5(2016):e13740), polymerase μ (Juarez et al., Nucleic acids research 34.16(2006):4572-4582; or McElhinny et al., Molecular cell 19.3(2005):357-366), or polymerases that, for example, induce template-independent activity by inserting elements of a template-independent polymerase (Juarez et al., Nucleic acids research 34.16(2006):4572-4582). In other DNA synthesis applications, the polymerase can be a template-dependent polymerase, i.e., a DNA-directed DNA polymerase (these terms are used interchangeably to refer to enzymes with activity 2.7.7.7 using the IUBMB nomenclature).

[0158] For RNA synthesis applications, tethered ribonucleoside triphosphates can be used. In these embodiments, RNA-specific nucleotransferases, such as *E. coli* poly(A) polymerase (IUBMB EC 2.7.7.19) or poly(U) polymerase, are particularly useful. RNA nucleotransferases may contain modifications, such as single-point mutations, that affect substrate specificity to a particular rNTP (Lunde et al., *Nucleic Acids Research* 40.19 (2012): 9815-9824.). In some embodiments, a very short tether between the RNA nucleotransferase and the ribonucleoside triphosphate can be used to induce high effective concentrations of the nucleoside triphosphate, thereby forcing the incorporation of rNTPs that may not be native substrates of the nucleotransferase.

[0159] Initial substrates used for de novo synthesis of nucleic acids.

[0160] Nucleic acid synthesis protocols using polymerase-tethered nucleoside triphosphates may require a nucleic acid substrate (or a molecule with similar properties) of at least 3-5 bases to initiate synthesis. This initial substrate nucleotide can then be nucleotide-extended to the desired product. In some embodiments, the initial substrate may be an oligonucleotide primer synthesized using a phosphoramidite method (as shown in Example 1). In some cases, a specific sequence of the initiating primer can be used for downstream applications of the synthesized nucleic acid. In some embodiments, the sequence of the initial substrate can be removed from the synthesized nucleic acid after synthesis, particularly if the initial substrate contains a cleavable bond near its 3' end. For example, if the initial substrate is a primer with a 3' deoxyuridine base, exposing the extended primer to a USER enzyme (i.e., a mixture of uracil DNA glycosidase and endonuclease VIII) will cleave the synthesized sequence from the initial substrate. However, other cleavable bonds can be used, such as bridging phosphate esters in primers that can be cleaved with silver or mercury ions.

[0161] In some implementations, a double-stranded DNA molecule can be used to initiate synthesis, particularly if it has a 3' overhang (as shown in Example 2). If the initial substrate is a linearized plasmid backbone, a DNA synthesis method can be used to extend the DNA molecule through one or more synthetic gene sequences, and the extended DNA can then be (optionally amplified and) circularized into a plasmid. Generally, the methods for nucleic acid synthesis described herein enable the de novo synthesis of DNA from native nucleic acid molecules; conversely, it is not possible to directly extend native nucleic acid molecules through a defined sequence using the phosphoramidite method.

[0162] The strategy of connecting adapters to nucleoside triphosphates that can be used for nucleic acid synthesis.

[0163] In some embodiments, cleavage of the linker attached to the nucleotide can result in the production of the native nucleotide upon cleavage. For example, the linker of the nitrobenzyl moiety of an amine containing an alkylated nucleobase can be photocleaved, for example as described below for exocyclic amines of cytosine and adenine, but applicable to any suitable nitrogen atom on any nucleobase:

[0164]

[0165] In other embodiments, the linker to the amine on the nucleobase via an amide bond can be cleaved by a suitable peptidase, for example as described below for exocyclic amines of cytosine and adenine, but applicable to any suitable amino group on any nucleobase:

[0166]

[0167] In such embodiments, the linker may contain additional atoms adjacent to the amide (included in R' above), which increase the peptidase's activity against the amide bond. In some embodiments, R' is a peptide or polypeptide.

[0168] In some implementations, cleavage of the linker during the deprotection step can leave scars that persist throughout the stepwise synthesis but are removed or further reduced once the stepwise synthesis is complete. This ligation strategy allows for the introduction of additional distance between the cleavable portion of the linker and the nucleotide, which can be used with certain (e.g., bulky) cleavable groups. Scarring can be used to prevent base pairing of the incorporated nucleotide, thereby preventing the formation of secondary structures during synthesis, for example, by preventing exocyclic amino groups from participating in base pairing, as discussed below. Once the synthesis of such a molecule is complete, a single “scar removal” step can be used to prevent the scars from interfering with downstream applications and to restore the base pairing ability of the nucleic acid.

[0169] For example, acyl scars left on the exocyclic amino groups of adenine, cytosine, and guanine after cleavage may hinder the formation of certain types of secondary structures during synthesis. Such scars can be quantitatively removed after synthesis using mild ammonia treatment (Schulhof et al., Nucleic Acids Research. 1987; 15(2):397-416), as follows:

[0170]

[0171] Examples of linkers using these groups are linkers containing a 2-((4-(dithioalkyl)butyryl)oxy)acetyl group to an exocyclic amino group containing a nucleobase (forming an amide), as described below for cytosine and adenine:

[0172]

[0173] The cleavage of disulfides (e.g., via DTT) can lead to the elimination of the 2-oxothiopental ring via intramolecular thiolactoneation, leaving a hydroxyacetyl scar. Another example of this strategy is attaching a linker containing a photocleavable group (e.g., NPPOC) to the hydroxyacetyl group of an exocyclic amino group of the nucleobase, as described below for cytosine and adenine:

[0174]

[0175] After photolysis of the linker, the bases still contain hydroxyacetyl (acyl) scars, which cannot be terminated by other nucleotide-polymerase conjugates for further elongation, but can eventually be removed by treatment with ammonia, as shown below:

[0176]

[0177] Other strategies for linking nucleoside triphosphates to polymerases following the principles described above can be used, and this will be apparent to those skilled in the art.

[0178] Template-independent polymerases can be highly tolerant to base modifications (e.g., for TdT, see Figeys et al. (Anal Chem. 1994 Dec 1; 66(23):4382-3) and Li et al. (Cytometry. 1995 Jun 1; 20(2):172-80.)), which makes them well tolerant to specific scarring in subsequent nucleic acid extension steps.

[0179] In some implementations, a connector with multiple cascaded cleavable groups can be used to increase the cleavage rate compared to a connector with a single cleavable group.

[0180] Reducing the inhibitory effect of the 3' terminal secondary structure

[0181] In some embodiments, modified nucleoside triphosphates with linked chemical moieties can be used, which prevent base pairing or the formation of undesirable secondary structures during synthesis. Such modifications can include, but are not limited to, N3-methylation of cytosine, N1-methylation of adenine, O6-methylation of guanine, and acetylation of exocyclic amines of guanine. Similar modifications have been shown to significantly enhance the rate of synthesis of dGTP homopolymers using TdT (Lefler and Bollum, Journal of Biological Chemistry 244.3 (1969): 594-601.). After synthesis, such base modifications can be simultaneously removed to restore base pairing for downstream applications. For example, as described above, N-acetylation of guanine can be removed by ammonia treatment. N3-methylation of cytosine and N1-methylation of adenine can be removed by the enzyme AlkB, and O6-methylation of guanine can be removed by the enzyme MGMT, as follows:

[0182]

[0183] In some implementations, de novo nucleic acid synthesis can be paused at an intermediate step, and complementary DNA synthesis can proceed, for example, via hybridization with suitable primers (e.g., random hexamers) extended using a template-dependent polymerase with nucleoside triphosphates. After complementary DNA synthesis, de novo DNA synthesis can be resumed, and the formation of secondary structures from the double-stranded portion of the nucleic acid can be prevented. In some cases, the complementary DNA synthesis step may include leaving a 3' overhang on the de novo synthesized nucleic acid to enable efficient subsequent extension via a polymerase-nucleotide conjugate.

[0184] In some implementations, a single-stranded binding protein (e.g., E. coli SSB) may be included in the extension reaction to prevent the formation of secondary structures in the synthesized nucleic acid.

[0185] Incorporation of non-natural or modified dNTPs.

[0186] Conjugates that can be used for de novo nucleic acid synthesis may contain nucleoside triphosphate analogs, including nucleotides that do not have base-pairing ability (e.g., non-base nucleotide analogs) or nucleoside triphosphates that have base-pairing ability different from that of natural nucleotides, such as deoxyinosine or nitroindole nucleoside triphosphates.

[0187] Automated de novo nucleic acid synthesis using polymerase-nucleotide conjugates

[0188] In some embodiments, nucleic acid molecules are immobilized on a solid support during synthesis, which can be washed and exposed to cyclic enzymes and buffers by an automated liquid handling device. Examples of solid supports include, but are not limited to: microtiter plates into which reagents can be dispensed and removed by a liquid handling robot; magnetic beads that can be magnetically separated from a suspension and then resuspended in new reagents in a microtitration; or the inner surface of a microfluidic device that can automatically dispense cyclic reagents to that location.

[0189] Applications of automated nucleic acid synthesis systems using polymerase and nucleoside triphosphate conjugates include the synthesis of 10-100 nt oligonucleotides for molecular biology applications such as PCR. Another application is the picomolar or femtomolar synthesis of 50-500 nt or longer oligonucleotides using inkjet-based liquid processing technology to produce DNA molecules, for example, as inputs for conventional DNA assembly methods (Kosuri and Church, Nature methods 11.5 (2014):499-507).

[0190] Single-molecule nucleic acid synthesis using fluorescent polymerase-nucleotide conjugates

[0191] In some implementations, the DNA synthesis method can be carried out in single-molecule form. In this method, the reaction chamber of an automated microfluidic device is loaded with a single DNA primer molecule, which is repeatedly extended into the desired sequence using a modified form of the aforementioned reaction cycle. Figure 5A In this system, a conjugate containing polymerase and nucleoside triphosphate is labeled with one or more reporter molecules (e.g., fluorophores), such that once the labeled conjugate molecule has been tethered to its nucleotide extension primer, it is then ligated to a solid vector via the primer. Figure 5B The grown DNA molecules then become fluorescent. After washing away the free conjugate molecules, the polymerase linked to the primers can be detected, for example, using fluorescence microscopy techniques such as total internal reflection fluorescence (TIRF) microscopy. Figure 5C After each extension attempt, the reaction chamber is washed and imaged. If extension is determined to have failed, it is retried with the same type of conjugate. Upon confirmation of successful extension, a deprotection agent is introduced into the reaction chamber to lyse the tethered labeled polymerase, thereby deprotecting the 3' end of the grown DNA molecule for subsequent extensions and rendering it non-fluorescent. Figure 5D If deprotection fails, the fluorescence signal will be retained, and the deprotection step will be retried. Figure 5EThese extension and deprotection checks prevent the introduction of deletion errors, which inevitably accumulate during large-scale reactions that do not reach 100% completion. The automated synthesizer executing the protocol will ultimately synthesize a DNA molecule with the desired sequence, which can then be amplified. Figure 6 An example of such a synthesizer is depicted. The synthesizer includes a PDMS device with input ports for reagents, including ports for each polymerase-nucleotide conjugate, a port for a wash buffer, a port for a deprotection buffer, and a port for an in situ amplification buffer. The input ports are connected to the reaction chamber where the synthesis takes place via microchannels (and optionally, computer-driven microvalves). The device also includes waste ports and output ports (for collecting the synthesized products) connected to the reaction chamber via microchannels. The device can be mounted on a microscope suitable for single-molecule imaging, for example, Figure 6 The objective lens of the TIRF microscope shown can be used to excite the fluorophore attached to the conjugate using a laser of a suitable wavelength, such as 532 nm. The emitted light can be collected by the objective lens and imaged on a suitable detector, such as an electron multiplication charge-coupled device (EMCCD) camera connected to a computer. The computer can execute the above synthesis scheme by (a) interpreting the signal from the detector using an algorithm, and (b) dispensing suitable reagents into the reaction chamber by driving microvalves or pumps inside or outside the microfluidic device.

[0192] A method for DNA sequencing using polymerase-nucleotide conjugates.

[0193] This article presents a method for nucleic acid sequencing using a template-dependent polymerase and a nucleoside triphosphate conjugate. This method is similar to sequencing by synthesis (SBS).

[0194] In some embodiments, the method uses "ACGT extension reagents," which comprise four conjugates having base-pairing capabilities equivalent to A, C, G, and T, wherein said conjugates are labeled with distinguishable markers, such as different fluorophores. In other embodiments, the method uses "four different extension reagents" in different containers, each container containing conjugates having base-pairing capabilities respectively equivalent to A, C, G, and T. In some embodiments, these four extension reagents may be labeled with fluorophores, for example, for detection.

[0195] In some embodiments, the method includes: (a) immobilizing a duplex containing primers and template nucleic acids on a vector; (b) exposing the duplex to “ACGT extension reagent” to extend the primers by nucleotides complementary to the template; (c) detecting the labeling of the ligated conjugate to infer the complementary bases of the template; (d) exposing the duplex to a deprotecting agent that cleaves the bond between the polymerase and the added nucleotide, leaving the duplex unlabeled; and (e) repeating steps (bd) 10 or more times to determine at least a portion of the sequence of the template molecule. An example of this method using a polymerase-conjugate with four distinguishable fluorophores is depicted in [the following text is missing from the original]. Figures 7A-7E middle.

[0196] In other embodiments, the method includes: (a) immobilizing a duplex containing primers and template nucleic acids on a vector; (b) exposing the duplex to a first extension reagent and extending the primers by nucleotides complementary to the template if the nucleotides of the conjugate are complementary; (c) detecting a label on the conjugate to infer whether extension has occurred; (d) exposing the nucleic acid to a deprotection reagent that cleaves the bond between the polymerase and the added nucleotide, leaving the duplex unlabeled; (e) repeating steps (bd) three more times with the remaining three extension reagents; and (f) repeating step (be) 10 or more times to determine at least a portion of the sequence of the template molecule.

[0197] In some embodiments, the detectable marker may be a fluorescent protein fused to the polymerase. In other embodiments, the detectable marker may be a quantum dot specifically linked to the polymerase.

[0198] In particular, in embodiments using four different extension reagents, the conjugate may be non-fluorescent or unlabeled. In such embodiments, extension can be detected by other signals of the extension reaction, such as H+. + Alternatively, the release of pyrophosphate may be detected. In other embodiments, the polymerase may be fused with a reporter enzyme such as luciferase or peroxidase, which can be detected when they generate light through a catalytic reaction. In other embodiments, the polymerase may be fused with detectable nanoparticles via scattered light. In other embodiments, additional unlabeled polymerases that have already extended nucleic acid conjugates may be detected by changes in surface plasmon resonance.

[0199] In particular, in implementations using conjugates with detectable markers, this method can be used to determine the sequence of a single molecule, i.e., "single-molecule sequencing".

[0200] In some embodiments, the method may additionally include an initial step of preparing 10, 100, 1000 or more copies of the template molecule, and then simultaneously applying the above method to all copies.

[0201] In any implementation of the nucleic acid sequencing method, the linker length between the nucleoside triphosphate of the conjugate and the polymerase can be selected to maximize the fidelity of nucleotide incorporation through the conjugate while minimizing the incorporation of mismatched bases relative to the template. Similarly, in any implementation, the divalent cations (e.g., Mg) in the extension reaction can be tuned. 2+ The concentration of ) is adjusted to maximize the fidelity of nucleotide incorporation through the conjugate.

[0202] In some implementations, the polymerase of the conjugate can be a polymerase with a “random binding sequence,” meaning that the primer-template duplex can bind to a catalytic site before or after the nucleoside triphosphate.

[0203] In other embodiments, the polymerase of the conjugate can be a polymerase with a defined binding sequence, i.e., the primer-template duplex must bind to the catalytic site prior to the nucleoside triphosphate. In such embodiments, the linker length between the nucleoside triphosphate of the conjugate and the polymerase can be selected to minimize inhibition of binding of the primer-template duplex to the conjugate, i.e., by using a linker longer than […]. or or The connector.

[0204] In any implementation, the conjugate may contain a reversible terminator nucleoside triphosphate.

[0205] Implementation Plan

[0206] Implementation Scheme 1. A conjugate comprising a polymerase and a nucleoside triphosphate, wherein the polymerase and the nucleoside triphosphate are covalently linked by a linker comprising a cleavable bond.

[0207] Implementation Scheme 2. The conjugate described in Implementation Scheme 1, wherein the polymerase is capable of catalyzing the addition of a nucleotide linked to the polymerase to the 3' end of a nucleic acid.

[0208] Implementation Scheme 3. The combination of any of the foregoing embodiments, wherein the polymerase is transmitted through a length of... The linker within the range is connected to a nucleoside triphosphate, and the length of the linker is sufficient for the nucleoside triphosphate to approach the active site of the polymerase.

[0209] Implementation Scheme 4. The conjugate described in any of the preceding embodiments, wherein the nucleoside triphosphate is linked to a cysteine ​​residue in the polymerase.

[0210] Implementation Scheme 5. The combination of any of the preceding embodiments, wherein the cleavable bond is a photocleavable bond or an enzyme-cleavable bond.

[0211] Implementation Scheme 6. The conjugate described in any of the preceding embodiments, wherein the polymerase is a DNA polymerase.

[0212] Implementation Scheme 7. The conjugate of any one of Implementation Schemes 1-5, wherein the polymerase is an RNA polymerase.

[0213] Implementation Scheme 8. The conjugate described in any of the preceding implementation schemes, wherein the polymerase is a template-independent polymerase.

[0214] Implementation Scheme 9. The conjugate of any one of Implementation Schemes 1-7, wherein the polymerase is a template-dependent polymerase.

[0215] Implementation Scheme 10. The conjugate of any of the preceding embodiments, wherein the nucleoside triphosphate or polymerase comprises a fluorescent label.

[0216] Implementation Scheme 11. The conjugate described in any of the preceding implementation schemes, wherein the nucleoside triphosphate is a deoxyribonucleoside triphosphate.

[0217] Implementation Scheme 12. The conjugate described in any of the preceding implementation schemes, wherein the nucleoside triphosphate is ribonucleoside triphosphate.

[0218] Implementation Scheme 13. A set of combinations according to any of the preceding embodiments, wherein the combinations correspond to G, A, T and C and are in different containers.

[0219] Implementation Scheme 14. A method for synthesizing nucleic acids, comprising:

[0220] Nucleic acid is incubated with a first conjugate, wherein the first conjugate is a conjugate as described in any of the foregoing embodiments, and the incubation is carried out under conditions in which a polymerase catalyzes the covalent addition of a nucleotide of the first conjugate to the 3' hydroxyl group of the nucleic acid to prepare an extended product.

[0221] Implementation Scheme 15. The method described in Implementation Scheme 14, wherein the nucleic acid is embolized to the vector.

[0222] Implementation Scheme 16. The method of Implementation Scheme 14 or 15, wherein the method includes, after adding nucleotides to nucleic acids, cleaving the cleavable bonds of the linker, thereby releasing polymerase from the elongation product.

[0223] Implementation Scheme 17. The method of Implementation Scheme 16, wherein the cleavable bond is an enzyme-cleavable bond or a photocleavable bond, and the cleavage includes exposing the extended product to an enzyme or to light.

[0224] Implementation Scheme 18. The method of Implementation Scheme 16 or 17, wherein the cleavage of the cleavable bond causes the added nucleotide to be deprotected to produce a deprotected extended product.

[0225] Implementation Scheme 19. The method of Implementation Scheme 18 further includes, after the added nucleotide deprotection:

[0226] The deprotected extension product is incubated with a second conjugate, wherein the second conjugate is the conjugate described in any one of embodiments 1-12, and the incubation is carried out under conditions in which a polymerase catalyzes the covalent addition of a nucleotide of the second conjugate to the 3' end of the deprotected extension product.

[0227] Implementation Scheme 20. The method of any one of Implementation Schemes 14-19, wherein the method comprises:

[0228] (a) Under conditions in which the polymerase catalyzes the covalent addition of the nucleotide of the first conjugate to the 3' hydroxyl group of the nucleic acid, the nucleic acid is incubated with the first conjugate of any one of embodiments 1-12 to prepare an extended product;

[0229] (b) Cleavage the cleavable bonds of the linker, thereby releasing the polymerase from the extension product and deprotecting the extension product;

[0230] (c) Under conditions in which polymerase catalyzes the covalent addition of the nucleotide of the second conjugate to the 3' end of the elongation product, the deprotected elongation product is incubated with the second conjugate of any one of embodiments 1-12 to prepare the second elongation product;

[0231] (d) Repeat steps (b)-(c) multiple times with the second extension product to produce extended nucleic acids with defined sequences.

[0232] Implementation Scheme 21. The method of Implementation Scheme 20, wherein the nucleotide is a reversible terminator, and wherein deprotection of the elongation product comprises removing the blocking group of the reversible terminator.

[0233] Implementation Scheme 22. The method of any one of Implementation Schemes 14-21, wherein the nucleic acid is an oligonucleotide.

[0234] Implementation Scheme 23. A sequencing method, comprising:

[0235] Incubate a bistrand containing primers and a template with a composition comprising a set of conjugates as described in embodiment 13, wherein the conjugates correspond to G, A, T (or U) and C and are distinguishably labeled;

[0236] Which nucleotide has been added to the primer can be detected by detecting the marker attached to the polymerase that has tethered the primer;

[0237] The extended product is deprotected by pyrolyzing the joint; and

[0238] Repeat the incubation, detection, and deprotection steps to obtain at least a portion of the template sequence.

[0239] Implementation Scheme 24. The method described in Implementation Scheme 23, wherein the sequencing method is a DNA sequencing method.

[0240] Implementation Scheme 25. The method described in Implementation Scheme 23, wherein the sequencing method is an RNA sequencing method.

[0241] Implementation Scheme 26. The method of any one of Implementation Schemes 23-25, wherein the nucleotide is a reversible terminator, and wherein the deprotection of the elongation product comprises the removal of the blocking group.

[0242] Implementation Scheme 27. A reagent kit, comprising:

[0243] Polymerase, which is modified to contain a single cysteine ​​residue on its surface; and

[0244] A group of nucleoside triphosphates, wherein each of the nucleoside triphosphates is attached to a thiol reactive group.

[0245] Implementation scheme 28. The reagent kit described in implementation scheme 27, wherein the nucleoside triphosphates correspond to G, A, T (or U) and C.

[0246] Implementation scheme 29. The reagent kit of any one of implementation schemes 27-28, wherein the nucleoside triphosphate is a reversible terminator.

[0247] Implementation Scheme 30. The reagent kit of any one of Implementation Schemes 27-29, wherein the nucleoside triphosphate comprises a length of [missing information]. Connectors within the specified range. Example

[0248] The various aspects of the teachings of this invention can be further understood from the following embodiments, which should not be construed as limiting the scope of the teachings of this invention in any way.

[0249] Example 1

[0250] Nucleotides are incorporated into the tethered nucleotides via polymerase-nucleotide conjugates using linkers that can be cleaved by reducing agents.

[0251] 1. Generation of polymerase (TdT) mutants with different ligation sites for the adapter.

[0252] The gBlock (Molecular Biotechnology 10.3 (1998): 199-208) encoding the rat TdT amino acid sequence used by Boulé et al. was ordered from IDT (Coralville, IA). This sequence was cloned into the pET19b vector, and the N-terminal his-tag of the vector was fused to the protein using isothermal assembly. QuickChange PCR was used to generate TdT mutants lacking all surface cysteine ​​residues, enabling the incorporation of the tethered dNTP (TdT5cysX). Cysteines at positions 188, 302, and 378 were mutated to alanine, and cysteines at positions 216 and 438 were mutated to serine (positions refer to their numbers in PDB structure 4I27). TdT mutants containing a surface cysteine ​​residue near the catalytic site were generated on TdT5cysX by QuickChange PCR. Cysteines were inserted at positions 188, 302, 180, or 253.

[0253] The following is the amino acid sequence (TdTwt) of the "wild-type" TdT protein used in this example prior to the cysteine ​​residue mutation:

[0254] MGHHHHHHHHHHSSGHIDDDDKHMSQYACQRRTTLNNHNQIFTDAFDILAENDEFRENEGPSLTFMRAASVLKSLPFTIISMKDIEGIPNLGDRVKSIIEEIIEDGESSAVKAVLNDERYKSFKLFTSVFGVGLKTSEKWFRMGFRTLSNIRSDKSLTFTRMQRAGFLYYEDLVSRVTRAEAEAVGVLVKEA VWASLPDAFVTMTGGFRRGKKTGHDVDFLITSPGATEEEEQQLLHKVISLWEHKGLLLYYDLVESTFEKLKLPSRKVDALDHFQKCFLILKLHHQRVDSDQSSWQEGKTWKAIRVDLVVCPYERRAFALLGWTGSRQFERDLRRYATHERKMIIDNHALYDKTKRIFLEAESEEEIFAHLGLDYIEPWERNA

[0255] (SEQ ID NO:1).

[0256] As described above, plasmids encoding six TdT mutants with different cysteine ​​residues (and therefore different linker sites) were generated:

[0257] (a) Plasmid encoding “wild-type” TdT with 7 cysteine ​​residues (TdTwt)

[0258] (b) Plasmid (TdT5cysX) encoding a TdT mutant that lacks surface cysteine ​​and contains only two buried cysteine ​​residues.

[0259] (c) Four plasmids encoding TdT mutants with two buried cysteine ​​residues plus one exposed surface cysteine ​​residue at different positions (188, 180, 253 and 302) are referred to in this paper as TdTcys188, TdTcys180, TdTcys253 and TdTcys302.

[0260] 2. Protein expression and purification of mutants.

[0261] TdT expression was performed on Rosetta-gami B(DE3)pLysS cells (Novagen) in LB medium containing antibiotics that target all four cell resistance markers (Kan, Cmp, Tet, and Carb, introduced via the pET19 vector). 400 mL of expression culture (1 / 20 volume) was seeded with 50 mL of overnight culture. Cells were grown at 37°C with shaking at 200 rpm until their OD reached 0.6. IPTG was added to a final concentration of 0.5 mM, and expression was performed at 30°C for 12 h. Cells were harvested by centrifugation at 8000 G for 10 min and resuspended in 20 mL of buffer A (20 mM Tris-HCl, 0.5 M NaCl, pH 8.3) + 5 mM imidazole. Cell lysis was performed by sonication followed by centrifugation at 15000 G for 20 min. The supernatant was applied to a gravity column containing 1 mL of Ni-NTA agarose (Qiagen). The column was washed with 20 volumes of buffer A + 40 mM imidazole, and the bound protein was eluted with 4 mL of buffer A + 500 mM imidazole. The protein was concentrated to ~0.15 mL using a Vivaspin 20 column (MWCO 10 kDa, Sartorius), and then purified using Pur-A-Lyzer. TM Dialyze the Mini 12000 tubes (Sigma) of the dialysis kit overnight with 200 mL TdT stock buffer (100 mM NaCl, 200 mM K2HPO4, pH 7.5).

[0262] All six TdT mutants with different cysteine ​​residues were expressed and purified.

[0263] 3. The tethered nucleoside triphosphate linkage to polymerase

[0264] To prepare the TdT-dUTP conjugate, the linker-nucleotide OPSS-PEG4-aa-dUTP was first synthesized and then reacted with TdT. OPSS-PEG4-aa-dUTP was synthesized by reacting amino-allyl dUTP (aa-dUTP) with the heterobifunctional crosslinking agent PEG4-SPDP. Figure 3A The reaction, containing 12.5 mM aa-dUTP, 3 mM PEG4-SPDP crosslinking agent, and 125 mM sodium bicarbonate (pH 8.3), was carried out in an 8 μL volume at room temperature for 1 hour. The reaction was quenched for 10 minutes by adding 1 μL of PBS containing 100 mM glycine. The buffer was adjusted to OPSS-labeled conditions by adding 1 μL of 10x TdT stock buffer, followed by adding 40 μL of 1x TdT stock buffer containing 70–100 μg of purified protein. The reaction of linker-nucleotides to TdT was carried out at room temperature for 13 hours. Free (i.e., unlinked) linker-nucleotides were removed using the Capturem His-labeled Purification Miniprep Kit (Clonetech). Purification resulted in protein concentrations between 0.2 and 0.4 μg / μL. The reaction was carried out in Pur-A-Lyzer. TM Dialyze for 4 hours in 100 mL of 1x TdT reaction buffer (NEB) in a Mini 12000 dialysis kit tube.

[0265] The scheme for preparing the polymerase-nucleotide conjugate is shown in Figure 3. First, aminoallyl dUTP (aa-dUTP) is reacted with the heterobifunctional amine-thiol crosslinking agent PEG4-SPDP (Figure A) to form a thiol-reactive linker-nucleotide OPSS-PEG4-aa-dUTP (Figure B). Then, OPSS-PEG4-aa-dUTP can be used to specifically label TdT at surface cysteine ​​residues (Figure C) via disulfide bond formation.

[0266] All six TdT mutants with different cysteine ​​residues were exposed to OPSS-PEG4-aa-dUTP tethering reaction.

[0267] (a)TdTwt contains five surface cysteine ​​residues, which result in labeling with up to five OPSS-PEG4-aa-dUTP moieties.

[0268] (b)TdT5cysX contains only two buried cysteine ​​residues but no surface cysteine ​​residues, and may not achieve substantial labeling by OPSS-PEG4-aa-dUTP.

[0269] (c) TdTcys188, TdTcys180, TdTcys253, and TdTcys302 have a single surface cysteine ​​residue that can be labeled with a single OPSS-PEG4-aa-dUTP moiety (at residues 188, 180, 253, and 302, respectively). These TdT mutants also contain two buried cysteine ​​residues in TdT5cysX, which may not have been labeled.

[0270] 4. The generation of trapezoidal stripes in extended product standards.

[0271] Trapezoidal bands of the linked incorporation product standard were generated by incorporating free OPSS-PEG4-aa-dUTP with TdT. The synthesis of OPSS-PEG4-aa-dUTP was carried out by mixing 6 μL of 50 mM aa-dUTP, 5 μL of 180 mM PEG4-SPDP, 5 μL of 1 M NaHCO3, and 4 μL of ddH2O. The reaction was incubated at room temperature for 1 hour, and an additional 5 μL of 180 mM PEG4-SPDP was added. After 1 hour, the reaction was quenched with 5 μL of PBS containing 100 mM glycine. To achieve different incorporation amounts, six TdT incorporation reactions were performed using free OPSS-PEG4-aa-dUTP as a substrate. The reaction consisted of 1.5 μL of 10x NEB TdT reaction buffer, 1.5 μL of NEB TdT CoCl2, 1.5 μL of 10 μM 5'-FAM-labeled 35-meric dT-oligonucleotide (5'-FAM-dT(35)), 1 μL of 10 mM OPSS-PEG4-aa-dUTP, and 4.5 μL of ddH2O, with varying TdT concentrations (100, 50, 25, 12.5, 6.3, and 3.13 units of NEB TdT in 5 μL of 1x NEB reaction buffer). The reaction was carried out at 37 °C for 5 minutes and terminated by adding 0.3 mM EDTA. Before running ladder-like bands on a polyacrylamide gel, the reaction product was mixed with an equal volume of 2x Novex Tris-glycine buffer + 1% v / v β-mercaptoethanol and heated to 95 °C for 5 minutes.

[0272] 5'-FAM-dT is extended by 0 to 5 or more OPSS-PEG4-aa-dUMP nucleotides. (60) Oligonucleotides. After reducing the ladder bands in the loaded dye, on a polyacrylamide gel (... Figure 8A The HS-PEG4-aa-dUTP extension product standard (the structure of an HS-PEG4-aa-dUTP extension product is drawn on lane B, labeled "L") was separated on lane B. Figure 3E(In Chinese). By comparing the migration of the extended product bands to the ladder bands, the ladder bands are used to identify the cleavage products of primer extension reactions with polymerase-nucleotide conjugates.

[0273] 5. Incorporate tethered nucleoside triphosphates into nucleic acids.

[0274] The OPSS-PEG4-aa-dUTP conjugates of TdTwt, TdT5cysX, TdTcys188, TdTcys180, TdTcys253 and TdTcys302 were reacted with 5'-FAM-dT(35). Figure 8A The reaction shown contained 1 μL of 5 μM 5'-FAM-dT(35), 17 μL of 1x TdT reaction buffer (NEB) containing the purified bound TdT variant, and 2 μL of 2.5 mM CoCl2. The reaction was carried out at 37 °C for 20 seconds, and then quenched by adding 33 mM EDTA. Figure 8B The reaction shown comprises 1.5 μL of 5 μM 5'-FAM-dT(35), 1 μL of 10x NEB reaction buffer, 1.5 μL of 2.5 mM CoCl2, 5 μL of 1x TdT buffer containing various TdT conjugates, and 6 μL of ddH2O. The reaction was carried out at 37 °C for 40 seconds and quenched by adding 33 mM EDTA. To prepare the reaction for gelation, the samples were mixed with equal volumes of 2x Novex Tris-glycine SDS sample buffer (ThermoScientific) or 2x Novex Tris-glycine buffer + 1% v / v 2-mercaptoethanol (BME). All samples were heated to 95 °C and held for 5 minutes before being run on SDS.

[0275] As shown in the chemical details in Figure 3, the TdT conjugate of OPSS-PEG4-aa-dUTP (Figure C) incorporates the tethered nucleotide into the primer, resulting in the TdT moiety being covalently linked to the extended primer (Figure D). For detection purposes, the primer can be labeled at its 5' end with a fluorescent dye such as 6-carboxyfluorescein (FAM). The primer-polymerase complex can be dissociated upon exposure to βME, which cleaves the disulfide bond between the incorporated nucleotide and TdT, releasing free TdT and the primer extended by dUMP containing the HS-PEG4-aa scar (Figure E).

[0276] To demonstrate that the polymerase-nucleotide conjugate adds its tethered nucleotide to the oligonucleotide, the TdT conjugate of OPSS-PEG4-aa-dUTP was incubated with a 5'FAM-labeled dT(35) primer. The reaction was carried out with a TdTwt conjugate having multiple tethered nucleotides, a TdT5cysX conjugate without tethered nucleotides, and a TdTcys302 conjugate having a single nucleotide tethered to a cysteine ​​residue at position 302. The reaction was stopped as described above, and the product was separated by SDS-PAGE. Figure 8A The TdTwt and TdTcys302 conjugates add their tethered nucleotides (multiple nucleotides) to the 3' end of the primers and become covalently linked, forming a polymerase-primer complex, as shown by the much slower migration of the bands on SDS-PAGE (lanes 3 and 11, respectively) compared to the primers (labeled "P": lanes 2, 6, and 10). Conversely, no change in migration was observed with TdT5cysX (lane 7), which does not contain the tethered nucleotides. Upon addition of the disulfide lysis reagent 2-mercaptoethanol, the primer-TdT complex dissociates, as shown by the migration through the recovery of the bands (labeled "B": lanes 4, 8, and 12 for TdTwt, TdT5cysX, and TdTcys302, respectively). Primer extension can be seen by referring to the trapezoidal bands of the product standards (labeled "L": lanes 1, 5, 9, and 13). The TdTwt conjugate incorporates multiple tethered nucleotides, producing a primer that extends through up to 5 scar dUMP nucleotides (lane 4). The TdTcys302 conjugate primarily extends the primer through a single scar dUMP nucleotide (lane 12), and the TdT5cysX conjugate does not extend the primer (lane 8).

[0277] These data show that the polymerase moiety of the polymerase-nucleotide conjugate can incorporate one or more tethered nucleoside triphosphates into the primer. They also show that conjugates labeled with a single nucleotide triphosphate can perform single primer extension and can prevent further primer extension by other polymerase-nucleotide conjugates, allowing nucleic acids to be extended by a single nucleotide.

[0278] To demonstrate that functional polymerase-nucleotide conjugates can be generated using multiple linker sites on the polymerase, OPSS-PEG4-aa-dUTP conjugates of TdT mutants with single surface cysteine ​​residues at positions 188, 302, 180, and 253 (TdTcys188, TdTcys302, TdTcys180, and TdTcys253, respectively) were used for primer extension via single nucleotides. The conjugates were exposed to 5' fluorescently labeled multi-dT primers. Products were isolated on SDS-PAGE. Figure 8BThe results showed that all four conjugates were able to incorporate the tethered nucleotide into the 3' end of the primer, as indicated by the formation of much slower migration bands corresponding to the polymerase-primer complex (lanes 4-7, respectively), compared to the primer bands (labeled "L": the fastest migration bands in lanes 1, 8, and 15). Upon cleavage of the linker by 2-mercaptoethanol (BME), the polymerase-primer complex dissociated, leaving primers that extended primarily through a single scarred dUMP nucleotide (lanes 11-14, respectively), as determined by comparison with the ladder-like bands of the product standards (lanes labeled "L").

[0279] These data suggest that the principle of tethering individual nucleotides to polymerases to achieve single nucleotide elongation of nucleic acids can be summarized by the linkage points on polymerases.

[0280] Example 2

[0281] Synthesize defined DNA sequences using polymerase-nucleotide conjugates employing photolytically cleavable linkers.

[0282] 1. Generation of MBP-TdT fusion protein with only one surface-exposed cysteine ​​residue (TdTcys).

[0283] The sequence encoding maltose-binding protein (MBP) was amplified from pMAL-c5X(NEB) and fused at the N-terminus to the TdTcys302 construct used in Example 1 using isothermal assembly. The resulting MBP-TdT fusion protein (referred to herein as TdTcys) was used throughout Example 2.

[0284] TdTcys protein sequence:

[0285] MGHHHHHHHHHHSSGHIDDDDKHMMKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQSGLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYENGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELVKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTNSSSNNNNNNNNNNLGIEGRISHMSMGGRDIVDGSEFSPSPVPGSQNVPAPAVKKISQYACQRRTTLNNYNQLFTDALDILAENDELRENEGSALAFMRASSVLKSLPFPITSMKDTEGIPSLGDKVKSIIEGIIEDGESSEAKAVLNDERYKSFKLFTSVFGVGLKTAEKWFRMGFRTLSKIQSDKSLRFTQMQKAGFLYYEDLVSCVNRPEAEAVSMLVKEAVVTFLPDALVTMTGGFRRGKMTGHDVDFLITSPEATEDEEQQLLHKVTDFWKQQGLLLYADILESTFEKFKQPSRKVDALDHFQKCFLILKLDHGRVHSEKSGQQEGKGWKAIRVDLVMSPYDRRAFALLGWTGSRQFERDLRRYATHERKMMLDNHALYDRTKRVFLEAESEEEIFAHLGLDYIEPWERNA

[0286] (SEQ ID NO:2).

[0287] 2. Protein expression and purification based on nickel affinity chromatography followed by anion exchange chromatography.

[0288] Unless otherwise stated, *E. coli* BL21(DE3) carrying pET19-TdTcys was grown in LB (Miller) containing 100 μg / mL carbenicillin with shaking at 200 RPM. In a 2 L flask without baffles, the overnight culture was diluted 1 / 60 to 400 mL of LB and grown at 37 °C until OD. 600 The concentration was adjusted to 0.40–0.45. The flasks were then cooled to room temperature for 45 minutes without shaking, followed by shaking at 15°C for 45 minutes. Protein expression was induced with IPTG (final concentration 1 mM), and cells were grown overnight at 15°C and harvested by centrifugation. All protein purification steps were performed at 4°C. Cells were lysed in buffer A (20 mM Tris-HCl, pH 8.3, 0.5 M NaCl) + 5 mM imidazole, and the lysates were subjected to nickel affinity chromatography using an imidazole gradient (buffer A + 5 mM imidazole to buffer A + 500 mM imidazole) (HisTrap FF 5 mL, GE Healthcare). Fractions of sufficient purity were combined and diluted 1:40 to 20 mM Tris-HCl, pH 8.3, and anion exchange chromatography was performed in 20 mM Tris-HCl using a gradient from 0 to 1 M NaCl (HiTrap Q HP 5 mL, GE Healthcare). Proteins were eluted in 200 mM NaCl. After adding 50% glycerol, the protein was stored at -20°C.

[0289] 3. Preparation of TdT-dNTP conjugates.

[0290] The propargyl amino-dNTP (pa-dNTP) was crosslinked with the photodegradable NHS carbonate-maleimide crosslinking agent BP-23354 at room temperature using a gentle vortex process. Figure 4AThe TdTcys-linker-dNTP conjugates (also known as TdT-dNTP conjugates) were coupled in 35 μL of a mixture containing 3.3 mM of various pa-dNTPs, 6.6 mM of linker, 66 mM of pH 7.5 KH₂PO₄, and 33 mM NaCl for 1 hour. The mixture was aliquoted into 7.5 μL aliquots, ground with ethyl acetate (~2 mL), and centrifuged at 15,000 g to form linker-nucleotide clusters. The supernatant was removed, and the clusters were dried in a Speedvac at room temperature for 8 minutes. The clusters were resuspended in 2.5 μL of water, and the linker-nucleotides were added to 20 μL of the TdTcys protein prepared as described above, along with 2.5 μL of pH 6.5 buffer (2 M KH₂PO₄, 1 M NaCl), and incubated at room temperature for 1 hour. The TdTcys-linker-dNTP conjugates were then purified using a 0.8 mL Pierce column with amylose resin (NEB). All reagents and buffers were pre-chilled on ice. 25 μL of the TdTcys linker-nucleotide binding reactant was diluted in 400 μL of buffer B (200 mM KH2PO4, pH 7.5, 100 mM NaCl) and aliquoted into two purification columns, each containing 250 μL of amylose resin in buffer B. After gently vortexing the protein for 10 min, the column was washed twice with buffer B, followed by two washes with 1x NEB TdT reaction buffer (50 mM potassium acetate, 20 mM Tris-acetate, 10 mM magnesium acetate, pH 7.9). The column was then washed with 500 μL of buffer, incubated on a shaker block for 1 min to mix the resin and buffer (800 RPM), and centrifuged at 50 g for 1 min. The resin was eluted twice by vortexing for 5 min in 150 μL of TdT reaction buffer + 10 mM maltose, followed by centrifugation. The eluents were combined and concentrated in a 30 kDa MWCO column, diluted 1:10 with TdT reaction buffer, and then concentrated to ~2.5 μg / μL.

[0291] Analogs of four nucleotides, dATP, dCTP, dGTP, and dTTP, were conjugated to photolytically cleavable cross-linking agents and tethered to TdT. Different polymerase-nucleotide conjugates were purified using amylose affinity chromatography.

[0292] 4. Capillary electrophoresis analysis and trapezoidal band generation.

[0293] Capillary electrophoresis (CE) analysis was performed throughout Example 2 on an ABI 3730xl DNA analyzer. A GeneScan Liz600 v1.0 (Thermo) was added to all samples as an internal size standard. Ladder bands (size standards) of the extended product of 5'-FAM-labeled 60-meric dT-oligonucleotides (5'-FAM-dT(60)) were generated by incorporating free pa-dNTPs with TdT. The reaction contained 100 nM 5'-FAM-dT(60), 100 μM of one type of pa-dNTP, 1x RBC, and 0.05 U / μL or 0.03 U / μL NEB TdT. The reaction was carried out at 37 °C. Aliquots were taken after 2, 5, and 10 minutes and quenched with EDTA to a final concentration of 33.3 mM. The quenched samples were then acetylated using NHS-acetate, purified Oligo Clean & Concentrator kit (“OCC”, Zymo Research) and analyzed by capillary electrophoresis. Samples with detectable peaks of 5'-FAM-dT(60) and +1 and +2pa-dNTP extensions were selected as size standards (trapezoidal bands).

[0294] 5. Demonstrate the two reaction cycles using PAGE and capillary electrophoresis.

[0295] Throughout the experiment, primer extension reactions were performed at 37°C for 2 minutes and quenched by adding an equal volume of 200 mM EDTA. All reactions contained 50 nM 5'-FAM-dT(60), 0.25 mg / mL TdTcys(-linker) / TdT-dCTP, and 1 x RBC (1 x NEB TdT reaction buffer, 0.25 mM cobalt). Photoinduced linker cleavage was performed on ice for 1 hour at 365 nm using a benchtop 2UV transilluminator (UVP, LLC). The measured irradiance was approximately 5 mW / cm². 2Dual-cycle experiments: Reactions containing TdT-dCTP conjugates and 5'-FAM-dT(60) were performed, and the products were photolyzed at 365 nm. The oligonucleotides (ZymoOCC) were then purified and reacted with TdT-dCTP again before the photolysis step. Aliquots were taken after two extension reactions (for PAGE) and two photolysis reactions (for PAGE and CE). For the control experiment ("unlinked control"), the TdT-dCTP conjugates were lysed on ice at 365 nm for 1 hour to produce an equimolar mixture of unlinked TdTcys(-linker) + pa-dCTP. The product was then reacted with 5'-FAM-dT(60). Aliquots were taken for PAGE and CE after quenching the reaction with EDTA. Sample preparation: All CE samples were acetylated with bicarbonate buffer containing 20 mM NHS-acetate prior to analysis. The samples were combined with SDS-loaded dyes and analyzed by PAGE to image the green fluorescence (5'FAM-labeled primers) of the gel, and the red fluorescence (total protein) after staining with Lumitein UV (Biotium).

[0296] Exposing 5'FAM-labeled oligonucleotide primers to TdT-dCTP conjugates forms a covalent complex containing DNA primers and proteins visible on SDS-PAGE. Figure 9A Irradiation of the complex with 365nm UV light cleaved the linker, thereby dissociating the complex and releasing the primer, which extends primarily through a single scar dCMP nucleotide. Figure 9B The product was exposed to fresh TdT-dCTP, and the primer-TdT complex was formed again, which was then dissociated by UV irradiation, releasing the primer that had now been extended by two nucleotides. Conversely, no primer-TdT complex formation was observed in the control reaction with TdT-dCTP before the addition of the DNA primer. Figure 9A Conversely, the control reaction produced a variety of primer extension products. Figure 9B This is consistent with the incorporation of free nucleotides catalyzed by TdT.

[0297] These data suggest that the process of extending a primer by a single nucleotide using a polymerase-nucleotide conjugate can be repeated to extend the primer by a defined sequence.

[0298] 6. Rapid single nucleotide incorporation via TdT-dCTP, TdT-dGTP, TdT-dTTP, and TdT-dATP.

[0299] Oligonucleotide elongation yields of the 1.5 mg / mL TdT-dNTP conjugate were measured at 8, 15, and 120 seconds. The reaction was carried out at 37 °C by adding 4.5 μL of the TdT-dNTP conjugate (2 mg / mL) to 1.5 μL of 5'-FAM-dT(60) (100 nM, final 25 nM) (both in 1x RBC). After rapid mixing, the 4.5 μL reactant was quenched in 18 μL of QS (94% HiDi formamide, 10 mM EDTA) after 8 or 15 seconds. The remaining reaction volume was quenched with 6 μL of QS after 2 minutes. All samples were irradiated at 365 nm for 30 minutes on a benchtop 2UV transilluminator (UVP, LLC) to cleave the linkers. The lysis products were diluted with washing buffer (0.67M NaH2PO4, 0.67M NaCl, 0.17M EDTA, pH 8) and captured onto DynaBeads M-280 streptavidin (Thermo) saturated with 5' biotinylated dA(60) oligonucleotides. After washing, the products were acetylated with bicarbonate buffer containing 100 mM NHS-acetate and eluted with 75% deionized formamide for CE.

[0300] Figure 10 CE data showed that primers were converted into single extended complexes within less than 20 seconds of incubation with TdT-dCTP, TdT-dGTP, TdT-dTTP, and TdT-dATP. These results demonstrate that polymerase-nucleotide conjugates can rapidly pass through a single nucleotide extension primer with excellent yield.

[0301] 7. Cyclic synthesis of defined DNA sequences.

[0302] Four repeated nucleic acid extensions and deprotections were performed using 0.25 mg / mL TdT-dNTP conjugate. The extension reaction was carried out at 37°C in 1x RBCs for 2 minutes, followed by quenching with an equal volume of quenching buffer (250 mM EDTA, 500 mM NaCl). Adapter cleavage was performed by irradiation at 365 nm. The first extension reaction consisted of TdT-dCTP and 50 nM oligoCl ( / 5Phos / UTGAAGAGCGAGAGTGAGTGA / iFluorT / CATTAAAGACGTGGGCCTGGAttt (SEQ ID NO:3), where / 5Phos / refers to 5' phosphorylation and / iFluorT / refers to the fluorescein-modified dT nucleotide base). After photolysis, the extension product (Zymo OCC) was purified, and the recovered DNA was used for the next extension step with TdT-dTTP. Two more cycles were performed using TdT-dATP, followed by TdT-dGTP, with the final product T-tailed using TdT and free dTTP+ddTTP at a 100:1 ratio. The tailed product was then amplified by PCR using HotStart Taq (NEB) with primers C2 (GTGCCGTGAGACCTGGCTCCTGACGATATGGATaagcttTGAAGAGCGAGAGTGAGTGA; SEQ ID NO:4) and C3 (AAAAgaattcAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA; SEQ ID NO:5) (PCR program: 98°C for 2 minutes, 49°C for 20 seconds, 68°C for 5 minutes, followed by 30 cycles of a three-step protocol: 98°C for 30 seconds, 49°C for 20 seconds, 68°C for 30 seconds). The PCR product was then inserted into the pUC19 plasmid using the EcoRI and HindIII sites introduced via the PCR primers. The plasmid was transformed into DH10B cells, grown overnight in LB cells, and single-colony plasmids were extracted and sequenced.

[0303] As described in detail above, the starting DNA molecule is subjected to four repeated reaction cycles, the tailing product is amplified by PCR, and the amplicon is cloned for sequencing. Figure 11A As shown. Of the 35 clones sequenced, 31 (89%) contained the complete 5'-CTAG-3' sequence (…). Figure 11B This means that the average step-by-step yield is 97%.

[0304] These data suggest that polymerase-nucleotide conjugates can be used in cyclic processes to write defined sequences of DNA with excellent stepwise yields.

[0305] Example 3

[0306] Incorporating free dNTPs into primers already tethered to polymerase

[0307] Experiments were performed using the TdT-dCTP conjugate prepared as described in Example 2. Capillary electrophoresis analysis was also performed as described in Example 2, using the same oligonucleotide ladder bands as a size reference. 5'-FAM-dT was incubated in 1x RBCs at 37°C with TdT-dCTP (0.25 mg / mL). 60 (50 nM) for 120 seconds to generate primer-polymerase complex, and aliquot the reaction mixture. One aliquot was quenched by adding EDTA to a final concentration of 100 mM. The other aliquot was diluted 10-fold with 1x RBC. The third aliquot was diluted 10-fold with 1x RBC containing free p-dCTP to a final concentration of 500 μM p-dCTP. After incubation at 37 °C for 60 seconds, both reactions were quenched by adding EDTA to a final concentration of 100 mM. All samples were subjected to photoinduced linker lysis for 1 hour on ice using a benchtop 2UV transilluminator (UVP, LLC) at 365 nm setting. Samples were acetylated using bicarbonate buffer containing 20 mM NHS-acetate, purified, and analyzed by CE.

[0308] During initial incubation of primers with TdT-dCTP, a primer-polymerase complex is formed, as observed by CE analysis, showing primer extension via a single nucleotide. Figure 12 No further extension of the primers was detected in the reaction, which was allowed to proceed for an additional 60 seconds. However, further extension of the tethered primers was observed in the reaction, which was allowed to proceed for an additional 60 seconds with the added free nucleoside triphosphate (pa-dCTP).

[0309] These results indicate that the nucleic acid-polymerase complex formed by tethering the polymerase-nucleotide conjugate nucleoside triphosphate can still incorporate free nucleoside triphosphate.

[0310] Example 4

[0311] Converting "scar" polynucleotides into natural DNA

[0312] Fluorescent primers were prepared using sodium bicarbonate buffer containing 4.5 mM NHS ester, followed by OCC-labeled amine oligonucleotides. Following the manufacturer's instructions (two-step protocol: 98°C denaturation for 10 seconds, 72°C annealing / extension for 1 minute), primers PA1 ( / 5Phos / cattaaagacgtgggcgtgga; SEQ ID NO:6) and PA2 (t*t*t / iUniAmM / tgtgaaatccttccctcgatcc; SEQ ID NO:7) were used with Phusion U (Thermo) to obtain a 639 nt DNA product containing deoxyuridine bases from the plasmid template pMal-c5x (NEB) via 35 cycles of PCR. Primer PA1 was 5' phosphorylated; primer PA2 contained an internal amino group ( / iUniAmM / ) labeled with Cy3 NHS (GE Healtcare) and began with two thiophosphate bonds (*), making it exonuclease resistant. PCR products were purified by 1% TAE-agarose gel electrophoresis, and ~6.7 μg of product was digested with 5 U λ exonuclease (NEB) in 100 μL of reaction at 37 °C for 20 min to separate the Cy3-labeled strand. The digested product was purified by OCC and then hybridized at ~1 μM with an equimolar amount of 5' FAM-labeled primer PA3 ( / 5AmMC6 / CAACACACCACCCACCCAACcgcagatgtccgctttctgg (SEQ ID NO:14); / 5AmMC6 / refers to the aminohexyl modification of the 5' phosphate) in 1x CutSmart buffer (NEB) by heating to 85 °C and cooling to 25 °C at 1 °C / min. N-acetylated propargylamino dNTPs were prepared by acetylation of 13 mM propargylamino dNTPs with 25 mM NHS acetate in 100 mM sodium bicarbonate buffer and quenched with glycine to a final concentration of 25 mM. Then, using 7.5 U of Klenow (NEB) at 37°C, primers were extended in 30 μL of reaction with 200 μM (each) N-acetylsalicylic acid dNTPs (reaction ii) or without any dNTPs (reaction i). After 1 hour, 3 μL of 2.5 mM (each) ddNTPs (Affymetrix) was added to both reactions and incubated for another 15 minutes, followed by polymerase inactivation by heating to 75°C for 20 minutes. The products were then immediately digested in 50 μL of reaction with 5 U of USER enzyme (NEB) at 37°C for 1 hour to remove dU-containing ssDNA templates. The digested products were purified by OCC, and USER digestion of the Cy3-labeled template with 5'FAM-labeled primers for propargyl amino dNTP-dependent extension and Cy3-labeled templates was confirmed by CE.The two products were then used as templates for complementary DNA synthesis (“read”), which was performed in 20 μL of 5 U Taq (Thermo) using 200 μM (each) of native dNTPs and 200 nM of 5'Cy3-labeled primer PA4 ( / 5AmMC6 / CGACTCACCTCACGTCCTCAtgtgaaatccttccctcgatcc; SEQ ID NO:8), heated to 95 °C for 2 min, followed by incubation at 45 °C. After 30 min, 1 μL of 2.5 mM (each) ddNTPs was added to both reactions and incubated for another 15 min, and the DNA products were purified. Then, equal volumes of the two read products were analyzed by qPCR on a CFX96 instrument (Bio-Rad) using Phusion HS II (Thermo) and 1x EvaGreen (Biotium) primers PA5 (ttttGAATTCCAACACACCACCCACCCAAC; SEQ ID NO:9) and PA6 (tttttAAGCTTCGACTCACCTCACGTCCTCA; SEQ ID NO:10), for 30 cycles at 98°C for 5 seconds, 67°C for 15 seconds, and 72°C for 30 seconds. Using the EcoRI and HindIII sites introduced by the PCR primers, the qPCR product from reaction ii was inserted into the pUC19 plasmid, and 81 clones were sequenced as described above.

[0313] The dNTP analog used in this example contains an alkyne amino group extending from the nucleobase (position 5 of pyrimidine and position 7 of 7-deazopurine), which is the same moiety left as scar tissue by the polymerase-nucleotide conjugate from Example 2. The alkyne amino moiety is further derivatized by N-acetylation. To demonstrate that DNA containing the scar bases prepared as in Example 2 can serve as a template for the precise synthesis of complementary DNA using native dNTPs via template-dependent polymerase, a DNA molecule containing 141 consecutive 3-acetamidopropynyl (i.e., N-acetylated alkyne amino) nucleotides was prepared in this example. This DNA product was isolated and used as a template for PCR ( Figures 13A-13C The PCR product was inserted into a plasmid, cloned into *E. coli*, and 81 colonies were sequenced. Five errors were found, meaning the error rate for synthesizing native DNA from a template modified with 3-acetaminopropynyl is approximately 6 x 10^6. -4 / nt.

[0314] The data shows that scar nucleic acids (in this case, polynucleotides with a portion of the propylamino scar that can be derived from Example 2) can be amplified by high-fidelity PCR, thereby producing natural DNA that can be used for biological applications.

[0315] Example 5.10 Synthesis of polymers

[0316] The 3' overhang of the double-stranded DNA molecule was extended 10 times using the TdT-dNTP conjugate prepared as in Example 2. Following the manufacturer's instructions (two-step protocol: 98°C for 10 seconds, 72°C for 45 seconds), primers T1 ( / 5Phos / GCAGCCAACTCAGCTTCTGCAGGGGCTTTGTTAGCAGCCGGATCCTC; SEQ ID NO: 11) and T2 (AAACAAGCGCTCATGAGCCAGAAATCTGGAGCCCGATCTTCCCCATCGG; SEQ ID NO: 12) were used with Phusion polymerase (Thermo) to prepare the double-stranded DNA used as the initial substrate from a ~350 bp PCR product derived from the pET19b plasmid. The PCR product was digested with PstI to generate a 3' overhang on one side, and then TdT was used to tail the resulting 3' overhang with ddTTP to prevent further extension of the 3' overhang. After tailing, the DNA was digested with BstXI to generate a 3' overhang at the other end of the amplicon for extension via polymerase-nucleotide conjugate. The digestion product was separated and purified by 2% agarose gel electrophoresis to obtain the initial substrate for extension via polymerase-nucleotide conjugate.

[0317] Extension reactions were performed for 90 seconds in 1x RBCs at 37°C using various polymerase-nucleotide conjugates at 1 mg / mL, followed by quenching with an equal volume of quenching buffer (250 mM EDTA, 500 mM NaCl). Adapter lysis was performed by irradiation at 365 nm. The first extension reaction contained ~40 nM of initial substrate. After each lysis step, the DNA product was purified (ZymoOCC), and the recovered DNA was used for the next extension step. The following conjugates were used for the extension steps: 1) TdT-dCTP, 2) TdT-dTTP, 3) TdT-dATP, 4) TdT-dCTP, 5) TdT-dTTP, 6) TdT-dGTP, 7) TdT-dATP, 8) TdT-dCTP, 9) TdT-dTTP, 10) TdT-dGTP. The ten-cycle product was T-tailed using TdT and free dTTP+ddTTP at a ratio of 100:1 and acetylated with bicarbonate buffer containing 20 mM NHS-acetate. The tailed product was then amplified by HotStart Taq (NEB) PCR using primers C3 and C4 (GTGCCGTGAGACCTGGCTCCTGACGAGGAtaagcttCTATAGTGAGTCGTATTAATTTCG; SEQ ID NO:13) (PCR program: 98°C for 2 min initial cycles, 49°C for 20 s, 68°C for 12 min, followed by a three-step protocol of 30 cycles: 98°C for 30 s, 49°C for 20 s, 68°C for 30 s). The PCR product was inserted into pUC19 cells using EcoRI and HindIII sites introduced via PCR primers. The plasmid was transformed into DH10B cells, grown overnight in LB, and single-colony plasmids were extracted and sequenced.

[0318] As described in more detail above, a polymerase-nucleotide conjugate was used to extend the double-stranded DNA template through 10 cycles, amplifying the synthetic product and cloning it for sequencing. Figure 14A Of the 32 clones sequenced, 13 (41%) contained the complete 5'-CTACTGACTG-3' sequence. Figure 14B This means that the average stepwise yield is 91%.

[0319] This result indicates that the DNA elongation cycle can be repeated multiple times to create nucleic acid molecules with the desired sequence and length. sequence list <110> Daniel Arlow S. Paruk (Sebastian) <120> Nucleic acid synthesis and sequencing using tethered nucleoside triphosphates <130> BERK‑339WO <150> 62354635 <151> 2016‑06‑24 <160> 24 <170> SIPOSequenceListing 1.0 <210> 1 <211> 384 <212> PRT <213> Philippine tarsier (Tarsius syrichta) <400> 1 Met Gly His His His His His His His His His His Ser Ser Gly His 1 5 10 15 Ile Asp Asp Asp Asp Lys His Met Ser Gln Tyr Ala Cys Gln Arg Arg 20 25 30 Thr Thr Leu Asn Asn His Asn Gln Ile Phe Thr Asp Ala Phe Asp Ile 35 40 45 Leu Ala Glu Asn Asp Glu Phe Arg Glu Asn Glu Gly Pro Ser Leu Thr 50 55 60 Phe Met Arg Ala Ala Ser Val Leu Lys Ser Leu Pro Phe Thr Ile Ile 65 70 75 80 Ser Met Lys Asp Ile Glu Gly Ile Pro Asn Leu Gly Asp Arg Val Lys 85 90 95 Ser Ile Ile Glu Glu Ile Ile Glu Asp Gly Glu Ser Ser Ala Val Lys 100 105 110 Ala Val Leu Asn Asp Glu Arg Tyr Lys Ser Phe Lys Leu Phe Thr Ser 115 120 125 Val Phe Gly Val Gly Leu Lys Thr Ser Glu Lys Trp Phe Arg Met Gly 130 135 140 Phe Arg Thr Leu Ser Asn Ile Arg Ser Asp Lys Ser Leu Thr Phe Thr 145 150 155 160 Arg Met Gln Arg Ala Gly Phe Leu Tyr Tyr Glu Asp Leu Val Ser Arg 165 170 175 Val Thr Arg Ala Glu Ala Glu Ala Val Gly Val Leu Val Lys Glu Ala 180 185 190 Val Trp Ala Ser Leu Pro Asp Ala Phe Val Thr Met Thr Gly Gly Phe 195 200 205 Arg Arg Gly Lys Lys Thr Gly His Asp Val Asp Phe Leu Ile Thr Ser 210 215 220 Pro Gly Ala Thr Glu Glu Glu Glu Gln Gln Leu Leu His Lys Val Ile 225 230 235 240 Ser Leu Trp Glu His Lys Gly Leu Leu Leu Tyr Tyr Asp Leu Val Glu 245 250 255 Ser Thr Phe Glu Lys Leu Lys Leu Pro Ser Arg Lys Val Asp Ala Leu 260 265 270 Asp His Phe Gln Lys Cys Phe Leu Ile Leu Lys Leu His His Gln Arg 275 280 285 Val Asp Ser Asp Gln Ser Ser Trp Gln Glu Gly Lys Thr Trp Lys Ala 290 295 300 Ile Arg Val Asp Leu Val Val Cys Pro Tyr Glu Arg Arg Ala Phe Ala 305 310 315 320 Leu Leu Gly Trp Thr Gly Ser Arg Gln Phe Glu Arg Asp Leu Arg Arg 325 330 335 Tyr Ala Thr His Glu Arg Lys Met Ile Ile Asp Asn His Ala Leu Tyr 340 345 350 Asp Lys Thr Lys Arg Ile Phe Leu Glu Ala Glu Ser Glu Glu Glu Ile 355 360 365 Phe Ala His Leu Gly Leu Asp Tyr Ile Glu Pro Trp Glu Arg Asn Ala 370 375 380 <210> 2 <211> 807 <212> PRT <213> Artificial Sequence <220> <221> CONFLICT <222> (1)..(807) <223> Synthetic Sequence <400> 2 Met Gly His His His His His His His His His His Ser Ser Gly His 1 5 10 15 Ile Asp Asp Asp Asp Lys His Met Met Lys Ile Glu Glu Gly Lys Leu 20 25 30 Val Ile Trp Ile Asn Gly Asp Lys Gly Tyr Asn Gly Leu Ala Glu Val 35 40 45 Gly Lys Lys Phe Glu Lys Asp Thr Gly Ile Lys Val Thr Val Glu His 50 55 60 Pro Asp Lys Leu Glu Glu Lys Phe Pro Gln Val Ala Ala Thr Gly Asp 65 70 75 80 Gly Pro Asp Ile Ile Phe Trp Ala His Asp Arg Phe Gly Gly Tyr Ala 85 90 95 Gln Ser Gly Leu Leu Ala Glu Ile Thr Pro Asp Lys Ala Phe Gln Asp 100 105 110 Lys Leu Tyr Pro Phe Thr Trp Asp Ala Val Arg Tyr Asn Gly Lys Leu 115 120 125 Ile Ala Tyr Pro Ile Ala Val Glu Ala Leu Ser Leu Ile Tyr Asn Lys 130 135 140 Asp Leu Leu Pro Asn Pro Pro Lys Thr Trp Glu Glu Ile Pro Ala Leu 145 150 155 160 Asp Lys Glu Leu Lys Ala Lys Gly Lys Ser Ala Leu Met Phe Asn Leu 165 170 175 Gln Glu Pro Tyr Phe Thr Trp Pro Leu Ile Ala Ala Asp Gly Gly Tyr 180 185 190 Ala Phe Lys Tyr Glu Asn Gly Lys Tyr Asp Ile Lys Asp Val Gly Val 195 200 205 Asp Asn Ala Gly Ala Lys Ala Gly Leu Thr Phe Leu Val Asp Leu Ile 210 215 220 Lys Asn Lys His Met Asn Ala Asp Thr Asp Tyr Ser Ile Ala Glu Ala 225 230 235 240 Ala Phe Asn Lys Gly Glu Thr Ala Met Thr Ile Asn Gly Pro Trp Ala 245 250 255 Trp Ser Asn Ile Asp Thr Ser Lys Val Asn Tyr Gly Val Thr Val Leu 260 265 270 Pro Thr Phe Lys Gly Gln Pro Ser Lys Pro Phe Val Gly Val Leu Ser 275 280 285 Ala Gly Ile Asn Ala Ala Ser Pro Asn Lys Glu Leu Ala Lys Glu Phe 290 295 300 Leu Glu Asn Tyr Leu Leu Thr Asp Glu Gly Leu Glu Ala Val Asn Lys 305 310 315 320 Asp Lys Pro Leu Gly Ala Val Ala Leu Lys Ser Tyr Glu Glu Glu Leu 325 330 335 Val Lys Asp Pro Arg Ile Ala Ala Thr Met Glu Asn Ala Gln Lys Gly 340 345 350 Glu Ile Met Pro Asn Ile Pro Gln Met Ser Ala Phe Trp Tyr Ala Val 355 360 365 Arg Thr Ala Val Ile Asn Ala Ala Ser Gly Arg Gln Thr Val Asp Glu 370 375 380 Ala Leu Lys Asp Ala Gln Thr Asn Ser Ser Ser Asn Asn Asn Asn Asn 385 390 395 400 Asn Asn Asn Asn Asn Leu Gly Ile Glu Gly Arg Ile Ser His Met Ser 405 410 415 Met Gly Gly Arg Asp Ile Val Asp Gly Ser Glu Phe Ser Pro Ser Pro 420 425 430 Val Pro Gly Ser Gln Asn Val Pro Ala Pro Ala Val Lys Lys Ile Ser 435 440 445 Gln Tyr Ala Cys Gln Arg Arg Thr Thr Leu Asn Asn Tyr Asn Gln Leu 450 455 460 Phe Thr Asp Ala Leu Asp Ile Leu Ala Glu Asn Asp Glu Leu Arg Glu 465 470 475 480 Asn Glu Gly Ser Ala Leu Ala Phe Met Arg Ala Ser Ser Val Leu Lys 485 490 495 Ser Leu Pro Phe Pro Ile Thr Ser Met Lys Asp Thr Glu Gly Ile Pro 500 505 510 Ser Leu Gly Asp Lys Val Lys Ser Ile Ile Glu Gly Ile Ile Glu Asp 515 520 525 Gly Glu Ser Ser Glu Ala Lys Ala Val Leu Asn Asp Glu Arg Tyr Lys 530 535 540 Ser Phe Lys Leu Phe Thr Ser Val Phe Gly Val Gly Leu Lys Thr Ala 545 550 555 560 Glu Lys Trp Phe Arg Met Gly Phe Arg Thr Leu Ser Lys Ile Gln Ser 565 570 575 Asp Lys Ser Leu Arg Phe Thr Gln Met Gln Lys Ala Gly Phe Leu Tyr 580 585 590 Tyr Glu Asp Leu Val Ser Cys Val Asn Arg Pro Glu Ala Glu Ala Val 595 600 605 Ser Met Leu Val Lys Glu Ala Val Val Thr Phe Leu Pro Asp Ala Leu 610 615 620 Val Thr Met Thr Gly Gly Phe Arg Arg Gly Lys Met Thr Gly His Asp 625 630 635 640 Val Asp Phe Leu Ile Thr Ser Pro Glu Ala Thr Glu Asp Glu Glu Gln 645 650 655 Gln Leu Leu His Lys Val Thr Asp Phe Trp Lys Gln Gln Gly Leu Leu 660 665 670 Leu Tyr Ala Asp Ile Leu Glu Ser Thr Phe Glu Lys Phe Lys Gln Pro 675 680 685 Ser Arg Lys Val Asp Ala Leu Asp His Phe Gln Lys Cys Phe Leu Ile 690 695 700 Leu Lys Leu Asp His Gly Arg Val His Ser Glu Lys Ser Gly Gln Gln 705 710 715 720 Glu Gly Lys Gly Trp Lys Ala Ile Arg Val Asp Leu Val Met Ser Pro 725 730 735 Tyr Asp Arg Arg Ala Phe Ala Leu Leu Gly Trp Thr Gly Ser Arg Gln 740 745 750 Phe Glu Arg Asp Leu Arg Arg Tyr Ala Thr His Glu Arg Lys Met Met 755 760 765 Leu Asp Asn His Ala Leu Tyr Asp Arg Thr Lys Arg Val Phe Leu Glu 770 775 780 Ala Glu Ser Glu Glu Glu Ile Phe Ala His Leu Gly Leu Asp Tyr Ile 785 790 795 800 Glu Pro Trp Glu Arg Asn Ala 805 <210> 3 <211> 46 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(46) <223> synthetic sequence <220> <221> misc_feature <222> (1)..(1) <223> 5' phosphorylation at the 5' end <220> <221> misc_feature <222> (1)..(1) <223> n is uracil <220> <221> misc_feature <222> (22)..(22) <223> n is a deoxythymidine nucleotide base modified with fluorescein. <400> 3 ntgaagagcg agagtgagtg ancattaaag acgtgggcct ggattt 46 <210> 4 <211> 59 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(59) <223> synthetic sequence <400> 4 gtgccgtgag acctggctcc tgacgatatg gataagcttt gaagagcgag agtgagtga 59 <210> 5 <211> 56 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(56) <223> synthetic sequence <400> 5 aaaagaattc aaaaaaaaaa aaaaaaaaaa aaaaaaaaaa aaaaaaaaaa aaaaaa 56 <210> 6 <211> twenty one <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(21) <223> synthetic sequence <220> <221> misc_feature <222> (1)..(1) <223> Phosphorylation at the 5' end <400> 6 cattaaagac gtgggcgtgg a 21 <210> 7 <211> 26 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(26) <223> synthetic sequence <220> <221> misc_feature <222> (4)..(4) <223> n is an internal amino group labeled with Cy3 NHS. <400> 7 tttntgtgaa atccttccct cgatcc 26 <210> 8 <211> 42 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(42) <223> synthetic sequence <220> <221> misc_feature <222> (1)..(1) <223> 5'-terminal 5'-phosphate aminohexyl modification <400> 8 cgactcacct cacgtcctca tgtgaaatcc ttccctcgat cc 42 <210> 9 <211> 30 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(30) <223> synthetic sequence <400> 9 ttttgaattc caacacaccacccacccaac 30 <210> 10 <211> 30 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(30) <223> synthetic sequence <400> 10 ttttaagctt cgactcacct cacgtcctca 30 <210> 11 <211> 47 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(47) <223> synthetic sequence <220> <221> misc_feature <222> (1)..(1) <223> 5' phosphorylation at the 5' end <400> 11 gcagccaact cagcttctgc aggggctttg ttagcagccg gatcctc 47 <210> 12 <211> 49 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(49) <223> synthetic sequence <400> 12 aaacaagcgc tcatgagcca gaaatctgga gcccgatctt ccccatcgg 49 <210> 13 <211> 60 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(60) <223> synthetic sequence <400> 13 gtgccgtgag acctggctcc tgacgaggat aagcttctat agtgagtcgt attaatttcg 60 <210> 14 <211> 40 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(40) <223> synthetic sequence <220> <221> misc_feature <222> (1)..(1) <223> 5'-terminal 5'-phosphate aminohexyl modification <400> 14 caacacacca cccacccaac cgcagatgtc cgctttctgg 40 <210> 15 <211> 11 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(11) <223> synthetic sequence <220> <221> misc_feature <222> (1)..(1) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (2)..(2) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (3)..(3) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (4)..(4) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (11)..(11) <223> OH group at the 3' end <400> 15 ctagtttttt t 11 <210> 16 <211> 11 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(11) <223> synthetic sequence <220> <221> misc_feature <222> (11)..(11) <223> OH group at the 3' end <400> 16 ctagtttttt t 11 <210> 17 <211> 11 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(11) <223> synthetic sequence <400> 17 tttctagttt t 11 <210> 18 <211> 10 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(10) <223> synthetic sequence <400> 18 ctactgactg 10 <210> 19 <211> 10 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(10) <223> synthetic sequence <220> <221> misc_feature <222> (1)..(1) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (1)..(1) <223> No other information <220> <221> misc_feature <222> (2)..(2) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (3)..(3) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (4)..(4) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (5)..(5) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (6)..(6) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (7)..(7) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (8)..(8) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (9)..(9) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (10)..(10) <223> OH group at the 3' end <220> <221> misc_feature <222> (10)..(10) <223> Connected to the -C triple bond CCH2NH2 <400> 19 ctactgactg 10 <210> 20 <211> 17 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(17) <223> synthetic sequence <220> <221> misc_feature <222> (1)..(1) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (2)..(2) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (3)..(3) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (4)..(4) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (5)..(5) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (6)..(6) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (7)..(7) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (8)..(8) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (9)..(9) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (10)..(10) <223> Connected to the -C triple bond CCH2NH2 <220> <221> misc_feature <222> (17)..(17) <223> OH group at the 3' end <400> 20 ctactgactg ttttttt 17 <210> twenty one <211> 17 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(17) <223> synthetic sequence <400> twenty one ctactgactg ttttttt 17 <210> twenty two <211> 17 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(17) <223> synthetic sequence <400> twenty two tttctactga ctgtttt 17 <210> twenty three <211> 11 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(17) <223> synthetic sequence <400> twenty three aaaaaaacta g 11 <210> twenty four <211> 17 <212> DNA <213> Artificial Sequence <220> <221> conflict <222> (1)..(17) <223> synthetic sequence <400> twenty four aaaaaaacag tcagtag 17

Claims

1. A conjugate comprising a nucleoside triphosphate, a linker, and a polymerase, the polymerase being capable of catalyzing the covalent addition of the nucleoside triphosphate to the 3' end of a nucleic acid, wherein the linker tethers the nucleoside triphosphate to the polymerase, and wherein the linker is connected to a nucleobase, sugar, or α-phosphate of the nucleoside triphosphate, and wherein the linker comprises a selectively cleavable bond; The polymerase mentioned therein is a template-independent polymerase.

2. The conjugate of claim 1, wherein the length of the linker is in the range of 4-100 Å, and wherein the length of the linker is sufficient to bring the nucleoside triphosphate close to the active site of the polymerase.

3. The conjugate of claim 1, wherein the nucleoside triphosphate is linked to a cysteine ​​residue in the polymerase.

4. The conjugate of claim 3, wherein cleavage of the linker releases the polymerase from the nucleoside triphosphate.

5. The conjugate of claim 4, wherein the cleavage of the linker leaves a scar on the nucleobase of the nucleoside triphosphate.

6. The conjugate of claim 5, wherein the scar on the nucleoside triphosphate is selectively removable.

7. The conjugate of claim 6, wherein the removal of the scar leaves naturally occurring nucleobases.

8. The conjugate of claim 1, wherein the selectively cleavable bond is a photo-cleavable bond or an enzyme-cleavable bond.

9. The conjugate of claim 1, wherein the nucleoside triphosphate comprises a nucleobase selected from adenine, cytosine, guanine, thymine, and uracil.

10. The conjugate of claim 1, wherein the polymerase is a DNA polymerase.

11. The conjugate of claim 1, wherein the polymerase is an RNA polymerase.

12. The conjugate of claim 1, wherein the nucleoside triphosphate or polymerase comprises a fluorescent label.

13. The conjugate of claim 1, wherein the nucleoside triphosphate is deoxyribonucleoside triphosphate.

14. The conjugate of claim 1, wherein the nucleoside triphosphate is riboside triphosphate.

15. A composition comprising a group of the complexes according to any one of claims 1-14, wherein the nucleoside triphosphates correspond to G, A, T and C.

16. A method for synthesizing nucleic acids, comprising: Nucleic acid is incubated with a first conjugate, wherein the first conjugate is any one of claims 1-14, and the incubation is carried out under conditions in which a polymerase catalyzes the covalent addition of the nucleoside triphosphate of the first conjugate to the 3' hydroxyl group of the nucleic acid to prepare an extended product.

17. The method of claim 16, wherein the nucleic acid is tethered to the vector.

18. The method of claim 17, wherein the method comprises, after adding nucleoside triphosphate to a nucleic acid, cleaving the linker to release polymerase from the elongation product.

19. The method of claim 18, wherein the linker comprises an enzyme-cleavable bond or a photo-cleavable bond, and the cleavage comprises exposing the elongation product to an enzyme or to light.

20. The method of claim 19, wherein the cleavage of the linker causes the added nucleoside triphosphate to be deprotected to produce a deprotected extended product.

21. The method of claim 20, further comprising, after the added nucleoside triphosphate is deprotected: The deprotected extension product is incubated with a second conjugate, wherein the second conjugate is the conjugate according to any one of claims 1-14, and the incubation is carried out under conditions in which a polymerase catalyzes the covalent addition of the nucleoside triphosphate of the second conjugate to the 3' end of the deprotected extension product.

22. The method of claim 21, wherein the method comprises: (a) Incubating the nucleic acid with the first conjugate under conditions in which the polymerase catalyzes the covalent addition of the nucleoside triphosphate of the first conjugate to the 3' hydroxyl group of the nucleic acid to prepare an extended product; (b) Cleavage the linker to release the polymerase from the elongation product and deprotect the elongation product; (c) The second extension product is prepared by incubating the deprotected extension product with the second conjugate under the condition that the polymerase catalyzes the covalent addition of the nucleoside triphosphate of the second conjugate to the 3' end of the extension product; (d) Repeat steps (b)-(c) multiple times with the second extension product to produce extended nucleic acids with defined sequences.

23. The method of claim 22, wherein the nucleoside triphosphate is a reversible terminator, and wherein deprotection of the extended product comprises removing the blocking group of the reversible terminator.

24. The method of any one of claims 16-23, wherein the nucleic acid is an oligonucleotide.