De novo stepwise template-independent synthesis of long polynucleotides
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- ANSA BIOTECHNOLOGIES INC
- Filing Date
- 2024-06-21
- Publication Date
- 2026-04-29
AI Technical Summary
Current de novo DNA synthesis methods, such as chemical synthesis and enzymatic template-dependent methods, face challenges in achieving high stepwise yield and low error rates for synthesizing long polynucleotides, particularly beyond 200 bases, due to side reactions and impurities that limit the length and purity of the synthesized products.
A template-independent de novo synthesis method involving nucleotide coupling to substrates with free hydroxyl groups, using a blocking group that is removed to allow further coupling, achieving long polynucleotides of up to 2000 nucleotides or more with an observed error rate of less than 1% per cycle, and a coupling rate of at least 10 nucleotides per hour.
This method significantly improves the stepwise yield and reduces errors, enabling the synthesis of long polynucleotides with high accuracy and efficiency, overcoming the limitations of existing methods by maintaining low error rates and achieving longer sequence lengths.
Smart Images

Figure US2024035137_26122024_PF_FP_ABST
Abstract
Description
DE NOVO STEPWISE TEMPLATE-INDEPENDENT SYNTHESIS OF LONG POLYNUCLEOTIDES CROSS-REFERENCE TORELATEDAPPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 509,522, filed on June 21, 2023, the contents of which is incorporated herein by reference in its entirety. SEQUENCE LISTING
[0002] This application contains a Sequence Listing, which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML file, created on June 14, 2024, is titled ABB-013WO_SL.xml and is 99.9 kilobytes in size. BACKGROUND
[0003] Polynucleotide synthesis includes creation of chains of nucleotides, which are building blocks of DNA and RNA. Standard de novo DNA synthesis performed today is based on the nucleoside phosphoramidite method (generally referred to as “chemical synthesis”) in which a desired sequence is synthesized by stepwise coupling of blocked monomers. These reactions are performed in organic solvents using highly-reactive activated monomers, and the conditions cause side reactions that damage the growing chain, limiting the yield of full-length product. The impurities produced can be difficult or impractical to separate from the desired oligonucleotide product, limiting the usefulness of the method for producing sequences longer than approximately 200 bases.
[0004] As an alternative, different enzymatic de novo DNA synthesis strategies using template-independent polymerases have more recently been developed, allowing an environmentally-friendly synthesis of longer DNA molecules than with chemical synthesis. However, stepwise yield and length have still not significantly improved. In some cases, stepwise yield can start to stall after the synthesized polynucleotide has reached a certain length. Even if the stepwise yield remains consistent, for polynucleotides up to 500 or 1000 nucleotides in length, a very high stepwise yield is needed to achieve a reasonable number of perfectly synthesized long polynucleotides via de novo stepwise synthesis.
[0005] What is needed, therefore, are improved methods of polynucleotide synthesis to achieve a high stepwise yield and / or low error rate and / or a long polynucleotide synthesis using stepwise de novo template-independent synthesis techniques.SUMMARY
[0006] Among other things, the present disclosure provides technologies (e.g., methods of synthesizing, polynucleotides, etc.) that represent advances and improvements in long polynucleotide synthesis. The disclosure is based, in part, upon methods of synthesis for long polynucleotides, including improved methods that achieve not only a high stepwise yield and / or lower error rate than previous methods, but do so in successfully producing long polynucleotides (e.g., about 500 to about 1000 nucleotides or more in length).
[0007] In some aspects, the present disclosure provides a method of template-independent de novo synthesis of a long single-stranded polynucleotide, comprising providing a substrate comprising a plurality of free hydroxyl groups linked to the substrate and suitable for nucleotide coupling; coupling a nucleotide to the plurality of free hydroxyl groups, wherein the nucleotide is linked to a blocking group; removing said blocking group from said coupled nucleotides; repeating steps (b) and (c) according to a predetermined nucleotide sequence (e.g., a reference sequence) to yield a plurality of de novo synthesized polynucleotides at least 500 nucleotides in length, wherein each coupling has an observed error rate of less than 1% as compared to said predetermined nucleotide sequence.
[0008] In some embodiments, the de novo synthesized long polynucleotides is at least 600, at least 700, at least 800, at least 900, or at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1800, at least 1900, or at least 2000 nucleotides in length.
[0009] In some embodiments, the de novo synthesized long polynucleotides from 500 to 1000 nucleotides in length, from 500 to 1500 nucleotides in length, from 500 to 2000 nucleotides in length, or from 500 to 2500 nucleotides in length.
[0010] In some embodiments, the observed error rate of each coupling is less than 0.9%, less than 0.8%, less than 0.7%, less than 0.6%, less than 0.5%, less than 0.4%, less than 0.3%, less than 0.2%, less than 0.1% per cycle, less than 0.09% per cycle, less than 0.08% per cycle, less than 0.07% per cycle, less than 0.06% per cycle, less than 0.05% per cycle, less than 0.04% per cycle, less than 0.03% per cycle, less than 0.02% per cycle, or less than 0.01% per cycle.
[0011] In some embodiments, said coupling step is performed in less than 120 seconds, less than 90 seconds, less than 60 seconds, less than 50 seconds, or less than 40 seconds, less than 30 seconds, less than 20 seconds, less than 15 seconds, less than 10 seconds, or less than 5 seconds.
[0012] In some embodiments, said template-independent polynucleotide synthesis to yield said de novo synthesized polynucleotides is performed at a rate of at least 10 nucleotides per hour, at least 15 nucleotides per hour, at least 20 nucleotides per hour, at least 25 nucleotides per hour, at least 30 nucleotides per hour, at least 40 nucleotides per hour, at least 50 nucleotides per hour, or at least 60 nucleotides per hour.
[0013] In some embodiments, said template-independent polynucleotide synthesis is performed at a rate of from 5 nucleotides per hour to 25 nucleotides per hour, from 10 nucleotides per hour to 30 nucleotides per hour, from 15 nucleotides per hour to 45 nucleotides per hour, or from 20 nucleotides per hour to 60 nucleotides per hour.
[0014] In some embodiments, said plurality of de novo synthesized polynucleotides on the substrate comprises at least 10 polynucleotides, at least 20 polynucleotides, at least 50 polynucleotides, at least 100 polynucleotides, at least 200 polynucleotides, at least 500 polynucleotides, at least 1,000 polynucleotides, at least 10,000 polynucleotides, at least 10 x 105polynucleotides, at least 10 x 106polynucleotides, at least 10 x 107polynucleotides, at least 10 x 108polynucleotides, at least 10 x 109polynucleotides, at least 10 x 1010polynucleotides, at least 10 x 1011polynucleotides, at least 10 x 1012polynucleotides, at least 10 x 1013polynucleotides, or at least 10 x 1014polynucleotides.
[0015] In some embodiments, the free hydroxyl groups are at the end of a plurality of starter oligonucleotides or growing polynucleotides attached to the substrate.
[0016] In some embodiments, the starter oligonucleotides comprise a single-stranded region at the 3’ end.
[0017] In some embodiments, the starter oligonucleotide is hybridized to an oligonucleotide bound to the substrate.
[0018] In some embodiments, the starter oligonucleotide is covalently linked to the substrate.
[0019] In some embodiments, said nucleotide coupling is performed enzymatically.
[0020] In some embodiments, said nucleotide coupling is catalyzed by a polymerase.
[0021] In some embodiments, said polymerase is a template-independent polymerase.
[0022] In some embodiments, said template-independent polymerase is covalently linked to said nucleotide.
[0023] In some embodiments, said template-independent polymerase is Terminal deoxynucleotidyl Transferase (TdT), or a variant thereof.
[0024] In some embodiments, said polymerase is an RNA polymerase.
[0025] In some embodiments, said blocking group is a template-independent polymerase linked to said nucleotide.
[0026] In some embodiments, said blocking group comprises cleaving a linker attaching said nucleotide to said template-independent polymerase.
[0027] In some embodiments, said blocking group is a 3'-O-blocking group.
[0028] In some embodiments, removing said blocking group comprises removing said 3'-O- blocking group from said nucleotide to leave a free 3' hydroxyl group.
[0029] In some embodiments, the blocking group is a 2' or 3' modification of the nucleotide.
[0030] In some embodiments, the 2' modification is selected from the group consisting of - H, -OH, -F, -OMe, -N3, -NH2, and -Ara.
[0031] In some embodiments, the 3' modification is selected from the group consisting of — H, -OH, -OCH2N3, -ONH2 and -Oallyl.
[0032] In some embodiments, the blocking group is a reversible terminator.
[0033] In some embodiments, the predetermined sequence has a GC content of at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95%.
[0034] In some embodiments, the predetermined sequence has an AT content of at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95%.
[0035] In some embodiments, said coupling is performed in the presence of phosphatase. In some embodiments, the phosphatase is an inorganic pyrophosphatase.
[0036] In some embodiments, the nucleotide linked to the blocking group is a nucleotidepolymerase conjugate and the conjugate has been treated with phosphatase.
[0037] In some embodiments, said coupling is performed in the presence of a divalent cation having a total concentration of divalent cations present in the reaction volume of said coupling no greater than about 500 pM.
[0038] In some embodiments, the total concentration of divalent cations present in the reaction volume is no greater than about 250 pM, about 125 pM or about 50 pM.
[0039] In some embodiments, the divalent cation present at the highest concentration in the reaction volume is cobalt (Co2+) or zinc (Zn2+).
[0040] In some embodiments, the at least one divalent cation is selected from Mg2+, Ca2+, Sr2+, Ba2+, Mn2+, Co2+, Fe2+, Ni2+, Cu2+, and Zn2+, or a combination thereof.
[0041] In some embodiments, the coupling reaction is performed in the absence of Mg2+.
[0042] In some embodiments, said nucleotide comprises one or more modifications to a hydrogen binding N or O on the nucleobase.
[0043] In some embodiments, said coupled nucleotide comprises one or more alkylated nucleobases after removal of said blocking group. In some such embodiments, the method further comprises contacting said de novo synthesized polynucleotides with an alkyl transferase.
[0044] In some embodiments, said alkyl transferase is from EC 2.1.1.63.
[0045] In some embodiments, said alkyl transferase is selected from an alkyl transferase listed in Table 1 or Table 2.
[0046] In some embodiments, said alkyl transferase is O6-alkylguanine DNA alkyltransferase.
[0047] In some embodiments, said alkyl transferase is AlkB.
[0048] In some embodiments, the alkylated nucleobase is represented by:wherein X is -C(R2)= or -N=; R1 is selected from the group consisting of C1-6 alkyl, C2-6 alkenyl, C1-6 alkynyl, and –(CH2)0-3Ph, wherein R1 is optionally substituted with 1-6 instances of R1a; each R1a is independently selected from halogen, C1-6 alkyl, -(CH2)0-3OR1b, -NO2, -N3, -OPO2OH, and –(CH2)0-3NHR1b; and each R1b is independently selected from hydrogen, C1-6 alkyl, -C(O)(C1-6 alkyl), C1-6 haloalkyl, -C(O)(C1-6 haloalkyl), and -CH2OAc; R2 is selected from the group consisting of hydrogen, optionally substituted C1-4 alkyl chain, wherein 1-2 methylene units is optionally and independently replaced with -O-, -N(Ra)-, -C(O)-, -S-, -S(O)-, -S(O)2-, or phenylene, wherein R2 is optionally substituted with 1-6 instances of R2a; each R2a is independently selected from halogen, C1-6 alkyl, -(CH2)0-3OR2b, -NO2, -N3, -OPO2OH, and –(CH2)0- 3NHR2b; and each R2b is independently selected from hydrogen and C1-6 alkyl.
[0049] In some embodiments, R1 is C1-4 alkyl.
[0050] In some embodiments, R1 is selected from the group consisting of methyl, ethyl, n- propyl, and n-butyl.
[0051] In some embodiments, R1 is selected from the group consisting of:.
[0052] In some embodiments, R1a is -OR1b.
[0053] In some embodiments, R1b is hydrogen.
[0054] In some embodiments, the alkylated nucleobase in the polynucleotide is selected from the group consisting of:
[0055] In some embodiments, said coupled nucleotide comprises one or more modifications to a base-pairing nitrogen or oxygen on the nucleobase after removal of said blocking group.
[0056] In some embodiments, said coupled nucleotide is represented by (B):(B) wherein R is a ribose polyphosphate or deoxyribose polyphosphate; Y is a nucleobase; L- R1 is a protecting group; wherein L is attached to a base-pairing nitrogen or oxygen of the nucleobase; and wherein R1 is selected from the group consisting of hydrogen, -OH, - N(Rb )2, and -SH, wherein each Rb is independently hydrogen or optionally substituted C1-6 alkyl.
[0057] In some embodiments, L is -Z-L1-L2-; Z is selected from the group consisting of a bond, -C(O)-, -C(O)CH2-, -C(O)C(RL)2-, -C(O)CH(RL)-, -C(O)O-, and -C(O)N(H)-;
[0058] L1 is selected from the group consisting of a bond, ,each RL is independently selected from the group consisting of halogen, hydroxyl, oxo, and optionally substituted C1-C3 alkyl, wherein 2 instances of R1 are optionally taken together with the intervening atom(s) to form a 3-6 membered carbocyclyl ring; L2 is selected from the group consisting of a bond, an optionally substituted C1-12 alkylene chain, C4-C20 polyethylene glycol, an optionally substituted C2-12 alkenylene chain, and an optionally substituted C2-12 alkynylene chain, wherein 1-6 methylene units are optionally and independently replaced with -O-, -N(Rb)-, -C(O)-, -S-, -S(O)-, -S(O)2-, or phenylene;
[0059] W is selected from the group consisting of -O-, -S-, -S(O)2-, and -N(Rb)-; each Ra is halogen, -Me, or -OMe; each Rb is independently hydrogen or C1-6 alkyl; and
[0060] n is 1 or 2; wherein L1 and Z are not both a bond.
[0061] In some embodiments, Z is a bond when L is attached to a base-pairing oxygen of the nucleobase.
[0062] In some embodiments, Z is selected from the group consisting of -C(O)-, -C(O)CH2- , -C(O)C(RL)2-, -C(O)CH(RL)-, -C(O)O-, and -C(O)N(H)-when L is attached to a base- pairing nitrogen of the nucleobase.
[0063] In some embodiments, L1 is, wherein L1 is optionally substituted with 1-4 instances of RL; n is 1 or 2; and W is selected from the group consisting of -O-, -S-, - S(O)2-, and -N(Rb)-.
[0064] In some embodiments, L1 is selected from the group consisting of,wherein L1 is optionally substituted with 1-4 instances of RL.
[0065] In some embodiments, each RL is independently optionally substituted C1-C3 alkyl, wherein 2 instances of R1 are optionally taken together with the intervening atom(s) to form a 3-6 membered carbocyclyl ring.
[0066] In some embodiments, RL is optionally substituted methyl.
[0067] In some embodiments, L1 is selected from the group consisting of.
[0069] In some embodiments, Z is -C(O)O-.
[0070] In some embodiments, Z is -C(O)N(H)-.
[0071] In some embodiments, Z is a bond.
[0072] In some embodiments, Z is -C(O)-.
[0073] In some embodiments, -Z-L1-L2-R1 is selected from the group consisting of
[0074] In some embodiments, L2 is an optionally substituted C1-12 alkylene chain, wherein 1-6 methylene units are optionally and independently replaced with -O-, -N(Rb)-, -C(O)-, -S- , -S(O)-, -S(O)2-, or phenylene.
[0075] In some embodiments, L2 is an optionally substituted C1-12 alkylene chain, wherein 1-6 methylene units are optionally and independently replaced with -O-.
[0076] In some embodiments, L2 is an optionally substituted C2-6 alkylene chain, wherein 1-3 methylene units are optionally and independently replaced with -O-.
[0077] In some embodiments, -Z-L1-L2-R1 is selected from the group consisting of
[0078] In some embodiments, said coupling comprises dipping said substrate into a solution comprising said nucleotide and a template-independent polymerase.
[0079] In some embodiments, the blocking group is a polymerase, and wherein the polymerase is linked to the nucleotide via a cleavable linker.
[0080] In some embodiments, the cleavable linker comprises an amino acid ester.
[0081] In some embodiments, the amino acid ester is attached to an amino acid.
[0082] In some embodiments, the amine group of the amino acid ester is bound to the amino acid.
[0083] In some embodiments, the cleavable linker comprises a peptide of at least 2, at least 3, at least 4, or at least 5 amino acids bound to the amine group of the amino acid ester.
[0084] In some embodiments, the amino acid or amino acids is selected from the group consisting of: alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine.
[0085] In some embodiments, the amino acid is glycine or the amino acids comprise glycine.
[0086] In some embodiments, the amino acid is a non-naturally occurring amino acid or the amino acids comprise a non-naturally occurring amino acid.
[0087] In some embodiments, the cleavable linker is bound to the alpha-phosphate, sugar, or nucleobase of the nucleotide.
[0088] In some embodiments, the amino acid ester is represented by:; wherein R1 and R1' are each independently selected from hydrogen and an optionally substituted C1-6 alkyl, or are optionally taken together with the atom on which they are attached to form an optionally substituted C3-C7 carbocyclic ring.
[0089] In some embodiments, the amino acid ester is represented by a compound selected from the group consisting of:
[0090] In some embodiments, the linker comprises the structure:wherein R1 and R1' are each independently selected from hydrogen and an optionally substituted C1-6 alkyl or are optionally taken together with the atom on which they are attached to form an optionally substituted C3-C7 carbocyclic ring; each R2 is an optionally substituted group independently selected from the group consisting of hydrogen, C1-6 alkyl, phenyl, C1-C6 carbocyclic ring and 3-7 heterocyclic ring; each R3 is hydrogen or optionally substituted C1-6 alkyl; and n is 1, 2, 3, 4 or 5.
[0091] In some embodiments, R3 is hydrogen.
[0092] In some embodiments, R2 is hydrogen.
[0093] In some embodiments, R2 is selected from the group consisting of hydrogen, -Me, - iso-Pr, -sec-butyl, iso-butyl, -CH2Ph, -CH2OH, -CH2SH, -CH2CH2SCH3, -CH2COOH, - CH2CH2COOH, -CH2CONH2, -CH2CH2CONH2, -CH2CH2, -CH2CH2NH2,
[0094] In some embodiments, n is 1.
[0095] In some embodiments, R1 and R1' are taken together to form an optionally substituted C3-C7 carbocyclic ring.
[0096] In some embodiments, R1 and R1' are taken together to form an optionally substituted C3 carbocyclic ring.
[0097] In some embodiments, the linker comprises the structure:.
[0098] In some embodiments, the nucleotide linked to the polymerase comprises the structure: Nuc—L1—L2—L3—Pol, wherein: Nuc is the nucleotide; Pol is the polymerase; L1 is a first portion of the linker connecting the nucleotide to L2; L2 is a second portion of the linker represented by:wherein R1 and R1' are each independently selected from an optionally substituted C1-6 alkyl, a halogen, or are optionally taken together with the atom on which they are attached to form an optionally substituted C3- C7 carbocyclic ring; each R2 is an optionally substituted group independently selected from the group consisting of hydrogen, C1-6 alkyl, phenyl, C1-C6 carbocyclic ring and 3-7 heterocyclic ring; each R3 is hydrogen or optionally substituted C1-6 alkyl; n is 0, 1, 2, 3, 4 or 5; wherein * indicates the attachment point of L2 to L1; and ** indicates the attachment point of L2 to L3; L2 is cleavable; and L3 is a linker connecting pol to L2.
[0099] In some embodiments, L1 is selected from the group consisting of a bond, an optionally substituted C1-12 alkylene chain, C4-C20 polyethylene glycol, an optionally substituted C2-12 alkenylene chain, and a C2-12 alkynylene chain, wherein 1-6 methylene units of L1 are optionally and independently replaced with -O-, -N(Rb)-, -N=C(H)-, -C(O)-, - S-, -S(O)-, -S(O)2-, optionally substituted phenylene, or optionally substituted cyclopropylene.
[0100] In some embodiments, L1 comprises:orTMS; each Ra is independently selected from the group consisting of halogen, hydroxyl, cyano, optionally substituted C1-6 alkyl, and optionally substituted C1-6 alkoxy.
[0101] In some embodiments, L2 comprises an amino acid ester selected from the group consisting of:
[0102] In some embodiments, L2 is represented by:.
[0103] In some embodiments, L1 is bound to the nucleobase of the nucleotide.
[0104] In some embodiments, L1 is bound to the nucleobase at an oxygen or nitrogen involved in base pairing.
[0105] In some embodiments, the nucleobase is selected from the group consisting of:
[0106] In some embodiments, L1 is bound to the sugar of the nucleotide.
[0107] In some embodiments, L1 is bound to a phosphate of the nucleotide.
[0108] In some embodiments, the phosphate is the alpha phosphate.
[0109] In some embodiments, the nucleotide is a ribonucleotide polyphosphate or a deoxyribonucleotide polyphosphate.
[0110] In some embodiments, the nucleotide is selected from the group consisting of: adenine, guanine, cytosine, uracil, and thymine.
[0111] In some embodiments, the polymerase is a template-independent polymerase.
[0112] In some embodiments, the polymerase is TdT.
[0113] In some embodiments, the linker is capable of being cleaved by a protease comprising esterase activity.
[0114] In some embodiments, the linker is capable of being cleaved by Proteinase K.
[0115] In some embodiments, said linker is capable of being cleaved at the ester group on L2, leaving a compound represented by Nuc-L1-OH after said cleavage.
[0116] In some aspects, the present disclosure provides substrate comprising a plurality of attached polynucleotides at least 500 nucleotides in length, wherein said plurality of polynucleotides are characterized by sequences generated by a stepwise template- independent polynucleotide synthesis having an observed error rate of less than 1% per cycle as compared to a predetermined nucleotide sequence.
[0117] In some embodiments, the polynucleotides are at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least1500, at least 1600, at least 1700, at least 1800, at least 1900, or at least 2000 nucleotides in length.
[0118] In some embodiments, the polynucleotides are from 500 to 1000 nucleotides in length, from 500 to 1500 nucleotides in length, from 500 to 2000 nucleotides in length, or from 500 to 2500 nucleotides in length.
[0119] In some embodiments, the observed error rate is less than 0.9%, less than 0.8%, less than 0.7%, less than 0.6%, less than 0.5%, less than 0.4%, less than 0.3%, less than 0.2%, less than 0.1% per cycle, less than 0.09% per cycle, less than 0.08% per cycle, less than 0.07% per cycle, less than 0.06% per cycle, less than 0.05% per cycle, less than 0.04% per cycle, less than 0.03% per cycle, less than 0.02% per cycle, or less than 0.01% per cycle as compared to said predetermined sequence.
[0120] In some embodiments, the observed error rate is greater than 0.001%, greater than 0.002%, greater than 0.005%, greater than 0.01%, greater than 0.02%, greater than 0.05%, or greater than 0.1% per cycle as compared to said predetermined sequence.
[0121] In some embodiments, the plurality of attached polynucleotides on the substrate comprises at least 10 polynucleotides, at least 20 polynucleotides, at least 50 polynucleotides, at least 100 polynucleotides, at least 200 polynucleotides, at least 500 polynucleotides, at least 1,000 polynucleotides, at least 10,000 polynucleotides, at least 10 x 105polynucleotides, at least 10 x 106polynucleotides, at least 10 x 107polynucleotides, at least 10 x 108polynucleotides, at least 10 x 109polynucleotides, at least 10 x 1010polynucleotides, at least 10 x 1011polynucleotides, at least 10 x 1012polynucleotides, at least 10 x 1013polynucleotides, or at least 10 x 1014polynucleotides.
[0122] In some embodiments, the polynucleotides comprise one or more nucleotides comprising a modification to a hydrogen binding N or O on the nucleobase. BRIEF DESCRIPTION OF THE DRAWINGS
[0123] The foregoing and other objects, features and advantages will be apparent from the following description of particular embodiments of the disclosure, as illustrated in the accompanying drawings in which like reference characters refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead placed upon illustrating the principles of various embodiments of the disclosure.
[0124] FIG. 1A depicts a scheme for two-step cyclic nucleic acid synthesis using polymerase-nucleotide conjugates. In the first step, a conjugate binds to a DNA molecule viaits linked dNTP moiety; in the second step the linkage between the polymerase and the elongated DNA molecule including the dNTP moiety is cleaved, deblocking the end of the DNA molecule for subsequent elongation.
[0125] FIG. 1B depicts a scheme for two-step cyclic nucleic acid synthesis using TdT- dNTP conjugates comprising a TdT molecule site-specifically linked to a dNTP via a cleavable linker.
[0126] FIG. 2 depicts (i) typical enzymatic DNA synthesis performed with an enzyme and free nucleotides with 3’ blocking groups, which can be inhibited by the formation of secondary structure during synthesis, and (ii) a diagram of improved conjugate-based synthesis provided herein, including use of polymerase-nucleotide conjugates comprising the polymerase linked to a base pairing N or O atom of the nucleobase. Upon cleavage of the polymerase from the nucleotide at each cycle, a scarred nucleotide comprising a portion of the linker (scar) at the N or O atom is retained in the polynucleotide, which inhibits secondary structure formation. The scar(s) can then be removed to generate a scarless polynucleotide.
[0127] FIG. 3 depicts a scheme for nucleic acid synthesis using TdT-dNTP conjugates where the dNTP has an O-alkyl modification separate from its binding site to the linker. Cyclic nucleic acid synthesis comprises cleavage of the linker to remove TdT from the added dNTP using a cleavage agent which leaves a scar. Once nucleic acid synthesis is complete, the O-alkyl group is removed from the nucleotide(s) using AGT.
[0128] FIG. 4 depicts a scheme for nucleic acid synthesis using TdT-dNTP conjugates where cleavage of the linker to remove TdT from the added dNTP using a cleavage agent results in a nucleotide comprising an O-alkyl scar from the cleavage. Once nucleic acid synthesis is complete, the O-alkyl group is removed from the nucleotide(s) using AGT.
[0129] FIG. 5A shows exemplary intramolecular cyclization reactions for removal of a scar or protecting group comprising an unsubstituted and substituted alkyl group where the alkyl group is attached to an amide group linked to a nitrogen on the nucleobase.
[0130] FIG. 5B shows exemplary intramolecular cyclization reactions for removal of a scar or protecting group comprising an unsubstituted and substituted alkyl group where the alkyl group is attached to carbamate group linked to a nitrogen on the nucleobase.
[0131] FIGS. 6A, 6B, and 6C are diagrams of exemplary unshielded nucleotides that could be present in a polymerase-nucleotide conjugate reagent. FIG. 6A shows an exemplary template-independent polymerase with an exemplary unshielded nucleotide (e.g., a deoxynucleoside triphosphate or “dNTP”) tethered in the wrong position.
[0132] FIG. 6B shows an exemplary unfolded template-independent polymerase with an exemplary tethered nucleotide (e.g., dNTP).
[0133] FIG. 6C illustrates an exemplary “free” (or untethered) nucleotide (e.g., dNTP) as present in an exemplary conjugate reagent. Such free dNTPs can be present in such polymerase-nucleotide conjugate reagent due to, e.g., cleavage of the linker between the nucleotide and the polymerase (e.g., due to instability) or, e.g., due to imperfect removal of free nucleotides from conjugates after conjugate synthesis. Each of the nucleotides in FIG. 6A, FIG. 6B, and FIG. 6C has a 5′ phosphate group accessible to catalytic removal via a phosphatase.
[0134] FIG. 6D shows an exemplary polymerase-nucleotide conjugate comprising an exemplary shielded nucleotide (e.g., dNTP). In this figure, the exemplary nucleotide is tethered in the catalytic site of a folded polymerase and is sterically hindered by the tethered polymerase from phosphatase cleavage at its 5′ phosphate.
[0135] FIG. 7 depicts a block diagram of a synthesis system that is suitable for exemplary embodiments.
[0136] FIG. 8 depicts a first illustrative configuration of the synthesis system.
[0137] FIG. 9 depicts a second illustrative configuration of the synthesis system.
[0138] FIG. 10 depicts an illustrative element and an illustrative well that are suitable for the exemplary embodiments.
[0139] FIG. 11 depicts illustrative cross-sections of element designs of exemplary embodiments.
[0140] FIG. 12 depicts a longitudinal view of illustrative element designs of exemplary embodiments.
[0141] FIG. 13A and 13B depict a floating elements design of an exemplary embodiment.
[0142] FIG. 14 depicts an illustrative reaction plate for use in exemplary embodiments.
[0143] FIG. 15 depicts a portion of an illustrative patterned surface for use in exemplary embodiments.
[0144] FIG. 16 depicts a flowchart of illustrative steps that may be performed in exemplary embodiments to synthesize polymers.
[0145] FIG. 17 depicts illustrative operations that may be performed in exemplary embodiments to perform hybridization of elements of an element array.
[0146] FIG. 18 depicts a flowchart of illustrative steps that may be performed in exemplary embodiments as part of the synthesis process.
[0147] FIG. 19 depicts a flowchart of illustrative steps that may be performed in exemplary embodiments to perform a cycle of the synthesis process.
[0148] FIG. 20 depicts an illustrative pattern of loading polymer extension solutions in wells of a reaction plate in exemplary embodiments.
[0149] FIG. 21 shows the results of extension reactions with TdT and (i) dGTP (left column) and (ii) O6-methyl dGTP (right column) performed for 30 seconds, 1 minute, 2 minutes, 4 minutes, or 8 minutes as described in Example 1. The starter is unmodified T35 oligo (SEQ ID NO: 48). The x-axis is the approximate oligo length in nucleotides and the y- axis is relative fluorescence of fluorescein at 517 nm.
[0150] FIG. 22 shows the results of extension reactions performed starting from a nucleotide comprising a 6 GC hairpin that i) is alkylated (left column) or ii) is unmodified (right column) performed for 3.8 seconds, 6.6 seconds, 11.6 seconds, 19.2 seconds, 29 seconds, 45 seconds, 67 seconds, and 139 seconds as described in Example 3.
[0151] FIG. 23 shows the results of a cyclic nucleotide synthesis reaction using TdT-dNTP conjugates using i) G nucleotides without alkylations in the O6 position (O7Et-G), ii) G nucleotides with alkylation in the O6 position only in two locations with the strongest predicted secondary structure, with the rest of the G nucleotides being without alkylations (O7Et-G & O6Bu-G), and iii) G nucleotides with alkylations in the O6 position (O6Bu-G), as described in Example 4.
[0152] FIG. 24 shows an agarose gel visualized with UV of the product of PCR amplification of a synthesized alkylated 50-mer that was treated or not treated with AGT for the time shown, as described in Example 5. A non-alkylated positive control is also shown for reference.
[0153] FIG. 25 shows the results of a cyclic nucleotide synthesis reaction TdT-dNTP conjugates of a 50-mer polyG homopolymer (SEQ ID NO: 49) using G nucleotides with alkylations in the O6 position, as described in Example 6.
[0154] FIG. 26 shows the results of a cyclic nucleotide synthesis reaction TdT-dNTP conjugates of a 40-mer sequence with an expected 15 bp hairpin structure G nucleotides with alkylations in the O6 position, as described in Example 7. FIG. 26 discloses SEQ ID NO: 35.
[0155] FIG. 27 shows a reaction scheme and to detect activity of human AGT on various O6-alkylated G nucleotides in a synthesized polynucleotide, and a gel showing the results of the assay for O6-methyl-G, O6-hydroxybutyl-G, O6-hydroxypropyl-G, N7-aminomethyl-O6- methyl-G, and O6-aminomethylbenzyl-G as described in Example 8.
[0156] FIGs. 28A and 28B show the results of dealklylation reactions of oligonucleotide comprising 3’ G nucleotide with an O6-allyl modified G nucleotide at the 3’ end, with resulting oligonucleotides measured by capillary electrophoresis. “Allyl-G” represents a control, untreated oligonucleotide with the O6-allyl modified G nucleotide at the 3’ end, while the remaining oligonucleotides were treated with the corresponding species of AGT as shown in FIGs. 28A and 28B.
[0157] FIGs. 29A and 29B show the result of polyG synthesis reactions performed without (“dGTP) and with an O6 alkyl modification at the 3’ end of a polyT starter oligo (FIG. 29A) and at the 3’ end of a polyC starter oligo (FIG. 29B).
[0158] FIG. 30 shows the results of incorporation of alkylated G and U nucleotides at the 3’ end of an oligonucleotide by a polymerase-nucleotide conjugate, followed by cleavage of the polymerase, leaving an alkyl scar on the G or U nucleotides.
[0159] FIGs. 31A-31D shows the successful dealkylation of the synthesized oligonucleotides from FIG. 30 by AGT, resulting in scarless conjugate-based oligonucleotide synthesis.
[0160] FIG. 32A is an HPLC chromatogram showing the trace of the photocleavable dGTP nucleotide, the starting material for the photocleavage experiment. The X-axis is represented in minutes. The peak eluting at 5.693 minutes is representative of the intact photocleavable dATP nucleotide.
[0161] FIG. 32B is an HPLC chromatogram showing the trace of native dGTP nucleotide, the expected product of the photocleavage reaction. The X-axis is represented in minutes. The peak eluting at 2.001 minutes is representative of the native dGTP nucleotide.
[0162] FIG. 32C is an HPLC chromatogram showing the trace of the reaction product of photocleavable dGTP after exposure to a 365 nm wavelength lamp for 120 minutes. The X- axis is represented in minutes.
[0163] FIG. 33A is an HPLC chromatogram showing the trace of the photocleavable dATP nucleotide, the starting material for the photocleavage experiment. The X-axis is represented in minutes. The peak eluting at 5.693 minutes is representative of the intact photocleavable dATP nucleotide.
[0164] FIG. 33B is an HPLC chromatogram showing the trace of the reaction product of photocleavable dATP after exposure to a 365 nm wavelength lamp for 120 minutes. The X- axis is represented in minutes. The peak eluting at 3.426 minutes is representative of the native dATP nucleotide.
[0165] FIG. 34A is a capillary electrophoresis pherogram showing UV exposed polynucleotide extension products following extension with a conjugate including TdT and a modified dGTP containing a photocleavable nitrobenzyl group.
[0166] FIG. 34B is a capillary electrophoresis pherogram showing polynucleotide extension products following extension with a conjugate including TdT and a modified dGTP containing a photocleavable nitrobenzyl group that were exposed to UV light.
[0167] FIG. 35 is a capillary electrophoresis pherogram showing polynucleotide extension products from conjugates with a linker attached to the N4 of cytosine and containing different removeable scars.
[0168] FIG. 36 is a capillary electrophoresis pherogram showing polynucleotide extension products from conjugates with a linker attached to the N6 of adenine and containing different removeable scars.
[0169] FIG. 37 is a capillary electrophoresis pherogram showing polynucleotide extension products from conjugates with a linker attached to the O4 of uracil and thymine and containing removeable scars. The X-axis is represented in approximate oligonucleotide length in number of nucleotides, and the Y-axis is represented in relative fluorescence.
[0170] FIG. 38 is a capillary electrophoresis pherogram showing polynucleotide extension products from conjugates with a linker attached to the O6 of guanine and containing different removeable scars.
[0171] FIG. 39 is a capillary electrophoresis pherogram showing incubation of polynucleotide extension at pH 7 or pH 8 for various times following extension with a conjugate including TdT and a modified dGTP containing an eliminable sulfone group. The first row contains a control ssDNA that serves as a marker for where a single natural guanine on the 3′ end of the starter oligo migrates. Hashed vertical markers showing the migration of natural guanine and +1 sulfone-guanine are indicated for clarity.
[0172] FIG. 40 is a capillary electrophoresis pherogram showing the results of a single nucleotide extension of a starter oligonucleotide with a dGTP comprising a sulfide scar (thioether) (top), an oxidation or the reaction product to form a sulfone scar on the dGTP (middle), and exposure of the sulfone dGTP to sodium hydroxide to remove the O-linked scar from dGTP.
[0173] FIG. 41A is a capillary electrophoresis pherogram showing the results of treatment of an oligonucleotide comprising an N6 Carbamate Sulfide A scarred nucleotide with base (50 mM NaOH) to remove the scar and convert the scarred nucleotide to a native adenine.
[0174] FIG. 41B is a capillary electrophoresis pherogram showing the results of treatment of an oligonucleotide comprising a N4 Carbamate Sulfide C scarred nucleotide with base (50mM NaOH) to remove the scar and convert the scarred nucleotide to a native cytosine.
[0175] FIG. 42 is a capillary electrophoresis pherogram showing the results of treatment of an oligonucleotide comprising an N6 Carbamate Ethyl A scarred nucleotide with base (50mM NaOH) to remove the scar and convert the scarred nucleotide to a native adenine.
[0176] FIG. 43 shows capillary electrophoresis results for i) an oligonucleotide comprising a N6-linked scarred adenine, and ii) the same oligonucleotide after 30 minute treatment with triethylamine (TEA), where the N6-linked scarred adenine are: N6 Carbamate Propyl A, N6 Carbamate Ethyl A, N6 Amide Propyl A, and N6 Amide Ethyl A.
[0177] FIG. 44 depicts an intramolecular cyclization reaction mechanism with kinetics impacted by the ring size for an N-linked carbamate scar or protecting group.
[0178] FIG. 45 shows the rate of the deprotection reaction of an N-linked scarred nucleotide incorporated into an oligonucleotide for the following scarred nucleotides: N6 Carbamate Ethyl A (large circles; Et-CO2-A), N6 Amide Propyl A (squares, Pr-CO-A), N4 Carbamate Ethyl C (triangles; Et-CO2-C) and N4 Carbamate (Methyl) Ethyl C (small circles, 2MeEt- CO2-C).
[0179] FIG. 46A shows the capillary electrophoresis analysis of uncontrolled oligonucleotide synthesis reactions using dG nucleotides to synthesize a G homopolymer on a 35T starter oligonucleotide (SEQ ID NO: 48). dGTP represents a synthesis performed with unmodified nucleotides. 06 Sulfone G and 06 Sulfide G represent syntheses performed with dGTP nucleotides modified to have a removable protecting group at the base pairing 06 atom of guanine. Oligo synthesis reactions were terminated at 30 seconds, 1 minute, 4 minutes, and 8 minutes to measure progress via capillary electrophoresis, as shown.
[0180] FIG. 46B shows the capillary electrophoresis analysis of uncontrolled oligonucleotide synthesis reactions using dG nucleotides to synthesize a G homopolymer on a 30C starter oligonucleotide (SEQ ID NO: 50). dGTP represents a synthesis performed with unmodified nucleotides. 06 Sulfone G and 06 Sulfide G represent syntheses performed with dGTP nucleotides modified to have a removable protecting group at the base pairing 06 atom of guanine. Oligo synthesis reactions were terminated at 30 seconds, 1 minute, 4 minutes, and 8 minutes to measure progress via capillary electrophoresis, as shown.
[0181] FIG. 47 shows the capillary electrophoresis analysis of uncontrolled oligonucleotide synthesis reactions using dA nucleotides to synthesize an A homopolymer on a 35T starter oligonucleotide (SEQ ID NO: 48). dATP represents a synthesis performed with unmodifiednucleotides. N6 Carbamate Ethyl A and N6 Carbamate Sulfide A represent syntheses performed with dATP nucleotides modified to have a removable protecting group at the base pairing N6 atom of adenine. Oligo synthesis reactions were terminated at 30 seconds, 1 minute, 4 minutes, and 8 minutes to measure progress via capillary electrophoresis, as shown.
[0182] FIG. 48 shows two amino acid ester dTTP analogs used for oligo synthesis and linker cleavage. One is based on a hydroxypropargyl scar (Linker 1) and the other on a smaller hydroxymethyl scar (Linker 2). The two amino acid ester dTTP analogs (Linkers 1 and 2, FIG. 48; synthesized by Jena Bioscience) were attached to cysteine-reactive crosslinkers and conjugated to TdT with the final structure shown in FIG. 48. Also shown are the alcohol-scarred cleavage products after ester cleavage of the linker.
[0183] FIG. 49(A-C) shows a plot of the kinetics of conjugate addition to an unscarred oligo (FIG. 49(A and B)) and to a hydroxymethyl scarred oligo (FIG. 49-C) FIG. 49-A: Natural DNA primer exposed to a dTTP conjugate comprising an ester linkage for 1 second results in -35% extension yield. FIG. 49-B: The oligo synthesis reaction proceeds to completion, with linker cleavage yielding a primer with a hydroxymethyl scar on the last base. FIG. 49-C. Exposure of the scarred primer to the dTTP conjugate comprising an ester linkage for 1 second again results in -35% extension yield.
[0184] FIG. 50 shows the results of a primer extension by TdT- dTTP conjugates based on Linker 1 or Linker 2 as measured by a gel shift assay on SDS-PAGE. An ssDNA primer was extended for 60s with 1) a Linker 1 conjugate, 2) a Linker 2 conjugate 3) a Linker 2 conjugate (replicates), 4) no conjugate. T / P: TdT / DNA primer complex. P: ssDNA primer.
[0185] FIG. 51 shows primer extension products as measured by capillary electrophoresis. Extension was performed by linker 2 conjugates stored overnight at the indicated pH, or in buffer only (negative control). Extension without insertion shows a peak at -58 nt. A peak indicating unwanted insertion (elongation products) is in some samples at -59 nt and indicates the presence of free dNTPs in the incubated conjugate.
[0186] FIG. 52 shows the results of an Enzymatic synthesis of lOOmer and 200mer dT oligos (SEQ ID NOs: 51 and 52, respectively) using the linker 2 dNTP conjugate as measured using capillary electrophoresis (part A). An enlarged view of the product distributions of the 100 mers from enzymatic synthesis (top) and chemical synthesis (bottom) synthesis as observed via capillary electrophoresis is shown in part B.
[0187] FIG. 53 shows the results of an extension of an oligonucleotide using TdT-dATP, - dCTP, -dGTP, and -dTTP conjugates comprising linker 6 as measured by capillaryelectrophoresis (Panel A), and a cleavage time course of a TdT-dTTP conjugate comprising linker 6 incorporated into an oligonucleotide and cleaved via proteinase K for 30-240 seconds, as measured by capillary electrophoresis (Panel B).
[0188] FIG. 54 shows structures for a linker nucleotide comprising a glycine amino acid ester (Gly-OMe-U) and an ACC amino acid ester (ACC-OMe-U) and the product of ester instability of both linkers (HOMe-U) (top), and a comparison of the intact (Gly-OMe-U or ACC-OMe-U) and hydrolized (HOMe-U) product after 60 minutes of exposure to a temperature of 45°C.
[0189] FIGs. 55A, 55B and 55C show a comparison of the linker cleavage efficiency of various TdT-nucleotide conjugates. Data shown is at the time point for 60 seconds of ProK treatment.
[0190] FIG. 56 shows a series of electropherographs characterizing the cleavage rate by Proteinase K (ProK) for illustrative linkers having an aminocyclopropyl carboxy ethyl group and either one (1XG) or two (2XG) glycines. The cleavage reactions were quenched after 15 seconds (s), 30 s, 60 s, 4 minutes (m), 8 m, or 16 m.
[0191] FIG. 57 shows the results of conjugate addition to the primer 3.8 seconds after addition of the conjugate for each of the TdT-nucleotide conjugates (L2 = ACC, Gly-ACC, or 2XGly-ACC).
[0192] FIG. 58 shows a plot of % ester hydrolysis for compounds 14-18 (ring expansion series linker nucleotides) after exposure to 50°C from 1 minute to 20 hours.
[0193] FIG. 59 shows the results of exposure to a temperature of 50°C for 1 hour, 4 hours, or overnight of an oligonucleotide extended with an Allyl G, ACC, AiB, AC4C, AC5C, or AC6C conjugate as measured by capillary electrophoresis to show proportion of intact and hydrolyzed products.
[0194] FIG. 60A shows results of an exemplary single nucleotide addition reaction onto an exemplary single- stranded DNA substrate using an A, C, T, or G polymerase conjugate in the presence (+Phos) or absence (-Phos) of phosphatase. The resulting synthesized oligonucleotides were analyzed by capillary electrophoresis. The x-axes show approximate nucleotide length of oligonucleotides and the y-axes indicate relative fluorescence at 517 nm. Reactions were terminated at the timepoints shown.
[0195] FIG. 60B shows an expanded view of the 21min 41s timepoint results from FIG.60A in present or absence of phosphatases. Specific nucleotides are indicated on each set of panels. Arrows designate the +2 additions.
[0196] FIG. 61 A shows graphical representations of results of capillary electrophoresis analysis of single nucleotide addition reactions onto a single- stranded DNA substrate using a T-polymerase conjugate in the presence of exemplary phosphatase variants from: B. taurus (Quick CIP, NEB), P. borealis (shrimp alkaline phosphatase, NEB), Antarctic bacterium TAB5 (Antarctic phosphatase, NEB), or E. coli (Takara Bio) phosphatase. Synthesis reactions were performed at room temperature (24 °C). A control synthesis reaction was performed without phosphatase. Reactions were terminated at the timepoints shown. The x- axes show relative electrophoretic migration of oligonucleotides (via approximate nucleotide length) and the y-axes indicate relative fluorescence at 517 nm.
[0197] FIG. 6 IB shows graphical representations of results of capillary electrophoresis analysis of an exemplary single nucleotide addition reaction onto a single- stranded DNA substrate using a T-polymerase conjugate in the presence of exemplary phosphatase variants: B. taurus (Quick CIP, NEB), P. borealis (shrimp alkaline phosphatase, NEB), Antarctic bacterium TAB5 (Antarctic phosphatase, NEB), or E. coli (Takara Bio) phosphatase. The synthesis reaction was performed at 37°C (plus and minus phosphatases) and terminated after 30 minutes. The arrow designates the expected size of +2 additions.
[0198] FIG. 62 shows graphical representations of results of an exemplary conjugate-based polynucleotide synthesis of an exemplary 50-mer polynucleotide , conducted in presence or absence of phosphatase, with resulting synthesized polynucleotides distinguished by size along the x-axis using a SeqStudio Genetic Analyzer. Peaks corresponding to the starter oligo and the correct 50-mer synthesis product are labeled.
[0199] FIGs. 63A-63D show a series of electropherograms showing results from analysis of products at different time points in an enzymatic polynucleotide extension reaction performed using a polymerase-nucleotide conjugate in a reaction buffer containing low cobalt acetate concentration (0.05 mM CoOAc) or a standard cobalt acetate concentration. FIG. 63E is a plot showing the quantification and analysis of products in FIGs. 63A-63D, and the associated calculated rates of reaction (kobs).
[0200] FIGs. 64A-64L show a series of electropherograms showing results from analysis of products at different time points in an enzymatic polynucleotide extension reaction performed using a polymerase-nucleotide conjugate in a reaction buffer containing a range of cobalt acetate concentrations (0.05 mM CoOAc, 0.125 mM CoOAc, 0.25 mM CoOAc, 0.75 mM CoOAc, 1.25 mM CoOAc, and 2.5 mM CoOAc). FIG. 64M is a plot showing the quantification and analysis of products in FIGs. 64A-64L, and the associated calculated rates of reaction (kobs).
[0201] FIGs. 65A-65L show a series of electropherograms showing results from analysis of products at different time points in an enzymatic polynucleotide extension reaction performed using a polymerase-nucleotide conjugate in a reaction buffer containing a range of zinc acetate (ZnOAc) concentrations (0.05 mM ZnOAc, 0.125 mM ZnOAc, 0.25 mM ZnOAc, 0.75 mM ZnOAc, 1.25 mM ZnOAc, and 2.5 mM ZnOAc). FIG. 65M is a plot showing the quantification and analysis of products in FIGs. 65A-65L, and the associated calculated rates of reaction (kobs).
[0202] FIGs. 66A-66L show a series of electropherograms showing results from analysis of products at different time points in an enzymatic polynucleotide extension reaction using a free polymerase and free nucleotide performed in a reaction buffer containing a range of cobalt acetate (CoOAc) concentrations (0.05 mM CoOAc, 0.125 mM CoOAc, 0.25 mM CoOAc, 0.75 mM CoOAc, 1.25 mM CoOAc, and 2.5 mM CoOAc). FIG. 66M is a plot showing the quantification and analysis of products in FIGs. 66A-66L.
[0203] FIG. 67 shows the percent of perfect polynucleotides at each step of synthesis of a 520 mer polynucleotide as measured by next generation sequencing.
[0204] FIG. 68 shows the percent of perfect polynucleotides at each step of synthesis of a 1005 mer polynucleotide as measured by next generation sequencing.
[0205] FIG. 69 shows charts characterizing the length and synthesis quality (stepwise yield) of oligonucleotides generated from the process described in Example 35.
[0206] FIG. 70 shows charts characterizing the length and synthesis quality (stepwise yield) of oligonucleotides greater than 1000 nucleotides in length and generated from the process described in Example 35.
[0207] FIG. 71 shows charts characterizing the length and synthesis quality of an ‘all5mer’ oligonucleotide sequence generated from the process described in Example 35.DETAILED DESCRIPTION
[0208] The details of various embodiments of the disclosure are set forth in the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and the drawings, and from the claims.Definitions
[0209] As described herein, compounds of the disclosure may contain “optionally substituted” moieties. In general, the term “substituted”, whether preceded by the term “optionally” or not, means that one or more hydrogens of the designated moiety are replacedwith a suitable substituent. Unless otherwise indicated, an “optionally substituted” group may have a suitable substituent at each substitutable position of the group, and when more than one position in any given structure may be substituted with more than one substituent selected from a specified group, the substituent may be either the same or different at every position. Combinations of substituents envisioned by this disclosure are preferably those that result in the formation of stable or chemically feasible compounds. The term “stable”, as used herein, refers to compounds that are not substantially altered when subjected to conditions to allow for their production, detection, and, in certain embodiments, their recovery, purification, and use for one or more of the purposes disclosed herein.
[0210] The term "alkyl" refers to a straight or branched full saturated hydrocarbon chain. Exemplary alkyl groups are methyl, ethyl, propyl, isopropyl, butyl, isobutyl, and tert-butyl.
[0211] The term "haloalkyl" refers to a straight or branched alkyl group that is substituted with one or more halogen atoms.
[0212] As described herein, compounds of the present disclosure may contain “optionally substituted” moieties. In general, the term “substituted”, whether preceded by the term “optionally” or not, means that one or more hydrogens of the designated moiety are replaced with a suitable substituent. Unless otherwise indicated, an “optionally substituted” group may have a suitable substituent at each substitutable position of the group, and when more than one position in any given structure may be substituted with more than one substituent selected from a specified group, the substituent may be either the same or different at every position. Combinations of substituents envisioned by this present disclosure are preferably those that result in the formation of stable or chemically feasible compounds. The term “stable”, as used herein, refers to compounds that are not substantially altered when subjected to conditions to allow for their production, detection, and, in certain embodiments, their recovery, purification, and use for one or more of the purposes disclosed herein.
[0213] Suitable monovalent substituents on a substitutable carbon atom of an “optionally substituted” group are independently halogen; —(CH2)0-4R∘; —(CH2)0-4OR∘; —O(CH2)0-4R∘, —O—(CH2)0-4C(O)OR∘; —(CH2)0-4CH(OR∘)2; —(CH2)0-4SR∘; —(CH2)0-4Ph, which may be substituted with R∘; —(CH2)0-4O(CH2)0-1Ph which may be substituted with R∘; —CH═CHPh, which may be substituted with R∘; —(CH2)0-4O(CH2)0-1-pyridyl which may be substituted with R∘; —NO2; —CN; —N3; —(CH2)0-4N(R∘)2; —(CH2)0-4N(R∘)C(O)R∘; —N(R∘)C(S)R∘; —(CH2)0-4N(R∘)C(O)NR∘2; —N(R∘)C(S)NR∘2; —(CH2)0-4N(R∘)C(O)OR∘; — N(R∘)N(R∘)C(O)R∘; —N(R∘)N(R∘)C(O)NR∘2; —N(R∘)N(R∘)C(O)OR∘; —(CH2)0-4C(O)R∘;—C(S)R∘; —(CH2)0-4C(O)OR∘; —(CH2)0-4C(O)SR∘; —(CH2)0-4C(O)OSiR∘3; —(CH2)0-4OC(O)R∘; —OC(O)(CH2)0-4SR∘, SC(S)SR∘; —(CH2)0-4SC(O)R∘; —(CH2)0-4C(O)NR∘2; — C(S)NR∘2; —C(S)SR∘; —SC(S)SR∘, —(CH2)0-4OC(O)NR∘2; —C(O)N(OR∘)R∘; — C(O)C(O)R∘; —C(O)CH2C(O)R∘; —C(NOR∘)R∘; —(CH2)0-4SSR∘; —(CH2)0-4S(O)2R∘; — (CH2)0-4S(O)2OR∘; —(CH2)0-4OS(O)2R∘; —S(O)2NR∘2; —(CH2)0-4S(O)R∘; — N(R∘)S(O)2NR∘2; —N(R∘)S(O)2R∘; —N(OR∘)R∘; —C(NH)NR∘2; —P(O)2R∘; —P(O)R∘2; — OP(O)R∘2; —OP(O)(OR∘)2; SiR∘3; —(C1-4 straight or branched alkylene)O—N(R∘)2; or — (C1-4straight or branched alkylene)C(O)O—N(R∘)2, wherein each R∘may be substituted as defined below and is independently hydrogen, C1-6 aliphatic, —CH2Ph, —O(CH2)0-1Ph, — CH2-(5-6 membered heteroaryl ring), or a 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur, or, notwithstanding the definition above, two independent occurrences of R∘, taken together with their intervening atom(s), form a 3-12-membered saturated, partially unsaturated, or aryl mono- or bicyclic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur, which may be substituted as defined below.
[0214] Suitable monovalent substituents on R∘(or the ring formed by taking two independent occurrences of R∘together with their intervening atoms), are independently halogen, —(CH2)0-2R●, -(haloR●), —(CH2)0-2OH, —(CH2)0-2OR●, —(CH2)0-2CH(OR●)2; — O(haloR●), —CN, —N3, —(CH2)0-2C(O)R●, —(CH2)0-2C(O)OH, —(CH2)0-2C(O)OR●, — (CH2)0-2SR●, —(CH2)0-2SH, —(CH2)0-2NH2, —(CH2)0-2NHR●, —(CH2)0-2NR●2, —NO2, — SiR●3, —OSiR●3, —C(O)SR●, —(C1-4 straight or branched alkylene)C(O)OR●, or — SSR●wherein each R●is unsubstituted or where preceded by “halo” is substituted only with one or more halogens, and is independently selected from C1-4 aliphatic, —CH2Ph, — O(CH2)0-1Ph, or a 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur. Suitable divalent substituents on a saturated carbon atom of R∘include ═O and ═S.
[0215] Suitable divalent substituents on a saturated carbon atom of an “optionally substituted” group include the following: ═O, ═S, ═NNR*2, ═NNHC(O)R*, ═NNHC(O)OR*, ═NNHS(O)2R*, ═NR*, ═NOR*, —O(C(R*2))2-3O—, or —S(C(R*2))2-3S—, wherein each independent occurrence of R* is selected from hydrogen, C1-6aliphatic which may be substituted as defined below, or an unsubstituted 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur. Suitable divalent substituents that are bound to vicinalsubstitutable carbons of an “optionally substituted” group include: —O(CR*2)2-3O—, wherein each independent occurrence of R* is selected from hydrogen, C1-6 aliphatic which may be substituted as defined below, or an unsubstituted 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.
[0216] Suitable substituents on the aliphatic group of R* include halogen, —R●, -(haloR●), —OH, —OR●, —O(haloR●), —CN, —C(O)OH, —C(O)OR●, —NH2, —NHR●, —NR●2, or —NO2, wherein each R●is unsubstituted or where preceded by “halo” is substituted only with one or more halogens, and is independently C1-4 aliphatic, —CH2Ph, —O(CH2)0-1Ph, or a 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.
[0217] Suitable substituents on a substitutable nitrogen of an “optionally substituted” group include —R†, —NR†2, —C(O)R†, —C(O)OR†, —C(O)C(O)R†, —C(O)CH2C(O)R†, — S(O)2R†, —S(O)2NR†2, —C(S)NR†2, —C(NH)NR†2, or —N(R†)S(O)2R†; wherein each R†is independently hydrogen, C1-6 aliphatic which may be substituted as defined below, unsubstituted —OPh, or an unsubstituted 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur, or, notwithstanding the definition above, two independent occurrences of R†, taken together with their intervening atom(s) form an unsubstituted 3-12-membered saturated, partially unsaturated, or aryl mono- or bicyclic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.
[0218] Suitable substituents on the aliphatic group of R†are independently halogen, —R●, - (haloR●), —OH, —OR●, —O(haloR●), —CN, —C(O)OH, —C(O)OR●, —NH2, —NHR●, — NR●2, or —NO2, wherein each R●is unsubstituted or where preceded by “halo” is substituted only with one or more halogens, and is independently C1-4 aliphatic, —CH2Ph, — O(CH2)0-1Ph, or a 5-6-membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.
[0219] The recitation of a listing of chemical groups in any definition of a variable herein includes definitions of that variable as any single group or combination of listed groups. The recitation of an embodiment for a variable herein includes that embodiment as any single embodiment or in combination with any other embodiments or portions thereof.
[0220] In an alternative embodiment, compounds described herein may also comprise one or more isotopic substitutions. For example, hydrogen may be2H (D or deuterium) or3H (T or tritium); carbon may be for example13C or14C; oxygen may be for example18O;nitrogen may be, for example,15N, and the like. In other embodiments, a particular isotope (e.g.,3H,13C,14C,18O, or15N) can represent at least 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or at least 99.9% of the total isotopic abundance of an element that occupies a specific site of the compound.
[0221] As used herein, the terms “about” and “approximately” refer to a value or composition that is within an acceptable error range for the particular value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, “about” or “approximately” can mean within one or more than one standard deviation per the practice in the art. Alternatively, “about” or “approximately” can mean a range of up to 10% (i.e., ±10%) or more depending on the limitations of the measurement system. For example, about 5 mg can include any number between 4.5 mg and 5.5 mg. Furthermore, particularly with respect to biological systems or processes, the terms can mean up to an order of magnitude or up to 5-fold of a value. When particular values or compositions are provided in the instant disclosure, unless otherwise stated, the meaning of “about” or “approximately” should be assumed to be within an acceptable error range for that particular value or composition. Also, where ranges and / or subranges of values are provided, the ranges and / or subranges can include the endpoints of the ranges and / or subranges.
[0222] The terms “nucleic acid”, “polynucleotide” and “oligonucleotide” and other related terms used herein are used interchangeably and refer to polymers of nucleotides and are not limited to any particular length. Nucleic acids include recombinant and chemically- synthesized forms. Nucleic acids can be isolated. Nucleic acids include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), analogs of the DNA or RNA generated using nucleotide analogs (e.g., peptide nucleic acids (PNA) and non-naturally occurring nucleotide analogs), and chimeric forms containing DNA and RNA. Nucleic acids can be single-stranded or double-stranded. Nucleic acids comprise polymers of nucleotides, where the nucleotides include natural or non-natural bases and / or sugars. Nucleic acids comprise naturally-occurring internucleosidic linkages, for example phosphodiester linkages. Nucleic acids can lack a phosphate group. Nucleic acids comprise non-natural internucleoside linkages, including phosphorothioate, phosphorothiolate, or peptide nucleic acid (PNA) linkages. In some embodiments, nucleic acids comprise a one type of polynucleotides or a mixture of two or more different types of polynucleotides
[0223] The term “operably linked” and “operably joined” or related terms as used herein refers to juxtaposition of components. The juxtapositioned components can be linked together covalently. For example, two nucleic acid components can be enzymatically ligated together where the linkage that joins together the two components comprises phosphodiester linkage. A first and second nucleic acid component can be linked together, where the first nucleic acid component can confer a function on a second nucleic acid component. For example, linkage between a primer binding sequence and a sequence of interest forms a nucleic acid library molecule having a portion that can bind to a primer. In another example, a transgene (e.g., a nucleic acid encoding a polypeptide or a nucleic acid sequence of interest) can be ligated to a vector where the linkage permits expression or functioning of the transgene sequence contained in the vector. In some embodiments, a transgene is operably linked to a host cell regulatory sequence (e.g., a promoter sequence) that affects expression of the transgene. In some embodiments, the vector comprises at least one host cell regulatory sequence, including a promoter sequence, enhancer, transcription and / or translation initiation sequence, transcription and / or translation termination sequence, polypeptide secretion signal sequences, and the like. In some embodiments, the host cell regulatory sequence controls expression of the level, timing and / or location of the transgene.
[0224] The terms “linked”, “joined”, “attached”, “appended” and variants thereof comprise any type of fusion, bond, adherence or association between any combination of compounds or molecules that is of sufficient stability to withstand use in the particular procedure. The procedure can include but are not limited to: nucleotide binding; nucleotide incorporation; de-blocking (e.g., removal of chain-terminating moiety); washing; removing; flowing; detecting; imaging and / or identifying. Such linkage can comprise, for example, covalent, ionic, hydrogen, dipole-dipole, hydrophilic, hydrophobic, or affinity bonding, bonds or associations involving van der Waals forces, mechanical bonding, and the like. In some embodiments, such linkage occurs intramolecularly, for example linking together the ends of a single-stranded or double-stranded linear nucleic acid molecule to form a circular molecule. In some embodiments, such linkage can occur between a combination of different molecules, or between a molecule and a non-molecule, including but not limited to: linkage between a nucleic acid molecule and a solid surface; linkage between a protein and a detectable reporter moiety; linkage between a nucleotide and detectable reporter moiety; and the like. Some examples of linkages can be found, for example, in Hermanson, G., “Bioconjugate Techniques”, Second Edition (2008); Aslam, M., Dent, A., “Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences” London: Macmillan (1998); Aslam, M.,Dent, A., “Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences”, London: Macmillan (1998).
[0225] When used in reference to nucleic acids, the terms “extend”, “extending”, “extension” and other variants, refers to incorporation of one or more nucleotides into a nucleic acid molecule (i.e. nucleotide coupling). Nucleotide incorporation comprises polymerization of one or more nucleotides into the terminal 3′ OH end of a nucleic acid strand (e.g., a nucleic acid primer), resulting in extension of the nucleic acid strand (e.g., starter oligo). Nucleotide incorporation can be conducted with natural nucleotides and / or nucleotide analogs.
[0226] The terms “cleavable linker” or “cleavable moiety” as used herein refers to a divalent or monovalent, respectively, moiety which is capable of being separated (e.g., detached, split, disconnected, hydrolyzed, a stable bond within the moiety is broken) into distinct entities. In embodiments, a cleavable linker is cleavable (e.g., specifically cleavable) in response to external stimuli (e.g., enzymes, nucleophilic / basic reagents, reducing agents, photo-irradiation, electrophilic / acidic reagents, organometallic and metal reagents, or oxidizing reagents).
[0227] Use of the term “cleavable linker” is not meant to imply that the whole linker is required to be removed. The cleavage site can be located at a position on the linker that ensures that part of the linker remains attached to the dye and / or substrate moiety after cleavage. Cleavable linkers may be, by way of non-limiting example, electrophilically cleavable linkers, nucleophilically cleavable linkers, photocleavable linkers, cleavable under reductive conditions (for example disulfide or azide containing linkers), oxidative conditions, cleavable via use of safety-catch linkers and cleavable by elimination mechanisms. The use of a cleavable linker to attach the dye compound to a substrate moiety ensures that the label can, if required, be removed after detection, avoiding any interfering signal in downstream steps.
[0228] In embodiments, the cleavable linker is cleaved by contacting the cleavable linker with a cleaving agent (e.g., a reducing agent). In embodiments, the cleaving agent is…
[0229] The term “polymerase-compatible cleavable moiety” and “polymerase-compatible cleavable linker” as used herein refers to a cleavable moiety or cleavable linker which does not interfere with the function of a polymerase (e.g., a DNA polymerase or modified DNA polymerase, in incorporating the nucleotide, to which the polymerase-compatible cleavable moiety is attached, to the 3′ end of the newly formed nucleotide strand). Methods fordetermining the function of a polymerase contemplated herein are described in B. Rosenblum et al. (Nucleic Acids Res. 1997 Nov 15; 25(22): 4500- 4504); and Z. Zhu et al. (Nucleic Acids Res. 1994 Aug 25; 22(16): 3418-3422), which are incorporated by reference herein in their entirety for all purposes. In embodiments the polymerase-compatible cleavable moiety does not decrease the function of a polymerase relative to the absence of the polymerase- compatible cleavable moiety. In embodiments, the polymerase-compatible cleavable moiety does not negatively affect DNA polymerase recognition. In embodiments, the polymerase- compatible cleavable moiety does not negatively affect (e.g., limit) the read length of the DNA polymerase.
[0230] As used herein, the term “nucleotide” refers to a molecule comprising a nucleoside and one or more phosphate groups. A “nucleoside” refers to a molecule comprising a nucleobase (e.g., adenine, thymine, cytosine, guanine, or uracil) and a five carbon sugar (e.g., ribose or 2'-deoxyribose). Exemplary nucleotides can be or comprise, without limitation, a nucleoside monophosphate, a nucleoside diphosphate, a nucleoside triphosphate, a nucleoside tetraphosphate, a nucleoside pentaphosphate, or a nucleoside hexaphosphate. As provided herein, TdT and TdT variants, can, in some embodiments, incorporate any nucleoside polyphosphate, including nucleotide analogs comprising modifications to the nucleobase.
[0231] As used herein, the term “nucleoside polyphosphate” is a “nucleotide” and may be called a “nucleotide polyphosphate.” For example, a “nucleotide triphosphate” and a “nucleoside triphosphate” both refer to a nucleotide comprising a nucleobase, a sugar, and a polyphosphate consisting of three linked phosphate groups.
[0232] As used herein, “non-termination” or “insertion” occurs when more than one nucleotide is added during a single step of a cyclic nucleotide extension. This can occur when an unshielded nucleotide with an uncleaved 5′ phosphate is added to an oligonucleotide.
[0233] As used herein, the term “phosphatase” refers to an enzyme capable of removing the 5′ phosphate of a nucleotide, especially a nucleotide that is unshielded as part of an improperly formed conjugate or is not tethered to a polymerase. When not referring to a specific phosphatase enzyme, as used in herein, phosphatase is meant to also include all phosphatase enzymes, engineered enzymes having phosphatase activity, or a functional fragment thereof, that is capable of removing one or more phosphate group(s) from a nucleotide. A phosphatase can also refer to any biomolecule (e.g., a polypeptide or ribozyme) capable of removing one or more phosphate group(s) from a nucleotide, including an engineered enzyme having phosphatase activity or functional fragments thereof
[0234] As used herein, the term “blocked nucleotide” or “shielded nucleotide” refers to a nucleotide that is sterically hindered by a tethered polymerase (or other entity or component such as, e.g., a blocking group) from a phosphatase capable of removing its 5′ phosphate. In some embodiments, such nucleotides are likely to inhibit subsequent nucleotide additions after having been added to an oligonucleotide and before removal of said tethered polymerase.
[0235] As used herein, the term “unblocked nucleotide” or “unshielded nucleotide” refers to a nucleotide that is not sterically hindered by a tethered polymerase (or other entity or component such as, e.g., a blocking group) from a phosphatase capable of removing its 5′ phosphate. In some embodiments, an unshielded nucleotide may be tethered to a polymerase, such as in a misfolded polymerase or tethered at an incorrect position. An unshielded nucleotide may be untethered (or free) from a polymerase. Unshielded nucleotides that have not been exposed to phosphatase are more likely to be erroneously added to a polynucleotide as an insertion after a shielded nucleotide has been properly added. Overview – Long Polynucleotide Synthesis
[0236] One of the major challenges in the field of polynucleotide synthesis is the do novo synthesis of long polynucleotide sequences. Because nucleotides are added one at a time during the synthesis process, it can be difficult to accurately and efficiently synthesize very long sequences without introducing errors. Even small stepwise error rates during synthesis can quickly accumulate to prevent synthesis of long sequences.
[0237] Described herein are reagents, systems, and methods for the de novo stepwise template-independent synthesis of long, high-quality polynucleotides, e.g., using enzymatic polynucleotide synthesis.
[0238] In particular, described herein are improved reagents, systems, and methods for stepwise polynucleotide synthesis that mitigate common errors during a synthesis reaction, such as an unwanted nucleotide insertion (e.g., due to addition of an incompletely blocked nucleotide to a growing polynucleotide strand, or premature removal of a blocking group during an extension reaction), or an unwanted nucleotide deletion (e.g., due to an incomplete reaction, such as failure to add a nucleotide during an extension reaction, or failure to remove a blocking group during a nucleotide deblocking reaction.).
[0239] In some embodiments, provided herein are modified nucleotides that inhibit secondary structure formation during enzymatic polynucleotide synthesis. Such secondary structure formation can inhibit extension of a growing polynucleotide by certainpolymerases, such as TdT. In some embodiments, provided herein are optimized reagent and cofactor conditions for improved enzymatic polynucleotide synthesis activity. Such improved reagents and methods inhibit deletions due to incomplete extensions, thereby improving stepwise yields of long polynucleotide synthesis.
[0240] In some embodiments, provided herein are reagents to inhibit unwanted insertions, such as phosphatase treatment of nucleotide-linker conjugates, which inhibits activity of unblocked nucleotides during the extension step.
[0241] In some embodiments, provided herein are improved linkers between blocking groups (e.g., TdT) and the nucleotide, that are stable during the extension step (to inhibit unwanted insertions), but quickly and completely cleave during the blocking group removal step (to inhibit unwanted deletions).
[0242] Also provided herein, in some embodiments, are improved systems for oligonucleotide synthesis, such as a synthesis surface that is dipped into the appropriate pre- prepared extension, blocking group removal, and wash buffers to improve synthesis speed and reagent delivery. Such systems can also act to improve overall stepwise yield and reaction cycle speed.
[0243] In some embodiments, provided herein are methods of de novo template- independent synthesis of a polynucleotide at least 500 bp in length with an observed error rate of less than 1% as compared to a predetermined sequence.
[0244] In some embodiments, the de novo synthesized long polynucleotides are at least 600, at least 700, at least 800, at least 900, or at least 1000 nucleotides in length. In some embodiments, the de novo synthesized long polynucleotides are from 500 to 1000 nucleotides in length.
[0245] In some embodiments, the observed error rate of each coupling during de novo template-independent synthesis of a long polynucleotide is less than 0.9%, less than 0.8%, less than 0.7%, less than 0.6%, less than 0.5%, less than 0.4%, less than 0.3%, less than 0.2%, or less than 0.1% per cycle.
[0246] In some embodiments, the observed stepwise yield of a de novo template- independent synthesis of a long polynucleotide is greater than 99%, greater than 99.1%, greater than 99.2%, greater than 99.3%, greater than 99.4%, greater than 99.5%, greater than 99.6%, greater than 99.7%, greater than 99.8%, or greater than 99.9%. In some embodiments, the observed stepwise yield of a de novo template-independent synthesis of a long polynucleotide is from 99% to 99.9%.
[0247] In some embodiments, each nucleotide coupling step of a de novo template- independent synthesis of a long polynucleotide is performed in less than 120 seconds, less than 90 seconds, less than 60 seconds, less than 50 seconds, or less than 40 seconds. In some embodiments, each nucleotide coupling step of a de novo template-independent synthesis of a long polynucleotide is performed at a rate of at least 10 nucleotides per hour, at least 15 nucleotides per hour, or at least 20 nucleotides per hour. In some embodiments, the nucleotide coupling to yield said de novo synthesized polynucleotides is performed at a rate of from 5 nucleotides per hour to 25 nucleotides per hour.
[0248] In some embodiments, the de novo template-independent polynucleotide synthesis generates at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 2000, or at least 5000 copies of each sequence.
[0249] The overall error rate or error rates for individual types of errors such as deletions, insertions, or substitutions for each oligonucleotide synthesized on the substrate, for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, 99.5%, or more of the oligonucleotides synthesized on the substrate, or for the substrate average may be at most or at most about 1:100, 1:500, 1:1000, 1:10000, 1:20000, 1:30000, 1:40000, 1:50000, 1:60000, 1:70000, 1:80000, 1:90000, 1:1000000, or less. The overall error rate or error rates for individual types of errors such as deletions, insertions, or substitutions for each oligonucleotide synthesized on the substrate, for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, 99.5%, or more of the oligonucleotides synthesized on the substrate, or the substrate average may fall between 1:100 and 1:10000, 1:500 and 1:30000. Those of skill in art, appreciate that the overall error rate or error rates for individual types of errors such as deletions, insertions, or substitutions for each oligonucleotide synthesized on the substrate, for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, 99.5%, or more of the oligonucleotides synthesized on the substrate, or the substrate average may fall between any of these values, for example 1:500 and 1:10000. The overall error rate or error rates for individual types of errors such as deletions, insertions, or substitutions for each oligonucleotide synthesized on the substrate, for at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, 99.5%, or more of the oligonucleotides synthesized on the substrate, or the substrate average may fall between any range defined by any of the values serving as endpoints of the range.
[0250] Desired predetermined sequences may be supplied by any method, typically by a user, e.g., a user entering data using a computerized system. In various embodiments, synthesized nucleic acids are compared against these predetermined sequences in some casesby sequencing at least a portion of the synthesized nucleic acids, e.g., using next-generation sequencing methods.
[0251] The polynucleotides can be released from the substrate by a variety of suitable methods as described in further details elsewhere herein and known in the art, for example by enzymatic cleavage, as is well known in that art. Examples of such enzymatic cleavage include, but are not limited to, the use of restriction enzymes such as MIyI, or other enzymes or combinations of enzymes capable of cleaving single or double-stranded DNA such as, but not limited to, Uracil DNA glycosylase (UDG) and DNA Endonuclease IV. Other methods of cleavage known in the art may also be advantageously employed in the present disclosure, including, but not limited to, chemical (base labile) cleavage of DNA molecules or optical (photolabile) cleavage from the surface. PCR or other amplification reactions can also be employed to generate building material for gene synthesis by copying the oligonucleotides while they are still anchored to the substrate. Methods of releasing polynucleotides are described in P.C.T. Patent Publication No. WO2007137242, and U.S. Pat. No. 5,750,672 which is herein incorporated by reference in its entirety. Enzymatic Polynucleotide Synthesis
[0252] In some embodiments, provided herein is a method of polynucleotide synthesis to generate polynucleotides of desired length and sequence according to the embodiments described herein. In some embodiments, the steps are performed by dipping a reaction surface comprising a bound synthesis initiator (which may also include previously added nucleotides) into a contained solution comprising the desired reagents. The method is amenable to multiplexed polynucleotide synthesis, such that a plurality of elements having an end with the reaction surface and bound synthesis initiator may be used and simultaneously dipped into a plurality of contained liquid reagents (such as in wells or droplets on a surface) aligned with the reaction surfaces. In some embodiments, the contained liquid reagent comprise a polymer extension solution with a nucleotide of a specific identity and polymerase capable of adding the nucleotide to the synthesis initiator.
[0253] In some embodiments, synthesis of a polynucleotide comprises adding blocked nucleotides stepwise to an oligonucleotide bound to the reaction surface on the element via the cycled steps of: addition of nucleotide comprising a blocking group (i.e., a blocked nucleotide) to a synthesis initiator or extended polynucleotide comprising previously added nucleotides, binding of the nucleotide to the end of the synthesis initiator or extended polynucleotide catalyzed by the polymerase, and removal of the blocking group from thenucleotide to allow addition of a subsequent nucleotide to the extended polynucleotide. These steps can be repeated until a desired polynucleotide sequence and length is synthesized.
[0254] The blocking group bound to the nucleotide (e.g., a polymerase or a reversible terminator) is a group capable of preventing addition of another nucleotide once the nucleotide has been added to the synthesis initiator or extended polynucleotide. After addition of the desired blocked nucleotide and removal of excess nucleotide during an extension cycle, the extended polynucleotide is immersed in a nucleotide deblocking solution capable of removing the blocking group from the nucleotide.
[0255] In some embodiments, the blocking group is the polymerase that catalyzes addition of the nucleotide to the surface-bound polynucleotide, wherein the polymerase is linked to the nucleotide (i.e., a nucleotide-polymerase conjugate). In this embodiment, the polymerase can sterically hinder addition of a subsequent nucleotide after addition of the blocked nucleotide to the polynucleotide. A monomer deblocking solution that removes the polymerase from the nucleotide can then be used to remove the blocking group, such as a linker cleavage solution.
[0256] In some embodiments, the blocking group is a reversible terminator bound to the nucleotide. A monomer deblocking solution that removes the reversible terminator from the nucleotide can then be used to remove the blocking group, such as a linker cleavage solution. In some embodiments, both a reversible terminator and a polymerase bound to the nucleotide may be used.
[0257] Both the nucleotide addition and blocking group removal steps may be quenched by immersing the extended polynucleotide in an appropriate reaction quenching solution, such as EDTA. In addition, washing steps may be used between steps by immersing the extended polynucleotide in a wash buffer. Conjugate-based Polynucleotide Synthesis
[0258] In some embodiments, the present disclosure includes use of TdT with free nucleotides that have a 3′ modification to enable single extensions. In some embodiments, the present disclosure also includes use of TdT with a tethered nucleotide (we call this polymerase-nucleotide conjugate). Linkage of the dNTP can occur via a tether to the nucleobase. In some embodiments, a nucleotide comprises an optionally substituted O-alkyl group. Additional tethered nucleotides can be found, e.g., in PCT PublicationWO2017 / 223517 “Nucleic Acid Synthesis and Sequencing Using Tethered Nucleoside Triphosphates,” the entirety of which is incorporated by reference.
[0259] Described herein is a method for the de novo synthesis of nucleic acids using conjugates comprising a polymerase and a nucleoside triphosphate. In some embodiments of the method, conjugates comprise the polymerase Terminal deoxynucleotidyl Transferase (TdT). In other embodiments, the method may employ conjugates comprising another template-independent polymerase.
[0260] FIG. 1A illustrates a typical process for the stepwise synthesis of a defined sequence using a template-independent polymerase. A nucleic acid that serves as an initial substrate for elongation (i.e., "starter molecule") is incubated with a first polymerase-nucleotide conjugate. Once the nucleic acid has been elongated by the tethered nucleotide of a conjugate, no further elongations occur because the conjugates implement a termination mechanism. In the second step of the process, the linker is cleaved to release the polymerase and reverse the termination mechanism, thus enabling subsequent elongations. The elongation products are then exposed to the second conjugate, and these two steps are iterated to elongate the nucleic acid by a defined sequence. FIG. 1B illustrates a synthesis procedure using a conjugate comprising TdT and a photocleavable linker. As described above, other strategies are available for the attachment and cleavage of the linker.
[0261] For RNA synthesis applications, tethered ribonucleoside triphosphates may be used. In these embodiments, an RNA specific nucleotidyl transferase, such as E. coli Poly(A) Polymerase (IUBMB EC 2.7.7.19) or Poly(U) Polymerase, among others, may be employed. The RNA nucleotidyl transferases can contain modifications, e.g., single point mutations, that influence the substrate specificity towards a specific rNTP (Lunde et al., Nucleic acids research 40.19 (2012): 9815-9824.). In some embodiments, a very short tether between an RNA nucleotidyl transferase and a ribonucleoside triphosphate may be used to induce a high effective concentration of the nucleoside triphosphate, thereby forcing incorporation of an rNTP that might not be the natural substrate of the nucleotidyl transferase.
[0262] When a conjugate comprising a polymerase and a nucleoside triphosphate is incubated with a nucleic acid, it preferentially elongates the nucleic acid using its tethered nucleotide (as opposed to using the nucleotide of another conjugate molecule). As described above, the polymerase then remains attached to the nucleic acid via its tether to the added nucleotide until exposed to some stimulus that causes cleavage of the linkage to the added nucleotide. In this situation, further extensions by polymerase- nucleotide conjugates are hindered due to "shielding" when: 1) the attached polymerase molecule hinders otherconjugates from accessing the 3′ OH of the extended DNA molecule and 2), other nucleoside triphosphates in the system are hindered from accessing the catalytic site of the polymerase that remains attached to the 3′ end of the extended nucleic acid. (The extent of shielding may be described as the extent to which both of these interactions are hindered.) To enable subsequent extensions, the linker tethering the incorporated nucleotide to the polymerase can be cleaved, releasing the polymerase from the nucleic acid and therefore re-exposing its 3′ OH group for subsequent elongation.
[0263] Methods for nucleic acid synthesis provided herein that employ the shielding effect to achieve termination comprise an extension step wherein a nucleic acid is exposed to conjugates preferentially in the absence of free (i.e., untethered) nucleoside triphosphates, because the termination mechanism of shielding may not prevent their incorporation into the nucleic acid.
[0264] In some embodiments, termination of further elongation may be "complete", meaning that after a nucleic acid molecule has been elongated by a conjugate, further elongations cannot occur during the reaction. In other embodiments, termination of further elongation may be "incomplete", meaning that further elongations can occur during the reaction but at a substantially decreased rate compared to the initial elongation, e.g., 100 times slower, or 1000 times slower, or 10,000 times slower, or more. Conjugates that achieve incomplete termination may still be used to extend a nucleic acid by predominantly a single nucleotide (e.g., in methods for nucleic acid synthesis and sequencing) when the reaction is stopped after an appropriate amount of time. In some embodiments, the reagent containing the conjugate may additionally contain polymerases without tethered nucleoside triphosphates, but those polymerases should not significantly affect the reaction because there are no free dNTPs in the mix.
[0265] Reagents based on conjugates employing the shielding effect to achieve termination preferentially only contain polymerase-nucleotide conjugates in which all polymerases remain folded in the active conformation. In some cases, if the polymerase moiety of a conjugate is unfolded, its tethered nucleoside triphosphate may become more accessible to the polymerase moieties of other conjugate molecules. In these cases, the unshielded nucleotides may be more readily incorporated by other conjugate molecules, circumventing the termination mechanism.
[0266] Polymerase-nucleotide conjugates employing the shielding effect to achieve termination are preferentially only labeled with a single nucleoside triphosphate moiety. Polymerase nucleotide conjugates labeled with multiple nucleoside triphosphates that canaccess the catalytic site can, in some cases, incorporate multiple nucleoside triphosphates into the same nucleic acid. Additional tethered nucleotides may therefore lead to additional, undesired nucleotide incorporations into a nucleic acid during a reaction. Furthermore, only one tethered nucleoside triphosphates can occupy the (buried) catalytic site of its polymerase at a time so the other tethered nucleoside triphosphate(s) may have an increasing accessibility to the polymerase moieties of other conjugate molecules, as discussed below.
[0267] Polymerase-nucleotide conjugates employing the shielding effect to achieve termination preferentially comprise as short of a linker as possible that still enables the nucleoside triphosphate to frequently access the catalytic site of its tethered polymerase molecule in a productive conformation, in order to enable fast incorporation of the nucleotide into a nucleic acid. Such conjugates may also preferentially employ an attachment position of the linker to the polymerase as close to the catalytic site as possible, enabling use of a shorter linker. The length of the linker will determine the maximum distance from the attachment point a tethered nucleoside triphosphate or a tethered nucleic acid can reach. A smaller distance may lead to a reduced accessibility of the tethered moiety to other polymerase- nucleotide molecules, as discussed below. In some embodiments, linkers are approximately 24 and 28 Å long. Shorter linkers, e.g., with lengths of 8-15 Å may increase shielding; longer linkers, e.g., linkers longer than 50 Å, 70 Å or 100 Å, may reduce shielding. The shielding effect may be influenced by a combination of factors including, but not limited to, to the structure of the polymerase, the length of the linker, the structure of the linker, the attachment position of the linker to the polymerase, the binding affinity of the nucleoside triphosphate to the catalytic site of the polymerase, the binding affinity of the nucleic acid to the polymerase, the preferred conformation of the polymerase, and the preferred conformation of the linker.
[0268] One contribution to shielding can be steric effects that block the 3' OH of a nucleic acid that has been elongated by a conjugate from reaching into the catalytic site of another conjugate's polymerase moiety. Steric effects may also hinder a tethered nucleoside triphosphate from reaching into the catalytic site of another polymerase-nucleotide conjugate molecule due to clashes between the conjugates that would occur during such approaches. These steric effects may result in complete termination if they completely block productive interactions between the tethered nucleoside triphosphate (or elongated nucleic acid) of one conjugate molecule with another conjugate molecule, or may result in incomplete termination if they only hinder such intermolecular interactions.
[0269] Another contribution to shielding arises from the binding affinity of the tethered nucleoside triphosphate to the catalytic site of the polymerase The tethered nucleosidetriphosphate of a conjugate will have a high effective concentration with respect to the catalytic site of its tethered polymerase so it may remain bound to that site much of the time. When the nucleoside triphosphate is bound to the catalytic site of its tethered polymerase molecule it is unavailable for incorporation by other polymerase molecules. Thus, tethering reduces the effective concentration of nucleoside triphosphates available for intermolecular incorporation (i.e., incorporation catalyzed by a polymerase molecule to which the nucleotide is not tethered). This shielding effect can enhance termination by reducing the rate by which a nucleic acid is elongated using the nucleoside triphosphate moiety of one conjugate molecule by the polymerase moiety of another conjugate molecule.
[0270] Another contribution to shielding arises from the binding affinity of the 3′ region of a nucleic acid molecule to the catalytic site of a polymerase molecule. After elongation by a conjugate, the nucleic acid is tethered to the conjugate via its 3′ terminal nucleotide and will have a high effective concentration with respect to the catalytic site of its tethered polymerase so it may remain bound to that site much of the time. When the nucleic acid is bound to the catalytic site of its tethered polymerase molecule it is unavailable for elongation by other conjugate molecules. This effect can enhance termination by reducing the rate by which a nucleic acid that has been elongated by a first conjugate is further elongated by other conjugate molecules.
[0271] In some embodiments, the polymerase-nucleotide conjugates comprise additional moieties that sterically hinder the tethered nucleoside triphosphate (or a tethered nucleic acid post-elongation) from approaching the catalytic sites of another conjugate molecule. Such moieties include polypeptides or protein domains that can be inserted into a loop of the polymerase, and those and other bulky molecules such as polymers that can be site- specifically ligated e.g., to an inserted unnatural amino acid or specific polypeptide tag.
[0272] In some embodiments, the linker is attached the 5 position of pyrimidines or the 7 position of 7-deazapurines. In other embodiments, the linker may be attached to an exocyclic amine of a nucleobase, e.g., by N-alkylating the exocyclic amine of cytosine with a nitrobenzyl moiety as discussed below. In other embodiments, the linker may be attached to any other atom in the nucleobase, sugar, or oc-phosphate, as will be apparent to those skilled in the art.
[0273] Certain polymerases have a high tolerance for modification of certain parts of a nucleotide, e.g., modifications of the 5 position of pyrimidines and the 7 position of purines are well-tolerated by some polymerases (He and Seela., Nucleic Acids Research 30.24(2002): 5485-5496.; or Hottin et al., Chemistry. 2017 Feb 10;23(9):2109-2118). In some embodiments, the linker is attached to these positions. Preparation of a nucleotide-polymerase conjugate
[0274] In some examples, a polymerase-nucleotide conjugate is prepared by first synthesizing an intermediate compound comprising a linker and a nucleoside triphosphate (referred to herein as a "linker-nucleotide"), and then this intermediate compound is attached to the polymerase.
[0275] In some examples, nucleosides with substitutions compared to natural nucleosides, e.g., pyrimidines with 5-hydroxymethyl or 5-propargylamino substituents, or 7- deazapurines with 7-hydroxymethyl or 7-propargylamino substituents may be useful starting materials for preparing linker- nucleotides. An exemplary set of nucleosides with 5- and 7- hydroxymethyl substituents that may be useful for preparing linker-nucleotides is shown below:
[0276] An exemplary set of nucleosides with 5- and 7-deaza-7-propargylamino substituents that may be useful for preparing linker-nucleotides is shown below:
[0277] These nucleosides are also commercially available as deoxyribonucleoside triphosphates.
[0278] In some embodiments, the tethered nucleotide may be specifically attached to a cysteine residue of the polymerase using a sulfhydryl-specific attachment chemistry. Possible sulfhydryl specific attachment chemistries include, but are not limited to ortho-pyridyl disulfide (OPSS), maleimide functionalities, 3- arylpropiolonitrile functionalities, allenamide functionalities, haloacetyl functionalities such as iodoacetyl or bromoacetyl, alkyl halides or perfluroaryl groups that can favorably react with sulfhydryls surrounded by a specific amino acid sequence (Zhang, Chi, et al. Nature chemistry 8, (2015) 120-128.). Other attachment chemistries for specific labeling of cysteine residues will be apparent to those skilled in theart or are described in the pertinent literature and texts (e.g., Kim, Younggyu, et al, Bioconjugate chemistry 19.3 (2008): 786-791.).
[0279] In other embodiments, the linker could be attached to a lysine residue via an amine - reactive functionality (e.g., NHS esters, Sulfo-NHS esters, tetra- or pentafluorophenyl esters, isothiocyanates, sulfonyl chlorides, etc.). In other embodiments, the linker may be attached to the polymerase via attachment to a genetically inserted unnatural amino acid, e.g., p- propargyloxyphenylalanine or p- azidophenylalanine that could undergo azide-alkyne Huisgen cycloaddition, though many suitable unnatural amino acids suitable for site-specific labeling exist and can be found in the literature (e.g., as described in Lang and Chin., Chemical reviews 114.9 (2014): 4764-4806.).
[0280] In other embodiments, the linker may be specifically attached to the polymerase N- terminus. In some embodiments, the polymerase is mutated to have an N-terminal serine or threonine residue, which may be specifically oxidized to generate an N-terminal aldehyde for subsequent coupling to e.g., a hydrazide. In other embodiments, the polymerase is mutated to have an N-terminal cysteine residue that can be specifically labeled with an aldehyde to form a thiazolidine. In other embodiments, an N-terminal cysteine residue can be labeled with a peptide linker via Native Chemical Ligation.
[0281] In other embodiments, a peptide tag sequence may be inserted into the polymerase that can be specifically labeled with a synthetic group by an enzyme, e.g., as demonstrated in the literature using biotin ligase, transglutaminase, lipoic acid ligase, bacterial sortase and phosphopantetheinyl transferase (e.g., as described in refs. 74-78 of Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).
[0282] In other embodiments, the linker is attached to a labeling domain fused to the polymerase. For example, a linker with a corresponding reactive moiety may be used to covalently label SNAP tags, CLIP tags, HaloTags and acyl carrier protein domains (e.g., as described in refs. 79-82 of Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).
[0283] In other embodiments, the linker is attached to an aldehyde specifically generated within the polymerase, as described in Carrico et al. (Nat. Chem. Biol. 3, (2007) 321 - 322). For example, after insertion of an amino acid sequence that is recognized by the enzyme formylglycine-generating enzyme (FGE) into the polymerase, it may be exposed to FGE, which will specifically convert a cysteine residue in the recognition sequence to formylglycine (i.e., producing an aldehyde). This aldehyde may then be specifically labeled with e.g., a hydrazide or aminooxy moiety of a linker.
[0284] In some embodiments, a linker may be attached to the polymerase via non-covalent binding of a moiety of the linker to a moiety fused to the polymerase. Examples of such attachment strategies include fusing a polymerase to streptavidin that can bind a biotin moiety of a linker, or fusing a polymerase to anti-digoxigenin that can bind a digoxigenin moiety of a linker. In some embodiments, site-specific labeling may lead to an attachment of the linker to the polymerase that may readily be reversed (e.g., an ortho-pyridyl disulfide (OPSS) group that forms a disulfide bond with a cysteine that can be cleaved using reducing agents, e.g., using TCEP), other attachment chemistries will produce permanent attachments.
[0285] In any embodiment, the polymerase may be mutated to ensure specific attachment of the tethered nucleotide to a particular location of the polymerase, as will be apparent to those skilled in the art. For example, with sulfhydryl-specific attachment chemistries such as maleimides or ortho-pyridyl disulfides, accessible cysteine residues in the wild-type polymerase may be mutated to a non-cysteine residue to prevent labeling at those positions. On this "reactive cysteine-free" background, a cysteine residue may be introduced by mutation at the desired attachment position. These mutations preferentially do not interfere with the activity of the polymerase.
[0286] Other strategies for site-specific attachment of synthetic groups to proteins will be apparent to those skilled in the art and are reviewed in literature, (e.g., Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).
[0287] In some examples, a polymerase-nucleotide conjugate is prepared by first synthesizing an intermediate compound comprising a linker and a nucleotide (referred to herein as a "linker-nucleotide"), and then this intermediate compound is attached to the polymerase.
[0288] A person of ordinary skill in the art will understand the conjugates and nucleotides disclosed herein can be prepared in a manner similar to the reaction schemes shown below. The synthetic approaches outlined in these reaction schemes may be illustrated for specific nucleotides. Similar synthetic approaches can be applied to related nucleotide analogs.
[0289] There are several known reactions and functional groups suitable for generating nucleotide polymerase conjugates having cleavable linkers conforming to those describe herein. For example, connection of the conjugate components can be achieved by the formation of a disulfide (forming a readily cleavable connection), formation of an amide, formation of an ester, protein-ligand linkage (e.g., biotin-streptavidin linkage), by alkylation (e.g., using a substituted iodoacetamide reagent) or forming adducts using aldehydes and amines or hydrazines
[0290] In some embodiments, the separate components of the conjugate comprise a site suitable for conjugation to facilitate conjugate synthesis (i.e., a conjugate group). Examples of such conjugate groups include but are limited to hydroxyl, ester, amine, carbonate, acetal, aldehyde, aldehyde hydrate, alkenyl, acrylate, methacrylate, acrylamide, active sulfone, hydrazide, thiol, alkanoic acid, acid halide, isocyanate, isothiocyanate, maleimide, vinylsulfone, dithiopyridine, vinylpyridine, iodoacetamide, epoxide, glyoxal, dione, mesylate, tosylate, and tresylate.
[0291] Further examples of conjugate groups include -NH2, -COOH, -COOCH3, -N- hydroxysuccinimide, and -maleimide. In some embodiments, the bioconjugate reactive group may be blocked (e.g., with a blocking group). Additional examples of bioconjugate reactive groups and the resulting bioconjugate reactive linkers can be found, e.g., in PCT Publication WO2021 / 226327, incorporated by reference in its entirety.
[0292] An exemplary reaction scheme for preparing linker nucleotide conjugates with a linker comprising an amino acid ester, including linkers described herein and variants thereof, is described here. In some embodiments, a desired nucleotide can be commercially obtained, and its hydroxyl groups protected by TBS before conjugation of the L1-OH group to an exocylic oxygen or amine on the nucleobase. Alternatively, an already modified nucleotide comprising the L1-OH group bound to the nucleotide can be obtained, such as an L1-OH group bound to the C5 of a pyrimidine or an L1-OH group bound to C7 of a 7 deazapurine. An Fmoc-protected amino acid ester is then coupled to the hydroxyl group of L1-OH, followed by hydroxyl group deprotection and triphosphorylation of the nucleoside.
[0293] To complete the linker, Fmoc is removed from the amino acid ester amine group and it is coupled to the rest of the linker including L3, which is capable of binding to a polymerase.
[0294] In some embodiments, the linker is bound to a portion of the nucleotide at an atom that is not involved in base pairing. In other embodiments, the linker is bound to the nucleobase of the nucleotide at an atom that is involved in base pairing. In some embodiments, the linker is considered to be at least the atoms that connect the polymerase to any atom in the monocyclic or polycyclic ring system bonded to the Γ position of the sugar (e.g., pyrimidine or purine or 7-deazapurine or 8-aza-7-deazapurine).
[0295] Certain polymerases have a high tolerance for modification of certain parts of a nucleotide, e.g., modifications of the 5 position of pyrimidines and the 7 position of purines are well-tolerated by some polymerases (He and Seela., Nucleic Acids Research 30.24 (2002): 54855496 ; or Hottin et al Chemistry 2017 Feb 10;23(9):2109 2118) In someembodiments, the linker is attached the 5 position of pyrimidines or the 7 position of 7- deazapurines. In other embodiments, the linker may be attached to an exocyclic amine of a nucleobase, e.g., by N-alkylating the exocyclic amine of cytosine with a nitrobenzyl moiety as discussed below.
[0296] In other embodiments, the linker is joined to the sugar or to the α-phosphate of the nucleotide. In some embodiments, the linker is jointed to the terminal phosphate of the nucleotide. In all embodiments, the linker used should be sufficiently long to allow the nucleotide to access the active site of the polymerase to which it is tethered. As will be described in greater detail below, the polymerase of a conjugate is capable of catalyzing the addition of the nucleotide to which it is linked onto the 3′ end of a nucleic acid.
[0297] Conjugation of nucleotides or other base-pairing moieties to linkers may be achieved by any means known in the art of chemical conjugation methods. Nucleotide bases can be obtained or modified to include an L1 portion of the linker. The rest of the linker can be attached to L1 using methods exemplified herein. Those skilled in the art will know or be able to determine appropriate methods for attaching linkers based on the reactivities of these bases.
[0298] In some embodiments, nucleotides containing base modifications that add a free amine group are contemplated for use in conjugation to linkers as described herein. Primary amines, for example, may be linked to the base in such a manner that they can be reacted with heterobifunctional polyethylene glycol (PEG) linkers to create a nucleotide containing a variable length PEG linker. Examples of such amine-containing nucleotides include 5- propargylamino-dNTPs, 5- propargylamino-NTPs, amino allyl-dNTPs, and amino allyl- NTPs.
[0299] In some embodiments, the tethered nucleotide may be specifically attached to a cysteine residue of the polymerase using a sulfhydryl-specific attachment chemistry. Possible sulfhydryl specific attachment chemistries include, but are not limited to ortho-pyridyl disulfide (OPSS), maleimide functionalities, 3- arylpropiolonitrile functionalities, allenamide functionalities, haloacetyl functionalities such as iodoacetyl or bromoacetyl, alkyl halides or perfluroaryl groups that can favorably react with sulfhydryls surrounded by a specific amino acid sequence (Zhang, Chi, et al. Nature chemistry 8, (2015) 120-128.). Other attachment chemistries for specific labeling of cysteine residues will be apparent to those skilled in the art or are described in the pertinent literature and texts (e.g., Kim, Younggyu, et al, Bioconjugate chemistry 19.3 (2008): 786-791.).
[0300] In other embodiments, the linker could be attached to a lysine residue via an amine - reactive functionality (e.g., NHS esters, Sulfo-NHS esters, tetra- or pentafluorophenyl esters, isothiocyanates, sulfonyl chlorides, etc.). In other embodiments, the linker may be attached to the polymerase via attachment to a genetically inserted unnatural amino acid, e.g., p- propargyloxyphenylalanine or p- azidophenylalanine that could undergo azide-alkyne Huisgen cycloaddition, though many suitable unnatural amino acids suitable for site-specific labeling exist and can be found in the literature (e.g., as described in Lang and Chin., Chemical reviews 114.9 (2014): 4764-4806.).
[0301] In other embodiments, the linker may be specifically attached to the polymerase N- terminus. In some embodiments, the polymerase is mutated to have an N-terminal serine or threonine residue, which may be specifically oxidized to generate an N-terminal aldehyde for subsequent coupling to e.g., a hydrazide. In other embodiments, the polymerase is mutated to have an N-terminal cysteine residue that can be specifically labeled with an aldehyde to form a thiazolidine. In other embodiments, an N-terminal cysteine residue can be labeled with a peptide linker via Native Chemical Ligation.
[0302] In other embodiments, a peptide tag sequence may be inserted into the polymerase that can be specifically labeled with a synthetic group by an enzyme, e.g., as demonstrated in the literature using biotin ligase, transglutaminase, lipoic acid ligase, bacterial sortase and phosphopantetheinyl transferase (e.g., as described in refs. 74-78 of Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).
[0303] In other embodiments, the linker is attached to a labeling domain fused to the polymerase. For example, a linker with a corresponding reactive moiety may be used to covalently label SNAP tags, CLIP tags, HaloTags and acyl carrier protein domains (e.g., as described in refs. 79-82 of Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).
[0304] In other embodiments, the linker is attached to an aldehyde specifically generated within the polymerase, as described in Carrico et al. (Nat. Chem. Biol. 3, (2007) 321 - 322). For example, after insertion of an amino acid sequence that is recognized by the enzyme formylglycine-generating enzyme (FGE) into the polymerase, it may be exposed to FGE, which will specifically convert a cysteine residue in the recognition sequence to formylglycine (i.e., producing an aldehyde). This aldehyde may then be specifically labeled with e.g., a hydrazide or aminooxy moiety of a linker.
[0305] In some embodiments, a linker may be attached to the polymerase via non-covalent binding of a moiety of the linker to a moiety fused to the polymerase. Examples of such attachment strategies include fusing a polymerase to streptavidin that can bind a biotinmoiety of a linker, or fusing a polymerase to anti-digoxigenin that can bind a digoxigenin moiety of a linker. In some embodiments, site-specific labeling may lead to an attachment of the linker to the polymerase that may readily be reversed (e.g., an ortho-pyridyl disulfide (OPSS) group that forms a disulfide bond with a cysteine that can be cleaved using reducing agents, e.g., using TCEP), other attachment chemistries will produce permanent attachments.
[0306] In any embodiment, the polymerase may be mutated to ensure specific attachment of the tethered nucleotide to a particular location of the polymerase, as will be apparent to those skilled in the art. For example, with sulfhydryl-specific attachment chemistries such as maleimides or ortho-pyridyl disulfides, accessible cysteine residues in the wild-type polymerase may be mutated to a non-cysteine residue to prevent labeling at those positions. On this "reactive cysteine-free" background, a cysteine residue may be introduced by mutation at the desired attachment position. These mutations preferentially do not interfere with the activity of the polymerase.
[0307] In some embodiments, the linker is specifically attached to an amino acid of the polymerase. In these cases, it is preferable to attach the linker to an amino acid at a position that can be mutated without loss of the polymerase activity, e.g., positions 180, 188, 253 or 302 of murine TdT (numbering as in the crystal structure PDB ID: 4127). It is preferable to not attach the linker to an amino acid involved in the catalytic activity of the polymerase to avoid interfering with catalysis. Residues known to be involved with catalysis and methods for determining if a residue is involved with catalysis (e.g., by site-specific mutagenesis) will be apparent to those skilled in the art and are reviewed in literature (e.g.. Joyce et al. (Journal of Bacteriology 177.22 (1995): 6321.) and Jara and Martinez (The Journal of Physical Chemistry B 120.27 (2016): 6504-6514.))
[0308] Other strategies for site-specific attachment of synthetic groups to proteins will be apparent to those skilled in the art and are reviewed in literature, (e.g., Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).
[0309] In any embodiment, the polymerase can be a template-independent polymerase, i.e., a terminal deoxynucleotidyl transferase or DNA nucleotidylexotransferase, which terms are used interchangeably to refer to an enzyme having activity 2.7.7.31 using the IUBMB nomenclature. A description of such enzymes can be found in Bollum, F.J.
[0310] Deoxynucleotide-polymerizing enzymes of calf thymus gland. V. Homogeneous terminal deoxynucleotidyl transferase. J. Biol. Chem. 246 (1971) 909-916; Gottesman, M.E. and Canellakis, E.S. The terminal nucleotidyltransferases of calf thymus nuclei. J. Biol.
[0311] Chem. 241 (1966) 4339-4352; and Krakow, J.S., Coutsogeorgopoulos, C. and Canellakis, E.S. Studies on the incorporation of deoxyribonucleic acid. Biochim. Biophys. Acta 55 (1962) 639-650, among others.
[0312] In some embodiments, for use with free RTdNTPs, the polymerase is mutated to improve addition of the modified nucleotide.
[0313] Any polymerase capable of extending a polynucleotide, incorporating a nucleotide into a polynucleotide, or incorporating a nucleotide analog into a polynucleotide is envisaged for use in the conjugates and methods described herein. In some embodiments, the polynucleotide is single stranded. In some embodiments, the polynucleotide is double stranded. In some embodiments, the polynucleotide is immobilized on a solid support.
[0314] Examples of DNA polymerases include polA, polB, polC, polD, polY, polX, reverse transcriptases (RT), and high-fidelity polymerases. In some instances, the polymerase is a modified polymerase. In some embodiments, the polymerase comprises 29, B103, GA-1, PZA, 15, BS32, M2Y, Nf, Gl, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, L17, ThermoSequenase®, 9°Nm™, Therminator™ DNA polymerase, Tne, Tma, Tfl, Tth, TIi, Stoffel fragment, Vent™ and Deep Vent™ DNA polymerase, KOD DNA polymerase, Tgo, JDF-3, Pfu, Taq, T7 DNA polymerase, T7 RNA polymerase, PGB-D, UlTma DNA polymerase, E. coli DNA polymerase I, E. coli DNA polymerase III, archaeal DP1EDP2 DNA polymerase II, 9°N DNA Polymerase, Taq DNA polymerase, Phusion® DNA polymerase, Pfu DNA polymerase, SP6 RNA polymerase, RB69 DNA polymerase, Avian Myeloblastosis Virus (AMV) reverse transcriptase, Moloney Murine Leukemia Virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, and SuperScript® III reverse transcriptase.
[0315] In some embodiments, the polymerase is DNA polymerase 1 -KI enow fragment, Vent polymerase, Phusion® DNA polymerase, KOD DNA polymerase, Taq polymerase, T7 DNA polymerase, T7 RNA polymerase, Therminator™ DNA polymerase, POLB polymerase, SP6 RNA polymerase, E. coli DNA polymerase I, E. coli DNA polymerase III, Avian Myeloblastosis Virus (AMV) reverse transcriptase, Moloney Murine Leukemia Virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, or SuperScript® III reverse transcriptase.
[0316] The polymerase molecules used in the methods described herein can be polymerase theta, a DNA polymerase, or any enzyme that can extend nucleotide chains. In some embodiments, the polymerase is tri29. In some embodiments, the polymerase is a protein with pockets that work around terminal phosphate groups, for example, a triphosphate group.In some embodiments, the described methods use TdT with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations to synthesize defined polynucleotides. In some embodiments, the described method uses TdT with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations to a surface-accessible amino acid residue. In some embodiments, the TdT is a variant of TdT. In some embodiments, the variant of TdT comprises a cysteine mutation. In some embodiments, the polymerase is mutated to improve addition of a modified nucleotide bound to the polymerase forming a conjugate. In some instances, the variant TdT comprises at least 70%, 80%, 90%, or 95% sequence identity to wild-type TdT.
[0317] In some embodiments, the described methods use polymerase theta with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations to synthesize defined polynucleotides. In some embodiments, the described method uses polymerase theta with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations to a surface-accessible amino acid residue. In some embodiments, the polymerase theta is a variant of polymerase theta. In some instances, the variant polymerase theta comprises at least 70%, 80%, 90%, or 95% sequence identity to wild-type polymerase theta. In some embodiments, the polymerase theta is encoded by POLQ.
[0318] Enzymes described herein (e.g., TdT), in some embodiments, comprise one or more unnatural amino acids. In some instances, the unnatural amino acid comprises: a lysine analogue; an aromatic side chain; an azido group; an alkyne group; or an aldehyde or ketone group. In some instances, the unnatural amino acid does not comprise an aromatic side chain. In some embodiments, the unnatural amino acid is selected from N6-azidoethoxy-carbonyl- L-lysine (AzK), N6-propargylethoxy-carbonyl-L-lysine (PraK), N6-(propargyloxy)- carbonyl- L-lysine (PrK), p-azido-phenylalanine, BCN-L-lysine, norbornene lysine, TCO- lysine, methyltetrazine lysine, allyloxycarbonyllysine, 2-amino-8-oxononanoic acid, 2- amino-8- oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo- L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p- propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-Dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p-acyl-L- phenyl alanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L- phenyl alanine, isopropyl-L-phenylalanine, O-allyltyrosine, O-methyl -L-tyrosine, O-4-allyl-L- tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L- phosphoserine, phosphonoserine, L-3-(2-naphthyl)alanine, 2-amino-3-((2-((3-(benzyloxy)-3- oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3- azidobenzyl)oxy)carbonyl)-L- lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L- lysine.
[0319] In some embodiments, the polymerase is a fusion protein. In some embodiments of the method, the fusion protein comprises maltose binding protein (MBP). In some embodiments, TdT is fused to other enzymes such as helicase.
[0320] In some embodiments, the polymerase comprises a template-independent polymerase. In some embodiments, the polymerase comprises a Pol-X family polymerase. In some embodiments, the polymerase comprises a Terminal deoxynucleotidyl Transferase (TdT), or a variant thereof. In some embodiments, the template-independent polymerase comprises a TdT or a variant thereof. In some embodiments, the TdT or variant thereof comprises a sequence sharing at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to SEQ ID NO: 1. In some embodiments of the method, the TdT comprises a sequence identical to SEQ ID NO: 1. In some embodiments, the TdT variant comprises one or more amino acid substitutions, insertions, or deletions to SEQ ID NO: 1.
[0321] >Terminal deoxynucleotidyl transferase (TdT): MGGRDIVDGSEFSPSPVPGSQNVPAPAVKKISQYACQRRTTLNNYNQLFTDALDILAENDEL RENEGSALAFMRASSVLKSLPFPITSMKDTEGIPSLGDKVKSIIEGIIEDGESSEAKAVLND ERYKSFKLFTSVFGVGLKTAEKWFRMGFRTLSKIQSDKSLRFTQMQKAGFLYYEDLVSCVNR PEAEAVSMLVKEAWTFLPDALVTMTGGFRRGKMTGHDVDFLITSPEATEDEEQQLLHKVTD FWKQQGLLLYADILESTFEKFKQPSRKVDALDHFQKCFLILKLDHGRVHSEKSGQQEGKGWK AIRVDLVMSPYDRRAFALLGWTGSRQFERDLRRYATHERKMMLDNHALYDRTKRVFLEAESE EEIFAHLGLDYIEPWERNA (SEQ ID NO: 1)
[0322] In some embodiments, template-independent polymerases having activity as described for E.C. class 2.7.7.31 are used. In some embodiments, the template-independent polymerase is a deoxynucleotidyl transferase or DNA nucleotidylexotransferase. A description of such enzymes can be found in Bollum, F.J. Deoxynucleotide-polymerizing enzymes of calf thymus gland. V. Homogeneous terminal deoxynucleotidyl transferase. J. Biol. Chem. 246 (1971) 909-916; Gottesman, M.E. and Canellakis, E.S. The terminal nucleotidyltransferases of calf thymus nuclei. J. Biol. Chem. 241 (1966) 4339-4352; and Krakow, J.S., Coutsogeorgopoulos, C. and Canellakis, E.S. Studies on the incorporation of deoxyribonucleic acid. Biochim. Biophys. Acta 55 (1962) 639-650, among others.
[0323] Additional polymerases with the ability to extend single stranded nucleic acids in the absence of template that can be used include, but are not limited to, Polymerase Theta (Kent et al., eLife 5 (2016): el3740.), polymerase mu (Juarez et al., Nucleic acids research 34.16 (2006): 4572-4582.; or McElhinny et all., Molecular cell 19.3 (2005): 357-366.) orpolymerases where template independent activity is induced, e.g. by the insertion of elements of a template independent polymerase (Juarez et al., Nucleic acids research 34.16 (2006): 4572-4582).
[0324] In other DNA synthesis applications, the polymerase can be a template-dependent polymerase i.e., a DNA-directed DNA polymerase (e.g., an enzyme having activity 2.7.7.7 using the IUBMB nomenclature) or an RNA-directed DNA polymerase. A description of such enzymes can be found in Richardson, A. Enzymatic synthesis of deoxyribonucleic acid. XIV. Further purification and properties of deoxyribonucleic acid polymerase of Escherichia coli. J. Biol. Chem. 239 (1964) 222-232; Schachman, A. Enzymatic synthesis of deoxyribonucleic acid. VIL Synthesis of a polymer of deoxyadenylate and deoxy thymidylate. J. Biol. Chem. 235 (1960) 3242-3249; and Zimmerman, B.K. Purification and properties of deoxyribonucleic acid polymerase from Micrococcus lysodeikticus. J. Biol. Chem. 241 (1966) 2035-2041.
[0325] In some embodiments, the polymerase comprises an RNA polymerase. In these embodiments, an RNA specific nucleotidyl transferase, such as E. coli Poly(A) Polymerase (IUBMB EC 2.7.7.19) or Poly(U) Polymerase, among others, may be employed. The RNA nucleotidyl transferases can contain modifications, e.g., single point mutations, which influence the substrate specificity towards a specific rNTP (Lunde et al., Nucleic acids research 40.19 (2012): 9815-9824.). In some embodiments, a very short tether between an RNA nucleotidyl transferase and a ribonucleotide may be used to induce a high effective concentration of the nucleotide, thereby forcing incorporation of an rNTP that might not be the natural substrate of the nucleotidyl transferase.Nucleotides
[0326] In some embodiments, a nucleotide is a ribose polyphosphate. In some embodiments, a ribose polyphosphate is selected from the group consisting of ribose triphosphate, ribose tetraphosphate, ribose pentaphosphate, and ribose hexaphosphate. In some embodiments, a ribose polyphosphate is a ribose triphosphate. In some embodiments, a ribose polyphosphate is a ribose hexaphosphate. In some embodiments, ribose polyphosphate is a ribose pentaphosphate. In some embodiments, ribose polyphosphate is a ribose tetraphosphate.
[0327] In some embodiments, a nucleotide is a deoxyribose polyphosphate. In some embodiments, a deoxyribose polyphosphate is selected from the group consisting of deoxyribose triphosphate, deoxyribose tetraphosphate, deoxyribose pentaphosphate, anddeoxyribose hexaphosphate. In some embodiments, a deoxyribose polyphosphate is a ribose triphosphate. In some embodiments, a deoxyribose polyphosphate is a deoxyribose hexaphosphate. In some embodiments, deoxyribose polyphosphate is a deoxyribose pentaphosphate. In some embodiments, deoxyribose polyphosphate is a deoxyribose tetraphosphate.
[0328] The term “nucleotides” and related terms refers to a molecule comprising an aromatic base, a five carbon sugar (e.g., ribose or deoxyribose), and at least one phosphate group. Canonical or non-canonical nucleotides are consistent with use of the term. The phosphate in some embodiments comprises a monophosphate, diphosphate, or triphosphate, or corresponding phosphate analog.
[0329] Nucleotides (and nucleosides) typically comprise a hetero cyclic base including substituted or unsubstituted nitrogen-containing parent heteroaromatic ring which are commonly found in nucleic acids, including naturally-occurring, substituted, modified, or engineered variants, or analogs of the same. Exemplary bases include, but are not limited to, purines and pyrimidines such as: 2-aminopurine, 2,6-diaminopurine, adenine (A), ethenoadenine, N6-A2-isopentenyladenine (6iA), N6-A2-isopentenyl-2-methylthioadenine (2ms6iA), N6-methyladenine, guanine (G), isoguanine, N2-dimethylguanine (dmG), 7- methylguanine (7mG), 2-thiopyrimidine, 6-thioguanine (6sG), hypoxanthine and 06- methylguanine; 7-deaza-purines such as 7-deazaadenine (7-deaza-A) and 7-deazaguanine (7- deaza-G); pyrimidines such as cytosine (C), 5-propynylcytosine, isocytosine, thymine (T), 4- thiothymine (4sT), 5,6-dihydrothymine, O4-methylthymine, uracil (U), 4-thiouracil (4sU) and 5,6-dihydrouracil (dihydrouracil; D); indoles such as nitroindole and 4-methylindole; pyrroles such as nitropyrrole; nebularine; inosines; hydroxymethylcytosines; 5- methycytosines; base (Y); as well as methylated, glycosylated, and acylated base moieties; and the like. Additional exemplary bases can be found in Fasman, 1989, in “Practical Handbook of Biochemistry and Molecular Biology”, pp. 385-394, CRC Press, Boca Raton, Fla.
[0330] Nucleotides (and nucleosides) typically comprise a sugar moiety, such as carbocyclic moiety (Ferraro and Gotor 2000 Chem. Rev. 100: 4319-48), acyclic moieties (Martinez, et al., 1999 Nucleic Acids Research 27: 1271-1274; Martinez, et al., 1997 Bioorganic & Medicinal Chemistry Fetters vol. 7: 3013-3016), and other sugar moieties (Joeng, et al., 1993 J. Med. Chem. 36: 2627-2638; Kim, et al., 1993 J. Med. Chem. 36: 30-7; Eschenmosser 1999 Science 284:2118-2124; and U.S. Pat. No. 5,558,991). The sugar moiety comprises: ribosyl; 2'-deoxyribosyl; 3 '-deoxyribosyl; 2',3'-dideoxyribosyl; 2', 3'-didehydrodideoxyribosyl; 2'-alkoxyribosyl; 2'-azidoribosyl; 2'-aminoribosyl; 2'- fluororibosyl; 2'-mercaptoriboxyl; 2'-alkylthioribosyl; 3 '-alkoxyribosyl; 3 '-azidoribosyl; 3'- aminoribosyl; 3 '-fluororibosyl; 3'-mercaptoriboxyl; 3 '-alkylthioribosyl carbocyclic; acyclic or other modified sugars.
[0331] In some embodiments, nucleotides comprise a chain of one, two or three phosphorus atoms where the chain is typically attached to the 5' carbon of the sugar moiety via an ester or phosphoramide linkage. In some embodiments, the nucleotide is an analog having a phosphorus chain in which the phosphorus atoms are linked together with intervening O, S, NH, methylene or ethylene. In some embodiments, the phosphorus atoms in the chain include substituted side groups including O, S or BH3. In some embodiments, the chain includes phosphate groups substituted with analogs including phosphoramidate, phosphorothioate, phosphordithioate, and O-methylphosphoroamidite groups.
[0332] In some embodiments, the polymerase of the conjugate may be covalently attached to oligonucleotides or nucleotides via the nucleotide base. For example, the nucleotide or oligonucleotide may have the polymerase attached to the C5 position of a pyrimidine base or the C7 position of a 7-deaza purine base through a linker moiety.
[0333] A nucleotide used in the present disclosure can also include native or non-native bases. In this regard a native deoxyribonucleic acid can have one or more bases selected from the group consisting of adenine, thymine, cytosine or guanine and a ribonucleic acid can have one or more bases selected from the group consisting of uracil, adenine, cytosine or guanine. Exemplary non-native bases that can be included in a nucleic acid, whether having a native backbone or analog structure, include, without limitation, inosine, xathanine, hypoxathanine, isocytosine, isoguanine, 5- methylcytosine, 5-hydroxymethyl cytosine, 2-aminoadenine, 6- methyl adenine, 6-methyl guanine, 2-propyl guanine, 2-propyl adenine, 2-thioLiracil, 2- thiothymine, 2-thiocytosine, 15 - halouracil, 15 -halocytosine, 5-propynyl uracil, 5-propynyl cytosine, 6-azo uracil, 6-azo cytosine, 6-azo thymine, 5-uracil, 4-thiouracil, 8-halo adenine or guanine, 8-amino adenine or guanine, 8- thiol adenine or guanine, 8-thioalkyl adenine or guanine, 8-hydroxyl adenine or guanine, 5-halo substituted uracil or cytosine, 7- methylguanine, 7-methyladenine, 8-azaguanine, 8-azaadenine, 7- deazaguanine, 7- deazaadenine, 3-deazaguanine, 3-deazaadenine or the like.
[0334] In some embodiments, the phosphorylated nucleoside (e.g., nucleotide) to be tethered to the polymerase is a nucleoside comprising at least one phosphate group. In some embodiments, the nucleoside comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or more than 9 phosphate groups. In some embodiments, the nucleoside comprises at least 3 phosphategroups. In some embodiments, the phosphorylated nucleoside is adenosine, cytidine, uridine, or guanosine, each of which comprises at least one phosphate group. In some embodiments, the phosphorylated nucleoside is a deoxy nucleoside comprising at least one phosphate group. In some embodiments, the phosphorylated nucleoside is a deoxynucleoside comprising at least 3 phosphate groups. In some embodiments, the deoxy nucleoside comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or more than 9 phosphate groups. In some embodiments, the phosphorylated nucleoside is deoxyadenosine, deoxycytidine, deoxythymidine, or deoxyguanosine, each of which comprises at least one phosphate group. In some embodiments, the phosphorylated nucleoside is a nucleoside triphosphate, such as dNTP. In some embodiments, the phosphorylated nucleoside is a nucleoside tetraphosphate, nucleoside pentaphosphate, a nucleoside hexaphosphate, a nucleoside heptaphosphate, nucleoside octaphosphate, or a nucleoside nonaphosphate. In some embodiments, the phosphorylated nucleoside is a nucleoside hexaphosphate. In some embodiments, the phosphorylated nucleoside is a nucleoside triphosphate. In some embodiments, the phosphorylated nucleoside is selected from the group consisting of deoxyadenosine triphosphate (dATP), deoxy guano sine triphosphate (dGTP), deoxy cytidine triphosphate (dCTP), deoxy thymidine triphosphate (dTTP), deoxyadenosine tetraphosphate, deoxyguanosine tetraphosphate, deoxycytidine tetraphosphate, deoxythymidine tetraphosphate, deoxyadenosine pentaphosphate, deoxyguanosine pentaphosphate, deoxycytidine pentaphosphate, deoxythymidine pentaphosphate, deoxyadenosine hexaphosphate, deoxyguanosine hexaphosphate, deoxycytidine hexaphosphate, deoxythymidine hexaphosphate, and any combination thereof.
[0335] In some embodiments, the nucleotides analogs described herein comprise a reversible terminator group, such as such as an O- azidomethyl or O-NH2 group on the 3' position of the sugar or an (alpha-tertbutyl-2- nitrobenzyl)oxymethyl group on the 5 position of pyrimidines or the 7 position of 7- deazapurines (for an overview see, e.g. Chen et al., Genomics, Proteomics & Bioinformatics 2013 11: 34-40). In these embodiments, the nucleotide analog prevents or hinders further elongation once incorporated into a nucleic acid to achieve controlled termination of synthesis. In some embodiments, when used as part of a conjugate, the RTdNTP- polymerase conjugates do not rely on the shielding effect to achieve termination, e.g.. when a 3' modified RTdNTP is tethered to the polymerase, the linker used may exceed 100 A or 200 A in length.Linker
[0336] In a conjugate, the linker is considered to be at least the atoms that connect the nucleotide to the polymerase. In some embodiments, the linker comprises atoms that connect the base, the sugar, or the oc-phosphate of a nucleotide to the polymerase. In some embodiments, the polymerase and the nucleotide are covalently linked and the distance between the linked atom of the nucleotide and the polymerase to which it is attached may be in the range of 4-100 A, e.g., 15-40 or 20-30 A, although this distance may vary depending on where the nucleoside triphosphate is tethered. In some embodiments, the linker may be a PEG or polypeptide linker, although, again, there is considerable flexibility on the type of linker used. In some embodiments, the linker should be joined to the base of the nucleotide at an atom that is not involved in base pairing. In such embodiments, the linker is considered to be at least the atoms that connect a Ca atom in the backbone of the polymerase to any atom in the monocyclic or polycyclic ring system bonded to the T position of the sugar (e.g.. pyrimidine or purine or 7-deazapurine or 8-aza-7-deazapurine). In other embodiments, the linker should be joined to the base of the nucleotide at an atom that is involved in base pairing. In other embodiments, the linker should be joined to the sugar or to the oc-phosphate of the nucleotide. In all embodiments, the linker used should be sufficiently long to allow the nucleoside triphosphate to access the active site of the polymerase to which it is tethered. As will be described in greater detail below, the polymerase of a conjugate is capable of catalyzing the addition of the nucleotide to which it is linked onto the 3' end of a nucleic acid.
[0337] As described above, the linker may be attached to various positions on the nucleotide, and a variety of cleavage strategies may be used. Those strategies may include, but are not limited to, the following examples:
[0338] In some embodiments, the linker may be cleaved by exposure to a reducing agent such as dithiothreitol (DTT). For example, a linker comprising a 4-(disulfaneyl)butanoyloxy- methyl group attached to the 5 position of a pyrimidine or the 7 position of a 7-deazapurine may be cleaved by reducing agents (e.g.. DTT) to produce a 4-mercaptobutanoyloxymethyl scar on the nucleobase. This scar may undergo intramolecular thiolactonization to eliminate a 2-oxothiolane, leaving a smaller hydroxymethyl scar on the nucleobase. An example of such a linker attached to the 5 position of cytosine is depicted below, but the strategy is applicable to any suitable nucleobase:
[0339] In other embodiments, the linker may be cleaved by exposure to light. For example, a linker comprising (2-nitrobenzyl)oxymethyl group may be cleaved with 365 nm light, leaving a hydroxymethyl scar, e.g., as depicted for cytosine below, but as is applicable to any suitable nucleobase:(where, e.g., R' ' =H or R' '=CH3 or R' =i-Bu.)
[0340] In other embodiments, the linker may comprise a 3-(((2- nitrobenzyl)oxy)carbonyl)aminopropynyl group that may be cleaved with 365 nm light release a nucleobase with a propargylamino scar. This strategy is applicable to any suitable nucleobase:
[0341] In other embodiments, the linker may comprise an acyloxymethyl group that may be cleaved with a suitable esterase to release a nucleobase with a hydroxymethyl scar, e.g.. as depicted for cytosine below, but as is applicable to any suitable nucleobase:
[0342] In such embodiments, the linker may comprise additional atoms (included in R' above) adjacent to the ester that increase the activity of the esterase towards the ester bond.
[0343] In other embodiments, the linker may comprise an N-acyl-aminopropynyl group that may be cleaved with a peptidase to release a nucleobase with propargylamino scar, e.g., as depicted for 5 -propargylamino cytosine below, but as is applicable to any suitable nucleobase:
[0344] In such embodiments, the linker may comprise additional atoms (included in R' above) adjacent to the amide that increase the activity of the peptidase towards the amide bond.
[0345] In some embodiments, a polymerase-nucleotide conjugate comprises a nucleotide linked to a polymerase using an enzymatically cleavable linker. In some embodiments, a polymerase-nucleotide conjugate comprising an enzymatically cleavable linker comprises a structure Nuc-L1-L2-Pol, wherein Nuc represents a nucleotide, pol represents a polymerase, and L'-L2represents an enzymatically cleavable linker. In some embodiments, L1represents a region of an enzymatically cleavable linker connecting the nucleotide to L2, L2represents a cleavable portion of an enzymatically cleavable linker. In some embodiments, L2also comprises a portion for connecting L2to Pol.
[0346] In some embodiments, an enzymatically cleavable linker comprises an amino acid ester moiety. In some embodiments, L2comprises an amino acid ester moiety. In some embodiments, the ester group of an amino acid ester moiety is cleavable by a protease comprising esterase activity. The ester of an amino acid of L2is attached to L1, which can also be referred to as a spacer or as a scar of a nucleotide after cleavage of the L2ester. In some embodiments, L2comprises attachment chemistry for polymerase conjugation. Insome embodiments, L2further comprises additional amino acids bound to the amine of the amino acid ester to serve as a protease substrate. In some embodiments, L2is optimized for ester stability to prevent spontaneous cleavage while retaining the ability to act as a suitable substrate for esterase activity of a protease comprising esterase activity.
[0347] In some embodiments, the amino acid ester comprises one or more substitutions at the alpha carbon, such as addition of an aliphatic or bulky substituent. In some embodiments, the amino acid ester is represented by:wherein R1and R1are each independently selected from an optionally substituted C1-3 alkyl, a halogen, or are optionally taken together with the atom on which they are attached to form an optionally substituted C3-C7 carbocyclic ring.
[0348] Exemplary L2linker structures with different substituents on the alpha carbon of the amino acid ester (e.g., to improve stability of the ester) are shown below (with L3representing a portion of the L2linker including attached to the polymerase):In some embodiments, L2comprises an amino acid ester adjacent to one or more amino acid residues. In some embodiments, the one or more amino acid residues are bound to the amine group of the amino acid ester.
[0349] In some embodiments, L2comprises or consists of:wherein R1and R1'are independently selected from hydrogen or an optionally substituted C1-3 alkyl or are taken together with the atom on which they are attached to form an optionally substituted C3-C7carbocyclic ring; each R3is an optionally substituted group independently selected from hydrogen, C1-6 alkyl, benzyl, -OH, -O(C1-6 alkyl), and -CN; each Rcis hydrogen or optionally substituted C1-6alkyl; and n is 1, 2, or 3.
[0350] In some embodiments, the one or more amino acids linked to the amine of the amino acid ester comprise L- or D- isomers of amino acid residues. The term “naturally-occurring amino acid” refer to Ala, Asp, Cys, Glu, Phe, Gly, His, He, Lys, Leu, Met, Asn, Pro, Gin, Arg, Ser, Thr, Val, Trp, Tyr, or citrulline. “D-” designates an amino acid having the “D” (dextrorotary) configuration, as opposed to the configuration in the naturally occurring (“L-”) amino acids. The amino acids described herein can be purchased commercially (Sigma Chemical Co., Advanced Chemtech) or synthesized using methods known in the art. In some embodiments, amino acids with non-natural or artificial side chains are linked to the amine of the amino acid ester.
[0351] The one or more amino acids included in the L2portion of the linker / bound to the amino acid ester can be selected, for example, to optimize protease binding and ester cleavage. A combinatorial library can be generated to test optimal cleavage activity, aminoacids can be chosen based on existing known peptide sequence targets for the protease. The protease comprising esterase activity can recognize the peptide portion of the linker and hydrolyzes the ester group of the amino acid ester of L2, resulting in removal of polymerase attached to the nucleotide via the linker, as disclosed herein.
[0352] If desired, a spacer can be used between the nucleotide and the linker, or between the linker and the label. Different lengths of spacers can be used in order to increase L2availability towards the protease / esterase and increase the efficiency and fidelity of polymerases. Exemplary spacers include, for example, polyethyleneglycol or other suitable spacers.
[0353] Examples of linkers comprising L2structures including an amino acid ester bound to one or more amino acid residues is shown below, with ‘L3’ representing a portion of the L2cleavable linker that is capable of binding to the polymerase:
[0354] There is considerable flexibility on the type of linker used for regions of the linker not associated with enzymatic cleavage disclosed herein. Examples of suitable linker structures may include, but are not limited to, carbon-chain linkers (e.g., C6, C12, C18, C24, etc.), peptide linkers (e.g., poly-glycine or poly-alanine ranging from about 1 residue to about 1,000 residues in length), or polyether linkers (e.g., PEG, PPG, PAG, PTMG from about 1 polyether unit to about 1,000 polyether units in length).
[0355] In some embodiments, the linker comprises a chain of atoms selected from C, N, O, S, Si, and P, preferably having 0-500 atoms, wherein L1covalently connects to Nuc and L2, and wherein L2is covalently attached to Pol. The atoms used in forming L1or including in L2(e.g., in L3shown above) may be combined in all chemically relevant ways, such as forming alkylene, alkenylene, and alkynylene, carbamates, carbonates, ethers, polyoxyalkylene, esters, amines, imines, polyamines, hydrazines, hydrazones, amides, ureas, semicarbazides, carbazides, alkoxyamines, alkoxylamines, urethanes, amino acids, peptides, acyloxylamines, hydroxamic acids, or combination above thereof.
[0356] In some embodiments, the linker comprises one or more carbon atoms, zero, one, or more oxygen atoms, zero, one or more nitrogen atoms, zero, one, or more sulfur atoms, or a combination thereof, in different embodiments. In some embodiments, the linker comprises,comprises about, comprises at least, comprises at least about, comprises at most, or comprises at most about, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or a range between any two of these values, carbon atom(s), oxygen atom(s), nitrogen atom(s), sulfur atom(s), or a combination thereof.
[0357] In some embodiments, the linker comprises a polymer, such as a homopolymer or a heteropolymer. In some embodiments, the linker comprise a plurality of repeat units. In some embodiments, the plurality of repeating units comprises identical repeating units. In some embodiments, the plurality of repeating units comprises two or more different repeating units. The plurality of repeating units can comprise a polyether such as paraformaldehyde, polyethylene glycol (PEG), polypropylene glycol (PPG), polyalkylene glycol (PAG), polytetramethylene glycol (PTMG), or a combination thereof. For example, the plurality of repeating units can comprise PEGix, PEG23, PEG24, or a combination thereof. The plurality of repeating units can comprise a polyalkylene, such as polyethene, polypropene, polybutene, or a combination thereof. In some embodiments, a repeating unit of the plurality of repeating units comprises no aromatic group. In some embodiments, a repeating unit of the plurality of repeating units comprises one or more aromatic groups.
[0358] In some embodiments, the linker comprises any number of basic chemical starting blocks. For example, linkers may comprise linear or branched alkyl, alkenyl, or alkynyl chains, or combinations thereof, that provide a useful distance between the nucleotide and polymerase, the nucleotide and L2, or the polymerase and the cleavable moiety of L2. For instance, amino-alkyl linkers, e.g., amino-hexyl linkers, have been used to attach linkers to nucleotide analogs, and are generally sufficiently rigid to maintain such distances. The longest chain of such linkers may include as many as 2 atoms, 3 atoms, 4 atoms, 5 atoms, 6 atoms, 7 atoms, 8 atoms, 9 atoms, 10 atoms, or even 11-35 atoms, or even 35-50 atoms. The linear or branched linker may also contain heteroatoms other than carbon, including, but not limited to, oxygen, sulfur, phosphate, and nitrogen. A polyoxyethylene chain (also commonly referred to as polyethyleneglycol, or PEG) is a preferred linker constituent due to the hydrophilic properties associated with polyoxyethylene. Insertion of heteroatom such as nitrogen and oxygen into the linkers may affect the solubility and stability of the linkers.
[0359] The linker may be rigid in nature or flexible. Rigid structures include laterally rigid chemical groups, e.g., ring structures such as aromatic compounds, multiple chemical bonds between adjacent groups, e.g., double or triple bonds, in order to prevent rotation of groups relative to each other, and the consequent flexibility that imparts to the overall linker. Thus, the degree of desired rigidity may be modified depending on the content of the linker, or the number of bonds between the individual atoms comprising the linker. Further, addition of ringed structures along the linker may impart rigidity. Ringed structures may include aromatic or non-aromatic rings. Rings may be anywhere from 3 carbons, to 4 carbons, to 5 carbons or even 6 carbons in size. Rings may also optionally include heteroatoms such as oxygen or nitrogen and also be aromatic or non-aromatic. Rings may additionally optionally be substituted by other alkyl groups and / or substituted alkyl groups.
[0360] Linkers that comprise ring or aromatic structures can include, for example aryl alkynes and aryl amides. Other examples of the linkers of the disclosure include oligopeptide linkers that also may optionally include ring structures within their structure.
[0361] In embodiments, the linker comprises is a C1-C10 alkylene chain, wherein 1-6 methylene units are optionally and independently replaced by -NH-, -O-, -C(O)-, -C(O)NH-, -NHC(O)-, -NHC(O)NH-, -C(O)O-, -OC(O)-, -SS-, optionally substituted cycloalkylene (e.g., C3-C8, C3-C6, or C5-C6), optionally substituted heterocycloalkylene (e.g., 3 to 8, 3 to 6, or 5 to 6 membered), optionally substituted arylene (e.g., C6-C10, C10, or phenylene), or substituted or unsubstituted heteroarylene (e.g., 5 to 10, 5 to 9, or 5 to 6 membered).
[0362] In some embodiments, L1is a bond, -NH-, -O-, -C(O)-, -C(O)NH-, -NHC(O)-, - NHC(O)NH-, -C(O)O-, -OC(O)-, -SS-, optionally substituted alkylene (e.g., C1- C20, C10- C20, C1-C8, C1-C6, or C1-C4), optionally substituted heteroalkylene (e.g., 2 to 20, 8 to 20, 2 to 10, 2 to 8, 2 to 6, or 2 to 4 membered), optionally substituted cycloalkylene (e.g., C3-C8, C3- C6, or C5-C6), optionally substituted heterocycloalkylene (e.g., 3 to 8, 3 to 6, or 5 to 6 membered), optionally substituted (e.g., C6-C10, C10, or phenylene), or optionally substituted (e.g., 5 to 10, 5 to 9, or 5 to 6 membered). In some embodiments, L1is optionally substitutedC1-C20 alkylene. In some embodiments, L1is optionally substituted 2 to 20 membered heteroalkylene. In some embodiments, L1is optionally substituted C3-C8 cycloalkylene. In some embodiments, L1is optionally substituted 3 to 8 membered heterocycloalkylene. In embodiments, L1is optionally substituted C6-C10arylene. In embodiments, L1is optionally substituted 5 to 10 membered heteroarylene.
[0363] In some embodiments, L1is substituted with 1-6 instances of RL. Each RLis independently selected from the group consisting of oxo, halogen, -CCI3, -CBr3, -CF3, -CI3, -CN, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, -NHNH2, -ONH2, - NHC(O)NHNH2, -NHC(O)NH2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -NHOH, -OCCl3, - OCF3, -OCBr3, -OCI2, -OCHCI2, -OCHBr2, -OCHb, -OCHF2, -N3, optionally substituted alkyl (e.g., C1-C20, C10-C20, C1-C8, C1-C6, or C1-C4), optionally substituted heteroalkyl (e.g., 2 to 20, 8 to 20, 2 to 10, 2 to 8, 2 to 6, or 2 to 4 membered), optionally substituted cycloalkyl (e.g., C3-C8, C3-C6, or C5-C6), optionally substituted heterocycloalkyl (e.g., 3 to 8, 3 to 6, or 5 to 6 membered), optionally substituted aryl (e.g., C6-C10, C10, or phenyl), and optionally substituted heteroaryl (e.g., 5 to 10, 5 to 9, or 5 to 6 membered).
[0364] In some embodiments, L1 acts as an attachment point to the nucleotide and includes a hydroxyl terminal group which binds to a portion of L2 during synthesis.
[0365] As described herein, in some embodiments, L1 is a scar that is enzymatically cleavable after cleavage of polymerase-nucleotide linker / removal of the L2-pol moiety.
[0366] In some embodiments, L2 comprises a bioconjugate group suitable for conjugation of L2 to the polymerase.
[0367] In some embodiments, the bioconjugate group is an N-hydroxysuccinimide ester (NHS) group. In some embodiments, the bioconjugate group is a maleimide group. The linker may then be covalently attached to the polymerase by reaction of the maleimide group with a cysteine residue of the polymerase.
[0368] In some embodiments, the polymerase may be operably linked to a linker moiety including a covalent or non-covalent bond; amino acid tag (e.g., poly-amino acid tag, poly- His tag, 6His-tag (SEQ ID NO: 53)); chemical compound (e.g., polyethylene glycol); protein-protein binding pair (e.g., biotin-avidin); affinity coupling; capture probes; or any combination of these. The linker moiety can be separate from or part of a polymerase variant.
[0369] The above described conjugates can be used in a method of nucleic acid synthesis. Nucleic acid synthesis can refer to synthesis, or generation of a product that is a nucleic acid molecule (e.g., a polynucleotide). Methods of nucleic acid synthesis can comprise stepwise synthesis, wherein nucleotides are inserted stepwise into a nucleic acid polymer or polynucleotide. By way of non-limiting example, one typical process for stepwise synthesis of a polynucleotide comprises adding nucleotides stepwise to a starter molecule (e.g., an initial oligonucleotide) via the cycled steps of: addition of a polymerase-nucleotide conjugate to an oligonucleotide under conditions suitable for covalently binding the nucleotide to the end of the oligonucleotide catalyzed by the polymerase. Successful incorporation of a nucleotide of a conjugate into an oligonucleotide can be referred to as an “extension” or “extension reaction,” which generates an “extension product.”
[0370] In some embodiments, this method comprises: incubating a nucleic acid with a first conjugate under conditions in which the polymerase catalyzes the covalent addition of the nucleotide of the first conjugate onto the 3' hydroxyl of the nucleic acid, to make an extension product. This reaction can be performed using a nucleic acid that is attached to a solid support or that is in solution, e.g., not tethered to a solid support. Addition of the conjugate to the nucleic acid results in a nucleic acid with an added 3’ group that is shielded by the linked polymerase, inhibiting subsequent addition of another nucleotide while the polymerase is attached. After elongation of the nucleic acid by the first desired nucleotide, the method comprises a deblocking (de-shielded) step wherein the cleavable linkage of the linker is cleaved, thereby releasing the polymerase from the extension product. Cleavage of the linker removes the polymerase to produce a deblocked extension product. Deblocking enables subsequent extension of the nucleic acid, and thus allows these steps to be repeated cyclically to produce an extension product of defined sequence. Specifically, in some embodiments, the method may further comprise, after deblocking: incubating the deblocked extension product with a second conjugate under conditions in which polymerase catalyzes the covalent addition of the nucleotide of the second conjugate onto the 3' end of the extension product.
[0371] In some embodiments, the method may involve (a) incubating a nucleic acid with a first conjugate under conditions in which the polymerase catalyzes the covalent addition of the nucleotide of the first conjugate (i.e., a single nucleotide) onto the 3' hydroxyl of the nucleic acid, to make an extension product; (b) cleaving the cleavable linkage of the linker, thereby releasing the polymerase from the extension product and deblocking the extension product; (c) incubating the deblocked extension product with a second conjugate of claim 1 under conditions in which the polymerase catalyzes the covalent addition of the nucleotide of the second conjugate onto the 3' end of the extension product, to make a second extension product; (d) repeating steps (b)-(c) on the second extension product multiple times (e.g., 2 to 100 or more times) to produce an extended oligonucleotide of a defined sequence. Steps (b) - (c) may be repeated as many times as necessary until an extension product of a defined sequence and length is synthesized. The end product may be 2-100 bases in length, although, in theory, the method can be used to produce products of any length, including greater than 200 bases or greater than 500 bases.
[0372] In some embodiments, methods of nucleic acid synthesis as provided herein are carried out in a reaction buffer composition. In some embodiments, the reaction buffer composition is an aqueous solution. In some embodiments, the reaction buffer compositioncomprises a set of components suitable for the stability of the polymerase, nucleotide, polymerase-nucleotide conjugates, starter molecule, nucleic acid molecule products, and any surface or matrix on which the methods disclosed herein are carried out. In some such embodiments, the reaction buffer composition comprises a set of components suitable for carrying out catalytic steps (e.g., polynucleotide polymerization performed by a polymerase) described in methods of nucleic acid synthesis in accordance with the present disclosure.
[0373] The conditions under which nucleic acid synthesis is carried out can be varied. For example, the amounts of times for carrying out each step in a stepwise nucleotide addition cycle can be varied to improve the purity of a plurality of products generated by the methods of nucleic acid synthesis described herein.
[0374] In some embodiments, methods of nucleic acid synthesis in accordance with the present disclosure generate a nucleic acid molecule product (i.e., a polynucleotide product). In some embodiments, the nucleic molecule product (i.e., polynucleotide product) has a target (i.e., pre-determined) sequence. A “target” or “pre-determined” sequence refers to a desired polynucleotide sequence that is intentionally produced by the method of nucleic acid synthesis. The pre-determined sequence can include any number of nucleotides comprising a nucleobase (e.g., adenine, thymine, guanine, cytosine, and / or uracil). In some embodiments, the nucleotide is a modified nucleotide (i.e., a nucleotide analog). In some embodiments, the nucleobase is a modified nucleobase. In some embodiments, the pre-determined sequence contains one or more designated positions which may be a random nucleobase. Inclusion of a position with a random nucleobase can be useful, for example, in introducing randomized mutation into a polynucleotide product.
[0375] In some embodiments, the present disclosure includes a method of synthesizing a polynucleotide comprising contacting a precursor polynucleotide with a conjugate comprising a nucleotide covalently linked to a polymerase via a cleavable linker, wherein said nucleotide comprises said blocked nucleobase. In some embodiments, the method of synthesizing a polynucleotide comprises cleaving a cleavable linker after addition of a nucleotide to a precursor polynucleotide. In some embodiments, the method of synthesizing a polynucleotide comprises repeating contacting, adding, and optionally cleaving steps described herein one or more times. In some embodiments, removal of one or more blocking groups described herein comprises contacting said polynucleotide with an enzyme capable of removing said one or more blocking groups from said blocked nucleobases. In some embodiments, a method of synthesizing a polypeptide comprising contacting saidpolynucleotide with two or more enzymes capable of removing said one or more blocking groups from said blocked nucleobases.
[0376] In some embodiments, synthesis of a polynucleotide comprises adding nucleotides stepwise to a starter molecule (e.g., an initial oligonucleotide) via the cycled steps of: addition of polymerase-nucleotide conjugate to an oligonucleotide, binding of the nucleotide to the 3’ end of the oligonucleotide catalyzed by the polymerase, and cleavage of the polymerase from the added nucleotide. These steps can be repeated until a desired polynucleotide is synthesized. As described herein, the use of nucleotides comprising protected nucleobases during polynucleotide synthesis helps to improve the efficiency and accuracy of synthesis by inhibiting secondary structure formation which can interfere with the addition of the incoming nucleotide by the polymerase during synthesis.
[0377] Although synthesis can be completed entirely with protected nucleotides, synthesis with a combination of unmodified and protected nucleotides can also be used effectively to improve polynucleotide synthesis. In some embodiments, only one of the four nucleotides added (e.g., from G or T) is protected during synthesis. In some embodiments, protected nucleotides are only added at targeted positions where secondary structure or ternary structure is predicted, which could interfere with synthesis. Such structures can be predicted based on the presence of complementary DNA regions in various ways and respective tools exist, such as the NUPACK algorithms (http: / / www.nupack.org / home / model). Thus, in some embodiments, synthesis of a completed polynucleotide where synthesis is improved can include the use of only 1 or 2 protected nucleotides. In some embodiments, about 5%, about 10%, about 20%, about 30%, about 50%, substantially all, or 100% of a specific nucleotide is incorporated into the polynucleotide in their protected version. In some embodiments, less than 5%, less than 10%, less than 20%, less than 30%, or less than 50% of a specific nucleotide is incorporated into the polynucleotide in its protected version. In some embodiments, more than 5%, more than 10%, more than 20%, more than 30%, or more than 50% of a specific nucleotide is incorporated into the polynucleotide in its protected version. In some embodiments, only protected guanine nucleotides are used in the nucleotide synthesis reaction. The removal of protecting groups in the terminal positions of a nucleic acid may be more challenging than the removal from internal DNA positions. Therefore, in some embodiments, nucleotide synthesis is performed such that the last and first 1, 2, or 3 positions of the synthesized nucleic acid does not comprise protected nucleotides.
[0378] A nucleic acid molecule product or polynucleotide product generated by the methods described herein can contain a plurality of products. In some embodiments, the plurality ofproducts comprises a nucleic acid molecule comprising the target (i.e., pre-determined) sequence. In some embodiments, the plurality of products comprises a nucleic acid molecule comprising a sequence that is not the target sequence. In some embodiments, the plurality of products comprises a nucleic acid molecule product comprising the target sequence and a nucleic acid molecule product that is not the target sequence. The “purity” of the plurality of products can refer to the ratio of the abundance of nucleic acid molecule products with the target sequence to the abundance of nucleic acid molecule products that do not have the target sequence. The purity of a product can be assessed by any number of methods known in the art for determining the sequence of a nucleic acid. Any suitable nucleic acid sequencing method can be used. For example, the product can be assessed, without limitation, by Sanger sequencing, next generation sequencing (e.g., Illumina sequencing), or long-read sequencing (e.g., small molecule, real-time sequencing (SMRT) and nanopore sequencing).
[0379] In some embodiments, a method of nucleic acid synthesis in accordance with the present disclosure produces a product having a purity between about 10% and about 99.99%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 10%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 10%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 20%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 30%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 40%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 50%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 60%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 70%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 80%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 90%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 95%. In some embodiments, the method of nucleic acid synthesis produces a product having a purity of at least 99%.
[0380] In any of the above-summarized embodiments, the nucleoside triphosphate may be a deoxyribonucleoside triphosphate or a ribonucleoside triphosphate. In some embodiments, a conjugate may comprise an RNA polymerase linked to a ribonucleoside triphosphate. In these embodiments, the nucleotide added to the nucleic acid may be a ribonucleotide. In other embodiments, a conjugate comprises an DNA polymerase linked to adeoxyribonucleoside triphosphate. In these embodiments, the nucleotide added to the nucleic acid may be a deoxyribonucleotide.
[0381] In some embodiments, the nucleotide is a nucleotide analog. In some embodiments, the nucleotide analog is a reversible terminator. Reversible terminators are known in the art for use in nucleic acid synthesis. Uses of reversible terminators in nucleic acid synthesis have been described previously; see, for example, WO 2021 / 122539 A1, WO 2018 / 215803 A1, WO 2021 / 094251 A1, and WO 2020 / 081985 A1.
[0382] In some embodiments, the nucleotide may comprise a reversible terminator (“RTdNTP”) and the deblocking step of the method further comprises removing the blocking group (e.g., removing the terminator group) from the added nucleotide to produce the deblocked extension product. Deblocking enables subsequent extension of the nucleic acid, and thus allows these steps to be repeated cyclically to produce an extension product of defined sequence.
[0383] A method of sequencing is also provided. These methods may comprise incubating a duplex comprising a primer and a template with a composition comprising a set of conjugates, wherein the conjugates correspond to G, A, T and C and are distinguishably labeled, e.g., fluorescently labeled; detecting which nucleotide has been added to the primer by detecting a label that is tethered to the polymerase that has added the nucleotide to the primer; deblocking the extension product by cleaving the linker; and repeating the incubation, detection and deblocking steps to obtain the sequence of at least part of the template.
[0384] The present disclosure describes a method of enzymatic polynucleotide synthesis using polymerase-nucleotide conjugates to control the iterative addition of a single nucleotide per cycle onto the 3′ hydroxyl terminus of a growing polynucleotide strand via the nucleotide-bound polymerase to perform polynucleotide synthesis. Such control is achieved through a so-called “shielding effect”. Shielding describes the steric hinderance that prevents the 3′ hydroxyl terminus that has been elongated by a conjugate from being accessed by another conjugate while the polymerase remains attached to the added nucleotide, as well as the preventing the polymerase tethered to the nucleotide at the 3′ terminus from accessing the nucleotides of other conjugates.
[0385] Described in PCT Publication No. WO2017 / 223517, incorporated by reference in its entirety herein, is a typical process for the stepwise synthesis of a defined sequence using a template-independent polymerase. A nucleic acid that serves as an initial substrate forelongation (i.e., "starter molecule") is incubated with a first polymerase-nucleotide conjugate. Once the nucleic acid has been elongated by the tethered nucleotide of a conjugate, no further elongations occur because the conjugates implement a termination mechanism. In the second step of the process, the linker is cleaved to release the polymerase and reverse the termination mechanism, thus enabling subsequent elongations. The elongation products are then exposed to the second conjugate, and these two steps are iterated to elongate the nucleic acid by a defined sequence. Also described in WO2017 / 223517 is a synthesis procedure using a conjugate comprising TdT and a photocleavable linker. As described above, other strategies are available for the attachment and cleavage of the linker.
[0386] An important step in this approach to polynucleotide synthesis is deblocking, or the removal of the tethered polymerase from the extended polynucleotide, making the 3′ terminus available for continued extension in the next cycle of synthesis. To be useful for polynucleotide synthesis, the removal of the tethered polymerase preferably occurs with rapid kinetics to reduce synthesis cycle time, while also being performed under benign conditions to prevent damage to the polynucleotide being synthesized. The removal of the tethered polymerase is also preferred to proceed to full completion, and to produce a cleavage product which does not impede continued extension or downstream applications of the complete DNA synthesis product. In some embodiments, the tether also allows for efficient conjugation of the nucleotide to the polymerase, and subsequently positions the nucleotide effectively within the active site to promote rapid incorporation to a free primer 3′ terminus.
[0387] Herein, we describe optimized cleavable linker designs used for the tethering of polymerases to nucleotides that are highly stable during storage and under oligo synthesis reaction conditions (before controlled linker cleavage), and enzymatically cleavable to completion in a short timeframe suitable for oligo synthesis.
[0388] Provided herein is a conjugate comprising a polymerase and a nucleotide linked via a linker that comprises an enzymatically cleavable linkage. The polymerase moiety of a conjugate can elongate a nucleic acid using its linked nucleotide (i.e., the polymerase can catalyze the attachment of a nucleotide to which it is joined onto a nucleic acid) and remains attached to the elongated nucleic acid via the linker until the linker is enzymatically cleaved.
[0389] In a conjugate, the linker comprises the atoms that connect the nucleotide to the polymerase. In some embodiments, the linker connects the base, the sugar, or the α- phosphate of a nucleotide to the polymerase. In some embodiments, the linker connects theterminal phosphate of a nucleotide to the polymerase. In some embodiments, the linker connects the nucleotide to the Cα atom in the backbone of the polymerase. In some embodiments, the polymerase and the nucleotide are covalently linked and the distance between the linked atom of the nucleotide and the polymerase to which it is attached may be in the range of 4-100Å, e.g., 15-40Å or 20-30Å, although this distance may vary depending on where the nucleotide is tethered. The linker used should be sufficiently long to allow the nucleotide to access the active site of the polymerase to which it is tethered. As will be described in greater detail below, the polymerase of a conjugate is capable of catalyzing the addition of the nucleotide to which it is linked onto the 3′ end of a nucleic acid.
[0390] Linkers contemplated herein are also of sufficient length and stability to allow efficient hydrolysis by enzymatic means. The number of carbons or atom in a linker, optionally derivatized by other functional groups, must be of sufficient length to allow either enzymatic cleavage of the polymerase from the nucleotide.
[0391] In certain aspects, a cleavable linker comprises an amino acid ester. In some aspects, an amino acid ester is the site of cleavage of the linker, thereby facilitating release of a polymerase upon exposure to an esterase or protease comprising esterase activity. A portion of the cleavable linker comprising the amino acid ester is referred to herein as the “L2” portion of the linker. L2can be designed and optimized for enzymatic cleavage by an esterase or protease comprising esterase activity, for example, by modifying the chemical group attached to the alpha carbon of the amino acid ester, or by including one or more amino acids adjacent to the amino acid ester as part of L2.
[0392] Described herein are polymerase-nucleotide conjugates comprising cleavable linkers that are highly stable and rapidly enzymatically cleavable by proteases comprising esterase activity. In some embodiments, a polymerase-nucleotide conjugate comprises a nucleotide linked to a polymerase using an enzymatically cleavable linker. In some embodiments, a polymerase-nucleotide conjugate comprising an enzymatically cleavable linker comprises a structure Nuc-L1-L2-L3-Pol, wherein Nuc represents a nucleotide, pol represents a polymerase, and L1-L2-L3represents an enzymatically cleavable linker. In some embodiments, L1represents a region of an enzymatically cleavable linker connecting the nucleotide to L2, L2represents a cleavable portion of an enzymatically cleavable linker, L3represents a region of an enzymatically cleavable linker connecting L2to Pol.
[0393] In some embodiments, an enzymatically cleavable linker comprises an amino acid ester moiety. In some embodiments, L2comprises an amino acid ester moiety. In someembodiments, the ester group of an amino acid ester moiety is cleavable by a protease comprising esterase activity. The ester of an amino acid of L2is attached to L1, which can also be referred to as a spacer or as a scar of a nucleotide after cleavage of the L2ester. L2is also attached to L3, the rest of the linker, which comprises attachment chemistry for polymerase conjugation. In some embodiments, L3can also include or be referred to as a spacer. In some embodiments, L2further comprises additional amino acids bound to the amine of the amino acid ester to serve as a protease substrate. As described herein, L2is optimized for ester stability to prevent spontaneous cleavage while retaining the ability to act as a suitable substrate for esterase activity of a protease comprising esterase activity.
[0394] In some embodiments, the linker is bound to the nucleotide at the nucleobase. In some embodiments, the linker is bound to the nucleotide at the sugar. In some embodiments, the linker is bound to the nucleotide at a 5′ phosphate group, wherein the nucleotide is any nucleoside polyphosphate. In some embodiments, the linker is bound to the alpha phosphate. In some embodiments, the linker is bound to the gamma, beta, delta, epsilon, zeta, eta, or theta phosphate. In some embodiments, the linker is bound to the terminal phosphate. In some embodiments, a linker of a conjugate may be attached to the 7-position of deaza dGTP or the 5-position of dTTP or dUTP.
[0395] Additional tethered nucleotides can be found, e.g., in PCT Publication WO2017 / 223517 “Nucleic Acid Synthesis and Sequencing Using Tethered Nucleoside Triphosphates,” the entirety of which is incorporated by reference.
[0396] In some embodiments, the tethered nucleotide may be specifically attached to a cysteine residue of the polymerase using a sulfhydryl-specific attachment chemistry. Possible sulfhydryl specific attachment chemistries include, but are not limited to ortho-pyridyl disulfide (OPSS), maleimide functionalities, 3- arylpropiolonitrile functionalities, allenamide functionalities, haloacetyl functionalities such as iodoacetyl or bromoacetyl, alkyl halides or perfluroaryl groups that can favorably react with sulfhydryls surrounded by a specific amino acid sequence (Zhang, Chi, et al. Nature chemistry 8, (2015) 120-128.). Other attachment chemistries for specific labeling of cysteine residues will be apparent to those skilled in the art or are described in the pertinent literature and texts (e.g., Kim, Younggyu, et al, Bioconjugate chemistry 19.3 (2008): 786-791.).
[0397] We have prepared and assayed several TdT-dNTP conjugates using different cleavable linkers to tether the nucleotide site-specifically to the polymerase. The use of a peptide bond in a linker results in a linker that can be cleaved by a protease. Cleavage of apeptide bond by a protease generates an amine and a carboxylic acid, both of which are charged under the buffer conditions that are typical for TdT activity. However, having charged functional groups persist on synthesized oligonucleotides can lead to deleterious effects during synthesis.
[0398] In contrast cleavage of an ester group generates an alcohol - a charge-neutral cleavage product. Herein, we demonstrate that nucleotides including scars containing such alcohols do not hinder conjugate-based oligonucleotide synthesis (see Example 2). We have also demonstrated that linkers including an amino acid ester can be cleaved enzymatically by a protease comprising esterase activity, such as Proteinase K (see Example 2).
[0399] Therefore, in some embodiments, L2comprises an amino acid ester. In some embodiments, the amino acid ester is the site of cleavage of the linker, facilitating the release of the polymerase from the nucleotide.
[0400] In addition, we initially observed that the ester group of a glycine amino acid ester in the linker could be unstable, resulting in spontaneous cleavage of the conjugate and unwanted nucleotide insertions during conjugate-based oligonucleotide synthesis (see Examples 2 and 3). However, addition of aliphatic or bulky substituents to the alpha carbon of the amino acid ester was observed to favorably improve stability of the adjacent ester (see Examples 4 and 6). In addition, substitution of atoms at the alpha carbon of the amino acid ester can affect hyperconjugation, resulting in an increase or decrease in the lability of the adjacent ester, as well as rate of cleavage by a protease comprising esterase activity. Thus, selection of a preferred substituent at the alpha carbon of the amino acid ester can be used to achieve an acceptable balance between stability and linker cleavage kinetics (see Example 6).
[0401] Therefore, in some embodiments, the amino acid ester comprises one or more substitutions at the alpha carbon, such as addition of an aliphatic or bulky substituent. In some embodiments, the amino acid ester is represented by:wherein R1and R1'are each independently selected from an optionally substituted C1-3 alkyl, a halogen, or are optionally taken together with the atom on which they are attached to form an optionally substituted C3-C7carbocyclic ring.
[0402] Exemplary L2linker structures with different substituents on the alpha carbon of the amino acid ester (e.g., to improve stability of the ester) are shown below:
[0403] Furthermore, we observed that the kinetics of cleavage of the ester group by a protease comprising esterase activity could be improved by including one or more amino acids in the L2moiety adjacent to the amino acid ester (see Example 5). In some embodiments, L2comprises an amino acid ester adjacent to one or more amino acid residues. In some embodiments, the one or more amino acid residues are bound to the amine group of the amino acid ester.
[0404] In some embodiments, L2comprises or consists of:wherein R1and R1'are independently selected from hydrogen or an optionally substituted C1-3 alkyl or are taken together with the atom on which they are attached to form an optionally substituted C3-C7carbocyclic ring; each R3is an optionally substituted group independently selected from hydrogen, C1-6 alkyl, benzyl, -OH, -O(C1-6 alkyl), and -CN; each Rcis hydrogen or optionally substituted C1-6alkyl; and n is 1, 2, or 3.
[0405] In some embodiments, the one or more amino acids linked to the amine of the amino acid ester comprise L- or D- isomers of amino acid residues. The term “naturally-occurring amino acid” refer to Ala, Asp, Cys, Glu, Phe, Gly, His, He, Lys, Leu, Met, Asn, Pro, Gin, Arg, Ser, Thr, Val, Trp, Tyr, or citrulline. “D-” designates an amino acid having the “D” (dextrorotary) configuration, as opposed to the configuration in the naturally occurring (“L-”) amino acids. The amino acids described herein can be purchased commercially (Sigma Chemical Co., Advanced Chemtech) or synthesized using methods known in the art. In some embodiments, amino acids with non-natural or artificial side chains are linked to the amine of the amino acid ester.
[0406] As discussed above, we observed that the composition of the linker (i.e., the peptide sequence of L2) has a significant impact on the rate of protease-mediated deblocking. Accordingly, various permutations of amino acids in L2could yield conjugates with faster addition and deblocking kinetics. Such linkers could include variations of amino acid identity and number of consecutive amino acids.
[0407] The one or more amino acids included in the L2portion of the linker / bound to the amino acid ester can be selected, for example, to optimize protease binding and ester cleavage. A combinatorial library can be generated to test optimal cleavage activity, amino acids can be chosen based on existing known peptide sequence targets for the protease. The protease comprising esterase activity can recognize the peptide portion of the linker and hydrolyzes the ester group of the amino acid ester of L2, resulting in removal of polymerase attached to the nucleotide via the linker, as disclosed herein.
[0408] If desired, a spacer can be used between the nucleotide and the linker, or between the linker and the label. Different lengths of spacers can be used in order to increase L2 availability towards the protease / esterase and increase the efficiency and fidelity ofpolymerases. Exemplary spacers include, for example, polyethyleneglycol or other suitable spacers.
[0409] Examples of linkers comprising L2structures including an amino acid ester bound to one or more amino acid residues is shown below:
[0410] There is considerable flexibility on the type of linker used for regions of the linker not associated with enzymatic cleavage disclosed herein (i.e., L1and L3). Examples of suitable linker structures may include, but are not limited to, carbon-chain linkers (e.g., C6, C12, C18, C24, etc.), peptide linkers (e.g., poly-glycine or poly-alanine ranging from about 1 residue to about 1,000 residues in length), or polyether linkers (e.g., PEG, PPG, PAG, PTMG from about 1 polyether unit to about 1,000 polyether units in length).
[0411] In some embodiments, L1or L3is a chain of atoms selected from C, N, O, S, Si, and P, preferably having 0-500 atoms, wherein L1covalently connects to Nuc and L2, and wherein L3covalently connects to L2and Pol. The atoms used in forming L1or L3may be combined in all chemically relevant ways, such as forming alkylene, alkenylene, and alkynylene, carbamates, carbonates, ethers, polyoxyalkylene, esters, amines, imines, polyamines, hydrazines, hydrazones, amides, ureas, semicarbazides, carbazides, alkoxyamines, alkoxylamines, urethanes, amino acids, peptides, acyloxylamines, hydroxamic acids, or combination above thereof.
[0412] In some embodiments, L1or L3comprises one or more carbon atoms, zero, one, or more oxygen atoms, zero, one or more nitrogen atoms, zero, one, or more sulfur atoms, or a combination thereof, in different embodiments. In some embodiments, L1or L3comprise, comprise about, comprise at least, comprise at least about, comprise at most, or comprise at most about, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000,8000, 9000, 10000, or a number or a range between any two of these values, carbon atom(s), oxygen atom(s), nitrogen atom(s), sulfur atom(s), or a combination thereof.
[0413] In some embodiments, L1or L3comprise a polymer, such as a homopolymer or a heteropolymer. In some embodiments, L1or L3comprise a plurality of repeat units. In some embodiments, the plurality of repeating units comprises identical repeating units. In some embodiments, the plurality of repeating units comprises two or more different repeating units. The plurality of repeating units can comprise a polyether such as paraformaldehyde, polyethylene glycol (PEG), polypropylene glycol (PPG), polyalkylene glycol (PAG), polytetramethylene glycol (PTMG), or a combination thereof. For example, the plurality of repeating units can comprise PEGix, PEG23, PEG24, or a combination thereof. The plurality of repeating units can comprise a polyalkylene, such as polyethene, polypropene, polybutene, or a combination thereof. In some embodiments, a repeating unit of the plurality of repeating units comprises no aromatic group. In some embodiments, a repeating unit of the plurality of repeating units comprises one or more aromatic groups.
[0414] In some embodiments, L1or L3comprises any number of basic chemical starting blocks. For example, linkers may comprise linear or branched alkyl, alkenyl, or alkynyl chains, or combinations thereof, that provide a useful distance between the nucleotide and polymerase, the nucleotide and L2, or the polymerase and L2. For instance, amino-alkyl linkers, e.g., amino-hexyl linkers, have been used to attach linkers to nucleotide analogs, and are generally sufficiently rigid to maintain such distances. The longest chain of such linkers may include as many as 2 atoms, 3 atoms, 4 atoms, 5 atoms, 6 atoms, 7 atoms, 8 atoms, 9 atoms, 10 atoms, or even 11-35 atoms, or even 35-50 atoms. The linear or branched linker may also contain heteroatoms other than carbon, including, but not limited to, oxygen, sulfur, phosphate, and nitrogen. A polyoxyethylene chain (also commonly referred to as polyethyleneglycol, or PEG) is a preferred linker constituent due to the hydrophilic properties associated with polyoxyethylene. Insertion of heteroatom such as nitrogen and oxygen into the linkers may affect the solubility and stability of the linkers.
[0415] The linker, including L1or L3, may be rigid in nature or flexible. Rigid structures include laterally rigid chemical groups, e.g., ring structures such as aromatic compounds, multiple chemical bonds between adjacent groups, e.g., double or triple bonds, in order to prevent rotation of groups relative to each other, and the consequent flexibility that imparts to the overall linker. Thus, the degree of desired rigidity may be modified depending on the content of the linker, or the number of bonds between the individual atoms comprising the linker. Further, addition of ringed structures along the linker may impart rigidity. Ringedstructures may include aromatic or non-aromatic rings. Rings may be anywhere from 3 carbons, to 4 carbons, to 5 carbons or even 6 carbons in size. Rings may also optionally include heteroatoms such as oxygen or nitrogen and also be aromatic or non-aromatic. Rings may additionally optionally be substituted by other alkyl groups and / or substituted alkyl groups.
[0416] Linkers that comprise ring or aromatic structures can include, for example aryl alkynes and aryl amides. Other examples of the linkers of the disclosure include oligopeptide linkers that also may optionally include ring structures within their structure.
[0417] In embodiments, L1or L3is a C1-C10 alkylene chain, wherein 1-6 methylene units are optionally and independently replaced by -NH-, -O-, -C(O)-, -C(O)NH-, -NHC(O)-, - NHC(O)NH-, -C(O)O-, -OC(O)-, -SS-, optionally substituted cycloalkylene (e.g., C3-C8, C3- C6, or C5-C6), optionally substituted heterocycloalkylene (e.g., 3 to 8, 3 to 6, or 5 to 6 membered), optionally substituted arylene (e.g., C6-C10, C10, or phenylene), or substituted or unsubstituted heteroarylene (e.g., 5 to 10, 5 to 9, or 5 to 6 membered).
[0418] In embodiments, L1or L3is a bond, -NH-, -O-, -C(O)-, -C(O)NH-, -NHC(O)-, - NHC(O)NH-, -C(O)O-, -OC(O)-, -SS-, optionally substituted alkylene (e.g., C1- C20, C10- C20, C1-C8, C1-C6, or C1-C4), optionally substituted heteroalkylene (e.g., 2 to 20, 8 to 20, 2 to 10, 2 to 8, 2 to 6, or 2 to 4 membered), optionally substituted cycloalkylene (e.g., C3-C8, C3- C6, or C5-C6), optionally substituted heterocycloalkylene (e.g., 3 to 8, 3 to 6, or 5 to 6 membered), optionally substituted (e.g., C6-C10, C10, or phenylene), or optionally substituted (e.g., 5 to 10, 5 to 9, or 5 to 6 membered). In embodiments, L1or L3is optionally substitutedC1-C20 alkylene. In some embodiments, L1or L3optionally substituted 2 to 20 membered heteroalkylene. In embodiments, L1or L3is optionally substituted C3-C8 cycloalkylene. In some embodiments, L1or L3is optionally substituted 3 to 8 membered heterocycloalkylene. In embodiments, L1or L3is optionally substituted C6-C10 arylene. In embodiments, L1or L3is optionally substituted 5 to 10 membered heteroarylene.
[0419] In some embodiments, L1is substituted with 1-6 instances of RL. In some embodiments, L3is substituted with 1-6 instances of RL. Each RLis independently selected from the group consisting of oxo, halogen, -CCI3, -CBr3, -CF3, -CI3, -CN, -OH, -NH2, - COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, -NHNH2, -ONH2, -NHC(O)NHNH2, -NHC(O)NH2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -NHOH, -OCCl3, -OCF3, -OCBr3, - OCI2, -OCHCI2, -OCHBr2, -OCHb, -OCHF2, -N3, optionally substituted alkyl (e.g., C1-C20, 8 to 20, 2 to C8, C3-C6,or C5-C6), optionally substituted heterocycloalkyl (e.g., 3 to 8, 3 to 6, or 5 to 6 membered), optionally substituted aryl (e.g., C6-C10, C10, or phenyl), and optionally substituted heteroaryl (e.g., 5 to 10, 5 to 9, or 5 to 6 membered).
[0420] In embodiments, L1or L3is -(CH2CH2O)b-. In embodiments, L1or L3is - CCCH2(OCH2CH2)a-NHC(O)-(CH2)c(OCH2CH2)b-. In embodiments, L1or L3is - CHCHCH2-NHC(O)-(CH2)c(0CH2CH2)b-. In embodiments, L1or L3is -CCCH2-NHC(O)- (CH2)c(OCH2CH2)b-. In embodiments, L1or L3is -CCCH2-. The symbol a is an integer from 0 to 8. In embodiments, a is 1. In embodiments, a is 0. The symbol b is an integer from 0 to 8. In embodiments, b is 1 or 2. In embodiments, b is an integer from 2 to 8. In embodiments, b is 1. The symbol c is an integer from 0 to 8. In embodiments, c is 3. In embodiments, c is 1. In embodiments, c is 2. In embodiments, L1 or L3 is independently a substituted or unsubstituted C1-C4 alkylene or substituted or unsubstituted 8 to 20 membered heteroalkylene.
[0421] L1 acts as an attachment point to the nucleotide and includes a hydroxyl terminal group which binds to a portion of L2 during synthesis.
[0422] In some embodiments, L1 is a scar that is enzymatically or chemically cleavable after cleavage of polymerase-nucleotide linker / removal of the L2-L3-pol moiety.
[0423] In some embodiments, L1 is selected from the group consisting of a bond, an optionally substituted C1-12 alkylene chain, C4-C20 polyethylene glycol, an optionally substituted C2-12alkenylene chain, and an optionally substituted C2-12alkynylene chain, wherein 1-4 methylene units of L1are optionally and independently replaced with -O-, - N(Rb)-, -C(O)-, -S-, -S(O)-, -S(O)2-, phenylene, cyclopropylene; wherein each Rbis independently hydrogen or optionally substituted C1-6 alkyl.
[0424] In some embodiments, L1 comprises:wherein each Rais independently selected from the group consisting of halogen, hydroxyl, cyano, optionally substituted C1-6 alkyl, and optionally substituted C1-6 alkoxy.
[0425] In some embodiments, L1is selected from the group consisting of:.
[0426] In some embodiments, L3 comprises a bioconjugate group suitable for conjugation of L3 to the polymerase.
[0427] In some embodiments, the bioconjugate group is an N-hydroxysuccinimide ester (NHS) group. In some embodiments, the bioconjugate group is a maleimide group. The linker may then be covalently attached to the polymerase by reaction of the maleimide group with a cysteine residue of the polymerase.
[0428] In some embodiments, the polymerase may be operably linked to a linker moiety including a covalent or non-covalent bond; amino acid tag (e.g., poly-amino acid tag, poly- His tag, 6His-tag (SEQ ID NO: 53)); chemical compound (e.g., polyethylene glycol); protein-protein binding pair (e.g., biotin-avidin); affinity coupling; capture probes; or any combination of these. The linker moiety can be separate from or part of a polymerase variant.
[0429] In some embodiments, the linker connecting the nucleotide and the polymerase comprises a saturated or unsaturated, substituted, or unsubstituted, straight, or branched carbon chain. The length of the linker can be different in different embodiments. The length of the linker may vary depending on the type of nucleotide and the polymerase. In some embodiments, the linker length in the enzyme linked nucleotide is different for each different nucleotide or nucleotide analog. In some embodiments, the linker has a length of, of about, of at least, of at least about, of at most, or of at most about, 19 Å, 20 Å, 21 Å, 22 Å, 23 Å, 24 Å, 25 Å, 26 Å, 27 Å, 28 Å, 29 Å, 30 Å, 31 Å, 32 Å, 33 Å, 34 Å, 35 Å, 36 Å, 37 Å, 38 Å, 39 Å, 40 Å, 41 Å, 42 Å, 43 Å, 44 Å, 45 Å, 46 Å, 47 Å, 48 Å, 49 Å, 50 Å, 51 Å, 52 Å, 53 Å, 54 Å, 55 Å, 56 Å, 57 Å, 58 Å, 59 Å, 60 Å, 61 Å, 62 Å, 63 Å, 64 Å, 65 Å, 66 Å, 67 Å, 68 Å, 69 Å, 70 Å, 71 Å, 72 Å, 73 Å, 74 Å, 75 Å, 76 Å, 77 Å, 78 Å, 79 Å, 80 Å, 81 Å, 82 Å, 83 Å, 84 Å, 85 Å, 86 Å, 87 Å, 88 Å, 89 Å, 90 Å, 91 Å, 92 Å, 93 Å, 94 Å, 95 Å, 96 Å, 97 Å, 98 Å, 99 Å, 100 Å, 200 Å, 300 Å, 400 Å, 500 Å, 600 Å, 700 Å, 800 Å, 900 Å, 1000 Å, or a number or a d thenucleotide are covalently linked, and the distance between the linked atom of the nucleotide and the polymerase is from about 4 Å to about 100 Å. In some embodiments, the distance between the linked atom of the nucleotide and the polymerase is about 5 Å to about 20 Å. In some embodiments, the distance between the linked atom of the nucleotide and the polymerase is about 20 Å to about 50 Å. In some embodiments, the distance between the linked atom of the nucleotide and the polymerase is about 50 Å to about 75 Å. In some embodiments, the distance between the linked atom of the nucleotide and the polymerase is about 75 Å to about 100 Å.
[0430] In some embodiments, the length of the linker will be defined as its persistence length, corresponding to the root-mean-square (RMS) distance between the ends of the linker as characterized by dynamic simulations, 2-D trapping experiments, or ab initio calculations based on statistical distributions of polymers in compact, collapsed, or fluid states as required by the solution, suspension, or fluid conditions present. In some embodiments, a linker may have a persistence length of at least 0.1, at least 0.2, at least 0.4, at least 1, at least 2, at least 4, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 700, or at least 1,000 nm, or a persistence length in a range defined by or comprising any two or more of these values. In some embodiments, a linker for connecting the nucleotide to the enzyme can have a persistence length of about 0.1 - 1,000 nm, 0.5 - 500 nm, 0.5 - 400 nm, 0.5 - 300 nm, 0.5 - 200 nm, 0.5 - 100 nm, 0.5 - 50 nm, 1 - 500 nm, 1 - 400 nm, 1 - 300 nm, 1 - 200 nm, 1 - 100 nm, 1-50 nm, 1.5 - 500 nm, 1.5 - 400 nm, 1.5 - 300 nm, 1.5 - 200 nm, 1.5 - 100 nm, 1.5 - 50 nm, 5 - 500 nm, 5 - 400 nm, 5 - 300 nm, 5 - 200 nm, 5 - 100 nm, or 5 - 50 nm. In some embodiments, the linker may have a persistence length of shorter than about 5, 10, 20, 30, 40, 50, 60, 80, 100, 200, 300, 400, 500, 700, or 1,000 nm. In some embodiments, linkers provided for one nucleotide may be longer or shorter than the linker provided for another nucleotide. In some embodiments, linkers provided for one polymerase may be longer or shorter than the linker provided for another polymerase.
[0431] In some embodiments, the conjugate is represented by
[0432] In some embodiments, a conjugate is represented by a structure of Formula (I) or (II):wherein L1is selected from the group consisting of an optionally substituted C1-6alkylene chain, an optionally substituted C2-6 alkenylene chain, and an optionally substituted C1-6 alkynylene chain, wherein 1-4 methylene units are optionally and independently replaced with -O-, - N(Ra)-, -C(O)-, -S-, -S(O)-, -S(O)2-, or phenylene; L2is a cleavable linker; L3is a linker connecting pol to L2each Rais independently hydrogen or C1-6alkyl; R2is hydrogen or methyl; R is a ribose polyphosphate or deoxyribose polyphosphate; and Pol is a polymerase.
[0433] In some embodiments, the conjugate is selected from the group consisting ofSecondary Structure Prevention and Scar Removal
[0434] One problem for de novo enzymatic polynucleotide synthesis can occur when a polynucleotide being synthesized begins to form secondary structures, inhibiting the extension reaction for a template-independent polymerase that has low activity on a duplex structure. Elevation of temperature has been explored to reduce secondary structure (Barthel et al., Enhancing Terminal Deoxynucleotidyl Transferase Activity on Substrates with 3′ Terminal Structures for Enzymatic De Novo DNA Synthesis, Genes 2020, 11(1), 102). However, elevated temperature induces damage to DNA (and quickly damages RNA), and secondary structure remains even in elevated temperatures suitable for synthesis. Synthesis reactions at elevated temperatures also require thermostable polymerases. Wild-type template independent polymerases such as Terminal deoxynucleotidyl Transferase (TdT) are not thermostable. Use of bases with exocyclic amines masked as azido groups has also been explored (Nuclera Nucleics PCT Publication WO2020229831A1). However, the unmasking reagent (TCEP) causes DNA damage, and there are doubts about the stability of the azido modification.
[0435] In some embodiments, provided herein are improved methods for enzymatic synthesis of long polynucleotides by cyclic stepwise extension using a template-independent polymerase and modified nucleotides to prevent secondary structure formation. As shown herein, secondary structure formation in the nascent chain, which may inhibit extension reactions, is suppressed by the use of modified nucleotides with protecting groups on a one or more oxygens or nitrogens of nucleobases that prevent secondary structure formation, such as Watson-Crick base pairing or G-quadruplex formation. After the synthesis is completed, the nucleobases can be converted back into their native form by removal of theprotecting group to facilitate further use and / or downstream processing of the synthesized polynucleotides.
[0436] In some embodiments, provided herein are improved base pair protecting groups on the nucleobase to inhibit secondary structure during polynucleotide synthesis, which can be efficiently removed after synthesis to leave a polynucleotide without modified nucleotides that might interfere with downstream applications. Such modified nucleotides are shown herein to be useful for both enzymatic polynucleotide synthesis with free nucleotides and for conjugate-based polynucleotide synthesis.
[0437] In some embodiments, provided herein are compositions and methods of oligonucleotide synthesis that inhibit secondary structure formation and improve oligonucleotide synthesis length and accuracy by providing protecting groups attached to the nucleobase of the synthesized polynucleotide that inhibit secondary structure formation. In some embodiments, these are provided as modified nucleotides incorporated into an oligonucleotide during enzymatic synthesis. In some embodiments, these are provided as part of a linker-nucleotide conjugate used during enzymatic oligonucleotide synthesis, such that the linker is attached to a base pairing nitrogen or oxygen atom on the nucleobase, and cleavage of the linker to separate the polymerase from the nucleotide during synthesis leaves a portion of the linker attached to a base pairing nitrogen or oxygen atom, which can act as a protecting group (also referred to herein as a “scar”) to inhibit secondary structure formation, as shown in FIG. 2.
[0438] In some embodiments, the present disclosure includes a method of synthesizing a polynucleotide, comprising: providing a polynucleotide comprising one or more protected nucleobases and removing one or more protecting groups from said protected nucleobases. In some embodiments, providing a polynucleotide comprises: contacting a precursor polynucleotide with a nucleotide comprising a protected nucleobase and a template- independent polymerase; and adding said nucleotide to the 3′ end of said precursor polynucleotide via said template-independent polymerase. In some embodiments, a method of synthesizing a polynucleotide further comprises repeating contacting and adding step one or more times.
[0439] As used herein, a “protected” nucleotide, as used herein, refers to a nucleotide that has biomolecule attached to a base pairing oxygen or nitrogen on the nucleobase. In some embodiments, the biomolecule inhibits hydrogen bonding of the oxygen or nitrogen to other nucleotides, such as in Watson-Crick base pairing, G-quadruplex formation, and other types of hydrogen bonding that can generate secondary structure. Thus, in some embodiments, thebiomolecule inhibits formation of secondary structure during oligonucleotide synthesis. A “scarred” nucleotide and a “protected” nucleotide can both refer to the same structure when a linker is bound to an oxygen or nitrogen on the nucleobase. Such that cleavage of the linker leaves a “scarred” nucleotide that is also a “protected” nucleotide. In some embodiments, a conjugate linker is attached to a base pairing oxygen or nitrogen to take advantage of the presence of a scar to provide a protected nucleotide to inhibit secondary structure formation during synthesis. Protected nucleotides can also refer to modified nucleotides using during oligonucleotide synthesis that are not part of a conjugate, but are useful to prevent secondary structure during oligonucleotide synthesis.
[0440] In some embodiments, the present disclosure includes a method of synthesizing a polynucleotide comprising contacting a precursor polynucleotide with a nucleotide and a polymerase, wherein said nucleotide comprises comprising a protecting group bound to a base pairing oxygen or nitrogen on the nucleobase. In some embodiments, the method of synthesizing a polynucleotide comprises removing a blocking group, such as a conjugated polymerase or a reversible terminator, after addition of a nucleotide to a precursor polynucleotide. In some embodiments, the method of synthesizing a polynucleotide comprises repeating contacting, adding, and optionally removing a blocking group described herein one or more times. In some embodiments, removal of one or more protecting groups described herein comprises exposing said polynucleotide to a chemical or photolytic condition capable of removing said one or more protecting groups from said protected nucleobases.
[0441] Although synthesis can be completed entirely with protected nucleotides, synthesis with a combination of unmodified and protected nucleotides can also be used effectively to improve polynucleotide synthesis. In some embodiments, only one of the four nucleotides added (e.g., from G or T) is protected during synthesis. In some embodiments, protected nucleotides are only added at targeted positions where secondary structure or ternary structure is predicted, which could interfere with synthesis. Such structures can be predicted based on the presence of complementary DNA regions in various ways and respective tools exist, such as the NUPACK algorithms (http: / / www.nupack.org / home / model). Thus, in some embodiments, synthesis of a completed polynucleotide where synthesis is improved can include the use of only 1 or 2 protected nucleotides. In some embodiments, about 5%, about 10%, about 20%, about 30%, about 50%, substantially all, or 100% of a specific nucleotide is incorporated into the polynucleotide in their protected version. In some embodiments, less than 5%, less than 10%, less than 20%, less than 30%, or less than 50% of a specificnucleotide is incorporated into the polynucleotide in its protected version. In some embodiments, more than 5%, more than 10%, more than 20%, more than 30%, or more than 50% of a specific nucleotide is incorporated into the polynucleotide in its protected version. In some embodiments, only protected guanine nucleotides are used in the nucleotide synthesis reaction. The removal of protecting groups in the terminal positions of a nucleic acid may be more challenging than the removal from internal DNA positions. Therefore, in some embodiments, nucleotide synthesis is performed such that the last and first 1, 2, or 3 positions of the synthesized nucleic acid does not comprise protected nucleotides.
[0442] In some embodiments, the present disclosure includes a method of synthesizing a polynucleotide comprising contacting a precursor polynucleotide with a conjugate comprising a nucleotide covalently linked to a polymerase via a cleavable linker, wherein cleavage of said nucleotide from said polymerase generates a scarred nucleobase. In some embodiments, the method of synthesizing a polynucleotide comprises cleaving a cleavable linker after addition of a nucleotide to a precursor polynucleotide. In some embodiments, the method of synthesizing a polynucleotide comprises repeating contacting, adding, and optionally cleaving steps described herein one or more times. In some embodiments, removal of one or more scars described herein comprises contacting said polynucleotide with a chemical of photolytic condition capable of removing said one or more scars from said scarred nucleobases.
[0443] In some embodiments, synthesis of a polynucleotide comprises adding nucleotides stepwise to a starter molecule (e.g., an initial oligonucleotide) via the cycled steps of: addition of polymerase-nucleotide conjugate to an oligonucleotide, binding of the nucleotide to the 3′ end of the oligonucleotide catalyzed by the polymerase, and cleavage of the polymerase from the added nucleotide. These steps can be repeated until a desired polynucleotide is synthesized. As described herein, the use of nucleotides comprising protected nucleobases during polynucleotide synthesis helps to improve the efficiency and accuracy of synthesis by inhibiting secondary structure formation which can interfere with the addition of the incoming nucleotide by the polymerase during synthesis.
[0444] Once a polymerase is cleaved from a nucleotide during conjugate synthesis, a portion of the linker may remain attached to the nucleotide, leaving a scar as compared to a naturally occurring nucleotide. This can negatively impact downstream use the synthesized polynucleotide, including amplification of the synthesized polynucleotide, or direct use for its intended application. Therefore, what are needed are improved conjugates and methodsof conjugate-based polynucleotide synthesis that include scars that can be removed during or after synthesis.
[0445] One advantage of the methods and compositions provided herein is that they allow scarless synthesis of a polynucleotide when using polymerase-nucleotide conjugates for synthesis. Preferred linker structures and methods of removing residual scars / protecting groups after oligonucleotide synthesis to leave a naturally occurring polynucleotide without scars are also provided herein.
[0446] As used herein, a “scarred” nucleotide refers to a nucleotide that has a portion of a linker still attached to the nucleotide after cleavage of the linker to release an attached biomolecule, such as a polymerase of a polymerase-nucleotide conjugate.
[0447] Accordingly, in some embodiments, provided herein is a method of synthesizing a polynucleotide using one or more conjugates, wherein cleavage of the polymerase from the nucleotide leaves scarred nucleotides in the polynucleotide. The resulting polynucleotide can then be treated suitable conditions as described herein to remove the scar from the modified nucleotides, resulting in a polynucleotide with unmodified nucleobases. In some embodiments, removal of one or more scars from the synthesized polynucleotide comprises contacting the polynucleotide with suitable conditions capable of removing said one or more scars from a scarred nucleobase.
[0448] In some embodiments, described herein is a method of synthesizing a polynucleotide, comprising: providing a polynucleotide comprising one or more scarred nucleobases and removing one or more scars from said scarred nucleobases. In some embodiments, providing a polynucleotide comprises: contacting a precursor polynucleotide with a nucleotide comprising a nucleobase linked to a template-independent polymerase; and adding said nucleotide to the 3′ end of said precursor polynucleotide via said template- independent polymerase. In some embodiments, a method of synthesizing a polynucleotide further comprises repeating contacting and adding step one or more times. Enzymatically Removeable Protecting Groups
[0449] In some embodiments, provided herein are modified nucleotides or nucleotide- linkers, than when the linker is cleaved, a protecting group or scar remains on the nucleobase that can be removed enzymatically.
[0450] Also provided herein are improved methods for synthesis of nucleic acids by cyclic extension using a template-independent polymerase. As shown herein, secondary structure formation in the nascent chain, which may inhibit extension reactions, is suppressed by theuse of modified nucleotides with methylation or other alkylations of the exocyclic oxygen of nucleobases that prevent Watson-Crick base pairing and / or other structures (such as G- quadruplex formation). After the synthesis is completed, the nucleobases can be converted back into their native form by enzymatic removal of the alkyl group.
[0451] The present disclosure includes a method of synthesizing a polynucleotide comprising one or more alkylated nucleobases, and removing one or more alkyl groups from said alkylated nucleobases. In some embodiments, inclusion of one or more alkylated nucleobases prevents base pairing or the formation of undesirable secondary structure during synthesis. In some embodiments, an alkylated nucleobase is described below in classes and subclasses herein. In some embodiments, providing a polynucleotide comprises: contacting a precursor polynucleotide with a nucleotide comprising an alkylated nucleobase and a template-independent polymerase; and adding said nucleotide to the 3′ end of said precursor polynucleotide via said template-independent polymerase. In some embodiments, a method of synthesizing a polynucleotide further comprises repeating contacting and adding step one or more times.
[0452] After completion of synthesis of a polynucleotide using one or more alkylated nucleotides, the resulting polynucleotide can then be treated with an alkyl transferase to remove the alkyl group bound to the nucleotides, resulting in a polynucleotide with unmodified nucleobases. In some embodiments, removal of one or more alkyl groups from the synthesized polynucleotide comprises contacting the polynucleotide with one or more enzymes capable of removing said one or more alkyl groups from an alkylated nucleobase. In some embodiments, the enzyme suitable for de-alkylation of the alkylated nucleobase is an alkyl transferase. In some embodiments, a suitable enzyme for de-alkylating the polynucleotide is an alkyl transferase from EC 2.1.1.63. In some embodiments, the alkyl transferase is an AGT (alkylguanine transferase) enzyme, e.g., O6-alkylguanine DNA alkyl transferase. In some embodiments the enzyme used to remove the alkyl group from the alkylated nucleobase is AlkB (E. coli), which is an alpha-ketoglutarate-dependent hydroxylase, which oxidatively dealkylates the DNA substrate.
[0453] FIG. 3 and FIG. 4 show a reaction including cleavage of a linker in a nucleotide-TdT conjugate and removal of an alkyl group from an alkylated nucleotide by an AGT enzyme. In FIG. 3, the linker is separate from the alkyl modification of the nucleotide. In FIG. 4, the TdT is linked to the alkyl modification of the nucleotide and the alkyl group remains after cleavage of the linker binding the nucleotide and the TdT.
[0454] Suitable enzymes for use in de-alkylating alkylated nucleobases after completion of synthesis can be determined by screening a set of enzymes known to be involved in a de- alkylation reaction of a nucleobase or closely related reaction. One example of an easy screening method is described herein in Example 8. Using such screening methods, one of ordinary skill in the art can identify suitable enzymes for de-alkylation and implementations of this DNA synthesis strategy.
[0455] For example, although O6-alkylguanine DNA alkyltransferase and AlkB are exemplified enzymes suitable for de-alkylation, a number of other enzymes from various species are closely related and could also be suitable for use in the de-alkylation of the synthesized nucleotides. Tables 1 and 2 below provides a list of alkyl transferases that could be suitable to de-alkylate polynucleotide synthesis products described herein. Table 1 – Alkyl transferase enzyme list (wild type)Table 2 – Alkyl transferase enzyme list (engineered variants)
[0456] In some embodiments, an alkylated nucleobase is represented by Formula (V) or (VI):wherein X is -C(R2)= or -N=; R1is selected from the group consisting of C1-6alkyl, C2-6alkenyl, C1-6alkynyl, and –(CH2)0-3Ph, wherein R1is optionally substituted with 1-6 instances of R1a; each R1ais independently selected from halogen, C1-6alkyl, -(CH2)0-3OR1b, -NO2, -N3, - OPO2OH, and –(CH2)0-3NHR1b; and each R1bis independently selected from hydrogen, C1-6 alkyl, -C(O)(C1-6 alkyl), C1-6 haloalkyl, -C(O)(C1-6haloalkyl), and -CH2OAc; R2is selected from the group consisting of hydrogen, optionally substituted C1-4alkyl chain, wherein 1-2 methylene units is optionally and independently replaced with -O-, -N(Ra)-, - C(O)-, -S-, -S(O)-, -S(O)2-, or phenylene, wherein R2is optionally substituted with 1-6 instances of R2a; each R2ais independently selected from halogen, C1-6alkyl, -(CH2)0-3OR2b, -NO2, -N3, - OPO2OH, and –(CH2)0-3NHR2b; and each R2bis independently selected from hydrogen and C1-6alkyl.
[0457] In some embodiments, R1is C1-4alkyl. In some embodiments, R1is selected from the group consisting to methyl, ethyl, n-propyl, and n-butyl. In some embodiments, R1is methyl. In some embodiments, R1is ethyl. In some embodiments, R1is n-propyl. In some embodiments, R1is n-butyl.
[0458] In some embodiments, R1is selected from the group consisting of:.
[0459] In some embodiments, R1is selected from the group consisting of:
[0460] In some embodiments, X is -C(R2)= or -N=. In some embodiments, X is -C(R2). In some embodiments, X is -N=.
[0461] In some embodiments, R2is selected from the group consisting of hydrogen or C1-C2 alkyl optionally substituted with -OH.
[0462] In some embodiments, an alkylated nucleobase is selected from the group consisting of:
[0463] In some embodiments, the linker is attached to a different position of the base than the alkylation.
[0464] In some embodiments, the linker is specifically attached to an amino acid of the polymerase. In these cases, it is preferable to attach the linker to an amino acid at a position that can be mutated without loss of the polymerase activity, e.g., positions 180, 188, 253 or 302 of murine TdT (numbering as in the crystal structure PDB ID: 4127). It is preferable to not attach the linker to an amino acid involved in the catalytic activity of the polymerase to avoid interfering with catalysis. Residues known to be involved with catalysis and methods for determining if a residue is involved with catalysis (e.g., by site-specific mutagenesis) will be apparent to those skilled in the art and are reviewed in literature (e.g., Joyce et al. (Journal of Bacteriology 177.22 (1995): 6321.) and Jara and Martinez (The Journal of Physical Chemistry B 120.27 (2016): 6504-6514.))
[0465] In some embodiments, a linker of a conjugate may be attached to the 7-position of deaza dGTP or the 5-position of dTTP or dUTP.
[0466] In some embodiments, the linker of the conjugate is cleaved to leave a scar. In such cases, scars may remain on the DNA after cleavage. In some embodiments, a scar comprises a hydroxyl. In some embodiments, a scar comprises an amine. In some embodiments, a scar comprises a hydroxylalkyl group.
[0467] In some embodiments, the linker of a conjugate is attached to an O-alkylated nucleobase at the alkyl group. In particular embodiments, the linker of a conjugate attached to an O-alkylated nucleobase is cleaved to leave a scar. In some embodiments, a scar is removed using an enzyme capable of removing said one or more alkyl groups from an alkylated nucleobase.
[0468] In some embodiments, a conjugate is represented by a structure of Formula (I) or (II):wherein L1is selected from the group consisting of an optionally substituted C1-6 alkylene chain, an optionally substituted C2-6alkenylene chain, and an optionally substituted C1-6alkynylene chain, wherein 1-4 methylene units are optionally and independently replaced with -O-, - N(Ra)-, -C(O)-, -S-, -S(O)-, -S(O)2-, or phenylene; L2is a cleavable linker; X is -N= or -C(H)=; each Rais independently hydrogen or C1-6 alkyl; R2is hydrogen or methyl; R is a ribose polyphosphate or deoxyribose polyphosphate; and Pol is a polymerase.
[0469] In some embodiments, the conjugate is represented by.
[0470] In some embodiments, L1is selected from the group consisting of.
[0471] In some embodiments, the conjugate is selected from the group consisting of
[0472] In some embodiments, a ribose polyphosphate is selected from the group consisting of ribose triphosphate, ribose tetraphosphate, ribose pentaphosphate, and ribose hexaphosphate. In some embodiments, a ribose polyphosphate is a ribose triphosphate. In some embodiments, a ribose polyphosphate is a ribose hexaphosphate. In some embodiments, ribose phosphate is a pentaphosphate. In some embodiments, ribose polyphosphate is a ribose tetraphosphate.
[0473] In some embodiments, the present disclosure includes a method of treating a polynucleotide synthesized with alkylated nucleobases, comprising: providing a polynucleotide comprising one or more alkylated nucleobases; and removing one or more alkyl groups from said one or more alkylated nucleobases.
[0474] In some embodiments, the present disclosure includes a method of synthesizing a polynucleotide comprising an alkylated nucleobase, comprising: contacting a precursor polynucleotide with a polymerase and a nucleotide comprising said alkylated nucleobase; adding said nucleotide to the 3' end of said precursor polynucleotide via said polymerase.
[0475] In some embodiments, the present disclosure includes a method of synthesizing a polynucleotide comprising contacting a precursor polynucleotide with a conjugate comprising a nucleotide covalently linked to a polymerase via a cleavable linker, wherein said nucleotide comprises said alkylated nucleobase. In some embodiments, the method of synthesizing a polynucleotide comprises cleaving a cleavable linker after addition of a nucleotide to a precursor polynucleotide. In some embodiments, the method of synthesizinga polynucleotide comprises repeating contacting, adding, and optionally cleaving steps described herein one or more times. In some embodiments, removal of one or more alkyl groups described herein comprises contacting said polynucleotide with an enzyme capable of removing said one or more alkyl groups from said alkylated nucleobases. In some embodiments, a method of synthesizing a polypeptide comprising contacting said polynucleotide with two or more enzymes capable of removing said one or more alkyl groups from said alkylated nucleobases.
[0476] In some embodiments, synthesis of a polynucleotide comprises adding nucleotides stepwise to a starter molecule (e.g., an initial oligonucleotide) via the cycled steps of: addition of polymerase-nucleotide conjugate to an oligonucleotide, binding of the nucleotide to the 3’ end of the oligonucleotide catalyzed by the polymerase, and cleavage of the polymerase from the added nucleotide. These steps can be repeated until a desired polynucleotide is synthesized. As described herein, the use of nucleotides comprising alkylated nucleobases during polynucleotide synthesis helps to improve the efficiency and accuracy of synthesis by inhibiting secondary structure formation which can interfere with the addition of the incoming nucleotide by the polymerase during synthesis.
[0477] Although synthesis can be completed entirely with alkylated nucleotides, synthesis with a combination of unmodified and alkylated nucleotides can also be used effectively to improve polynucleotide synthesis. In some embodiments, only one of the four nucleotides added (e.g., from G or T) is alkylated during synthesis. In some embodiments, alkylated nucleotides are only added at targeted positions where secondary structure or ternary structure is predicted, which could interfere with synthesis. Such structures can be predicted based on the presence of complementary DNA regions in various ways and respective tools exist, such as the NUPACK algorithms (http: / / www.nupack.org / home / model). Thus, in some embodiments, synthesis of a completed polynucleotide where synthesis is improved can include the use of only 1 or 2 alkylated nucleotides. In some embodiments, about 5%, about 10%, about 20%, about 30%, about 50%, substantially all, or 100% of a specific nucleotide is incorporated into the polynucleotide in their alkylated version. In some embodiments, less than 5%, less than 10%, less than 20%, less than 30%, or less than 50% of a specific nucleotide is incorporated into the polynucleotide in its alkylated version. In some embodiments, more than 5%, more than 10%, more than 20%, more than 30%, or more than 50% of a specific nucleotide is incorporated into the polynucleotide in its alkylated version. In some embodiments, only alkylated guanine nucleotides are used in the nucleotide synthesis reaction. The removal of alkylations in the terminal positions of a nucleic acid withalkyl transferases may be more challenging than the removal from internal DNA positions. Therefore, in some embodiments, nucleotide synthesis is performed such that the last and first 1, 2, or 3 positions of the synthesized nucleic acid does not comprise alkylated nucleotides.
[0478] In some embodiments, the nucleotides analogs described herein comprise a reversible terminator group, such as such as an O- azidomethyl or O-NH2 group on the 3' position of the sugar or an (alpha-tertbutyl-2- nitrobenzyl)oxymethl group on the 5 position of pyrimidines or the 7 position of 7- deazapurines (for an overview see, e.g. Chen et al., Genomics, Proteomics & Bioinformatics 2013 11: 34-40). In these embodiments, the nucleotide analog prevents or hinders further elongation once incorporated into a nucleic acid to achieve controlled termination of synthesis. In some embodiments, when used as part of a conjugate, the RTdNTP- polymerase conjugates do not rely on the shielding effect to achieve termination, e.g., when a 3' modified RTdNTP is tethered to the polymerase, the linker used may exceed 100 A or 200 A in length.
[0479] O6-alkylguanine-DNA alkyltransferase (A GT) irreversibly transfers an alkyl group from its substrate, a modified nucleotide. Described herein are alkyl modified nucleotides useful during synthesis, where the alkyl group can be removed by AGT. In some embodiments, this removal converts a ‘scarred’ nucleotide to a naturally occurring nucleotide or nucleobase in a synthesized oligonucleotide. Several forms of the enzyme can be used considered provided they have similar properties in reacting with an alkyl group substrate (such as human, murine, rat, a chimera, or other species of AGT). The As used herein, the alkyl group refers to any group that can act as a substrate and is removed from the nucleotide by an AGT enzyme.
[0480] In the present disclosure, O6-alkylguanine-DNA alkyltransferase also includes variants of a wild-type AGT which may differ by virtue of one or more amino acid substitutions, deletions, or additions, but which still retain the property of transferring a label present on a substrate to the AGT part of the fusion protein. AGT variants may be obtained by chemical modification using techniques well known to those skilled in the art. AGT variants may preferably be produced using protein engineering techniques known to the skilled person and / or using molecular evolution to generate and select new O6-aIkylguanine- DNA alkyltransferases. Such techniques are e.g., saturation mutagenesis, error prone PCR to introduce variations anywhere in the sequence, DNA shuffling used after saturation mutagenesis and / or error prone PCR, or family shuffling using genes from several species.
[0481] In some embodiments, an alkylated nucleobase is a nucleobase of formula (I) or formula (II):whereinR1is selected from the group consisting of Ci-6 alkyl, C2-6 alkenyl, C1-6 alkynyl, and -(CH2)o- 3PI1, wherein R1is optionally substituted with 1-6 instances of Rla; each Rlais independently selected from halogen, C1-6 alkyl, -(CH2)o-30Rlb, -NO2, -N3, - OPO2OH, and -(CH2)o-3NHRlb; and each Rlbis independently selected from hydrogen, C1-6 alkyl, -C(O)(Ci-6 alkyl), C1-6 haloalkyl, -C(O)(Ci-6haloalkyl), and -CthOAc;R2is hydrogen or methyl;X is -N= or -C(H)=; andR is a ribose polyphosphate or deoxyribose polyphosphate.
[0482] In some embodiments, X is -C(H)= or -N=. In some embodiments, X is -C(H)=. In some embodiments, X is -N=.
[0483] In some embodiments, R1is selected from the group consisting of C1-6 alkyl, C2-6 alkenyl, C1-6 alkynyl, and -(CH2)o-3Ph, wherein R1is optionally substituted with 1-6 instances of Rla. In some embodiments, R1is selected from the group consisting of C2-6 alkyl, C2-6 alkenyl, C2-6 alkynyl, and -(CH2)o-3Ph, wherein R1is optionally substituted with 1- 6 instances of Ra. In some embodiments, R1is selected from the group consisting of C1-3 alkyl, C2-4 alkenyl, C2-4 alkynyl, and -CH2PI1, wherein R1is optionally substituted with one instance of Ra. In some embodiments, R1is selected from the group consisting of methyl, ethyl, C4 alkenyl, C4 alkynyl, and -CH2PI1, wherein R1is optionally substituted with one instance of Ra. In some embodiments, R1is selected from the group consisting of ethyl, C4 alkenyl, C4 alkynyl, and -CH2PI1, wherein R1is optionally substituted with one instance of Ra. In some embodiments, R1is selected from the group consisting of C1-C3 alkyl, whereinR1is optionally substituted with 1-3 instances of Ra. In some embodiments, R1is selected from the group consisting of C1-C3 alkyl, wherein R1is optionally substituted with 1-3 instances of Ra.
[0484] In some embodiments, R1is selected from the group consisting of
[0485] In some embodiments, R1is selected from the group consisting of methyl, ethyl, and n-propyl optionally substituted with 1-3 instances of Ra.
[0486] In some embodiments, R1is methyl.
[0487] In some embodiments, the nucleotide is selected from the group consisting ofChemically Removeable Protecting Groups
[0488] In some embodiments, suitable conditions for removal of a protecting group from a protected nucleobase include exposing a protected nucleobase to light. In some embodiments, light is ultraviolet light. In some embodiments, ultraviolet light has a wavelength of about 365 nm.
[0489] In some embodiments, suitable conditions for removal of a protecting group include treating a protected nucleobase with an acidic oxidizing solution. In some embodiments, suitable conditions for removal of a protecting group include treating a protected nucleobase with a solution of nitrous acid.
[0490] In some embodiments, suitable conditions for removal of a protecting group include treating a protected nucleobase with a reducing agent. In some embodiments, a reducing agent is a phosphorous-based reducing agent. In some embodiments, a phosphorous-based reducing agent is PPI13 or TCEP. In some embodiments, a phosphorous- based reducing agent is TCEP.
[0491] In some embodiments, suitable conditions for removal of a protecting group comprises the step of treating a protected nucleobase with an oxidizing agent. In some embodiments, an oxidizing agent is capable of oxidizing a thioether (i.e., sulfide) to a sulfone. In some embodiments, an oxidizing agent is selected from the group consisting of H2O2, mCPBA, and TAPC. In some embodiments, conditions for removal of a protecting group further comprise an alkaline (or high pH) environment. In some embodiments the alkaline environment comprises a pH of about 8. In some embodiments, the alkaline environment is one that has a pH higher than the pH at polynucleotide synthesis.
[0492] In some embodiments, suitable conditions for removal of a protecting group comprises the step of treating a protected nucleobase with a base. In some embodiments, a base is selected from the group consisting of NH4OH, KOH, NaOH, KOMe, NaOMe, and KOtBu.
[0493] In some embodiments, suitable conditions for removal of a protecting group comprises the step of treating a protected nucleobase with conditions to remove an allyl group. In some embodiments, suitable conditions for removal of a protecting group comprises the step of treating a protected nucleobase with Pd. In some embodiments, suitable conditions for removal of a protecting group comprises the step of treating a protected nucleobase with Pd(OAc), Pd2(dba)3, and Pd2(pmdba)3.
[0494] Described herein are removeable protecting group or scar structures that can be chemically removed from a nitrogen or oxygen of a nucleobase, leaving the natural nucleotide. For a polymerase-nucleotide conjugate, a linker connecting the nucleotide to the polymerase is attached to the nucleotide at a nitrogen or oxygen on the nucleobase, such that cleavage of the linker leaves the removeable protecting group on the nitrogen or oxygen. In some embodiments, the protecting group / scar structures described herein remain attached to the nitrogen or oxygen on the nucleobase during synthesis but are removed before downstream processing or use of the newly synthesized oligonucleotide. In some embodiments, a nucleotide comprising a removeable protecting group structure described herein bound to a nitrogen or oxygen atom of the nucleobase (with a structure corresponding to a removeable scar after cleavage of a conjugate linker) can be used for oligonucleotide synthesis.
[0495] In some embodiments, the polymerase-nucleotide conjugates comprise a linker attached to the nucleobase represented by the following structure:where Y is a nucleobase; L is a linker attached to a base-pairing nitrogen or oxygen of the nucleobase; R is a ribose polyphosphate or deoxyribose polyphosphate; and Pol is a polymerase.
[0496] The linker L can be represented by the formula Z-L1-L2-L3, where Z-L '-L2represents a portion of the linker that remains after cleavage of the linker to separate thepolymerase from the nucleotide, and L3is a linker that can be cleaved from L2, and which is attached to the polymerase.
[0497] Also described herein are scarred nucleotides and protected nucleotides. Scarred nucleotides are nucleotides comprising the remainder of a linker, i.e., a “scar” after cleavage of the linker to separate the polymerase from the nucleotide. Protected nucleotides are nucleotides that comprise a protecting group bound to a base pairing oxygen or nitrogen of the nucleobase. The protected nucleotides have a protecting group structure that corresponds with the structure of the scarred nucleotides.
[0498] In some embodiments, the scarred nucleotides or protected nucleotides have a nucleobase represented by the followings structure:where Y is a nucleobase; L-R1is a protecting group or scar; L is a moiety attached to a basepairing nitrogen or oxygen of the nucleobase; and R1is selected from the group consisting of hydrogen, -OH, - N(Rb)2, and -SH, wherein each Rbis independently hydrogen or optionally substituted Ci-6 alkyl. L for protected nucleotides / scarred nucleotides can be represented by the formula Z-L'-L2, which corresponds to the same Z-L'-L2formula in the conjugate structure. Exemplary removeable scars / protecting groups comprising Z-L1-L2-R1and conditions for removal from the nucleotide are also described herein. In some embodiments, removal is achieved by exposure of a photocleavable scar or protecting group to the corresponding wavelength of light. In some embodiments, removal is achieved by an appropriate set of chemical conditions, such as a basic environment.
[0499] In some embodiments, Z is selected from the group consisting of a bond, -C(O)-, - C(O)CH2-, -C(O)C(RL)2-, -C(O)CH(RL)-, -C(O)O-, and -C(O)N(H)-;L1is selected from the group consisting of a bond,wherein L1is optionally substituted with 1-4 instances of RL; each RLis independently selected from the group consisting of halogen, hydroxyl, oxo, and optionally substituted Ci- C3 alkyl, wherein 2 instances of R1are optionally taken together with the intervening atom(s)to form a 3-6 membered carbocyclyl ring; L2is selected from the group consisting of a bond, an optionally substituted C1-12 alkylene chain, C4-C20 polyethylene glycol, an optionally substituted C2-12 alkenylene chain, and an optionally substituted C2-12 alkynylene chain, wherein 1-6 methylene units are optionally and independently replaced with -O-, -N(Rb)-, - C(O)-, -S-, -S(O)-, -S(O)2-, or phenylene; W is selected from the group consisting of -O-, -S-, -S(O)2-, and -N(Rb)-; each Rais halogen, -Me, or -OMe; each Rbis independently hydrogen or optionally substituted C1-6 alkyl; and n is 1 or 2; wherein L1and Z are not both a bond.O-Linked Scars / Protecting Groups
[0500] In some embodiments, a conjugate comprising a linker with a removable scar is bound to an oxygen on the nucleobase of the nucleotide. In some embodiments, a scar or a protecting group is bound to an oxygen on the nucleobase of the nucleotide. In some embodiments, the oxygen is a base pairing oxygen on the nucleobase.
[0501] In some embodiments, the linker, scar or protecting group is bound to the 04 oxygen of uracil or adenine, or to the 06 oxygen of guanine. Examples of chemically removeable scars or protecting groups are described below:
[0502] In some embodiments, a sulfonyl group is used in a protecting group or scar that can be removed from the nucleobase. An O-linked sulfonyl group can be removed by a suitable base, such as NH4OH, in a beta-elimination reaction.
[0503] In some embodiments, a linker of a conjugate attached to an oxygen or nitrogen on the nucleobase comprises a sulfonyl group. In some embodiments, a scar remaining after cleavage of the linker comprises a sulfonyl group. In some embodiments, a protected nucleotide comprises an O-linked protecting group comprising a sulfonyl group. In some embodiments, a scar or protecting group comprising a sulfonyl group is removed from the nucleotide upon exposure to an appropriate base. In some embodiments, a suitable base is a strong base. In some embodiments, a suitable base is a hydroxide with a suitable counterion. In some embodiments, a suitable base is selected from the group consisting of NaOH, NH4OH, KOH, KOtBu, NaOMe, and KOMe.
[0504] In some embodiments,
[0505] In some embodiments,
[0506] Illustrated below are exemplary Z1-L1-L2 structures in conjugates with a linker comprising a sulfonyl group. These Z1-L1-L2 structures can also be used for scarred nucleotides (after linker cleavage) or protected nucleotides for synthesis. Exemplary conjugates comprising a sulfonyl group attached to an oxygen on the nucleobase (that is retained as a scar on the nucleotide after cleavage of the linker) and can be removed are shown below:
[0507] Further Z-L1-L2 structures in exemplary nucleobases representative of scarred or protected nucleobases are shown below:
[0508] The present disclosure includes a method of preparing a polynucleotide comprising treating a scarred nucleobase comprising a sulfonyl group with a suitable base. Without being bound by any particular theory, treatment of a scarred nucleobase comprising a sulfonyl group with a suitable base results in beta-elimination:
[0509] In some embodiments, a conjugate or nucleotide comprising a sulfonyl group can be prepared as outlined in Scheme 1 :
[0510] In some embodiments, a thioether group is used in a protecting group or scar that can be removed from the nucleobase. A thioether group can be removed by exposure to a suitable nucleophile.
[0511] In some embodiments, a linker of a conjugate attached to an oxygen or nitrogen on the nucleobase comprises a thioether. In some embodiments, a scar remaining after cleavage of the linker comprises a thioether. In some embodiments, a protected nucleotide comprises an O-linked or N-linked protecting group comprising a thioether. In some embodiments, a scar or protecting group comprising a thioether is removed upon exposure to a suitable oxidant followed by exposure to a suitable base.
[0513] In some embodiments,
[0514] Illustrated below are exemplary Z1-L1-L2 structures in conjugates with a linker comprising a nitrobenzyl photocleavable group. These Z1-L1-L2 structures can also be used for scarred nucleotides (after linker cleavage) or protected nucleotides for synthesis.Exemplary conjugates comprising a photocleavable group attached to an oxygen on the nucleobase (that is retained as a scar on the nucleotide after cleavage of the linker) and can be removed are shown below:
[0515] Further Z-L1-L2 structures in exemplary nucleobases representative of scarred or protected nucleobases are shown below:
[0516] In some embodiments, the present disclosure includes a method of preparing a polynucleotide comprising removal of a scar or protecting group comprising a thioether group and attached to a nucleobase via oxidation to generate a sulfonyl group and treatment with a suitable base to remove the sulfonyl group.
[0517] For example, removal of a sulfonyl group bound to an oxygen of the nucleobase can be accomplished through exposure to an oxidant followed by exposure to a base as follows:Oxidation of Elimination of sulfone sulfide to sulfone to native DNA
[0518] In some embodiments, a conjugate or nucleotide comprising a thioethyl can be prepared as outlined in Scheme 2Scheme 2
[0519] In some embodiments, a cyanoethyl group is used in a protecting group or scar that can be removed from the nucleobase. An O-linked cyanoethyl group can be removed by a suitable base, such as NH4OH.
[0520] In some embodiments, a linker comprises a cyanoethyl group. In some embodiments, a scar comprises a cyanoethyl group. In some embodiments, a scar comprises a cyanoethyl group that is removed upon exposure to an appropriate base. In some embodiments, a suitable base is a strong base. In some embodiments, a suitable base is a hydroxide with a suitable counterion. In some embodiments, a suitable base is selected from the group consisting of NaOH, NH4OH, KOH, KOtBu, NaOMe, and KOMe.
[0521] In some embodiments,
[0522] In some embodiments, a nucleotide is selected from the group consisting of
[0523] Illustrated below are exemplary removeable Z1-L1-L2 structures attached to an oxygen of a protected nucleobase with a linker comprising a cyanoethyl group. These Zl- L1-L2 structures can be used for conjugates that leave scarred nucleotides (after linker cleavage) or protected nucleotides for synthesis:
[0524] In some embodiments, a protected or scarred nucleobase is selected from the group consisting of:
[0525] The present disclosure includes a method of preparing a polynucleotide comprising treating a scarred nucleobase comprising a cyanoethyl group with a suitable base. Without being bound by any particular theory, treatment of a scarred nucleobase comprising a cyanoethyl group with a suitable base results in beta-elimination:
[0526] In some embodiments, a conjugate or nucleotide comprising a cyanoethyl group can be prepared as outlined in Scheme 3:Scheme 3
[0527] In some embodiments, an allyl group is used in a protecting group or scar that can be removed from the nucleobase. In some embodiments, a linker comprises an allyl group. In some embodiments, a scar comprises an allyl group. In some embodiments, a scar comprises an allyl group that is removed upon exposure to an appropriate transition metal catalyst. In some embodiments, an appropriate transition metal catalyst is a palladium catalyst. In some embodiments, an appropriate transition metal catalyst is selected from the group consisting of Pd2(dba)3, Pd2pmdba)3, PdCl2, Pd(TFA)2, and Na2PdCl4.
[0528] In some embodiments,
[0529] In some embodiments,
[0530] Illustrated below are exemplary removeable Z1-L1-L2 structures attached to an oxygen of a protected nucleobase with a linker comprising a cyanoethyl group. These Zl-L1-L2 structures can be used for conjugates that leave scarred nucleotides (after linker cleavage) or protected nucleotides for synthesis:
[0531] The present disclosure includes a method of preparing a polynucleotide comprising treating a scarred nucleobase comprising an allyl group with a transition metal catalyst and, optionally, one or more suitable ligand. In some embodiments, a suitable ligand is P(PhSO3Na)3. Without being bound by any particular theory, treatment of a scarred nucleobase comprising an allyl group with a suitable transition metal catalyst results in catalytic deallylation:
[0532] In some embodiments, a conjugate or nucleotide comprising an allyl group can be prepared as outlined in Scheme 4:Scheme 4
[0533] In some embodiments, an azide is used in a protecting group or scar that can be removed from the nucleobase. In some embodiments, a linker comprises an azide. In some embodiments, a scar comprises an azide. In some embodiments, a scar comprises an azide that is removed upon exposure to a suitable reductant. In some embodiments, a suitable reductant is a phosphine. In some embodiments, a suitable reductant is TCEP.
[0535] Illustrated below are exemplary Z1-L1-L2 structures in conjugates with a linker comprising an azide. These Z1-L1-L2 structures can also be used for scarred nucleotides (after linker cleavage) or protected nucleotides for synthesis. Exemplary conjugatescomprising a photocleavable group attached to an oxygen on the nucleobase (that is retained as a scar on the nucleotide after cleavage of the linker) and can be removed are shown below:
[0536] Further Z-L1-L2 structures in exemplary nucleobases representative of scarred or protected nucleobases are shown below:
[0537] The present disclosure includes a method of preparing a polynucleotide comprising treating a scarred nucleobase comprising an azidyl group with a suitable reductant. Without being bound by any particular theory, treatment of a scarred nucleobase comprising an azidyl group with a suitable reductant results in removal of the scar:
[0538] In some embodiments, an oxime is used in a protecting group or scar that can be removed from the nucleobase. In some embodiments, a linker comprises an oxime. In some embodiments, a scar comprises an oxime. In some embodiments, a scar comprises an oxime that is removed upon exposure to suitable nucleophile and subsequently to a suitable acid. In some embodiments, a suitable nucleophile is HoNOtBu. In some embodiments, a suitable acid is a HONO.
[0540] Illustrated below are exemplary Z1-L1-L2 structures in conjugates with a linker comprising an oxime. These Z1-L1-L2 structures can also be used for scarred nucleotides (after linker cleavage) or protected nucleotides for synthesis. Exemplary conjugates comprising an oxime attached to an oxygen on the nucleobase (that is retained as a scar on the nucleotide after cleavage of the linker) and can be removed are shown below:
[0541] Further Z-L1-L2 structures in exemplary nucleobases representative of scarred or protected nucleobases are shown below:
[0542] The present disclosure includes a method of preparing a polynucleotide comprising treating a scarred nucleobase comprising an oximyl group with a suitable nucleophile and, subsequently a suitable acid. In Without being bound by any particular theory, removal of an oximyl group can be accomplished as shown below:
[0543] In some embodiments, a silyl group is used in a protecting group or scar that can be removed from the nucleobase. In some embodiments, a linker comprises a silyl group. In some embodiments, a scar comprises a trimethylsilyl group. In some embodiments, a linker comprises a silyl group. In some embodiments, a scar comprises a trimetylsilyl group. In some embodiments, a scar comprises a silyl group that is removed upon exposure to a halogen source. In some embodiments, a scar comprises a silyl group that is removed upon exposure to ZnBr2. In some embodiments, a scar comprises a silyl group that is removed upon exposure to suitable fluoride source. In some embodiments, a suitable fluoride source is selected from the group consisting of KF, TBAF, and TBAT.N-Linked Scars / Protecting Groups
[0544] In some embodiments, a conjugate comprising a linker with a removable scar is bound to a nitrogen on the nucleobase of the nucleotide. In some embodiments, a scar or a protecting group is bound to a nitrogen on the nucleobase of the nucleotide. In some embodiments, the nitrogen is a base pairing nitrogen on the nucleobase.
[0545] In some embodiments, the linker, scar or protecting group is bound to the N1 or N2 nitrogen of guanine, the N4 nitrogen of cytosine, or the N6 nitrogen of adenine. Examples of chemically removeable scars or protecting groups are described below:
[0546] In some embodiments, Z is carbamate, which is attached to a nitrogen of the nucleobase. When Z is carbamate, an O-linked protecting group or scar structure described above (-L1 or -L1-L2) can be attached to carbamate. Exposure to conditions for removal ofthese protecting groups or scars from carbamate as described above also results in removal of the carbamate.
[0547] O-linked removeable scars / protecting groups described above can also be attached to a nitrogen on the nucleobase by using an intermediate structure (e.g., Z), such as a carbamate. For example, a carbamate moiety can be attached to a nitrogen on the nucleobase, and the groups / moieties (-L1 or -L1-L2) described above can be attached to the oxygen of the carbamate. Under conditions suitable for removal of these protecting groups or scars from the carbamate also results in release of the carbamate from the nitrogen of the nucleobase.
[0548] Exemplary corresponding structures for removeable moieties discussed above bound to a carbamate attached to a nitrogen of a nucleobase are illustrated below:
[0549] Illustrated below are exemplary Z1-L1-L2 structures in conjugates with a linker comprising a sulfonyl group bound to a nitrogen atom via a carbamate group. These Zl-Ll- L2 structures can also be used for scarred nucleotides (after linker cleavage) or protected nucleotides for synthesis. Exemplary conjugates comprising a sulfonyl group attached to a nitrogen (via a carbamate) on the nucleobase (that is retained as a scar on the nucleotide after cleavage of the linker) and can be removed are shown below:
[0550] After cleavage of the linker, removal of the remaining scar can be achieved by exposing the scarred nucleotide to the appropriate base.
[0551] An exemplary synthesis of an N-linked carbamate sulfone nucleotide is shown inScheme 5:
[0552] Without wishing to be bound by theory, an exemplary reaction mechanism for removal of the above described scars and protecting groups is by beta-elimination, as shown below:
[0553] A similar mechanism applies to protecting groups or scars attached to a carbamate that have an electron withdrawing group:
[0554] For example, a cyanoethyl scar / protecting group attached to a carbamate:Thus, suitable scars or protecting groups include N-linked carbamates bound to L1-L2 structures having an electron withdrawing group.
[0555] Illustrated below are exemplary Z1-L1-L2 structures in conjugates with a linker comprising a thioether group bound to a nitrogen atom via a carbamate group. These Zl-Ll- L2 structures can also be used for scarred nucleotides (after linker cleavage) or protected nucleotides for synthesis. Exemplary conjugates comprising a thioether group attached to a nitrogen (via a carbamate) on the nucleobase (that is retained as a scar on the nucleotide after cleavage of the linker) and can be removed are shown below:
[0556] After cleavage of the linker, removal of the remaining scar can be achieved by exposing the scarred nucleoeide to an oxidant, followed by exposure to basic conditions.
[0557] An exemplary synthesis of an N-linked carbamate thioether nucleotide is shown in Scheme 6:
[0558] In some embodiments, a linker comprises an acyl group linked to a nitrogen of a nucleobase. In some embodiments, a scar or protecting group comprises an acyl group linked to a nitrogen of a nucleobase. In some embodiments, for the structure Z-L1-L2, Z is an acyl group. In some embodiments, the scar or protecting group comprising an acyl group is removed upon exposure to an appropriate nucleophile or base. In some embodiments, a scar comprises a group that is removed upon exposure to an appropriate nucleophile or base. In some embodiments, the acyl group is a carbamate or amide. An exemplary reaction mechanism for removing a scar or protecting group comprising an acyl group (carbamate when X is O and amide with X is C) with a nucleophile is shown below:
[0559] An exemplary reaction mechanism for removing a scar or protecting group comprising an acyl group (carbamate when X is O and amide with X is C) with a base is shown below:
[0560] Thus, in some embodiments, when Z is an amide group or carbamate, a large number of suitable L1-L2 structures can be used that are compatible with removal of the scar or leaving group from the nitrogen of the nucleotide via exposure to a nucleophile or base. Intramolecular Cyclization
[0561] In some embodiments, a linker comprises a carbamate group. In some embodiments, a scar comprises a carbamate group. In some embodiments, a scar comprises a carbamate group that is removed upon exposure to an appropriate base. In some embodiments, a scar comprises a group that is removed upon exposure to an appropriate base. In some embodiments, a suitable base is a strong base. In some embodiments, a suitable base is a hydroxide with a suitable counterion. In some embodiments, a suitable base is selected from the group consisting of NaOH, NH4OH, KOH, KOtBu, NaOMe, and KOMe.
[0562] In some embodiments, -Z-L1-L2-R1iswhere X is O, CH2, or NH; n is 1 or 2; and R1is hydrogen, -OH, - N(Rb)2, and -SH, wherein each Rbis independently hydrogen or optionally substituted C1-6alkyl; and wherein the hydrogen atoms bound to the carbon atoms between the carbonyl and the R1 group are each optionally substituted with a C1-3 alkyl group.
[0563] In some embodiments, -Z-L1-L2-R1is selected from the group consisting of,
[0564] Without being bound by any particular theory, the results from Example 11 suggest that exposure of a scarred or protected nucleobase comprising a carbamate or amide group linked to an alkyl group of a preferred length results in the cyclization and subsequent removal of the scar, which is more preferred than the nucleophilic or base removal mechanism described above.
[0565] A reaction mechanism for intramolecular cyclization of an N-linked carbamate-L1-Thorpe-Ingold effect
[0566] Without being bound by any particular theory, the results from Example 12 suggest that removal of a scar via intramolecular cyclization can be improved by exploiting the Thorpe-Ingold Effect. For example, the introduction of substituent groups on L1can increase the rate of removal of a scar. For example, a scar comprising single methyl substituent may be removed faster than an unsubstituted scar. Exemplary intramolecular cyclization reactions for removal of a scar or protecting group comprising an unsubstituted and substituted alkyl group are shown in FIG. 5 A (alkyl group attached to an amide group linked to a nitrogen on the nucleobase), and FIG. 5B (alkyl group attached to a carbamate group linked to a nitrogen on the nucleobase. As shown in Example 12, Scar elimination by cyclization can be further accelerated by gem-dimethyl substitution (Thorpe-Ingold Effect). Therefore, addition and removal of methyl substitutions impacts the rate of the reaction, and can be changed depending on the desired stability and efficiency of removal needed.
[0567] In some embodiments, to improve the removal efficiency of the scar or protected nucleobase, alkyl substituents occur on carbon groups at any position between the carbonyl group and the R1group, except for at the carbon next to R1.
[0568] In some embodiments, -Z-L1-L2-R1is selected from the group consisting of
[0569] In some embodiments, a conjugate is selected from the group consisting of
[0570] In some embodiments, a scarred nucleobase is selected from the group consisting ofSynthetic Schemes
[0571] In some embodiments, a conjugate or nucleotide comprising carbamate or amide (Z) bound to a nitrogen of a nucleotide and linked to an L1-L2-R1 or an L1-L2-L3 can be prepared as outlined below:Scheme 9: Synthesis ofN4 Carbamate Ethyl C, N4 Carbamate (Methyl) Ethyl C, and N4Carbamate (Dimethyl) Ethyl C. y p32 R, R’= H33 R = H, R’= Me34 R, R’= Me,
[0572] General protocol for carbamate formation: 3 ',5'- diTBS protected nucleosides (dA and dC) (1 eq) was dried by repeated co-ev aporation with pyridine, toluene and thendissolved in 1,2-dichloroethane (IM solution). CDI (1.6 eq) was added to the solution and the mixture was stirred for 20 h) at reflux. Alcohol (2 eq) dissolved in 1,2-dichloroethane was added and then the solution was stirred for 20 h ((3h-5h for C-nucleoside) at reflux. The solution was cooled to room temperature and washed with a saturated aqueous solution of NaHCO3. The aqueous solution was then backextracted with CH2C12, and all the collected organic solution was dried over Na2SO4, filtered, and concentrated. The crude was purified by silica gel column chromatography using a 0-50% EtOAC / hexane gradient to afford carbamates 2 and 7.
[0573] General protocol for Benzyl group deprotection (Compound 23 / 24 / 25): 10% Pd / C (10% w / w) was added to a solution of compound 20 / 21 / 22 (1 mmol) in MeOH (3 mL). The inside air was replaced with th balloon by three vacuum / th cycles. The reaction mixture was stirred at room temperature until the thin layer chromatography monitoring indicated the complete consumption of the starting material. Then, the reaction mixture was passed through a Celite pad or membrane filter using EtOAc to remove Pd / C. The filtrate was concentrated in vacuo to give deprotected product.
[0574] General protocol for esterification reaction (Compound 3 / 8 / 26 / 27 / 28): To a solution of the compound 2 / 7 (1 equiv) in dry DCM (0.2 M) was added DMAP (0.6 eq) and the Fmoc amino acid (1.2 eq) under inert atmosphere at ambient temperature. The reaction mixture was cooled to 0° and then EDC’HCl (1.2 eq) was added slowly. The reaction mixture was allowed to stir for 16 hr at ambient temperature. The reaction solution was then diluted with additional DCM and extracted with water. The aqueous layer was then washed with more DCM (2x). The combined organic layer was then washed with sat. ammonium bicarb, dried over Na2SO4, and concentrated. The crude material was purified by chromatography on a silica gel column using a 10-45% EtOAC / hexane gradient.
[0575] General procedure for TBS deprotection (compound 4 / 9 / 14 / 15 / 16 / 29 / 30 / 31): To a 0.1 M solution of protected nucleosides (3 / 8) in dry THF at 0 °C, was added 3HF-TEA (10 equiv) dropwise. The mixture was stirred for 16 hr while warming to ambient temperature. The reaction was quenched with the addition of a few drops of MeOH, and the solvents were removed under reduced pressure. The residue was diluted with DCM and washed with water. The aqueous layer was again extracted with DCM (lx). The combined organic layer was dried over anhydrous Na2SO4, filtered, and concentrated. The compound was purified by chromatography on a silica gel column using a 1-12% MeOH / DCM gradient. All the compounds were obtained as white foamy solids.
[0576] General protocol for the synthesis of triphosphates (Compound 5 and 10): Nucleoside analogue (4 / 9 / 14 / 15 / 16, 50 mg, 1 eq) was placed into a 10 mL round bottom flask equipped with a stir bar and tetrabutylammonium pyrophosphate was placed in a separate 5 mL conical tube. The two flasks were placed in a vacuum desiccator with P2O5 and allowed to dry under vacuum for at least 16 hr. Additionally, molecular sieves and three small round bottom flasks were placed in a drying oven for at least 16 hr. Two small flasks from the oven were charged with molecular sieves and flame-activated under vacuum. While these were cooling, the other small flask was attached to a Hickman distillation apparatus and flame dried. Upon cooling, the first two flasks were backfilled with nitrogen. Trimethyl phosphate and tributyl amine were then placed over the molecular sieves in the initial two flasks for drying. The Hickman distillation apparatus was then used to freshly distill POCI3. The vacuum desiccator was purged with N2 gas, and the flasks inside were then transferred to nitrogen balloons or the Schlenk line. Trimethyl phosphate (40 eq) was added to the nucleoside and the mixture was cooled at -5 °C. To this nucleoside mixture was added dry tributyl amine (3 eq) followed by POCI3 (2.1 eq) slowly via micro syringe. The combined mixture was stirred at -5 °C. for 45 mins. After 45 min, the reaction mixture was treated with a mixture of tributylamine pyrophosphate (5 eq, 0.5 M in dry acetonitrile) and tributyl amine (6 eq). After 1 hour, the mixture was treated with triethylammonium bicarbonate (0.5 M, 1:2 of the total reaction volume) and allowed to stir at ambient temperature for 1 hour.
[0577] Fmoc Deprotection (compound 5 / 10 / 32 / 33 / 34): Then, this reaction mixture was further treated with N-Methyl piperidine (1 / 5111of total reaction volume) and stirred for 90 mins at ambient temperature followed by extraction with dichloromethane (2X). The aqueous layer was then purified by reverse phase HPLC (0.1 M triethylammonium acetate buffer / Acetonitrile, 4-47%, 0-15 min, flow 5 ml min'1). Product containing fractions were pooled and lyophilized to provide desired product as a triethylammonium salt. The resulting solid was reconstituted in RNase Free DI water for further experiments.
[0578] For Boc deprotection (Compound 17 / 18 / 19): After the reaction quench using triethyl ammonium bicarbonate (0.5 M, 1:2 of the total reaction volume) and allowed to stir at ambient temperature for 1 hour, the reaction mixture was extracted with dichloromethane (2X). The aqueous layer was then purified by reverse phase HPLC (0.1 M triethylammonium acetate buffer / Acetonitrile, 4-47%, 0-15 min, flow 5 ml min'1). The fraction containing triphosphate was pooled and lyophilized to give triphosphate with linker- amine-Boc which was separated into 4 eppendorf tubes. To each eppendorf tube were added MeOH (10 pL) and TFA (20 pL). After vortex, the mixture was kept at RT for 2 min and Et2<D (500 pL) wasadded. After vortexing, the mixture was cooled at -20 ° C for 10 min. After centrifuge (10,000 rpm, 5 min), the supernatant was decanted. The ether wash procedure was repeated twice. The resulting solid was reconstituted in RNase Free DI water for further experiments.
[0579] Provided below are exemplary reaction schemes for preparing a conjugate or nucleotide comprising a removeable scar or protecting group attached to a nitrogen on the nucleobase for each nucleotide:Scheme 11: General synthetic scheme for the access of cleavable amide and carbamate modified nucleotides attached to the N2 positions of G:
[0580] Scheme 12: General synthetic scheme for the access of cleavable amide and carbamate modified nucleotides attached to the N1 position of G:
[0581] Scheme 13: General synthetic scheme for the access of cleavable amide and carbamate modified nucleotides attached to the N6 position of A
[0582] Scheme 14: General synthetic scheme for the access of cleavable amide and carbamate modified nucleotides attached to the N4 position of C
[0583] Scheme 15: General synthetic scheme for the access of cleavable amide and carbamate modified nucleotides attached to the N3 position of U / TPhotocleavable Protecting Groups
[0584] In some embodiments, a photocleavable group is used in a protecting group or scar that can be removed from the nucleobase. A photocleavable group can be removed by exposure to appropriate wavelength(s) of light to leave a native nucleotide. In some embodiments, the wavelength of light is ultraviolet light (uv). In some embodiments, a linker of a conjugate attached to an oxygen or nitrogen on the nucleobase comprises a photocleavable group. In some embodiments, a scar after cleavage of the linker comprises a photocleavable group. In some embodiments, a nucleotide comprises a protecting group comprising a photocleavable group.
[0585] In some embodiments, L1comprises a photocleavable group. In some embodiments, L1comprises an optionally substituted 2-nitrobenzyl group. In some embodiments, L1comprises a 5-methoxy-2-nitrobenzyl group. In some embodiments, L1iswherein, each Rais independently selected from the group consisting of halogen, -Me, and - OMe.
[0586] In some embodiments, L1is
[0587] In some embodiments, a conjugate comprising a nucleotide bound to a photocleavable group is represented by:wherein the linker is attached to the nucleotide at an oxygen or nitrogen of the nucleobase.
[0588] In some embodiments, a protected nucleotide or scarred nucleotide comprising a nucleotide bound to a photocleavable group is represented by:wherein the nitrobenzyl group is attached to the nucleotide at an oxygen or nitrogen of the nucleobase.
[0589] Illustrated below are exemplary Z1-L1-L2 structures in conjugates with a linker comprising a nitrobenzyl photocleavable group. These Z1-L1-L2 structures can also be used for scarred nucleotides (after linker cleavage) or protected nucleotides for synthesis.Exemplary conjugates comprising a photocleavable group attached to an oxygen on the nucleobase (that is retained as a scar on the nucleotide after cleavage of the linker) and can be removed are shown below:
[0590] In some embodiments, a protected or scarred nucleotide comprising a photocleavable group is selected from the group consisting of:.
[0591] In some embodiments, a protected nucleobase comprising a photocleavable protecting group or scar bound to an oxygen on the nucleobase is selected from the group consisting of:wherein R1is selected from the group consisting of hydrogen, -OH, - N(Rb )2, and -SH; and each Rb is independently hydrogen or optionally substituted C1-6 alkyl.
[0592] In some embodiments, a protected nucleobase is selected from the group consisting of:, ,
[0593] In some embodiments, the present disclosure includes a method of preparing a polynucleotide comprising removal of a scar or protecting group attached to a nucleobase via photolysis. In some embodiments, photolysis comprises exposure of a polynucleotide to ultraviolet light. In some embodiments photolysis comprises expose of a polynucleotide to light with a wavelength of about 365 nm.
[0594] For example, removal of a scar can be accomplished through exposure of a nucleobase to light with a wavelength of about 365 nm:
[0595] In some embodiments, a conjugate or nucleotide comprising a photocleavable group can be prepared as outlined in Scheme 16: Scheme 16
[0596] In some embodiments, the photocleavable linker can also be attached to a nitrogen of a nucleobase. An exemplary scheme for preparing an N-linked photocleavable group attached to adenine is outlined in Scheme 17:
[0597] Examples of conjugates comprising N-linked carbamate + photocleavable nitrobenzyl protecting groups are shown below. Upon cleavage of L3, the nucleotides will have an N-linked scar that can be cleaved via photolytic cleavage.
[0598] In some embodiments, a linker of a conjugate may be attached to an oxygen or nitrogen of the nucleobase. In particular embodiments, the linker of a conjugate attached to an oxygen or nitrogen of the nucleobase is cleaved to leave a scar, which can also be a protecting group that inhibits secondary structure formation. In some embodiments, a scar is removed using a chemical or photolytic condition capable of removing said one or more protecting groups from a protected nucleobase. Linker attachment to Polymerase
[0599] In some embodiments, L3 comprises a bioconjugate group suitable for conjugation of L3 to the polymerase.
[0600] In some embodiments, the bioconjugate group is an N-hydroxysuccinimide ester (NHS) group. In some embodiments, the bioconjugate group is a maleimide group. The linker may then be covalently attached to the polymerase by reaction of the maleimide group with a cysteine residue of the polymerase.
[0601] In some embodiments, the polymerase may be operably linked to a linker moiety including a covalent or non-covalent bond; amino acid tag (e.g., poly-amino acid tag, poly- His tag, 6His-tag (SEQ ID NO: 53)); chemical compound (e.g., polyethylene glycol); protein-protein binding pair (e.g., biotin-avidin); affinity coupling; capture probes; or any combination of these. The linker moiety can be separate from or part of a polymerase variant. Phosphatase-treatment of conjugates to reduce insertions
[0602] Among other things, the present disclosure provides polymerase nucleotide conjugates. As is understood to those in the art, there are various challenges associated with precise and accurate polynucleotide synthesis. For example, among other things, polymerases can erroneously catalyze covalent addition of nucleotides, which may result in the addition of more than one nucleotide per step when using polymerase-nucleotide conjugates in a controlled, step-wise nucleic acid synthesis (e.g., insertion or non-termination). Technologies provided herein, including combining such conjugates with phosphatases (e.g., providing a polymerase-nucleotide conjugate in the presence of a phosphatase) overcome such challenges. These technologies help to achieve more accurate and precise stepwise additions,reducing errors as compared to previously described synthesis approaches (e.g., those conducted in the absence of phosphatases).
[0603] In some embodiments, the conjugates are provided in the presence of a phosphatase. In some embodiments, the present disclosure provides a conjugate reagent comprising a plurality of polymerase-nucleotide conjugates, wherein a polymerase and a nucleotide are linked via a linker. In some embodiments, the linker is cleavable.
[0604] In some embodiments, the conjugate reagent exists in the presence of phosphatases.
[0605] In some embodiments, the polymerase-nucleotide conjugates are combined with a template (e.g., a start oligo or initial oligonucleotide) in the presence of a phosphatase.
[0606] A typical process for stepwise synthesis of a polynucleotide comprises adding individual nucleotides step-wise to a starter oligo (i.e., an initial oligonucleotide) via cyclical steps. For example, in some embodiments, the steps comprise: addition of a polymerase- nucleotide conjugate to an oligonucleotide, covalent addition of the nucleotide to the 3′ end of the oligonucleotide catalyzed by the polymerase, and cleavage of the polymerase from the added nucleotide. These steps can be repeated until a desired elongated polynucleotide is synthesized such that the elongated polynucleotide has a length one or more nucleotides longer than the polynucleotide prior to the steps being repeated one or more times.
[0607] Among other things, provided herein are methods of nucleic acid synthesis. In some embodiments, a method of nucleic acid synthesis comprises a step of contacting (e.g., incubating) a conjugate reagent comprising polymerase-nucleotide conjugates (e.g., a plurality of polymerase-nucleotide conjugates) in the presence of one or more phosphatases. In some embodiments, the nucleotides in the plurality are the same nucleotides (e.g., A, G, T, or C, etc.). In some embodiments, the nucleotides are different nucleotides (e.g., A, G, T, and / or C, etc.) In some such embodiments, synthesis conducted in the presence of a phosphatase is improved in one or more ways (e.g., more precise, more efficient, more accurate) as compared to the same synthesis in the absence of a phosphatase.
[0608] In some embodiments, the synthesis performed in the presence of a phosphatase prevents addition of unshielded nucleotides to a nucleic acid. The methods provided herein comprise a step of contacting (e.g., incubating) a conjugate reagent comprising a plurality of polymerase-nucleotide conjugates with a phosphatase, wherein there is a reduction in rates of processes that lead to addition of more than one nucleotide per step when using polymerase- nucleotide conjugates in nucleic acid synthesis (e.g. non-termination leading to an additional nucleotide insertion) as compared to synthesis without phosphatase or without treatment of the conjugate reagent with phosphatase.
[0609] Achieving precisely one nucleotide addition in each step of a stepwise nucleic acid synthesis is essential for producing accurate synthesis of longer oligonucleotides. It remains a challenge in the industry to conduct these additions with precision and accuracy. For example, stepwise nucleic acid synthesis using polymerase-nucleotide conjugates may be susceptible to insertions and / or non-termination resulting in the addition of more than one nucleotide to a nucleic acid in a single step of a cyclic nucleotide extension.
[0610] An unshielded nucleotide is not sterically-hindered or is only partially sterically- hindered by a tethered polymerase from phosphatase cleavage at its 5′ phosphate. During oligonucleotide synthesis, a polymerase can erroneously catalyze the covalent addition of the unshielded nucleotide, which may result in the addition of more than one nucleotide per step when using polymerase-nucleotide conjugates in nucleic acid synthesis (e.g., insertion or non-termination). Technologies provided herein help overcome this challenge to achieve accurate and precise stepwise addition with reduced errors as compared to previously described synthesis approaches.
[0611] In some embodiments, non-termination may occur when an unshielded nucleotide with an uncleaved 5′ phosphate is added to an oligonucleotide. As provided herein, in some embodiments, a phosphatase hydrolyzes a 5’ phosphate (e.g., a terminal 5’ phosphate) of a nucleotide (e.g., of a nucleotide triphosphate, etc.). In some embodiments, the terminal 5’ phosphate is on an α-phosphate, β-phosphate, χ-phosphate, δ-phosphate, ε-phosphate, φ- phosphate, or γ-phosphate of the nucleotide. In some embodiments, a phosphatase as disclosed herein hydrolyzes a 5′ phosphate (e.g., a terminal 5’ phosphate) of a nucleotide in a polymerase-nucleotide conjugate or of a free nucleotide. In some embodiments, a phosphatase as disclosed herein hydrolyzes 5’ phosphates of nucleotides, e.g., one or more 5′ terminal phosphate(s) of nucleotides in a plurality of polymerase-nucleotide conjugates, and prevents the hydrolyzed nucleotide from addition to the nucleic acid during oligonucleotide synthesis. In some embodiments, a phosphatase hydrolyzes a 5′ phosphate of an unshielded nucleotide of a polymerase-nucleotide conjugate. In some such embodiments, the hydrolysis of the 5’ phosphate prevents the unshielded nucleotide from addition to the nucleic acid during oligonucleotide synthesis. In some embodiments, a phosphatase as disclosed herein hydrolyzes one or more 5′ phosphate(s) of unshielded nucleotides in a plurality of polymerase-nucleotide conjugates and prevents said unshielded nucleotides from addition to the nucleic acid during oligonucleotide synthesis. In some embodiments, a phosphatase as disclosed herein hydrolyzes one or more 5′ phosphate(s) of one or more free nucleotides in acomposition comprising one or more polymerase-nucleotide conjugates and prevents the free nucleotides from addition to the nucleic acid during oligonucleotide synthesis. In some embodiments, a phosphatase hydrolyzes a 5′ phosphate of one or more free nucleotides present in a composition comprising one or more polymerase-nucleotide conjugates. In some such embodiments, the hydrolysis of the 5’ phosphate prevents the one or more free nucleotides from addition to the nucleic acid during oligonucleotide synthesis.
[0612] The presence of an unshielded nucleotide in a conjugate reagent can lead to non- termination (i.e., insertion) in oligonucleotide synthesis. In some embodiments, an unshielded nucleotide is less likely to inhibit subsequent nucleotide addition after having been added to an oligonucleotide during nucleic acid synthesis.
[0613] In some embodiments, a fraction of nucleotides in a plurality of polymerase- nucleotide conjugates are not shielded by a polymerase. In some embodiments, a polymerase-nucleotide conjugate comprises an unshielded nucleotide. In some embodiments, a polymerase molecule in a polymerase-nucleotide conjugate does not sterically hinder access of a phosphatase to the 5′ phosphate of a tethered nucleotide. In some embodiments, a tethered nucleotide is an unshielded nucleotide. In some embodiments, an unshielded nucleotide in a polymerase-nucleotide conjugate is not sterically hindered by a tethered polymerase from a phosphatase capable of removing its 5′ phosphate. In some embodiments, removing the 5′ phosphate (e.g., terminal 5’ phosphate) of a nucleotide in a polymerase- nucleotide conjugate prevents the nucleotide from addition to the nucleic acid during nucleic acid synthesis. In some embodiments, a phosphatase hydrolyzes the 5′ phosphate (e.g., terminal 5’ phosphate) of an unshielded nucleotide in a polymerase-nucleotide conjugate. In some such embodiments, more than one 5’ terminal phosphate is removed, for example, wherein a 5’ terminal phosphates are removed serially, i.e., from a first nucleotide, then a second nucleotide, etc. Thus, in some such embodiments, one or more 5’ terminal phosphates may be removed, though in a given nucleotide, a single 5’ terminal phosphate exists and is removed, upon which point a different phosphate becomes the 5’ terminal phosphate of a nucleotide having at least one 5’ terminal phosphate.
[0614] In some embodiments, an unshielded nucleotide is part of an improperly formed conjugate. In some embodiments, an improperly formed conjugate comprises a mis-folded polymerase, a polymerase in which a nucleotide is attached in the wrong position, and / or a polymerase in which multiple nucleotides are attached. In some embodiments, a nucleotide is free or untethered from a polymerase due to instability or imperfect purification. In some embodiments, a free or untethered nucleotide is an unshielded nucleotide.
[0615] Lack of shielding of a nucleotide in a polymerase-nucleotide conjugate can occur due to a number of processes during preparation of polymerase-nucleotide conjugates or during the addition reaction itself. Nucleotides that are not attached to a polymerase in a composition comprising a polymerase-nucleotide conjugate are considered unshielded nucleotides. Non-limiting examples of processes that may result in a polymerase-nucleotide conjugate comprising an unshielded nucleotide include: spontaneous cleavage of a linker between a nucleotide and a polymerase (see, e.g., linkers in FIGs. 6A-6D), unfolding of a polymerase (see, e.g., FIG. 6B), a polymerase having a nucleotide attached in the wrong position (see, e.g., FIG. 6A), a polymerase comprising multiple attached nucleotides (i.e., on a single polymerase), or an untethered, free nucleotide (see, e.g., FIG. 6C). In some embodiments, spontaneous cleavage of a linker between a nucleotide and a polymerase may occur due to instability, resulting in free nucleotides in the conjugate reagent.
[0616] In some embodiments, a fraction of nucleotides in a plurality of polymerase- nucleotide conjugates are shielded by a polymerase. A non-limiting example of a shielded nucleotide includes a nucleotide that is tethered in the catalytic site of a correctly folded polymerase (see, e.g., exemplary schematic in FIG. 6D). In some embodiments, a polymerase-nucleotide conjugate comprises a shielded nucleotide. In some embodiments, a polymerase molecule in a polymerase-nucleotide conjugate sterically hinders access of a phosphatase to the 5′ phosphate (e.g., the 5’ terminal phosphate) of a tethered nucleotide. In some embodiments, the 5’ phosphate (e.g., the 5’ terminal phosphate) can be, for example, on an α-phosphate, β-phosphate, χ-phosphate, δ-phosphate, ε-phosphate, φ-phosphate, or γ- phosphate. In some embodiments, a tethered nucleotide is a shielded nucleotide. In some embodiments, a shielded nucleotide in a polymerase-nucleotide conjugate is sterically hindered by a tethered polymerase from a phosphatase capable of removing its 5′ phosphate. In some embodiments, a phosphatase is unable to hydrolyze the 5′ phosphate (e.g., 5’ terminal phosphate) of a shielded nucleotide in a polymerase-nucleotide conjugate. Phosphatase
[0617] A phosphatase typically uses water to cleave a phosphoric acid monoester into a phosphate ion and an alcohol. A phosphatase enzyme catalyzes the hydrolysis of its substrate.
[0618] The 5′ phosphate of a nucleotide in a polymerase-nucleotide conjugate is necessary for addition of the nucleotide to an oligonucleotide by a polymerase. In some embodiments,removal of the 5′ phosphate of a nucleotide in a polymerase-nucleotide conjugate prevents the nucleotide from addition to the nucleic acid during oligonucleotide synthesis.
[0619] Disclosed herein are methods comprising adding a phosphatase to a conjugate reagent to hydrolyze the 5′ phosphate group of a nucleotide in a polymerase-nucleotide conjugate. In some embodiments, a phosphatase removes a phosphate moiety from an unshielded nucleotide in a polymerase-nucleotide conjugate. In some embodiments, the methods comprise adding a phosphatase capable of hydrolyzing a 5′ phosphate group of an unshielded nucleotide to a polymerase-nucleotide conjugate.
[0620] Any suitable phosphatase, engineered enzyme having phosphatase activity, or a functional fragment thereof for the methods described herein is contemplated by the disclosure. In some embodiments, the phosphatase is a nucleotidase. Enzymes having phosphatase activity are included in the enzyme class E.C 3.1.3.-., hydrolases acting on ester bonds, e.g., a phosphoric monoester hydrolase. However, enzymes having suitable phosphatase activity (such as apyrase) may be found in other enzyme classes. In some embodiments, a phosphatase may optionally be capable of hydrolyzing an inorganic phosphate substrate, e.g., pyrophosphate.
[0621] In some embodiments, the phosphatase is immobilized to a solid support. In some embodiments, the phosphatase is a fusion protein. In some embodiments, the phosphatase comprises a detectable label. In some embodiments, the phosphatase is a recombinant polypeptide. In some embodiments, the phosphatase is a wild type phosphatase. In some embodiments, the wild type phosphatase is isolated from the organism in which it is natively expressed.
[0622] Alkaline phosphatases (ALP, ALKP, ALPase, Alk Phos), or basic phosphatases, are plasma membrane-bound glycoproteins that catalyze the hydrolysis of phosphate monoesters and are optimally active at alkaline pH environments. Alkaline phosphatases are homodimeric protein enzymes of 86 kilodaltons. Each monomer contains five cysteine residues, two zinc atoms, and one magnesium atom crucial to its catalytic function.
[0623] Non-limiting examples of alkaline phosphatases include: B. taurus (Quick calf- intestinal alkaline phosphatase, or CIP, NEB), Pandalus borealis (shrimp alkaline phosphatase, NEB), Antarctic bacterium TAB5 (Antarctic phosphatase, NEB), and E. coli (Takara Bio) phosphatase. Additional non-limiting examples of alkaline phosphatases include: placental alkaline phosphatase (PLAP) and human-intestinal alkaline phosphatase.
[0624] Non-alkaline phosphatases may be acid phosphatases. A non-limiting example of a non-alkaline phosphatases is tartrate resistant acid phosphatase.
[0625] Illustrative amino acid sequences encoding phosphatases for use in the methods described herein are shown, without limitation, in Table 3. Table 3. Exemplary Alkaline Phosphatase SequencesPolynucleotide Synthesis
[0626] In some embodiments, a method of synthesizing a polynucleotide comprises contacting (e.g., incubating) a polymerase-nucleotide conjugate with a nucleic acid, wherein a polymerase of the polymerase-nucleotide conjugate elongates the nucleic acid using its tethered nucleotide.
[0627] As described above, disclosed herein are methods of nucleic acid synthesis comprising the step of contacting (e.g., incubating) a conjugate reagent comprising a plurality of polymerase-nucleotide conjugates with a phosphatase. In some embodiments, contacting a conjugate reagent comprising a plurality of polymerase-nucleotide conjugates with a phosphatase occurs before or during cyclic extension reactions. In some embodiments, the presence of a phosphatase in a stepwise method of nucleic acid synthesis reduces non- terminations and processes that lead to addition of more than one nucleotide per step. In some embodiments, use of a conjugate reagent treated with (e.g., incubated with) a phosphatase reduces non-terminations when the conjugate reagent is used in a stepwise method of nucleic acid synthesis as compared to an untreated conjugate reagent.
[0628] Polymerase-nucleotide conjugates may be stored together with a phosphatase and remain in the system (also during the DNA extension reaction) or a phosphatase may be removed, for a certain incubation period before the conjugates are added to the DNA, and upon initiation of the DNA addition reaction.
[0629] In some embodiments a phosphatase is incubated with a conjugate reagent comprising a plurality of polymerase-nucleotide conjugates. In some embodiments, incubation of a conjugate reagent with a phosphatase is performed before contacting a sample with the conjugate reagent. In some embodiments, a phosphatase is removed from a conjugate reagent prior to contacting a sample with the conjugate reagent. In some embodiments, incubation of a conjugate reagent with a phosphatase is performed after contacting a sample with the conjugate reagent.
[0630] The concentration of phosphatase in contact or incubated with the polymerase- nucleotide conjugate (e.g., a conjugate reagent) can be expressed in, for example, a stoichiometric ratio of phosphatase to conjugate fold increase relative to the conjugate concentration, units of activity of phosphatase, molarity, or mg / mL.
[0631] Any suitable stoichiometric ratio of conjugate to phosphatase can be used in the methods described herein. In some embodiments, the stoichiometric ratio of conjugate to phosphatase is from about 1:1 to about 1:500. In some embodiments, the stoichiometric ratio of conjugate to phosphatase is about 1:1, about 1:2, about 1:3, about 1:4, about 1:5, about 1:10, about 1:15, about 1:20, about 1:25, about 1:30, about 1:35, about 1:40, about 1:45, about 1:50, about 1:55, about 1:60, about 1:65, about 1:70, about 1:75, about 1:80, about 1:85, about 1:90, about 1:95, about 1:100, about 1:105, about 1:110, about 1:115, about 1:120, about 1:125, about 1:130, about 1:135, about 1:140, about 1:145, about 1:150, about 1:155, about 1:160, about 1:165, about 1:170 about 1:175, about 1:180, about 1:185, about 1:190, about 1:195, about 1:200, about 1:225, about 1:250, about 1:275, about 1:300, about 1:325, about 1:350, about 1:375, about 1:400, about 1:425, about 1:450, about 1:475, or about 1:500.
[0632] Any suitable stoichiometric ratio of phosphatase to conjugate can be used in the methods described herein. In some embodiments, the stoichiometric ratio of phosphatase to conjugate is from about 1:1 to about 1:500. In some embodiments, the stoichiometric ratio of conjugate to phosphatase is about 1:1, about 1:2, about 1:3, about 1:4, about 1:5, about 1:10, about 1:15, about 1:20, about 1:25, about 1:30, about 1:35, about 1:40, about 1:45, about 1:50, about 1:55, about 1:60, about 1:65, about 1:70, about 1:75, about 1:80, about 1:85, about 1:90, about 1:95, about 1:100, about 1:105, about 1:110, about 1:115, about 1:120, about 1:125, about 1:130, about 1:135, about 1:140, about 1:145, about 1:150, about 1:155, about 1:160, about 1:165, about 1:170 about 1:175, about 1:180, about 1:185, about 1:190, about 1:195, about 1:200, about 1:225, about 1:250, about 1:275, about 1:300, about 1:325, about 1:350, about 1:375, about 1:400, about 1:425, about 1:450, about 1:475, or about 1:500.
[0633] Any suitable phosphatase concentration can be used in the methods described herein. In some embodiments, the phosphatase concentration is from about 0.01 mg / mL to about 10.5 mg / mL. In some embodiments, the phosphatase concentration is about 0.1 mg / mL, about 0.15 mg / mL, about 0.25 mg / mL, about 0.5 mg / mL, about 0.75 mg / mL, about 1 mg / mL, about 1.25 mg / mL, about 1.5 mg / mL, about 1.75 mg / mL, about 2 mg / mL, about 2.25 mg / mL, about 2.5 mg / mL, about 2.75 mg / mL, about 3, about 3.25 mg / mL, about 3.5mg / mL, about 3.75 mg / mL, about 4 mg / mL, about 4.25 mg / mL, about 4.5 mg / mL, about 4.75 mg / mL, about 5, about 5.25 mg / mL, about 5.5 mg / mL, about 5.75 mg / mL, about 6 mg / mL, about 6.25 mg / mL, about 6.5 mg / mL, about 6.75 mg / mL, about 7, about 7.25 mg / mL, about 7.5 mg / mL , about 7.75 mg / mL, about 8 mg / mL, about 8.25 mg / mL, about 8.5 mg / mL, about 8.75 mg / mL, about 9, about 9.25 mg / mL, or about 9.5 mg / mL , about 9.75 mg / mL, about 10 mg / mL, about 10.25 mg / mL, or about 10.5 mg / mL.
[0634] Any suitable fold increase of phosphatase over conjugate can be used in the methods described herein. In some embodiments, the fold increase of phosphatase over conjugate is from about 2-fold to about 500-fold. In some embodiments, the fold increase of phosphatase over conjugate is about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 10-fold, about 15-fold, about 20-fold, about 25-fold, about 30-fold, about 35-fold, about 40-fold, about 45- fold, or about 50-fold, about 55-fold, about 60-fold, about 65-fold, about 70-fold, about 75- fold, about 80-fold, about 85-fold, about 90-fold, about 95-fold, about 100-fold, about 105- fold, about 110-fold, about 115-fold, about 120-fold, about 125-fold, about 130-fold, about 135-fold, about 140-fold, about 145-fold, or about 150-fold, about 155-fold, about 160-fold, about 165-fold, about 170-fold, about 175-fold, about 180-fold, about 185-fold, about 190- fold, about 195-fold, about 200-fold, about 225-fold, about 250-fold, about 275-fold, about 300-fold, about 325-fold, or about 350-fold, about 375-fold, about 400-fold, about 425-fold, about 450-fold, about 475-fold, or about 500-fold.
[0635] Any suitable fold increase of conjugate over phosphatase can be used in the methods described herein. In some embodiments, the fold increase of conjugate over phosphatase is from about 2-fold to about 500-fold. In some embodiments, the fold increase of conjugate over phosphatase is about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 10-fold, about 15-fold, about 20-fold, about 25-fold, about 30-fold, about 35-fold, about 40-fold, about 45-fold, or about 50-fold, about 55-fold, about 60-fold, about 65-fold, about 70-fold, about 75-fold, about 80-fold, about 85-fold, about 90-fold, about 95-fold, about 100-fold, about 105-fold, about 110-fold, about 115-fold, about 120-fold, about 125-fold, about 130- fold, about 135-fold, about 140-fold, about 145-fold, or about 150-fold, about 155-fold, about 160-fold, about 165-fold, about 170-fold, about 175-fold, about 180-fold, about 185- fold, about 190-fold, about 195-fold, about 200-fold, about 225-fold, about 250-fold, about 275-fold, about 300-fold, about 325-fold, or about 350-fold, about 375-fold, about 400-fold, about 425-fold, about 450-fold, about 475-fold, or about 500-fold.
[0636] Any suitable concentration of phosphatase can be used in the methods described herein. In some embodiments, the concentration of phosphatase is from about 0.5 µM toabout 500 µM. In some embodiments, the concentration of phosphatase is about 0.5 µM, about 1 µM, about 2 µM, about 5 µM, about 10 µM, about 15 µM, about 20 µM, about 25 µM, about 30 µM, about 35 µM, about 40 µM, about 45 µM, about 50 µM, about 55 µM, about 60 µM, about 65 µM, about 70 µM, about 80 µM, about 85 µM, about 90 µM, about 95 µM, about 100 µM, about 105 µM, about 110 µM, about 115 µM, about 120 µM, about 125 µM, about 130 µM, about 135 µM, about 140 µM, about 145 µM, about 150 µM, about 155 µM, about 160 µM, about 165 µM, about 170 µM, about 180 µM, about 185 µM, about 190 µM, about 195 µM, about 200 µM, about 225 µM, about 250 µM, about 275 µM, about 300 µM, about 325 µM, about 350 µM, about 375 µM, about 400 µM, about 425 µM, about 450 µM, about 475 µM, or about 500 µM.
[0637] In some embodiments, presence of a phosphatase in a stepwise method of nucleic acid synthesis reduces non-terminations and processes that lead to addition of more than one nucleotide per step. Polynucleotides or nucleic acids generated in the methods described herein are said to contain an insertion if a non-termination event has occurred. In some embodiments, nucleic acid synthesis in the presence of a phosphatase reduces the rate of non- terminations by about 50% to about 100% compared to nucleic acid synthesis in the absence of a phosphatase. In some embodiments, rates of non-terminations are reduced by about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% compared to nucleic acid synthesis in the absence of a phosphatase. In some embodiments, the total amount of nucleic acid synthesis product with insertions generated in the presence of a phosphatase is less than about 5%, less than about 4%, less than about 3%, less than about 2%, less than about 1%, less than about 0.5%, less than about 0.1%, less than about 0.05%, or less than about 0.01%. In some embodiments, nucleic acid synthesis product generated in the presence of a phosphatase is absent of nucleic acid synthesis product with insertions.
[0638] Disclosed herein, in some embodiments, are methods of synthesizing a polynucleotide comprising a pre-determined sequence, comprising contacting a conjugate reagent comprising a plurality of polymerase-nucleotide conjugates with a phosphatase. In some embodiments, contacting comprises incubating the conjugate reagent with the phosphatase. In some embodiments, the method generates a heterogeneous population of polynucleotide products comprising the pre-determined sequence. The heterogeneous population of polynucleotide products comprising the pre-determined sequence can be referred to as an “end product.”
[0639] In some embodiments, contacting the conjugate reagent comprising a plurality of polymerase-nucleotide conjugates with a phosphatase prevents insertion of nucleotides (i.e., non-terminations), such that these insertions are absent from the pre-determined sequence. In some embodiments, an end product comprises nucleic acids, a percentage of which comprise a target sequence and a percentage of which do not comprise a target sequence. In some embodiments, an end product comprises less than about 99%, less than about 95%, less than about 90%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 55%, less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, less than about 1%, less than about 0.5%, less than about 0.1%, less than about 0.05%, or less than about 0.01% of a polynucleotide comprising a sequence that is not the pre-determined sequence (that is, not the “target” sequence) as compared to a polynucleotide comprising a sequence that is a predetermined (“target”) sequence. In some embodiments, the end product is substantially absent of a polynucleotide comprising a sequence that is not the pre-determined (“target”) sequence.
[0640] Any suitable method known in the art can be used for analyzing the end product. The end product can be assessed by analyzing nucleic acid synthesis products any time following the initiation of an extension reaction (e.g., reaction time course). Analyses can be performed by, for example, capillary electrophoresis (CE) as previously demonstrated (Smith and Nelson. Curr Protoc Nucleic Acid Chem. Chapter 10: Unit 10.9. 2003; Durney et al. Anal Bioanal Chem. 407:6923-6938. 2015). CE can separate and report abundance of polynucleotide products with single nucleotide resolution. The relative abundance of each nucleic acid product generated by the methods of nucleic acid synthesis provided herein can be analyzed by CE. By comparing the abundance of the starting material (i.e., initial polynucleotide or oligonucleotide to which a nucleotide is being incorporated) and the extension products, it is possible to determine the extent to which the extension reaction is completed. The change over time of the starting material and extended species is indicative of the turnover rate, as described herein. This approach to determine turnover rate has been demonstrated previously (Palluk et al. Nat Biotech. 36(7):645-650. 2018). Alternatively, analysis of nucleic acid synthesis products can be performed using reverse-phase high- performance liquid chromatography (RP-HPLC) as described previously (Jensen and Davis. Biochemistry. 57(12):1821-1832. 2018).
[0641] CE and RP-HPLC may also be used to determine the purity of each species in a nucleic acid synthesis product by determining the area under the curve for peaks in the electropherograms and chromatograms for CE and RP-HPLC, respectively. Any suitable software package suitable for fitting curves to electropherograms and chromatograms and calculating area under the curve (AUC) may be used to determine the abundance of each polynucleotide product in a plurality of nucleotide products.
[0642] Any suitable polynucleotide sequencing method can be used for analysis (e.g., analysis of a sequence of a polynucleotide product, e.g., intermediate product, e.g., end product, etc.). For example, the sequencing method can be long-read sequencing, next generation sequencing, short-read sequencing, shotgun sequencing, sanger sequencing, high throughput sequencing, sequencing by synthesis, sequencing by ligation, sequencing by hybridization, and / or sequencing by mass spectrometry. Sequencing can be suitable for identifying the pre-defined sequence in an end product. Sequencing can be suitable for determining the percentage of sequences in an end product that are not the pre-defined sequence.
[0643] Any method known in the art for analysis of size of an end product can be used to determine the purity of the end product. For example, gel electrophoresis and / or mass spectrometry can be used to determine purity of an end product.
[0644] Disclosed herein are compositions comprising conjugates comprising nucleotides attached to a polymerase, wherein the purity of nucleotides shielded by a linked polymerase of the composition, as compared to total nucleotides in the composition, is greater than about 80%, about 85%, about 90%, about 95%, about 97%, about 98%, about 99%, about 99.5%, or about 99.9% shielded nucleotides. In some embodiments, the purity of nucleotides shielded by a linked polymerase is substantially free of impurities.
[0645] Low Divalent Cation Concentrations to Improve Stepwise Yield
[0646] The present disclosure provides insights based on a surprising discovery that rate of extension reactions using polymerase-nucleotide conjugates in polynucleotide synthesis can be increased by reducing the concentration of divalent cation concentrations present in the reactions as compared to standard divalent cation concentrations used in polynucleotide synthesis reactions using a free polymerase and free polynucleotide.
[0647] The disclosure provides compositions and methods related to the discovery. The details of various embodiments of the compositions and methods are set forth in the disclosure. Other features, objects, and advantages of the compositions and methods disclosed herein will be apparent from the description and the drawings, and from the claims.
[0648] Disclosed herein are methods of nucleic acid synthesis. Nucleic acid synthesis can refer to synthesis, or generation of a product that is a nucleic acid molecule (i.e., a polynucleotide). The methods of nucleic acid synthesis can comprise stepwise synthesis, wherein nucleotides are inserted stepwise into a nucleic acid polymer or polynucleotide. A typical process for stepwise synthesis of a polynucleotide comprises adding nucleotides stepwise to a starter molecule (e.g., an initial oligonucleotide) via the cycled steps of: addition of a polymerase and a nucleotide (e.g., a polymerase-nucleotide conjugate) to an oligonucleotide and covalently incorporating the nucleotide to the 3′ end of the oligonucleotide catalyzed by the polymerase. Successful incorporation of a nucleotide to an oligonucleotide can be referred to as an “extension” or “extension reaction”.
[0649] In some embodiments, the polymerase and nucleotide are linked together (i.e., tethered) to form a conjugate (i.e., a polymerase-nucleotide conjugate). In a stepwise synthesis using the conjugate, the tethered nucleotide is covalently incorporated (i.e., is added) into the 3′ end of the oligonucleotide, which is catalyzed by the tethered polymerase. The tethered polymerase can stay tethered to the nucleotide following covalent incorporation. Covalent incorporation of a nucleotide can be referred to as an extension or an addition. The tethered polymerase can be cleaved from the inserted nucleotide to expose the 3′ end of the oligonucleotide. These steps can be repeated to synthesize a desired polynucleotide. The desired polynucleotide can have a pre-determined (i.e., pre-defined or target) sequence.
[0650] As will be understood to those of skill in the art, methods provided herein, such as nucleic acid synthesis, are performed in a reaction volume. Given context, the contents of the reaction volume may change before, after, and during the synthesis. For example, in some embodiments, the reaction volume comprises one or more of a buffer, polynucleotide, polymerase-nucleotide conjugate, nucleotide starter molecule, synthesis products, phosphatases, etc. In some such embodiments, a reaction volume may be prepared comprising only select components (e.g., a buffer and nucleotide starter molecule), and one or more additional components may be added to the volume at one or more subsequent times. In some embodiments, all components for a given synthesis may be added substantially simultaneously. In some embodiments, one or more components of a synthesis reaction maybe pre-treated (e.g., pretreatment of a polymerase-nucleotide conjugate with a phosphatase) prior to being included in a reaction volume for a nucleic acid synthesis reaction.
[0651] Among other things, methods of nucleic acid synthesis disclosed herein are carried out in a reaction buffer composition. The reaction buffer composition is an aqueous solution. The reaction buffer composition comprises a set of components suitable for the stability of the polymerase, nucleotide, polymerase-nucleotide conjugates, st...
Claims
CLAIMS 1. A method of template-independent de novo synthesis of a long single-stranded polynucleotide, comprising: a) providing a substrate comprising a plurality of free hydroxyl groups linked to the substrate and suitable for nucleotide coupling; b) coupling a nucleotide to the plurality of free hydroxyl groups, wherein the nucleotide is linked to a blocking group; c) removing said blocking group from said coupled nucleotides; d) repeating steps (b) and (c) according to a predetermined nucleotide sequence (e.g., a reference sequence) to yield a plurality of de novo synthesized polynucleotides at least 500 nucleotides in length, wherein each coupling has an observed error rate of less than 1% as compared to said predetermined polynucleotide sequence.
2. The method of claim 1, wherein the de novo synthesized long polynucleotides are at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1800, at least 1900, or at least 2000 nucleotides in length.
3. The method of claim 1, wherein the de novo synthesized long polynucleotides are from 500 to 1000 nucleotides in length, from 500 to 1500 nucleotides in length, from 500 to 2000 nucleotides in length, or from 500 to 2500 nucleotides in length.
4. The method of claim 1, wherein the observed error rate of each coupling is less than 0.9%, less than 0.8%, less than 0.7%, less than 0.6%, less than 0.5%, less than 0.4%, less than 0.3%, less than 0.2%, less than 0.1% per cycle, less than 0.09% per cycle, less than 0.08% per cycle, less than 0.07% per cycle, less than 0.06% per cycle, less than 0.05% per cycle, less than 0.04% per cycle, less than 0.03% per cycle, less than 0.02% per cycle, or less than 0.01% per cycle.
5. The method of claim 1, wherein said coupling step is performed in less than 120 seconds, less than 90 seconds, less than 60 seconds, less than 50 seconds, less than 40 seconds, less than 30 seconds, less than 20 seconds, less than 15 seconds, less than 10 seconds, or less than 5 seconds.
6. The method of claim 1, wherein said template-independent polynucleotide synthesis is performed at a rate of at least 10 nucleotides per hour, at least 15 nucleotides per hour, atleast 20 nucleotides per hour, at least 25 nucleotides per hour, at least 30 nucleotides per hour, at least 40 nucleotides per hour, at least 50 nucleotides per hour, or at least 60 nucleotides per hour.
7. The method of claim 1, wherein said template-independent polynucleotide synthesis is performed at a rate of from 5 nucleotides per hour to 25 nucleotides per hour, from 10 nucleotides per hour to 30 nucleotides per hour, from 15 nucleotides per hour to 45 nucleotides per hour, or from 20 nucleotides per hour to 60 nucleotides per hour.
8. The method of claim 1, wherein said plurality of de novo synthesized polynucleotides on the substrate comprises at least 10 polynucleotides, at least 20 polynucleotides, at least 50 polynucleotides, at least 100 polynucleotides, at least 200 polynucleotides, at least 500 polynucleotides, at least 1,000 polynucleotides, at least 10,000 polynucleotides, at least 10 x 105polynucleotides, at least 10 x 106polynucleotides, at least 10 x 107polynucleotides, at least 10 x 108polynucleotides, at least 10 x 109polynucleotides, at least 10 x 1010polynucleotides, at least 10 x 1011polynucleotides, at least 10 x 1012polynucleotides, at least 10 x 1013polynucleotides, or at least 10 x 1014polynucleotides.
9. The method of claim 1, wherein the free hydroxyl groups are at the end of a plurality of starter oligonucleotides or growing polynucleotides attached to the substrate.
10. The method of claim 9, wherein the starter oligonucleotides comprise a single- stranded region at the 3’ end.
11. The method of claim 9, wherein the starter oligonucleotide is hybridized to an oligonucleotide bound to the substrate.
12. The method of claim 9, wherein the starter oligonucleotide is covalently linked to the substrate.
13. The method of claim 1, wherein said nucleotide coupling is performed enzymatically.
14. The method of claim 1, wherein said nucleotide coupling is catalyzed by a polymerase.
15. The method of claim 14, wherein said polymerase is a template-independent polymerase.
16. The method of claim 15, wherein said template-independent polymerase is covalently linked to said nucleotide.
17. The method of claim 15 or 16, wherein said template-independent polymerase is Terminal deoxynucleotidyl Transferase (TdT), or a variant thereof.
18. The method of claim 14, wherein said polymerase is an RNA polymerase.
19. The method of claim 1, wherein said blocking group is a template-independent polymerase linked to said nucleotide.
20. The method of claim 19, wherein removing said blocking group comprises cleaving a linker attaching said nucleotide to said template-independent polymerase.
21. The method of claim 1, wherein said blocking group is a 3′-O-blocking group.
22. The method of claim 21, wherein removing said blocking group comprises removing said 3′-O-blocking group from said nucleotide to leave a free 3′ hydroxyl group.
23. The method of claim 1, wherein the blocking group is a 2′ or 3′ modification of the nucleotide.
24. The method of claim 23, wherein the 2′ modification is selected from the group consisting of -H, -OH, -F, -OMe, -N3, -NH2, and -Ara.
25. The method of claim 23, wherein the 3′ modification is selected from the group consisting of --H, -OH, -OCH2N3, -ONH2 and -Oallyl.
26. The method of claim 1, wherein the blocking group is a reversible terminator.
27. The method of claim 1, wherein the predetermined sequence has a GC content of at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95%.
28. The method of claim 1 or 27, wherein the predetermined sequence has an AT content at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95%.
29. The method of claim 1, wherein said coupling is performed in the presence of phosphatase.
30. The method of claim 29, wherein said phosphatase is an inorganic pyrophosphatase.
31. The method of claim 1, wherein the nucleotide linked to the blocking group is a nucleotide-polymerase conjugate, and wherein the conjugate has been treated with phosphatase.
32. The method of claim 1, wherein said coupling is performed in the presence of a divalent cation, wherein a total concentration of divalent cations present in the reaction volume of said coupling is no greater than about 500 μM.
33. The method of claim 32, wherein the total concentration of divalent cations present in the reaction volume is no greater than about 250 μM, about 125 μM or about 50 μM.
34. The method of claim 32, wherein the divalent cation present at the highest concentration in the reaction volume is cobalt (Co2+) or zinc (Zn2+).
35. The method of claim 32, wherein the at least one divalent cation is selected from Mg2+, Ca2+, Sr2+, Ba2+, Mn2+, Co2+, Fe2+, Ni2+, Cu2+, and Zn2+, or a combination thereof.
36. The method of claim 32, wherein the coupling reaction is performed in the absence of Mg2+.
37. The method of claim 1, wherein said nucleotide comprises one or more modifications to a hydrogen binding N or O on the nucleobase.
38. The method of claim 1, wherein said coupled nucleotide comprises one or more alkylated nucleobases after removal of said blocking group.
39. The method of claim 38, further comprising contacting said de novo synthesized polynucleotides with an alkyl transferase.
40. The method of claim 39, wherein said alkyl transferase is from EC 2.1.1.
63.
41. The method of claim 39, wherein said alkyl transferase is selected from an alkyl transferase listed in Table 1 or Table 2.
42. The method of claim 39, wherein said alkyl transferase is O6-alkylguanine DNA alkyltransferase.
43. The method of claim 39, wherein said alkyl transferase is AlkB.
44. The method of claim 38, wherein the alkylated nucleobase is represented by:wherein X is -C(R2)= or -N=; R1is selected from the group consisting of C1-6 alkyl, C2-6 alkenyl, C1-6 alkynyl, and – (CH2)0-3Ph, wherein R1is optionally substituted with 1-6 instances of R1a; each R1ais independently selected from halogen, C1-6 alkyl, -(CH2)0-3OR1b, -NO2, -N3, -OPO2OH, and –(CH2)0-3NHR1b; and each R1bis independently selected from hydrogen, C1-6alkyl, -C(O)(C1-6alkyl), C1-6haloalkyl, -C(O)(C1-6haloalkyl), and -CH2OAc; R2is selected from the group consisting of hydrogen, optionally substituted C1-4 alkyl chain, wherein 1-2 methylene units is optionally and independently replaced with -O-, - N(Ra)-, -C(O)-, -S-, -S(O)-, -S(O)2-, or phenylene, wherein R2is optionally substituted with 1-6 instances of R2a; each R2ais independently selected from halogen, C1-6 alkyl, -(CH2)0-3OR2b, -NO2, -N3, -OPO2OH, and –(CH2)0-3NHR2b; and each R2bis independently selected from hydrogen and C1-6 alkyl.
45. The method of claim 38, wherein R1is C1-4 alkyl.
46. The method of claim 39, wherein R1is selected from the group consisting of methyl, ethyl, n-propyl, and n-butyl.
47. The method of claim 38, wherein R1is selected from the group consisting of:.
48. The method of claim 47, wherein R1ais -OR1b.
49. The method of claim 48, wherein R1bis hydrogen.
50. The method of claim 38, wherein the alkylated nucleobase in the polynucleotide is selected from the group consisting of:, , , ,51. The method of claim 1, wherein said coupled nucleotide comprises one or more modifications to a base-pairing nitrogen or oxygen on the nucleobase after removal of said blocking group.
52. The method of claim 1, wherein said coupled nucleotide is represented by (B):(B) wherein R is a ribose polyphosphate or deoxyribose polyphosphate; Y is a nucleobase; L-R1is a protecting group; wherein L is attached to a base-pairing nitrogen or oxygen of the nucleobase; and wherein R1is selected from the group consisting of hydrogen, -OH, - N(Rb)2, and - SH, wherein each Rbis independently hydrogen or optionally substituted C1-6 alkyl.
53. The method of claim 52, wherein: L is -Z-L1-L2-;Z is selected from the group consisting of a bond, -C(O)-, -C(O)CH2-, -C(O)C(RL)2-, -C(O)CH(RL)-, -C(O)O-, and -C(O)N(H)-; L1is selected from the group consisting of a bond, ,wherein L1is optionally substituted with 1-4 instances of RL; each RLis independently selected from the group consisting of halogen, hydroxyl, oxo, and optionally substituted C1-C3 alkyl, wherein 2 instances of R1are optionally taken together with the intervening atom(s) to form a 3-6 membered carbocyclyl ring; L2is selected from the group consisting of a bond, an optionally substituted C1-12alkylene chain, C4-C20 polyethylene glycol, an optionally substituted C2-12 alkenylene chain, and an optionally substituted C2-12 alkynylene chain, wherein 1-6 methylene units are optionally and independently replaced with -O-, -N(Rb)-, -C(O)-, -S-, -S(O)-, -S(O)2-, or phenylene; W is selected from the group consisting of -O-, -S-, -S(O)2-, and -N(Rb)-; each Rais halogen, -Me, or -OMe; each Rbis independently hydrogen or C1-6alkyl; and n is 1 or 2; wherein L1and Z are not both a bond.
54. The method of claim 53, wherein Z is a bond when L is attached to a base-pairing oxygen of the nucleobase.
55. The method of claim 53, wherein Z is selected from the group consisting of -C(O)-, - C(O)CH2-, -C(O)C(RL)2-, -C(O)CH(RL)-, -C(O)O-, and -C(O)N(H)-when L is attached to a base-pairing nitrogen of the nucleobase.
56. The method of any of claims 53-55, wherein L1is, wherein L1is optionally substituted with 1-4 instances of RL; n is 1 or 2; and W is selected from the group consisting of -O-, -S-, -S(O)2-, and -N(Rb)-.
57. The method of claim 56, wherein L1is selected from the group consisting ofwherein L1is optionally substituted with 1-4 instances of RL.
58. The method of any of claims 53-57, wherein each RLis independently optionally substituted C1-C3 alkyl, wherein 2 instances of R1are optionally taken together with the intervening atom(s) to form a 3-6 membered carbocyclyl ring.
59. The method of claim 58, wherein RLis optionally substituted methyl.
60. The method of any of claims 53-57, wherein L1is selected from the group consisting of61. The method of any of claims 53-60, wherein Z is -C(O)O-.
62. The method of any of claims 53-60, wherein Z is -C(O)N(H)-.
63. The method of any of claims 53-60, wherein Z is a bond.
64. The method of any of claims 53-60, wherein Z is -C(O)-.
65. The method of any of claims 53-60, wherein -Z-L1-L2-R1is selected from the group consisting of66. The method of any of claims 53-60, wherein L2is an optionally substituted C1-12 alkylene chain, wherein 1-6 methylene units are optionally and independently replaced with - O-, -N(Rb)-, -C(O)-, -S-, -S(O)-, -S(O)2-, or phenylene.
67. The method of claim 66, wherein L2is an optionally substituted C1-12 alkylene chain, wherein 1-6 methylene units are optionally and independently replaced with -O-.
68. The method of claim 67, wherein L2is an optionally substituted C2-6alkylene chain, wherein 1-3 methylene units are optionally and independently replaced with -O-.
69. The method of claim 53-60, wherein -Z-L1-L2-R1is selected from the group consisting of70. The method of claim 1, wherein said coupling comprises dipping said substrate into a solution comprising said nucleotide and a template-independent polymerase.
71. The method of claim 1, wherein the blocking group is a polymerase, and wherein the polymerase is linked to the nucleotide via a cleavable linker.
72. The method of claim 71, wherein the cleavable linker comprises an amino acid ester.
73. The method of claim 72, wherein the amino acid ester is attached to an amino acid.
74. The method of claim 73, wherein the amine group of the amino acid ester is bound to the amino acid.
75. The method of claim 74, wherein the cleavable linker comprises a peptide of at least 2, at least 3, at least 4, or at least 5 amino acids bound to the amine group of the amino acid ester.
76. The method of any one of claims 72-75, wherein the amino acid or amino acids is selected from the group consisting of: alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine.
77. The method of any one of claims 72-75, wherein the amino acid is glycine or the amino acids comprise glycine.
78. The method of any one of claims 72-75, wherein the amino acid is a non-naturally occurring amino acid or the amino acids comprise a non-naturally occurring amino acid.
79. The method of any one of claims 71-78, wherein the cleavable linker is bound to the alpha-phosphate, sugar, or nucleobase of the nucleotide.
80. The method of any one of claims 72-79, wherein the amino acid ester is represented by:wherein R1and R1'are each independently selected from hydrogen and an optionally substituted C1-6alkyl, or are optionally taken together with the atom on which they are attached to form an optionally substituted C3-C7 carbocyclic ring.
81. The method of claim 80, wherein the amino acid ester is represented by a compound selected from the group consisting of:
82. The method of any one of claims 72-79, wherein the linker comprises the structure:wherein R1and R1'are each independently selected from hydrogen and an optionally substituted C1-6 alkyl or are optionally taken together with the atom on which they are attached to form an optionally substituted C3-C7 carbocyclic ring; each R2is an optionally substituted group independently selected from the group consisting of hydrogen, C1-6alkyl, phenyl, C1-C6carbocyclic ring and 3-7 heterocyclic ring; each R3is hydrogen or optionally substituted C1-6 alkyl; and n is 1, 2, 3, 4 or 5.
83. The method of claim 82, wherein R3is hydrogen.
84. The method of claim 82 or 83, wherein R2is hydrogen.
85. The method of claim 82 or 83, wherein R2is selected from the group consisting of hydrogen, -Me, -iso-Pr, -sec-butyl, iso-butyl, -CH2Ph, -CH2OH, -CH2SH, -CH2CH2SCH3, - CH2COOH, -CH2CH2COOH, -CH2CONH2, -CH2CH2CONH2, -CH2CH2,CH2CH2NH2,86. The method of any one of claims 82-85, wherein n is 1.
87. The method of any one of claims 82-86, wherein R1and R1'are taken together to form an optionally substituted C3-C7 carbocyclic ring.
88. The method of claim 87, wherein R1and R1'are taken together to form an optionally substituted C3carbocyclic ring.
89. The method of any one of claims 72-79, wherein the linker comprises the structure:.
90. The method of claim 72, wherein the nucleotide linked to the polymerase comprises the structure: Nuc—L1—L2—L3—Pol, wherein: Nuc is the nucleotide; Pol is the polymerase; L1 is a first portion of the linker connecting the nucleotide to L2; L2 is a second portion of the linker represented by:wherein R1and R1'are each independently selected from an optionally substituted C1-6 alkyl, a halogen, or are optionally taken together with the atom on which they are attached to form an optionally substituted C3-C7 carbocyclic ring; each R2is an optionally substituted group independently selected from the group consisting of hydrogen, C1-6 alkyl, phenyl, C1-C6 carbocyclic ring and 3-7 heterocyclic ring; each R3is hydrogen or optionally substituted C1-6 alkyl; n is 0, 1, 2, 3, 4 or 5; wherein * indicates the attachment point of L2 to L1; and ** indicates the attachment point of L2 to L3; wherein L2is cleavable; and L3is a linker connecting pol to L2.
91. The method of claim 90, whereinL1is selected from the group consisting of a bond, an optionally substituted C1-12alkylene chain, C4-C20 polyethylene glycol, an optionally substituted C2-12 alkenylene chain, and a C2- 12 alkynylene chain, wherein 1-6 methylene units of L1are optionally and independently replaced with -O-, -N(Rb)-, -N=C(H)-, -C(O)-, -S-, -S(O)-, -S(O)2-, optionally substituted phenylene, or optionally substituted cyclopropylene.
92. The method of claim 91, wherein L1comprises:wherein each Rais independently selected from the group consisting of halogen, hydroxyl, cyano, optionally substituted C1-6alkyl, and optionally substituted C1-6alkoxy.
93. The method of any one of claims 90-92, wherein L2comprises an amino acid ester selected from the group consisting of:
94. The method of claim 93, wherein L2is represented by:.
95. The method of any one of claims 90-94, wherein L1is bound to the nucleobase of the nucleotide.
96. The method of claim 95, wherein L1is bound to the nucleobase at an oxygen or nitrogen involved in base pairing.
97. The method of claim 96, wherein the nucleobase is selected from the group consisting of:
98. The method of any one of claims 90-94, wherein L1is bound to the sugar of the nucleotide.
99. The method of one of claims 90-94, wherein L1is bound to a phosphate of the nucleotide.
100. The method of claim 99, wherein the phosphate is the alpha phosphate.
101. The method of any one of the above claims, wherein the nucleotide is a ribonucleotide polyphosphate or a deoxyribonucleotide polyphosphate.
102. The method of any one of the above claims, wherein the nucleotide is selected from the group consisting of: adenine, guanine, cytosine, uracil, and thymine.
103. The method of any one of the above claims, wherein the polymerase is a template- independent polymerase.
104. The method of claim 103, wherein the polymerase is TdT.
105. The method of any one of the above claims, wherein the linker is capable of being cleaved by a protease comprising esterase activity.
106. The method of claim 105, wherein the linker is capable of being cleaved by Proteinase K.
107. The method of claim 105 or 106, wherein said linker is capable of being cleaved at the ester group on L2, leaving a compound represented by Nuc-Ll-OH after said cleavage.
108. A substrate comprising a plurality of attached polynucleotides at least 500 nucleotides in length, wherein said plurality of polynucleotides are characterized by sequences generated by a stepwise template-independent polynucleotide synthesis having an observed error rate of less than 1% per cycle as compared to a predetermined nucleotide sequence.
109. The substrate of claim 108, wherein the polynucleotides are at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1800, at least 1900, or at least 2000 nucleotides in length.
110. The method of claim 108, wherein the polynucleotides are from 500 to 1000 nucleotides in length, from 500 to 1500 nucleotides in length, from 500 to 2000 nucleotides in length, or from 500 to 2500 nucleotides in length.
111. The method of claim 108, wherein the observed error rate is less than 0.9%, less than 0.8%, less than 0.7%, less than 0.6%, less than 0.5%, less than 0.4%, less than 0.3%, less than 0.2%, less than 0.1% per cycle, less than 0.09% per cycle, less than 0.08% per cycle, less than 0.07% per cycle, less than 0.06% per cycle, less than 0.05% per cycle, less than 0.04% per cycle, less than 0.03% per cycle, less than 0.02% per cycle, or less than 0.01% per cycle as compared to said predetermined sequence.
112. The method of claim 108, wherein the observed error rate is greater than 0.001%, greater than 0.002%, greater than 0.005%, greater than 0.01%, greater than 0.02%, greater than 0.05%, or greater than 0.1% per cycle as compared to said predetermined sequence.
113. The method of claim 108, wherein said plurality of attached polynucleotides on the substrate comprises at least 10 polynucleotides, at least 20 polynucleotides, at least 50 polynucleotides, at least 100 polynucleotides, at least 200 polynucleotides, at least 500 polynucleotides, at least 1,000 polynucleotides, at least 10,000 polynucleotides, at least 10 x 105polynucleotides, at least 10 x 106polynucleotides, at least 10 x 107polynucleotides, at least 10 x 108polynucleotides, at least 10 x 109polynucleotides, at least 10 x 1010polynucleotides, at least 10 x 1011polynucleotides, at least 10 x 1012polynucleotides, at least 10 x 1013polynucleotides, or at least 10 x 1014polynucleotides.
114. The method of claim 108, wherein said polynucleotides comprise one or more nucleotides comprising a modification to a hydrogen binding N or O on the nucleobase.