Substrate cleavage for nucleic acid synthesis
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-04-06
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for synthesizing and cleaving polynucleotides face challenges such as low yields, harsh conditions, and damage to newly synthesized polynucleotides, particularly when large numbers are synthesized on small devices, complicating analysis and resulting in mixed oligo pools.
A method involving enzymatic and chemical approaches to cleave polynucleotides from a substrate using enzymes and linkers, including photocleavable, acid-labile, and base-labile linkers, with controlled conditions to ensure precise and efficient release.
Enables independent cleavage of polynucleotides from a substrate, allowing access to specific sequences for various applications, reducing damage and improving yield and precision in polynucleotide synthesis and analysis.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] (cross reference) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 328,688, filed April 7, 2022, and U.S. Provisional Patent Application No. 63 / 479,672, filed January 12, 2023, which applications are incorporated by reference in their entireties. [Background technology]
[0002] Biomolecule-based information storage systems, such as DNA-based ones, have large storage capacities and long-term stability. However, the generation of biomolecules for information storage requires scalable, automated, highly accurate and efficient systems.
[0003] (Incorporated by reference) All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Summary of the Invention [Means for solving the problem]
[0004] Provided herein is a method for cleaving polynucleotides, the method comprising: (a) synthesizing a plurality of polynucleotides, each of which comprises one or more bases susceptible to enzymatic cleavage; (b) exposing the plurality of polynucleotides to one or more enzymes; and (c) treating the plurality of polynucleotides in aqueous base at a temperature between about 55 degrees Celsius and 75 degrees Celsius. In some examples, exposing the plurality of polynucleotides to the one or more enzymes comprises exposing the plurality of polynucleotides to a first enzyme of the one or more enzymes. In some examples, exposing the plurality of polynucleotides to the one or more enzymes further comprises exposing the plurality of polynucleotides to a second enzyme of the one or more enzymes. In some examples, the first enzyme and the second enzyme are different enzymes. In some examples, synthesizing comprises enzymatic synthesis or chemical synthesis. In some examples, synthesizing comprises synthesizing the plurality of polynucleotides on a solid support. In some examples, the plurality of polynucleotides are attached to a surface of the solid support via a supported linker. In some examples, the supported linker comprises a stilt. In some examples, the support comprises thymidine. In some examples, the one or more bases comprise deoxyuracil. In some examples, the one or more enzymes comprise one or more of uracil DNA glycosylase, apurinic apyrimidinic site (AP) endonuclease, alkylpurine glycosylase C and D, OGG1, NTH1, NEIL1-3, endonuclease V, or endonuclease VII. In some examples, the plurality of polynucleotides is treated in aqueous base for about 1 hour. In some examples, the temperature is about 65 degrees Celsius. In some examples, the plurality of polynucleotides encodes digital information. In some examples, the digital information comprises text, audio, or visual information.
[0005] Further provided herein is a method for cleaving a polynucleotide, the method comprising: (a) synthesizing a plurality of polynucleotides on a surface of a solid support, the plurality of polynucleotides being attached to the surface via a support linker; and (b) irradiating the plurality of polynucleotides. In some examples, synthesizing comprises enzymatic synthesis or chemical synthesis. In some examples, the support linker comprises a support. In some examples, the support comprises a thymidine. In some examples, the support linker comprises a photocleavable linker. In some examples, the photocleavable linker comprises an ortho-nitrobenzyl-based linker, a phenacyl linker, an alkoxybenzoin linker, a chromarene complex linker, an NpSSMpact linker, or a pivaloyl glycol linker. In some examples, the photocleavable linker is cleaved by irradiating the support linker at about 312 nm, 365 nm, or 405 nm. In some examples, the photocleavable linker is irradiated for about 1 minute to about 15 minutes. In some examples, the plurality of polynucleotides encodes digital information. In some examples, the digital information includes text, audio, or visual information.
[0006] Provided herein is a method for synthesizing a polynucleotide, comprising: a) contacting a polynucleotide with a complex according to the formula: ALB (Formula I) During the ceremony, A comprises a polymerase, B comprises a nucleotide, Methods are provided, including: (b) extending the polynucleotide by addition of a nucleotide, the addition of the nucleotide resulting in a cleavage between the chemical linker and the nucleotide; and (c) cleaving the polymerase from the polynucleotide, the cleavage leaving no part of the linker on the polynucleotide. Further provided herein are methods further comprising cleaving the polynucleotide from the solid support. In some examples, the method further comprises cleaving the polynucleotide from the solid support using a chemical reaction. In some examples, the cleavage of the polynucleotide is independently addressable. In some examples, the chemical reaction comprises acid, base, or electrochemistry. In some examples, the method further comprises generating acid at a region of the surface. In some examples, the acid is generated by applying a potential to a solution containing a mixture of benzoquinone and hydroquinone, or a mixture of their derivatives. In some examples, the support linker includes an aldol, tetrahydrofuran, or trityl group. In some examples, the method further includes generating a base at a region of the surface. In some examples, the base is generated by applying a potential to a solution containing (1) an arene or heteroarene, and (2) a protic solvent. In some examples, the arene or heteroarene includes one or more of substituted or unsubstituted azobenzene, hydrabenzene, azophenanthrene, azonaphthalene, and azopyridine. In some examples, the protic solvent includes an alcohol. In some examples, the base is generated by applying a potential to a solution containing unsubstituted, 1,6 or 2,7 disubstituted phenazine, or tetrasubstituted phenazine with their respective corresponding hydrophenazine compounds. In some examples, the arene or heteroarene includes a phenol group, a cresol group, or a catechol group.In some examples, the arene or heteroarene comprises an amine. In some examples, the arene or heteroarene is substituted with one or more of trifluoromethylsulfonyl, hexafluoropropyl, trifluoromethyl, pentafluorophenyl, and nitrophenyl. In some examples, the arene or heteroarene is substituted with one or more halogens. In some examples, the supported linker comprises an ester. In some examples, the supported linker is cleaved by beta elimination. In some examples, the supported linker comprises an electron withdrawing group. In some examples, the electron withdrawing group comprises a sulfone, fluorine(s), nitro group, sulfonyl, or cyano. In some examples, the supported linker comprises a latent nucleophile. In some examples, the supported linker comprises a levulinyl group. In some examples, the supported linker comprises hydroquinone-O,O-diacetic acid (Q-linker). In some examples, the supported linker comprises an alkyl-substituted silane. In some examples, the method further comprises an electrochemical reaction. In some examples, the supported linker comprises a redox active group. In some examples, the supported linker comprises a metal center. In some examples, the metal center comprises any one of metals in groups 8-10 of the periodic table. In some examples, the supported linker comprises an organoborane. In some examples, the supported linker comprises an aryl or alkyl sulfonate. In some examples, the supported linker comprises a ligand. In some examples, the support comprises a ligand binding agent. In some examples, the method comprises cleaving the polynucleotide from the solid support with an enzyme. In some examples, the supported linker comprises a support. In some examples, the support comprises thymidine. In some examples, the supported linker comprises uracil. In some examples, the carrying linker comprises one or more of 3-methyladenine, 8-oxoguanine, oxoinosine, 2,6-diamino-4-hydroxy-5-formamidopyrimidine (FapyG), 4,6-diamino-5-formamidopyrimidine (FapyA), 5-hydroxyuracil, 5-hydroxymethyluracil, and 5-formyluracil.In some examples, the enzyme comprises one or more of uracil DNA glycosylase, apurinic apyrimidinic site (AP) endonuclease, alkylpurine glycosylase C and D, OGG1, NTH1, NEIL1-3, endonuclease V, or endonuclease VII. In some examples, the method further comprises treating the polynucleotide with aqueous base, heating the polynucleotide, or a combination thereof. In some examples, heating the polynucleotide comprises heating at a temperature of about 55-75 degrees Celsius. In some examples, the supported linker comprises one or more ribonucleosides. In some examples, the one or more ribonucleosides comprise protecting groups at one or both of the 2'OH and 3'OH positions. In some examples, the protecting groups comprise acetyl, benzoyl, trimethylsilyl, TBDMS, TOM, or levulinyl. In some examples, the enzyme comprises RNase H. In some examples, the method further comprises hybridizing a complementary or partially complementary polynucleotide to the support linker. In some examples, the enzyme comprises one or more of thymidine DNA glycosylase (TDG) and methylated CpG binding domain protein 4 (MBD4). In some examples, the enzyme comprises one or more of BamHI, EcoRI, EcoRV, HindIII, and HaeIII. In some examples, steps a)-c) are repeated to generate an extended polynucleotide. In some examples, the extended polynucleotide comprises at least about 10 nucleotides. In some examples, the polymerase is a template-independent polymerase. In some examples, the polymerase is terminal deoxynucleotidyl transferase (TdT) or polymerase theta. In some examples, the chemical linker is an acid-labile linker, a base-labile linker, a pH-sensitive linker, an amine-to-thiol cross-linker, a thiomaleamic acid linker, or a photocleavable linker. In some examples, the photocleavable linker is selected from the group consisting of an ortho-nitrobenzyl-based linker, a phenacyl linker, an alkoxybenzoin linker, a chromarene complex linker, an NpSSMpact linker, a pivaloyl glycol linker, and any combination thereof.In some examples, the chemical linker is selected from the group consisting of a silyl linker, an alkyl linker, a polyether linker, a polysulfonyl linker, a polysulfoxide linker, and any combination thereof. In some examples, the nucleotide comprises at least three phosphate groups. In some examples, the nucleotide is selected from the group consisting of a nucleoside triphosphate, a nucleoside tetraphosphate, a nucleoside pentaphosphate, a nucleoside hexaphosphate, a nucleoside heptaphosphate, a nucleoside octaphosphate, a nucleoside nonaphosphate, and any combination thereof. In some examples, the nucleotide is selected from the group consisting of deoxyadenosine triphosphate (dATP), deoxyguanosine triphosphate (dGTP), deoxycytidine triphosphate (dCTP), deoxythymidine triphosphate (dTTP), deoxyadenosine tetraphosphate, deoxyguanosine tetraphosphate, deoxycytidine tetraphosphate, deoxythymidine tetraphosphate, deoxyadenosine pentaphosphate, deoxyguanosine pentaphosphate, deoxycytidine pentaphosphate, deoxythymidine pentaphosphate, deoxyadenosine hexaphosphate, deoxyguanosine hexaphosphate, deoxycytidine hexaphosphate, deoxythymidine hexaphosphate, and any combination thereof. In some examples, the polynucleotide encodes digital information. In some examples, the digital information comprises text, audio, or visual information.
[0007] Provided herein is a method for synthesizing a polynucleotide, comprising: (a) contacting a polynucleotide with a complex according to the formula: ALB (Formula I) During the ceremony, A comprises a polymerase, B comprises a nucleotide, Methods are provided, including (a) L comprising a chemical linker that covalently bonds the polymerase to the terminal phosphate group of the nucleotide, the polymerase being configured to catalyze the covalent addition of a nucleotide to the 3' hydroxyl of the polynucleotide and subsequent extension of the polynucleotide from the surface of the solid support, the polynucleotide being attached to the surface via the supported linker, and (b) cleaving the polymerase from the polynucleotide, the cleavage not leaving a portion of the linker on the polynucleotide. Further provided herein are methods further comprising cleaving the polynucleotide from the solid support. Further provided herein are methods further comprising cleaving the polynucleotide from the solid support with an enzyme. Further provided herein are methods, wherein the supported linker comprises a support. Further provided herein are methods, wherein the support comprises a thymidine. Further provided herein are methods, wherein the supported linker comprises a uracil. Further provided herein is a method in which the supported linker comprises one or more of 3-methyladenine, 8-oxoguanine, oxoinosine, 2,6-diamino-4-hydroxy-5-formamidopyrimidine (FapyG), 4,6-diamino-5-formamidopyrimidine (FapyA), 5-hydroxyuracil, 5-hydroxymethyluracil, and 5-formyluracil. Further provided herein is a method in which the enzyme comprises one or more of uracil DNA glycosylase, apurinic apyrimidinic site (AP) endonuclease, alkylpurine glycosylase C and D, OGG1, NTH1, NEIL1-3, and endonuclease V. Further provided herein is a method in which the supported linker comprises one or more ribonucleosides. Further provided herein is a method in which the one or more ribonucleosides comprise a protecting group at one or both of the 2'OH and 3'OH positions. Further provided herein are methods, wherein the protecting group comprises acetyl, benzoyl, trimethylsilyl, TBDMS, TOM, or levulinyl. Further provided herein are methods, wherein the enzyme comprises RNase H. Further provided herein are methods, further comprising hybridizing a complementary or partially complementary polynucleotide to the supported linker.Further provided herein is a method, wherein the enzyme comprises one or more of thymidine DNA glycosylase (TDG) and methylated CpG binding domain protein 4 (MBD4). Further provided herein is a method, wherein the enzyme comprises one or more of BamHI, EcoRI, EcoRV, HindIII, and HaeIII. Further provided herein is a method, wherein steps a)-b) are repeated to generate an extended polynucleotide. Further provided herein is a method, wherein the extended polynucleotide comprises at least about 10 nucleotides. Further provided herein is a method, wherein the polymerase is a template-independent polymerase. Further provided herein is a method, wherein the polymerase is terminal deoxynucleotidyl transferase (TdT) or polymerase theta. Further provided herein is a method, wherein the chemical linker is an acid-labile linker, a base-labile linker, a pH-sensitive linker, an amine to thiol cross-linker, a thiomaleamic acid linker, or a photocleavable linker. Further provided herein is a method in which the photocleavable linker is selected from the group consisting of an ortho-nitrobenzyl-based linker, a phenacyl linker, an alkoxybenzoin linker, a chromarene complex linker, an NpSSMpact linker, a pivaloyl glycol linker, and any combination thereof. Further provided herein is a method in which the chemical linker is selected from the group consisting of a silyl linker, an alkyl linker, a polyether linker, a polysulfonyl linker, a polysulfoxide linker, and any combination thereof. Further provided herein is a method in which the nucleotide comprises at least three phosphate groups. Further provided herein is a method in which the nucleotide is selected from the group consisting of a nucleoside triphosphate, a nucleoside tetraphosphate, a nucleoside pentaphosphate, a nucleoside hexaphosphate, a nucleoside heptaphosphate, a nucleoside octaphosphate, a nucleoside nonaphosphate, and any combination thereof.Further provided herein is a method, wherein the nucleotide is selected from the group consisting of deoxyadenosine triphosphate (dATP), deoxyguanosine triphosphate (dGTP), deoxycytidine triphosphate (dCTP), deoxythymidine triphosphate (dTTP), deoxyadenosine tetraphosphate, deoxyguanosine tetraphosphate, deoxycytidine tetraphosphate, deoxythymidine tetraphosphate, deoxyadenosine pentaphosphate, deoxyguanosine pentaphosphate, deoxycytidine pentaphosphate, deoxythymidine pentaphosphate, deoxyadenosine hexaphosphate, deoxyguanosine hexaphosphate, deoxycytidine hexaphosphate, deoxythymidine hexaphosphate, and any combination thereof. In some examples, the polynucleotide encodes digital information. In some examples, the digital information comprises text, audio, or visual information. [Brief description of the drawings]
[0008] [Figure 1A] 1 is a first exemplary scheme for cleaving a supported linker to release a surface-bound nucleic acid. Deprotection of the anomeric hydroxyl group opens the ribose ring, followed by beta-elimination to release the polynucleotide. [Figure 1B] 2 is a second exemplary scheme for cleaving a supported linker to release a surface-bound nucleic acid. Formation of a cyclophosphate at the 2'OH displaces the 5'OH of the polynucleotide, releasing the polynucleotide. [Figure 2A] 1 shows the addition of uracil phosphoramidites to surface-attached thymine pillars. After the enzymatic synthesis step that adds the additional base, an enzyme(s) is used to cleave the synthesized polynucleotide from the surface. [Figure 2B] 1 shows the addition of a protected ribonucleic acid to a surface-attached thymine pillar. After the enzymatic synthesis step that adds an additional base, an enzyme (e.g., base or RNase) is used to cleave the synthesized polynucleotide from the surface. [Diagram 3] 1 is an exemplary workflow for nucleic acid-based information storage according to some embodiments. [Figure 4] FIG. 1 illustrates an example of a computer system according to some embodiments. [Diagram 5] FIG. 1 is a block diagram illustrating the architecture of a computer system according to some embodiments. [Figure 6] FIG. 1 illustrates a network configured to incorporate multiple computer systems, multiple mobile phones and personal data assistants, and network attached storage (NAS). [Figure 7] FIG. 1 is a block diagram of a multi-processor computer system using a shared virtual address memory space according to some embodiments. [Figure 8A] 1 shows an exemplary mechanism for enzymatic cleavage of a polynucleotide according to some embodiments. In some examples, a polynucleotide (A) contains deoxyuracil, which can be cleaved using uracil deglycosylase (B) followed by endonuclease VIII (C). In some examples, the polynucleotide is exposed to one or more enzymes, followed by treatment with aqueous base, heating, or both (C). [Figure 8B] 8A-8C are exemplary LCMS chromatograms from the process depicted in FIG. 8A according to some embodiments. Exposure of polynucleotide (A) to uracil deglycosylase and endonuclease VIII can yield a combination of products B and C, as shown in FIG. 8A (FIG. 8B top). The top chromatogram shows the amount of interaction versus acquisition time (min). Subsequent treatment with aqueous base and heat can increase the yield of product C (FIG. 8B bottom). The bottom chromatogram shows intensity versus time (min). [Figure 9A] 1 shows an exemplary mechanism for cleavage of a photolabile linker on a polynucleotide according to some embodiments, in some embodiments, the photolabile linker is an ortho-nitrobenzyl-based linker that can be cleaved by irradiation at a wavelength of about 365 nm. [Figure 9B]9A-9C are exemplary LCMS chromatograms for various exposure times to irradiation of the polynucleotide shown in FIG. 9A according to some embodiments. Chromatograms are shown for exposure times of 3 minutes (FIG. 9B top), 5 minutes (FIG. 9B bottom), 10 minutes (FIG. 9C top), and 15 minutes (FIG. 9C bottom). Each chromatogram shown in FIG. 9B-9C shows the amount of interaction versus acquisition time (minutes). [Figure 9C] 9A-9C are exemplary LCMS chromatograms for various exposure times to irradiation of the polynucleotide shown in FIG. 9A according to some embodiments. Chromatograms are shown for exposure times of 3 minutes (FIG. 9B top), 5 minutes (FIG. 9B bottom), 10 minutes (FIG. 9C top), and 15 minutes (FIG. 9C bottom). Each chromatogram shown in FIG. 9B-9C shows the amount of interaction versus acquisition time (minutes). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] definition
[0010] Throughout this disclosure, various embodiments are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiment. Thus, the description of a range should be considered to have specifically disclosed all possible subranges within that range, and individual numerical values to the tenth of the unit of the lower limit, unless the context clearly indicates otherwise. For example, the description of a range such as 1-6 should be considered to have specifically disclosed subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, and individual values within that range, such as 1.1, 2, 2.3, 5, 5.9. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may be independently included in the smaller ranges and are encompassed within the scope of the disclosure, subject to any specifically excluded limits in the stated range. Where the stated range includes one or both of the limits, ranges excluding one or both of those included limits are also included in the disclosure, unless the context clearly indicates otherwise.
[0011] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit any embodiment. As used herein, the singular forms "a," "an," and "the" are intended to include the plural unless the context clearly indicates otherwise. Furthermore, it will be understood that the terms "comprise" and / or "comprising" as used herein specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0012] Unless otherwise stated or clear from the context, as used herein, the term "about" in reference to a numerical value or range of numerical values is understood to mean the stated numerical value and + / - 10% of that numerical value, or 10% below the lower limit and 10% above the upper limit of the stated range.
[0013] As used herein, the term "symbol" generally refers to a representation of a unit of digital information. Digital information may be divided or converted into one or more symbols. In one example, a symbol may be a bit, which may have a numerical value. In some examples, a symbol may have a value of "0" or "1." In some examples, digital information may be represented as an array of symbols or a string of symbols. In some examples, the array of symbols or the string of symbols may include binary data.
[0014] Unless otherwise indicated, the term "nucleic acid" as used herein includes single-stranded molecules as well as double-stranded or triple-stranded nucleic acids. In double-stranded or triple-stranded nucleic acids, the nucleic acid strands need not be coextensive (i.e., a double-stranded nucleic acid need not be double-stranded along the entire length of both strands). Nucleic acid sequences are listed in the 5' to 3' direction unless otherwise indicated. The methods described herein provide for the production of isolated nucleic acids. The methods described herein further provide for the production of isolated and purified nucleic acids. A "nucleic acid" as referred to herein can comprise at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, or more bases in length. Further provided herein are methods for synthesizing any number of nucleotide sequences encoding polypeptide segments, including sequences encoding non-ribosomal peptides (NRPs), non-ribosomal peptide synthetase (NRPS) modules and synthetic variants, polypeptide segments of other modular proteins such as antibodies, polypeptide segments from other protein families, including non-coding DNA or RNA, such as regulatory sequences, such as promoters, transcription factors, enhancers, siRNAs, shRNAs, RNAi, miRNAs, small nucleolar RNAs derived from microRNAs, or any functional or structural DNA or RNA unit of interest.Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, intergenic DNA, sequence positions defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), small nucleolar RNA, ribozymes, complementary DNA (cDNA), which is a DNA representation of messenger RNA (mRNA), usually obtained by reverse transcription or amplification of mRNA, DNA molecules produced by synthesis or amplification, genomic DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. cDNAs encoding genes or gene fragments referred to herein may include at least one region encoding an exon sequence without intervening intron sequences in the genomic equivalent. cDNAs described herein may be produced by de novo synthesis.
[0015] Provided herein are methods and compositions for generating polynucleotides. Also provided herein are methods and compositions for cleaving or removing polynucleotides. Polynucleotides may also be referred to as oligonucleotides or oligos.
[0016] Polynucleotide Synthesis
[0017] Polynucleotide synthesis often occurs at the surface of a substrate, such as a discrete locus. Once synthesis is complete, polynucleotides are often cleaved from the surface of the substrate. However, cleavage methods often present challenges, such as low yields, harsh conditions / reagents, and damage to the newly synthesized polynucleotides. Furthermore, large numbers of sequences can be synthesized on devices that are too small to individually cleave the polynucleotides by chemical means. This can complicate analysis and result in mixed oligo pools if all synthesized sequences are cleaved at once.
[0018] Provided herein are compositions and methods that allow for cleavage of polynucleotides from a substrate. In some examples, compositions and methods allow for independent cleavage of polynucleotides from a substrate. Independent cleavage of polynucleotides from a substrate may be performed on a surface that includes addressable sequence locations. Independent cleavage of polynucleotides allows access to specific sequences for different applications (e.g., access to different gene fragments) from the same chip. In some examples, these methods are used in combination with chemical or enzymatic polynucleotide synthesis. Polynucleotides are attached to a substrate or solid support surface via a linker in some examples. The linker may be referred to as a supported linker. In some examples, the methods and compositions provided herein cleave the supported linker to release the polynucleotide. In some examples, the polynucleotide is released into solution. In some examples, chemical or enzymatic methods are used to cleave the supported linker. In some examples, electrochemical methods are used to cleave the supported linker (e.g., acid generation).
[0019] Provided herein are compositions and methods for improving cleavage of polynucleotides from surfaces. In some examples, these methods are used in combination with chemical or enzymatic polynucleotide synthesis. In some examples, the polynucleotide is attached to the surface of a substrate or solid support via a supported linker. In some examples, the methods and compositions provided herein cleave the supported linker to release the polynucleotide into solution. In some examples, chemical or enzymatic methods are used to cleave the supported linker. In some examples, the enzymatic methods used to cleave the supported linker include exposing the supported linker to one or more enzymes (e.g., at least one, two, or three enzymes). Exposing the supported linker to one or more enzymes may be performed sequentially.
[0020] Provided herein are compositions and methods in which a polynucleotide is attached to a surface via a supported linker. In some examples, the supported linker comprises a stilt. In some examples, the stilt comprises one or more thymidines. In some examples, the stilt comprises 1-10 thymidines. In some examples, a uracil is attached to the 3' end of the stilt. In some examples, the desired sequence is enzymatically synthesized from the uracil. In some examples, the synthesized polynucleotide is treated with uracil DNA glycosylase to excise the base leaving an aldehyde anomeric carbon. In some examples, the resulting sugar is treated with a weak base to cleave the strand, leaving 5' and 3' phosphate strands. Alternatively, after base excision, the strand is cleaved by treatment with an apurinic apyrimidinic site (AP) endonuclease. AP classes I-IV are used in some examples to generate alternating phosphorylation or non-phosphorylation of the 3' and 5' ends of the cleaved strand.
[0021] Base excision repair (BER) enzymes may be used for different endogenous targets. In some examples, the carrying linker includes one or more bases configured for removal by BER. In some examples, the carrying linker includes one or more of 3-methyladenine, 8-oxoguanine, oxoinosine, 2,6-diamino-4-hydroxy-5-formamidopyrimidine (FapyG), 4,6-diamino-5-formamidopyrimidine (FapyA), 5-hydroxyuracil, 5-hydroxymethyluracil, and 5-formyluracil. These bases are incorporated using phosphoramidite chemistry, in some examples, with phosphoramidites containing labile base protecting groups that may be cleaved before enzymatic synthesis begins. In some examples, alkylpurines are further excised by alkylpurine glycosylases C and D (AlkC, AlkD). In some examples, bifunctional DNA glycosylases are used. In some examples, the bifunctional glycosylase includes OGG1, NTH1, NEIL1-3, and homologs thereof. In some examples, the use of a bifunctional glycosylase eliminates the need for a second enzymatic treatment. In some examples, the carrying linker includes inosine. In some examples, endonuclease V is used to cleave the inserted inosine. In some examples, uracil deglycosylase is used to cleave the inserted inosine. In some examples, uracil deglycosylase followed by endonuclease VII is used to cleave the inserted inosine.
[0022] In some examples, the polynucleotides are exposed to one or more enzymes and then treated with aqueous base and / or heat for a period of time. In some examples, the aqueous base is NH 3 / CH 3 NH 2In some embodiments, the predetermined time is about 1 hour. In some embodiments, the predetermined time is about 5 minutes, 10 minutes, 15 minutes, 20 minutes, 30 minutes, 45 minutes, 1 hour, 1.5 hours, 2 hours, 3 hours, 4 hours, or 5 hours. In some embodiments, the predetermined time is up to about 5 minutes, 10 minutes, 15 minutes, 20 minutes, 30 minutes, 45 minutes, 1 hour, 1.5 hours, 2 hours, 3 hours, 4 hours, or 5 hours. In some embodiments, the predetermined time is at least about 5 minutes, 10 minutes, 15 minutes, 20 minutes, 30 minutes, 45 minutes, 1 hour, 1.5 hours, 2 hours, 3 hours, 4 hours, or 5 hours. In some embodiments, the predetermined time is about 5-10 minutes, 5-15 minutes, 5-20 minutes, 5-30 minutes, 5 minutes to 1 hour, 10-15 minutes, 10-20 minutes, 10-30 minutes, 10-45 minutes, 10 minutes to 1 hour, 15-20 minutes, 15-30 minutes, 15-45 minutes, 15 minutes to 1 hour, 20-30 minutes, 20-45 minutes, 20 minutes to 1 hour, 30-45 minutes, 30 minutes to 1 hour, 30 minutes to 2 hours, 30 minutes to 3 hours, 45 minutes to 1 hour, 45 minutes to 2 hours, 45 minutes to 3 hours, 1-2 hours, 1-3 hours, 1-4 hours, 1-5 hours, 2-3 hours, 2-4 hours, 2-5 hours, 3-4 hours, 3-5 hours, or 4-5 hours. In some examples, the heating is at a temperature of about 30 to 90 degrees Celsius. In some examples, the heating is to a temperature of about 55-75 degrees Celsius. In some embodiments, the temperature is about 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90 degrees Celsius. In some embodiments, the temperature is at least about 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90 degrees Celsius. In some embodiments, the temperature is up to about 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90 degrees Celsius. In some embodiments, the temperature is about 30-50, 30-60, 30-70, 30-80, 40-60, 40-70, 40-80, 40-90, 45-65, 45-75, 45-85, 50-70, 50-80, 50-90, 55-75, 55-85, 60-80, 60-90, 65-85, or 70-90 degrees Celsius.In some examples, the plurality of polynucleotides is treated with aqueous base and heated at a temperature for a duration (or a predetermined period of time) as provided herein.
[0023] In some instances, the site at which cleavage occurs is further away from the start of enzymatic synthesis. In some instances, the site at which cleavage occurs is about 1, 2, 3, 4, 5, 10, 15, 20, 25, or about 30 bases from the start of enzymatic synthesis. In some instances, the site at which cleavage occurs is at least 1, 2, 3, 4, 5, 10, 15, 20, 25, or at least 30 bases from the start of enzymatic synthesis. In some examples, the site at which cleavage occurs is about 1 to 2, 1 to 3, 1 to 4, 1 to 5, 1 to 10, 1 to 15, 1 to 20, 1 to 25, 1 to 30, 2 to 3, 2 to 4, 2 to 5, 2 to 10, 2 to 15, 2 to 20, 2 to 25, 2 to 30, 3 to 4, 3 to 5, 3 to 10, 3 to 15, 3 to 20, 3 to 25, 3 to 30, 4 to 5, 4 to 10, 4 to 15, 4 to 20, 4 to 25, 4 to 30, 5 to 10, 5 to 15, 5 to 20, 5 to 25, 5 to 30, 10 to 15, 10 to 20, 10 to 25, 10 to 30, 15 to 20, 15 to 25, 15 to 30, 20 to 25, 20 to 30, or 25 to 30 bases from the start position of enzymatic synthesis.
[0024] RNA nucleotides may be incorporated into the supported linkers described herein. In some examples, the supported linker includes an RNA nucleoside at the 3' end of the support. In some examples, treatment of this DNA / RNA hybrid with basic conditions generates a 3'-cyclic phosphate on the support and a 5'-OH on the enzymatically synthesized strand. In some examples, a complementary strand to the region surrounding the excision site is used for enzymatic cleavage. Hybridization of a DNA complement to the support region may selectively cleave a specific enzymatically synthesized sequence using a restriction endonuclease. In some examples, endonucleases include BamHI, EcoRI, EcoRV, HindIII, HaeIII, and the like. In some examples, partially complementary polynucleotides are used. In some examples, mispairings can be introduced in this way, providing a T:G mispair that is excised by thymidine DNA glycosylase (TDG) and / or methylated CpG binding domain protein 4 (MBD4). In some embodiments, multiple RNA bases may be added to the ends of the support. In some instances, addition of a DNA complement to the RNA region in the presence of RNase H cleaves the synthesized nucleic acid from the surface. Any uncleaved RNA that remains is, in some instances, subsequently removed enzymatically or by incubation under basic conditions. In some instances, the RNA nucleoside comprises a 5' protecting group. In some instances, the RNA nucleoside comprises a 3' protecting group. In some instances, the RNA nucleoside comprises a 3' protecting group and a 5' protecting group. In some cases, the protection comprises benzoyl, trimethylsilyl, TBDMS, TOM, or levulinyl. In some cases, the protection is selected from the group consisting of benzoyl, trimethylsilyl, TBDMS, TOM, and levulinyl.
[0025] The supported linkers described herein may include nucleotide analogs that are recognized by specific enzymes. In some examples, the supported linkers include nucleotide analogs. In some examples, the supported linkers include deoxyuridine or 8-oxodeoxyguanosine that are recognized by specific glycosylases (e.g., uracil deoxyglycosylase, followed by endonuclease VIII, and 8-oxoguanine DNA glycosylase, respectively). In some embodiments, cleavage by glycosylases and / or endonucleases may require a double-stranded DNA substrate. In some embodiments, the supported linker comprises a base analogue cleavable by endonuclease III, including, but not limited to, urea, thymine glycol, methyltartonyl urea, alloxan, uracil glycol, 6-hydroxy-5,6-dihydrocytosine, 5-hydroxyhydantoin, 5-hydroxycytosine, trans-1-carbamoyl-2-oxo-4,5-dihydroxyoxyimidazolidine, 5,6-dihydrouracil, 5-hydroxycytosine, 5-hydroxyuracil, 5-hydroxy-6-hydrouracil, 5-hydroxy-6-hydrothymine, 5,6-dihydrothymine. In some embodiments, the supported linker comprises a base analogue cleavable by formamidopyrimidine DNA glycosylase, including but not limited to 7,8-dihydro-8-oxoguanine, 7,8-dihydro-8-oxoinosine, 7,8-dihydro-8-oxoadenine, 7,8-dihydro-8-oxonebularine, 4,6-diamino-5-formamidopyrimidine, 2,6-diamino-4-hydroxy-5-formamidopyrimidine, 2,6-diamino-4-hydroxy-5-N-methylformamidopyrimidine, 5-hydroxycytosine, 5-hydroxyuracil. In some embodiments, the supported linker comprises a base analogue cleavable by hNeil1, including but not limited to guanidinohydantoin, spiroiminodihydantoin, 5-hydroxyuracil, thymine glycol.In some embodiments, the supported linker comprises a base analogue cleavable by thymine DNA glycosylase, including, but not limited to, 5-formylcytosine and 5-carboxycytosine. In some embodiments, the supported linker comprises a base analogue cleavable by human alkyladenine DNA glycosylase, including, but not limited to, 3-methyladenine, 3-methylguanine, 7-methylguanine, 7-(2-chloroethyl)-guanine, 7-(2-hydroxyethyl)-guanine, 7-(2-ethoxyethyl)-guanine, 1,2-bis-(7-guanyl)ethane, 1,N. 6 -Ethenoadenine, 1,N 2 -Ethenoguanine, N 2 ,3-Ethenoguanine, N 2 Examples of suitable linkers include, but are not limited to, 3-ethanoguanine, 5-formyluracil, 5-hydroxymethyluracil, and hypoxanthine. In some embodiments, the supported linker comprises a 5-methylcytosine that is cleavable by 5-methylcytosine DNA glycosylase.
[0026] The polynucleotide may be cleaved from the solid support using a chemically reactive agent. In some embodiments, the supported linker is a disulfide bond that can be cleaved by a reducing agent. In some embodiments, the disulfide supported linker is cleaved using β-mercaptoethanol (βME). In some embodiments, the supported linker is a base-cleavable bond, such as an ester (e.g., a succinate ester). In some embodiments, the supported linker is a base-cleavable linker that can be cleaved, for example, using ammonia or trimethylamine. In some embodiments, the supported linker is a quaternary ammonium salt that can be cleaved, for example, using diisopropylamine. In some embodiments, the supported linker is a urethane that can be cleaved by a base, such as, for example, aqueous sodium hydroxide.
[0027] In some embodiments, the supported linker is an acid-cleavable linker. In some embodiments, the supported linker is a benzyl alcohol derivative. In some embodiments, the acid-cleavable linker can be cleaved using trifluoroacetic acid. In some embodiments, the supported linker is a teicoplanin aglycone that can be cleaved by treatment with trifluoroacetic acid and base. In some embodiments, the supported linker is an acetal or thioacetal that can be cleaved, for example, by trifluoroacetic acid. In some embodiments, the supported linker is a thioether that can be cleaved, for example, by hydrogen fluoride or cresol. In some embodiments, the supported linker is a sulfonyl group that can be cleaved, for example, by trifluoromethanesulfonic acid, trifluoroacetic acid, or thioanisole. In some embodiments, the supported linker contains a nucleophilic cleavage site, such as a phthalimide that can be cleaved, for example, by treatment with hydrazine. In some embodiments, the supported linker can be an ester that can be cleaved, for example, by aluminum trichloride.
[0028] In some embodiments, the supported linker is a phosphorothionate, which can be cleaved by silver or mercury ions. In some embodiments, the supported linker can be a diisopropyldialkoxysilyl group, which can be cleaved by fluoride ions. In some embodiments, the supported linker can be a diol, which can be cleaved by sodium periodate. In some embodiments, the supported linker can be an azobenzene, which can be cleaved by sodium dithionate.
[0029] In some embodiments, the support linker is a photocleavable linker. In some embodiments, the photocleavable linker is an ortho-nitrobenzyl-based linker, a phenacyl linker, an alkoxybenzoin linker, a chromarene complex linker, an NpSSMpact linker, or a pivaloyl glycol linker. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker with a wavelength of about 300-500 nm. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker with a wavelength of about 300-400, 300-450, 300-500, 350-370, 350-400, 350-450, 350-500, 400-420, 400-450, or 400-500 nm. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker with a wavelength of about 312 nm. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker at about 365 nm. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker at about 405 nm. In some embodiments, the photocleavable linker is irradiated for about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes. In some embodiments, the photocleavable linker is irradiated for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes. In some embodiments, the photocleavable linker is irradiated for up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes. In some embodiments, the photocleavable linker is irradiated for about 1-3, 1-5, 1-8, 1-10, 2-4, 2-6, 2-8, 2-10, 3-5, 3-7, 3-9, 3-10, 4-6, 4-8, 4-10, 5-8, 5-10, 6-8, 6-10, 7-9, 7-10, 8-10, or 9-10 minutes.
[0030] In some embodiments, the supporting linker is selected from the group consisting of a silyl linker, an alkyl linker, a polyether linker, a polysulfonyl linker, a polysulfoxide linker, and any combination thereof.
[0031] The supported linker may be used to independently cleave one or more polynucleotides from a surface. In some embodiments, the supported linker is cleaved by generation of acid (e.g., electrochemical acid generation) at a region of the surface. The region may include a location on a feature or sequence of the solid support. In some embodiments, the region is addressable on the solid support. In some embodiments, the acid is generated by applying an electric potential to a solution. In some embodiments, the supported linker is cleaved by generation of base at a region of the surface. In some embodiments, the supported linker is reduced or oxidized to release the biomolecule (e.g., polynucleotide) from the region of the surface. In some examples, the surface is a surface of a solid support provided herein.
[0032] The acid may be generated by applying an electric potential to the solution. In some embodiments, the solution includes a mixture of benzoquinone and / or hydroquinone, or derivatives thereof. In some embodiments, the linker includes an acid-labile linker. The acid-labile linker may be any of those described herein. In some embodiments, the acid-labile linker includes an aldol, a tetrahydrofuran, a trityl group, a chlorotrityl group, a hydroxytrityl group, or other acid-labile protecting group such as a hydrazone, a carbonate ester, a cis-aconityl, an azidomethyl-methyl maleic anhydride linker, a Rink amide linker, a FMOC-PAL linker, a pyrophosphate linker, or any combination thereof.
[0033] A linker (e.g., a supported linker) may be cleaved by the generation of a base at a region of a surface. In some examples, the surface is a surface of a solid support provided herein. The region may include a location on a feature or sequence of the solid support. In some embodiments, the region is addressable on the solid support. A potential may be applied to the solution to reverse the polarity, which may result in the generation of a base when a potential is applied to another solution.
[0034] The base may be generated by applying an electric potential to the solution. In some embodiments, the base is generated using a solution comprising (1) an arene or heteroarene, (2) a protic solvent, or a combination thereof. In some embodiments, the arene or heteroarene comprises a substituted or unsubstituted azobenzene, hydrabenzene, azophenanthrene, azonaphthalene, azopyridine, or any combination thereof. In some embodiments, the solution comprises an azo compound. In some embodiments, the azo compound comprises an aromatic heterocycle. In some embodiments, the solution comprises a hydrazo compound (e.g., hydrazobenzene).
[0035] In some embodiments, the base is generated in a solution containing a phenazine. In some embodiments, the phenazine is unsubstituted. In some embodiments, the phenazine is a 1,6 or 2,7 disubstituted phenazine. In some embodiments, the phenazine is tetrasubstituted. In some embodiments, the solution contains the corresponding hydrophenazine compound.
[0036] In some embodiments, the protic solvent can include an alcohol. In some embodiments, the alcohol is a primary alcohol, a secondary alcohol, or a tertiary alcohol. In some embodiments, the protic solvent is deprotonated. In some embodiments, deprotonation of the protic solvent generates a chemical species that can initiate cleavage of a biomolecule (e.g., a polynucleotide) from the surface of the solid support. In some embodiments, the protic solvent includes one or more compounds. In some embodiments, the one or more compounds include an arene or a heteroarene.
[0037] In some embodiments, the arene or heteroarene comprises a phenol group, a cresol group, a catechol group, or any combination thereof. In some embodiments, the arene or heteroarene comprises an amine. In some embodiments, the pKa of the amino proton of the arene or heteroarene is manipulated by substitution. In some embodiments, the arene or heteroarene is substituted with trifluoromethylsulfonyl, hexafluoropropyl, trifluoromethyl, pentafluorophenyl, nitrophenyl, or any combination thereof. In some embodiments, the arene or heteroarene is substituted with one or more halogens. In some embodiments, the one or more halogens comprise F, Cl, Br, I, or any combination thereof. In some embodiments, the one or more halogens manipulate the pKa of the compound.
[0038] In some embodiments, the linker comprises an ester.
[0039] In some embodiments, the linker is cleaved by beta-elimination. In some embodiments, the linker is cleaved similar to decyanoethylation of the phosphate backbone in phosphoramidite chemistry. In some embodiments, the linker comprises an electron-withdrawing group (EWG). In some embodiments, the EWG comprises a sulfone, fluorine(s), nitro group, sulfonyl, cyano, or any combination thereof.
[0040] In some embodiments, the linker comprises a latent nucleophile. In some embodiments, the latent nucleophile generates a nucleophile upon activation. In some embodiments, activation of the nucleophile causes the linker to self-cleave. In some embodiments, activation of the nucleophile causes cleavage of a biomolecule (e.g., a polynucleotide) from the surface of the solid support.
[0041] In some embodiments, the linker comprises a levulinyl group.
[0042] In some embodiments, the linker comprises hydroquinone-O,O-diacetic acid (Q-linker).
[0043] In some embodiments, the linker comprises an alkyl-substituted silane, hi some embodiments, the alkyl-substituted silane is cleaved by electrochemical generation of an alkoxide.
[0044] The linker may be reduced or oxidized to release the biomolecule (e.g., polynucleotide) from the surface of the solid support. In some embodiments, the linker comprises a redox active group. In some embodiments, the linker comprises a metal center. In some embodiments, the metal center comprises any one of metals in groups 8-10 of the periodic table. In some embodiments, the metal center is catalytically promoting. In some embodiments, the metal center is tethered. In some embodiments, the metal center is untethered.
[0045] In some embodiments, the linker comprises an organoborane. In some embodiments, the linker is cleaved by oxidative elimination followed by reductive elimination. In some embodiments, the linker comprises an aryl, an alkylsulfonate, or a combination thereof. In some embodiments, the aryl or alkylsulfonate oxidatively adds to the metal center.
[0046] In some embodiments, the linker comprises a ligand. In some embodiments, the linker comprises a transition metal complex. In some embodiments, the transition metal complex undergoes oxidation or reduction. In some embodiments, the oxidation or reduction results in a conformational change that results in the release of the ligand-modified biomolecule (e.g., a polynucleotide). In some embodiments, the biomolecule is tethered to the surface of the solid support by linkage. In some embodiments, the biomolecule is released by a deprotonation reaction. In some embodiments, the biomolecule is released by unmasking a ligand that has a lower dissociation constant for the metal center. In some embodiments, the metal center or a complex comprising the metal center is immobilized on a surface. In some embodiments, the metal center or a complex comprising the metal center is free floating in solution.
[0047] In some embodiments, the supported linker comprises an aldol, tetrahydrofuran, chlorotrityl group, hydroxytrityl group, or other acid labile protecting group such as hydrazone, carbonate ester, cis-aconityl, azidomethyl-methyl maleic anhydride linker, Rink amide linker, FMOC-PAL linker, pyrophosphate linker, or any combination thereof. In some embodiments, the supported linker comprises an ester. In some embodiments, the supported linker is cleaved by beta elimination. In some embodiments, the supported linker comprises an electron withdrawing group (EWG). In some embodiments, the EWG comprises a sulfone, fluorine(s), nitro group, sulfonyl, cyano, or any combination thereof. In some embodiments, the supported linker comprises a latent nucleophile. In some embodiments, the supported linker comprises a levulinyl group. In some embodiments, the supported linker comprises hydroquinone-O,O-diacetic acid (Q-linker). In some embodiments, the supported linker comprises an alkyl substituted silane. In some embodiments, the alkyl substituted silane is cleaved by electrochemical generation of an alkoxide. In some embodiments, the supported linker comprises a redox active group. In some embodiments, the supported linker comprises a metal center. In some embodiments, the metal center comprises any one of metals in groups 8-10 of the periodic table. In some embodiments, the supported linker comprises an organoborane. In some embodiments, the supported linker comprises an aryl, alkylsulfonate, or combinations thereof. In some embodiments, the linker comprises a ligand.
[0048] Enzymes may be used for the synthesis of polynucleotides. Terminal deoxynucleotidyl transferase (TdT) is a polymerase that adds deoxynucleotidyl triphosphates (dNTPs) to the 3' end of single-stranded DNA. Disclosed herein is a method for enzymatically synthesizing polynucleotides using TdT. A two-step method is used to extend polynucleotides using TdT-dNTP conjugates consisting of TdT molecules site-specifically labeled with dNTPs via a cleavable linker. The synthesis cycle includes two steps. (1) In the extension step, a DNA primer is exposed to an excess of TdT-dNTP conjugates. When a tethered nucleotide is incorporated into the 3' end of the primer, the conjugates become covalently bound, thereby preventing extension by other TdT-dNTP molecules. Each TdT molecule is conjugated to a single dNTP molecule incorporated into the primer. (2) The deprotection step inactivates excess TdT-dNTP conjugates and cleaves the linkage between the incorporated nucleoside and TdT. Cleavage of the TdT releases the primer for further extension. The two-step process can be repeated to generate defined sequences.
[0049] Described herein are methods for synthesizing polynucleotides that involve the use of conjugates according to the formula: ALB (Formula I) wherein A comprises a polymerase, B comprises a nucleotide, and L comprises a chemical linker that covalently links the polymerase to the terminal phosphate group of the nucleotide, and the polymerase is configured to catalyze the covalent addition of the nucleotide to the 3' hydroxyl of the polynucleotide and subsequent extension of the polynucleotide. After extension of the polynucleotide, the polynucleotide may be cleaved using methods and compositions described herein. In some embodiments, using the compositions and methods described herein, cleavage does not leave a portion of the linker on the polynucleotide. In some examples, the chemical linker and the supported linker are different.
[0050] In some embodiments, the polymerase is site-specifically conjugated to the terminal phosphate group of the phosphorylated nucleoside to form a tethered molecule. The phosphorylated nucleoside is referred to as a nucleotide in some embodiments. When the polymerase incorporates the tethered phosphorylated nucleoside into the primer, the polymerase remains covalently attached to the terminal phosphate group at the 3' end of the primer via the linker, blocking further extension by other polymerase conjugates. The linker can then be cleaved to deprotect the 3' end of the primer for subsequent extension. This process can be repeated to extend the polynucleotide to a desired length and sequence. In some examples, extending the polynucleotide includes incorporating a nucleotide. In some examples, incorporating the nucleotide causes spontaneous cleavage between the linker and the nucleotide, releasing the polymerase, the linker, or both. In some examples, the polymerase is released from the extended polynucleotide after condensation. In some instances, cleavage and release of the polymerase linker 5P occurs spontaneously upon reaction to the 3' end of the polynucleotide.
[0051] In some embodiments, the phosphorylated nucleoside (e.g., nucleotide) to be tethered to the polymerase is a nucleoside that comprises at least one phosphate group. In some embodiments, the nucleoside comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more phosphate groups. In some embodiments, the nucleoside comprises at least three phosphate groups. In some embodiments, the phosphorylated nucleoside is adenosine, cytidine, uridine, or guanosine, each of which comprises at least one phosphate group. In some embodiments, the phosphorylated nucleoside is a deoxynucleoside that comprises at least one phosphate group. In some embodiments, the phosphorylated nucleoside is a deoxynucleoside that comprises at least three phosphate groups. In some embodiments, the deoxynucleoside comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more phosphate groups. In some embodiments, the phosphorylated nucleoside is deoxyadenosine, deoxycytidine, deoxythymidine, or deoxyguanosine, each of which includes at least one phosphate group. In some embodiments, the phosphorylated nucleoside is a nucleoside triphosphate, such as a dNTP. In some embodiments, the phosphorylated nucleoside is a nucleoside tetraphosphate, a nucleoside pentaphosphate, a nucleoside hexaphosphate, a nucleoside heptaphosphate, a nucleoside octaphosphate, or a nucleoside nonaphosphate. In some embodiments, the phosphorylated nucleoside is a nucleoside hexaphosphate. In some embodiments, the phosphorylated nucleoside is a nucleoside triphosphate.In some embodiments, the phosphorylated nucleoside is selected from the group consisting of deoxyadenosine triphosphate (dATP), deoxyguanosine triphosphate (dGTP), deoxycytidine triphosphate (dCTP), deoxythymidine triphosphate (dTTP), deoxyadenosine tetraphosphate, deoxyguanosine tetraphosphate, deoxycytidine tetraphosphate, deoxythymidine tetraphosphate, deoxyadenosine pentaphosphate, deoxyguanosine pentaphosphate, deoxycytidine pentaphosphate, deoxythymidine pentaphosphate, deoxyadenosine hexaphosphate, deoxyguanosine hexaphosphate, deoxycytidine hexaphosphate, deoxythymidine hexaphosphate, and any combination thereof.
[0052] The methods described herein can use polynucleotides enzymatically synthesized using solid phase supports. In some embodiments, the methods of the disclosure can synthesize polynucleotides in wells of a multi-well plate, such as, for example, a 96-well plate or a 384-well plate. In some embodiments, the methods of the disclosure can synthesize polynucleotides using a non-swelling or low-swelling solid phase support. In some embodiments, the methods of the disclosure can synthesize polynucleotides using controlled pore glass (CPG) or microporous polystyrene (MPPS). In some embodiments, the methods of the disclosure can synthesize polynucleotides on CPG treated with a surface coating material. In some embodiments, the methods of the disclosure can synthesize polynucleotides on CPG treated with (3-aminopropyl)triethoxysilane (3-aminopropyl CPG). In some embodiments, the methods of the disclosure can synthesize polynucleotides on long chain amino alkyl (LCAA) CPG. In some embodiments, the methods of the disclosure can synthesize polynucleotides using CPGs with average pore sizes of about 500, about 1000, about 1500, about 2000, or about 3000 Å.
[0053] Provided herein are various surfaces for enzymatically synthesizing polynucleotides. In some embodiments, the surface comprises one or more reverse phosphoramidites. In some embodiments, the surface comprises a linker attached to the surface. In some embodiments, the linker is attached to the surface after treatment with diethylamine. In some embodiments, the surface comprises dT.
[0054] In some embodiments, the surface comprises at least one hydrophilic polymer, which in various embodiments includes polyethylene glycol (PEG), poly(vinyl alcohol) (PVA), poly(vinylpyridine), poly(vinylpyrrolidone) (PVP), poly(acrylic acid) (PAA), polyacrylamide, poly(N-isopropylacrylamide) (PNIPAM), poly(methyl methacrylate) (PMA), poly(2-hydroxyethyl methacrylate) (PHEMA), poly(oligo(ethylene glycol) methyl ether methacrylate) (POEGMA), polyglutamic acid (PGA), polylysine, polyglucosides, streptavidin, and dextran. In some embodiments, the surface comprises polyethylene glycol (PEG).
[0055] In some embodiments, the surface comprises a siloxane monomer or a siloxane polymer. In some embodiments, the siloxane monomer or the siloxane polymer comprises an epoxide functional group. In some embodiments, the siloxane monomer or the polymer thereof comprises one or more monomers selected from (3-glycidylpropyl)trimethoxysilane (GPTMS), diethoxy(3-glycidyloxypropyl)methylsilane, 3-glycidoxypropyldimethoxymethylsilane, 2-(3,4-epoxycyclohexyl)ethyltriethoxysilane, 2-(3,4-epoxycyclohexyl)ethyltrimethoxysilane, or combinations thereof. In some embodiments, the siloxane monomer is GPTMS. In some embodiments, the siloxane monomer is diethoxy(3-glycidyloxypropyl)methylsilane. In some embodiments, the siloxane monomer is 3-glycidoxypropyldimethoxymethylsilane. In some embodiments, the siloxane monomer is 2-(3,4-epoxycyclohexyl)ethyltriethoxysilane. In some embodiments, the siloxane monomer is 2-(3,4-epoxycyclohexyl)ethyltrimethoxysilane.
[0056] In some embodiments, the surface is selected from the group consisting of heptadecafluorodecyltrichlorosilane, poly(tetrafluoroethylene), octadecyltrichlorosilane, methyltrimethoxysilane, nonafluorohexyltrimethoxysilane, vinyltriethoxysilane, paraffin wax, ethyltrimethoxysilane, propyltrimethoxysilane, glass, poly(chlorotrifluoroethylene), polypropylene, poly(propylene oxide), polyethylene, trifluoropropyltrimethoxysilane, 3-(2-aminoethyl)aminopropyltrimethoxysilane, polystyrene, p-tolyltrimethoxy ...
[0033] The present invention may comprise any of the above-mentioned compounds, including but not limited to, methoxysilane, cyanoethyltrimethoxysilane, aminopropyltriethoxysilane, acetoxypropyltrimethoxysilane, poly(methyl methacrylate), poly(vinyl chloride), phenyltrimethoxysilane, chloropropyltrimethoxysilane, mercaptopropyltrimethoxysilane, glycidoxypropyltrimethoxysilane, poly(ethylene terephthalate), copper (dry), poly(ethylene oxide), aluminum, nylon 6 / 6, iron (dry), glass, soda lime (dry), titanium dioxide (anatase), iron(III) oxide, tin oxide, or combinations thereof.
[0057] Provided herein are various supports for enzymatically synthesized polynucleotides. In some embodiments, the polynucleotides described herein are synthesized on one or more solid supports. Exemplary solid supports include, for example, slides, beads, chips, particles, strands, gels, sheets, tubes, spheres, containers, capillaries, pads, slices, films, plates, polymers, or microfluidic devices. Furthermore, the solid supports may be biological, non-biological, organic, inorganic, or combinations thereof. For supports that are substantially planar, the supports may be physically separated into regions, for example, by trenches, grooves, wells, or chemical barriers (e.g., hydrophobic coatings, etc.). The supports may also include physically separated regions incorporated into the surface, optionally spanning the entire width of the surface. Supports suitable for improved oligonucleotide synthesis are further described herein. In some embodiments, polynucleotides are provided on solid supports for use in microfluidic devices, for example, as part of a PCA reaction chamber. In some embodiments, the polynucleotides are introduced into the microfluidic device after synthesis. In some embodiments, the solid phase support is part of or incorporated into a flow cell assembly.
[0058] Provided herein is a device for polynucleotide synthesis. The device can include an addressable solid support for independently cleaving one or more polynucleotides. In some examples, the device can include an addressable region or array location where polynucleotides are synthesized. In some examples, the addressable region or array location is in fluid communication with solvents and other reagents for polynucleotide synthesis and / or subsequent cleavage of one or more polynucleotides from the solid support.
[0059] A solid support for polynucleotide synthesis can include multiple sites (e.g., spots) or locations for synthesis. In some examples, the solid support can be used for storage of polynucleotides. In some examples, the solid support includes up to or about 10,000 x 10,000 locations within an area. In some examples, the solid support includes about 1000-20,000 x about 1000-20,000 locations within an area. In some examples, the solid support includes at least or about 10, 30, 50, 75, 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 12,000, 14,000, 16,000, 18,000, 20,000 by at least or about 10, 30, 50, 75, 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 12,000, 14,000, 16,000, 18,000, 20,000 locations. In some examples, the area is at most 0.25, 0.5, 0.75, 1.0, 1.25, 1.5, or 2.0 square inches. In some examples, the solid support comprises addressable array locations having a pitch of at least or about 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5, 6, 7, 8, 9, 10 um, or more than 10 um. In some examples, the solid support comprises addressable array locations having a pitch of about 5 um. In some examples, the solid support comprises addressable array locations having a pitch of about 2 um. In some examples, the solid support comprises addressable array locations having a pitch of about 1 um. In some examples, the solid support comprises addressable array locations having a pitch of about 0.2 um. In some examples, the solid phase carrier comprises addressable array locations having a pitch of about 0.2 um to about 10 um, about 0.2 to about 8 um, about 0.5 to about 10 um, about 1 um to about 10 um, about 2 um to about 8 um, about 3 um to about 5 um, about 1 um to about 3 um, or about 0.5 um to about 3 um.In some examples, the solid support comprises addressable array locations having a pitch of about 0.1 um to about 3 um. In some examples, the solid support comprises addressable array locations having a pitch of at least about 0.01, 0.02, 0.025, 0.03, 0.04, 0.05, 0.1, 0.15, .02, 0.25, 0.30, 0.35, 0.4, 0.45, 0.5, 0.6, 0.7, 0.8, 0.9, 1 um, or greater than 1 um. In some examples, the solid support comprises addressable array locations having a pitch of about 0.5 um. In some examples, the solid support comprises addressable array locations having a pitch of about 0.2 um. In some examples, the solid support comprises addressable array locations having a pitch of about 0.1 um. In some examples, the solid support comprises addressable array locations having a pitch of about 0.02 um. In some examples, the solid phase carrier comprises addressable array locations having a pitch of about 0.02 um to about 1 um, about 0.02 to about 0.8 um, about 0.05 to about 0.1 um, about 0.1 um to about 1 um, about 0.2 um to about 0.8 um, about 0.3 um to about 0.5 um, about 0.1 um to about 0.3 um, or about 0.05 um to about 0.3 um. In some examples, the solid phase carrier comprises addressable array locations having a pitch of about 0.01 um to about 0.3 um.
[0060] Chemical reactions used in polynucleotide synthesis and / or subsequent cleavage of one or more polynucleotides can be controlled using electrochemistry. In some examples, electrochemical reactions are controlled by any energy source, such as light, heat, radiation, electricity, etc. For example, electrodes are used to control chemical reactions as all or part of the locations on a discrete array on a surface. The electrodes, in some examples, are charged by applying a potential to the electrodes to control one or more chemical steps in polynucleotide synthesis. In some examples, the electrodes are addressable. Any number of chemical steps described herein are controlled by one or more electrodes in some examples. Electrochemical reactions can include oxidation, reduction, acid / base chemistry, or other reactions controlled by electrodes. In some examples, the electrodes generate electrons or protons that are used as reagents for chemical transformations. The electrodes, in some examples, directly generate a reagent, such as an acid. In some examples, the acid is a proton. The electrodes, in some examples, directly generate a reagent, such as a base. Acids or bases are often used to cleave protecting groups or to affect the kinetics of various polynucleotide synthesis reactions, for example, by adjusting the pH of the reaction solution. Electrochemically controlled polynucleotide synthesis reactions, in some instances, include redox-active metals or other redox-active organic materials. In some instances, metal or organic catalysts are used in these electrochemical reactions. In some instances, acids are generated by oxidation of quinones.
[0061] Control of chemical reactions includes, but is not limited to, electrochemical generation of reagents. Chemical reactivity may be indirectly affected by biophysical changes in substrates or reagents via an electric field (or gradient) generated by electrodes. In some examples, substrates include, but are not limited to, nucleic acids. In some examples, an electric field is generated that repels certain reagents or substrates away from or attracts them toward an electrode or surface. Such an electric field is generated in some examples by applying a potential to one or more electrodes. For example, negatively charged nucleic acids are repelled away from a negatively charged electrode surface. In some examples, this repelling or attraction of polynucleotides or other reagents by a local electric field results in the movement of polynucleotides or other reagents into or out of a region of a synthesis device or structure. In some examples, electrodes generate an electric field that repels polynucleotides away from a synthesis surface, structure, or device. In some examples, electrodes generate an electric field that attracts polynucleotides toward a synthesis surface, structure, or device. In some examples, protons are repelled away from the positively charged surface, limiting contact of the protons with the substrate or portions thereof. In some examples, repelling or attractive forces are used to allow or block entry of reagents or substrates to specific regions of the synthesis surface. In some examples, nucleoside monomers are prevented from contacting the polynucleotide chain by applying an electric field near one or both components. Such an arrangement allows gating of certain reagents, which may eliminate the need for protecting groups when controlling the concentration of the reagent and / or substrate or the rate of contact between the reagent and / or substrate. In some examples, unprotected nucleoside monomers are used for polynucleotide synthesis. Alternatively, applying an electric field near one or both components promotes contact of the nucleoside monomer with the polynucleotide chain. Furthermore, applying an electric field to the substrate may change the reactivity or conformation of the substrate. In an exemplary application, an electric field generated by electrodes is used to prevent polynucleotides at adjacent sequence positions from interacting. In some examples, the substrate is a polynucleotide, optionally attached to a surface.Application of an electric field, in some instances, changes the three-dimensional structure of the polynucleotide. Such changes include folding or unfolding of various structures, such as helices, hairpins, loops, or other three-dimensional nucleic acid structures. Such changes are useful for manipulating nucleic acids inside wells, channels, or other structures. In some instances, an electric field is applied to the nucleic acid substrate to prevent secondary structures. In some instances, the electric field eliminates the need for linkers or attachment to solid supports during polynucleotide synthesis.
[0062] Conventional electrochemical acid generation methods often require voltages that are greater than high density transistor devices (e.g., CMOS) can tolerate. In some instances, excessive voltages result in unstable currents and reduced fidelity of deprotection during polynucleotide synthesis. In some instances, the methods described herein are configured to operate at voltages less than 2 volts. In some instances, the methods described herein are configured for voltages of 2.00 volts or less, 1.95 volts or less, 1.9 volts or less, 1.85 volts or less, 1.80 volts or less, 1.75 volts or less, 1.70 volts or less, 1.65 volts or less, 1.60 volts or less, or 1.50 volts or less. In some instances, the methods described herein are configured for voltages of 0.1-2, 0.1-1.5, 1-1.9, 1-1.8, 1-1.7, 1-1.6, or 1-1.5 volts. In some instances, the compositions described herein allow for reduced concentrations of redox compounds compared to conventional methods. In some examples, the compositions described herein allow for a reduction in the concentration of additives, such as a reduction or elimination of the concentration of bases. In some examples, the compositions described herein allow for a reduction in the concentration of additives, such as a reduction or elimination of the concentration of amine bases (e.g., 2,6-lutidine).
[0063] Provided herein are devices for enzymatically synthesized polynucleotides that include layers of material. Such devices may include any number of layers of material, including conductive, semiconductive, or insulating materials. In some examples, the various layers of such devices are combined to form an addressable solid support. The layers or surfaces of such devices may be in fluid communication with solvents, solutes, or other reagents used during polynucleotide synthesis. Further described herein are devices that include multiple surfaces. In some examples, the surfaces include features for polynucleotide synthesis in close proximity to conductive materials. In some examples, the devices described herein include 1, 2, 5, 10, 50, 100, or even thousands of surfaces per device. In some examples, a voltage is applied to one or more layers of the devices described herein to facilitate polynucleotide synthesis. In some examples, a voltage is applied to one or more layers of the devices described herein to facilitate steps of polynucleotide synthesis, such as deprotection. Different layers on different surfaces of different devices are often voltage-applied for different times or at different voltages. For example, a positive voltage is applied to a first layer and a negative voltage is applied to a second layer of the same or a different device. In some examples, one or more layers on different devices are powered while other layers are disconnected from ground. In some examples, the base layer includes additional circuitry such as a complementary metal oxide semiconductor (CMOS) device. In some examples, the various layers of one or more devices are connected laterally through routing and / or vertically using vias. In some examples, the various layers of one or more devices are connected laterally through routing and / or vertically using vias to the CMOS layer. In some examples, the various layers of one or more devices are connected to the CMOS device via wire bonds, pogo pin contacts, or through silicon vias (TSVs).
[0064] The substrates, solid supports, or devices described herein may be fabricated from a variety of materials suitable for the disclosed methods and compositions described herein. In certain embodiments, the materials from which the substrates / solid supports of those included in the present disclosure are fabricated exhibit low levels of oligonucleotide binding. In some circumstances, materials that are transparent to visible and / or ultraviolet light may be used. Sufficiently conductive materials may be utilized, such as materials that can form a uniform electric field across all or a portion of the substrates / solid supports described herein. In some embodiments, such materials may be connected to electrical ground. In some instances, the substrates or solid supports may be thermally conductive or insulating. These materials may be chemically and heat resistant to support chemical and biochemical reactions, such as a series of oligonucleotide synthesis reactions. For flexible materials, materials of interest may include both modified and unmodified nylon, nitrocellulose, polypropylene, and the like. For rigid materials, specific materials of interest include glass, fused silica, silicon, plastics (e.g., polytetrafluoroethylene, polypropylene, polystyrene, polycarbonate, and blends thereof, and the like), and metals (e.g., gold, platinum, and the like). The substrate, solid support or reactor may be made from a material selected from the group consisting of silicon, polystyrene, agarose, dextran, cellulose polymers, polyacrylamide, polydimethylsiloxane (PDMS), and glass.
[0065] In various embodiments, surface modification is used to chemically and / or physically alter a surface by additive or subtractive processes to alter one or more chemical and / or physical properties of the substrate surface or selected sites or regions of the substrate surface. For example, surface modification may include (1) changing the wettability of the surface, (2) functionalizing the surface, i.e., providing, modifying or substituting surface functional groups, (3) defunctionalizing the surface, i.e., removing surface functional groups, (4) otherwise altering the chemical composition of the surface (e.g., etching), (5) increasing or decreasing the surface roughness, (6) providing a coating on the surface, i.e., providing a coating that exhibits wettability different from that of the surface, and / or (7) depositing particulates on the surface.
[0066] Described herein are methods for enzymatically synthesizing polynucleotides. In some embodiments, the methods include using a chain-extending enzyme. In some examples, the chain-extending enzyme is a polymerase. In some examples, the polymerase is a template-independent polymerase. In some examples, the polymerase is an RNA polymerase or a DNA polymerase. In some examples, the polymerase is a DNA polymerase. Examples of DNA polymerases include polA, polB, polC, polD, polY, polX, reverse transcriptase (RT), and high-fidelity polymerases. In some examples, the polymerase is an engineered polymerase.
[0067] In some embodiments, the polymerase includes Φ29, B103, GA-1, PZA, Φ15, BS32, M2Y, Nf, G1, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, L17, ThermoSequenase®, 9°Nm™, Therminator™ DNA polymerase, Tne, Tma, TfI, Tth, TIi, Stoffel fragment, Vent® and Deep Vent® DNA polymerase, KOD DNA polymerase, Tgo, JDF-3, Pfu, Taq, T7 DNA polymerase, T7 RNA polymerase, PGB-D, UlTma DNA polymerase, E. coli DNA polymerase I, E. coli DNA polymerase III, archaeal DP1I / DP2 DNA polymerase II, 9°N DNA polymerase, Taq DNA polymerase, Phusion® DNA polymerase, Pfu DNA polymerase, SP6 RNA polymerase, RB69 DNA polymerase, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, and SuperScript® III reverse transcriptase.
[0068] In some embodiments, the polymerase is DNA polymerase 1 - Klenow fragment, Vent polymerase, Phusion® DNA polymerase, KOD DNA polymerase, Taq polymerase, T7 DNA polymerase, T7 RNA polymerase, Therminator™ DNA polymerase, POLB polymerase, SP6 RNA polymerase, E. coli DNA polymerase I, E. coli DNA polymerase III, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, SuperScript® II reverse transcriptase, or SuperScript® III reverse transcriptase.
[0069] The polymerase molecule used in the methods described herein can be polymerase theta, DNA polymerase, or any enzyme capable of extending a nucleotide chain. In some embodiments, the polymerase is tri29. In some embodiments, the polymerase is a protein with a pocket that functions around a terminal phosphate group, e.g., a triphosphate group.
[0070] In some embodiments, the described methods use TdT with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations to synthesize defined polynucleotides. In some embodiments, the described methods use TdT with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations in surface accessible amino acid residues. In some embodiments, the TdT is a variant of TdT. In some embodiments, the variant of TdT includes a cysteine mutation (e.g., NTT-1). In some embodiments, the variant of TdT is NTT-1, NTT-2, or NTT-3. In some examples, the variant TdT has at least 70%, 80%, 90%, or 95% sequence identity to wild-type TdT.
[0071] In some embodiments, the described methods use polymerase theta with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations to synthesize defined polynucleotides. In some embodiments, the described methods use polymerase theta with 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations in surface accessible amino acid residues. In some embodiments, the polymerase theta is a variant of polymerase theta. In some examples, the variant polymerase theta has at least 70%, 80%, 90%, or 95% sequence identity to wild-type polymerase theta. In some embodiments, the polymerase theta is encoded by POLQ.
[0072] The enzymes described herein (e.g., TdT), in some embodiments, comprise one or more unnatural amino acids. In some examples, the unnatural amino acids comprise a lysine analog, an aromatic side chain, an azide group, an alkyne group, or an aldehyde or ketone group. In some examples, the unnatural amino acids do not comprise an aromatic side chain. In some embodiments, the unnatural amino acids are selected from the group consisting of N6-azidoethoxy-carbonyl-L-lysine (Azk), N6-propargylethoxy-carbonyl-L-lysine (Prak), N6-(propargyloxy)-carbonyl-L-lysine (PrK), p-azidophenylalanine (pAzF), BCN-L-lysine, norbornene lysine, TCO-lysine, methyltetrazine lysine, allyloxycarbonyl lysine, 2-amino-8-oxononanoic acid, 2-Amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, m-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propargyloxyphenylalanine, p-propargyl-phenylalanine, 3-methylphenylalanine, L-dopa, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L -phenylalanine, p-acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphonotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphonoserine, L-3-(2-naphthyl)-L-pyrrole, L-phosphoserine ... 2-amino-3-((2-((3-(benzyloxy)-3-oxopropyl)amino)ethyl)selanyl)propanoic acid, 2-amino-3-(phenylselanyl)propanoic, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine.
[0073] In some embodiments, the enzymes described herein are fused to one or more other enzymes, for example, TdT is fused to another enzyme, such as a helicase.
[0074] Provided herein are various linkers for conjugating an enzyme or other nucleic acid (e.g., polymerase) binding moiety to one or more base pairing moieties, such as modified nucleotides, during enzymatic synthesis of polynucleotides. Conjugation of nucleotides or other base pairing moieties to linkers may be accomplished by any means known in the art of chemical conjugation methods. For example, nucleotides containing base modifications that add free amine groups are envisioned for use in conjugation to the linkers described herein. Primary amines may be attached to bases in such a manner that they can be reacted with, for example, heterobifunctional polyethylene glycol (PEG) linkers to generate nucleotides containing PEG linkers of variable length that will still be properly bound to the enzyme active site. Examples of such amine-containing nucleotides include 5-propargylamino-dNTPs, 5-propargylamino-NTPs, aminoallyl-dNTPs, and aminoallyl-NTPs.
[0075] In some embodiments, the amine-containing nucleotides are suitable for conjugation with PEG-based linkers. The PEG linker may vary in length, e.g., from 1 to 1000, 1 to 500, 1 to 11, 1 to 100, 1 to 50, or 1 to 10 subunits. In some embodiments, the PEG linker comprises fewer than 100 subunits. In some embodiments, the PEG linker comprises more than 100 subunits. In some embodiments, the PEG linker comprises more than 500 subunits. In some embodiments, the PEG linker comprises more than 1000 subunits. In some examples, a suitable PEG linker (or branch thereof) may comprise at least 10 subunits, at least 20 subunits, at least 30 subunits, at least 40 subunits, at least 50 subunits, at least 60 subunits, at least 70 subunits, at least 80 subunits, at least 90 subunits, at least 100 subunits, at least 200 subunits, at least 300 subunits, at least 400 subunits, at least 500 subunits, at least 600 subunits, at least 700 subunits, at least 800 subunits, at least 900 subunits, or at least 1,000 subunits. In some examples, the PEG linker (or branches thereof) includes up to 1,000 subunits, up to 900 subunits, up to 800 subunits, up to 700 subunits, up to 600 subunits, up to 500 subunits, up to 400 subunits, up to 300 subunits, up to 200 subunits, up to 100 subunits, up to 90 subunits, up to 80 subunits, up to 70 subunits, up to 60 subunits, up to 50 subunits, up to 40 subunits, up to 30 subunits, up to 30 subunits, or up to 10 subunits.Any of the lower and upper limits described in this paragraph may be combined to form a range within the disclosure, e.g., in some examples, a suitable PEG linker (or branch thereof) may contain from about 90 subunits to about 400 subunits.
[0076] In some embodiments, the linker (e.g., a PEG linker) has an apparent average molecular weight as measured by mass spectrometry, electrophoresis, size exclusion chromatography, reverse phase chromatography, or any other means known in the art for estimating or measuring the molecular weight of a polymer. In some examples, the apparent average molecular weight of the linker selected for conjugation may be less than about 1,000 Da, less than about 2,000 Da, less than about 3,000 Da, less than about 4,000 Da, less than about 5,000 Da, less than about 7,500 Da, less than about 10,000 Da, less than about 15,000 Da, less than about 20,000 Da, less than about 50,000 Da, less than about 100,000 Da, or less than about 200,000 Da. In some examples, the apparent average molecular weight of the linker selected for conjugation may be greater than about 1,000 Da, greater than about 2,000 Da, greater than about 3,000 Da, greater than about 4,000 Da, greater than about 5,000 Da, greater than about 7,500 Da or more, greater than about 10,000 Da, greater than about 15,000 Da, greater than about 20,000 Da, greater than about 50,000 Da, greater than about 100,000 Da, or greater than about 200,000 Da.
[0077] Examples of other suitable linkers may include, but are not limited to, poly-T and poly-A oligonucleotide chains (e.g., ranging from about 1 base to about 1,000 bases in length), peptide linkers (e.g., polyglycine or polyalanine ranging from about 1 residue to about 1,000 residues in length), or carbon chain linkers (e.g., C6, C12, C18, C24, etc.).
[0078] In some embodiments, the linker contains an N-hydroxysuccinimide ester (NHS) group. In some embodiments, the linker contains a maleimide group. In some embodiments, the linker contains an NHS group and a maleimide group. The NHS group of the linker may then react with a primary amine of a nucleotide or other base pairing moiety to form a covalent bond without modifying or destroying the maleimide group. Such functionalized nucleotides may then be covalently attached to enzymes by reaction of the maleimide group with a cysteine residue of the enzyme.
[0079] Linkage of nucleotides can be achieved by disulfide formation (forming an easily cleavable linkage), amide formation, ester formation, protein-ligand linkages (e.g., biotin-streptavidin linkages), alkylation (e.g., using substituted iodoacetamide reagents), or adduct formation using aldehydes and amines or hydrazines.
[0080] In some embodiments, the linker contains, for example, a maltose group, a biotin group, an O2-benzylcytosine group or a derivative thereof, an O6-benzylguanine group, or an O6-benzylguanine derivative.The NHS group of the linker can then react with a primary amine on the nucleotide to form a covalent bond without modifying or destroying the maltose group, the biotin group, the O2-benzylcytosine group or a derivative thereof, an O6-benzylguanine group, or an O6-benzylguanine derivative.Such functionalized nucleotides can then be covalently or non-covalently bound to enzymes by reaction of the maltose group, the biotin group, the O2-benzylcytosine group or a derivative thereof, an O6-benzylguanine group, or an O6-benzylguanine derivative with a suitable functional group or binding partner attached to the enzyme.
[0081] Since branched PEG molecules allow for simultaneous coupling of proteins, dye(s) and nucleotide(s), multiple aspects of the compositions described herein may be present in a single reagent. Examples of suitable branched PEG molecules include, but are not limited to, PEG molecules that contain at least 4 branches, at least 8 branches, at least 16 branches, or at least 32 branches. Alternatively, it is envisioned that each individual element may be provided separately.
[0082] The length of the linker may vary depending on the type of nucleotide (or other base-pairing moiety) and enzyme (or other nucleic acid-binding moiety). In some examples, the enzyme-binding nucleotide should have a length effective to allow the nucleotide or nucleotide analog to pair with a complementary nucleotide while preventing the nucleotide or nucleotide analog from being incorporated at the 3' end of the polynucleotide. In some examples, the length of the linker in the enzyme-binding nucleotide varies for different nucleotides or nucleotide analogs. In some examples, the length of the linker is defined as the persistence length corresponding to the root mean square (RMS) distance between the two ends of the linker characterized by molecular dynamics simulations, 2D trapping experiments, or ab initio calculations. Such simulations, experiments, and calculations can be based on the statistical distribution of the polymer in a compact state, a collapsed state, or a fluid state depending on the existing solution, suspension, or fluid conditions. In some examples, the linker has a persistence length of 0.1 to 1,000 nm, 0.6 to 500 nm, or 0.6 to 400 nm. In some examples, the linker may have a persistence length defined by 0.6, 3.1, 12.7, 22.3, 31.8, 47.7, 95.5, 190.9, 381.8, 763.8 nm, or 989.5 nm, or a range defined by any two or more of these values, or a range including any two or more of these values. In some examples, the linker provided for one nucleotide may be longer or shorter than the linker provided for another nucleotide.For example, in some instances, dTTP may be attached to a nucleic acid binding moiety via a longer linker than the linker used to tether dGTP, or vice versa.
[0083] In some examples, the linker for connecting the nucleotide to the enzyme is about 0.1-1,000 nm, 0.5-500 nm, 0.5-400 nm, 0.5-300 nm, 0.5-200 nm, 0.5-100 nm, 0.5-50 nm, 0.6-500 nm, 0.6-400 nm, 0.6-300 nm, 0.6-200 nm, 0.6-100 nm, 0.6-50 nm, 1-5 The persistence length may be 1 to 500 nm, 1 to 400 nm, 1 to 300 nm, 1 to 200 nm, 1 to 100 nm, 1.5 to 500 nm, 1.5 to 400 nm, 1.5 to 300 nm, 1.5 to 200 nm, 1.5 to 100 nm, 1.5 to 50 nm, 1 to 50 nm, 5 to 500 nm, 5 to 400 nm, 5 to 300 nm, 5 to 200 nm, 5 to 100 nm, or 5 to 50 nm. In some examples, the linker may have a persistence length of about 0.1, 0.5, 0.6, 1.0, 1.5, 1.8, 2.0, 2.5, 3.0, 3.1, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 12.7, 22.3, 31.8, 47.7, 95.5, 190.9, or 381.8 nm, or a range defined by, or including, any two or more of these values. In some examples, the linker may have a persistence length of greater than about 0.1, 0.5, 0.6, 1.0, 1.5, 1.8, 2.0, 2.5, 3.0, 3.1, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 12.7, 22.3, 31.8, 47.7, 95.5, 190.9, or 381.8 nm. In some examples, the linker may have a persistence length of less than about 5, 10, 20, 30, 40, 50, 60, 80, 100, 200, 300, 400, 500, 700, or 1,000 nm. In some examples, the linker may have a persistence length of 0.1, 0.2, 0.4, 1, 2, 4, 10, 20, 30, 40, 50, 60, 80, 100, 200, 300, 400, 500, 700, or 1,000 nm, or a range defined by, or including, any two or more of these values.
[0084] The polymerase molecules of the present disclosure can be site-specifically conjugated to the terminal phosphate group of a nucleoside to form a tethered molecule via a chemical linker. In some embodiments, the chemical linker is an acid-labile linker. In some embodiments, the chemical linker is a base-labile linker. In some embodiments, the chemical linker can be cleaved by irradiation. In some embodiments, the chemical linker can be cleaved by an enzyme, such as, for example, a peptidase or an esterase. In some embodiments, the chemical linker is a pH-sensitive linker. In some embodiments, the chemical linker is an amine to thiol cross-linker, such as PEG4-SPDP. In some embodiments, the chemical linker is a thiomaleamic acid linker. In some embodiments, the chemical linker is a silane. In some embodiments, the chemical linker is cleavable using pH or fluoride ions.
[0085] A polymerase chemically bound to a nucleotide can be cleaved using a chemically reactive agent. In some embodiments, the chemical linker is a disulfide bond that can be cleaved by a reducing agent. In some embodiments, the disulfide chemical linker is cleaved using β-mercaptoethanol (βME). In some embodiments, the chemical linker is a base-cleavable bond, such as an ester (e.g., a succinate ester). In some embodiments, the chemical linker is a base-cleavable linker that can be cleaved using ammonia or trimethylamine. In some embodiments, the chemical linker is a quaternary ammonium salt that can be cleaved using diisopropylamine. In some embodiments, the chemical linker is a urethane that can be cleaved by a base, such as aqueous sodium hydroxide.
[0086] In some embodiments, the chemical linker is an acid-cleavable linker. In some embodiments, the chemical linker is a benzyl alcohol derivative. In some embodiments, the acid-cleavable linker can be cleaved using trifluoroacetic acid. In some embodiments, the chemical linker is a teicoplanin aglycone that can be cleaved by treatment with trifluoroacetic acid and base. In some embodiments, the chemical linker is an acetal or thioacetal that can be cleaved by trifluoroacetic acid. In some embodiments, the chemical linker is a thioether that can be cleaved by hydrogen fluoride or cresol. In some embodiments, the chemical linker is a sulfonyl group that can be cleaved by trifluoromethanesulfonic acid, trifluoroacetic acid, or thioanisole. In some embodiments, the chemical linker contains a nucleophile-cleavable site, such as phthalimide, that can be cleaved by treatment with hydrazine. In some embodiments, the chemical linker can be an ester that can be cleaved with aluminum trichloride.
[0087] In some embodiments, the chemical linker is a Weinreb amide, which can be cleaved by lithium aluminum hydride. In some embodiments, the chemical linker is a phosphorothionate, which can be cleaved by silver or mercury ions. In some embodiments, the chemical linker can be a diisopropyldialkoxysilyl group, which can be cleaved by fluoride ions. In some embodiments, the chemical linker can be a diol, which can be cleaved by sodium periodate. In some embodiments, the chemical linker can be an azobenzene, which can be cleaved by sodium dithionate.
[0088] In some embodiments, the chemical linker is a photocleavable linker. In some embodiments, the photocleavable linker is an ortho-nitrobenzyl-based linker, a phenacyl linker, an alkoxybenzoin linker, a chromarene complex linker, an NpSSMpact linker, or a pivaloyl glycol linker. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker at about 300-500 nm. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker at about 300-400, 300-450, 300-500, 350-370, 350-400, 350-450, 350-500, 400-420, 400-450, or 400-500 nm. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker at about 312 nm. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker at about 365 nm. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker at about 405 nm. In some embodiments, the photocleavable linker is irradiated for about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes. In some embodiments, the photocleavable linker is irradiated for at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes. In some embodiments, the photocleavable linker is irradiated for up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes. In some embodiments, the photocleavable linker is irradiated for about 1-3, 1-5, 1-8, 1-10, 2-4, 2-6, 2-8, 2-10, 3-5, 3-7, 3-9, 3-10, 4-6, 4-8, 4-10, 5-8, 5-10, 6-8, 6-10, 7-9, 7-10, 8-10, or 9-10 minutes.
[0089] In some embodiments, the chemical linker is selected from the group consisting of a silyl linker, an alkyl linker, a polyether linker, a polysulfonyl linker, a polysulfoxide linker, and any combination thereof.
[0090] In some embodiments, the linker is cleaved by an enzyme. In some embodiments, the enzyme is a protease, esterase, glycosylase, or peptidase. In some embodiments, the cleaving enzyme cleaves a bond within the polymerase. In some embodiments, the cleaving enzyme cleaves the linked nucleoside directly.
[0091] Provided herein are methods for enzymatically synthesizing polynucleotides, comprising using various buffers. In some embodiments, the buffers are used in the coupling reaction, the deprotection reaction, the wash solutions, or combinations thereof. In some embodiments, the buffers are sodium cacodylate, Tris-HCl, MgCl 2 , ZnSO 4 , sodium acetate, or a combination thereof.
[0092] The enzymatic methods described herein can be used to synthesize biopolymers. Biopolymers include, but are not limited to, polynucleotides and oligonucleotides. The polynucleotide sequences described herein may include DNA or RNA, unless otherwise specified. In some examples, the polynucleotide includes RNA. In some examples, the RNA includes short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), double-stranded RNA (dsRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), or heterogeneous nuclear RNA (hnRNA). In some examples, the RNA includes shRNA. In some examples, the RNA includes miRNA. In some examples, the RNA includes dsRNA. In some examples, the RNA includes tRNA. In some examples, the RNA includes rRNA. In some examples, the RNA includes hnRNA. In some examples, the polynucleotide is a phosphorodiamidate morpholino oligomer (PMO), which is a short single-stranded polynucleotide analogue built on a backbone of morpholine rings connected by phosphorodiamidate bonds. In some examples, the RNA comprises an siRNA. In some examples, the polynucleotide comprises an siRNA.
[0093] In some embodiments, the polynucleotide is about 8 to about 50 nucleotides in length. In some embodiments, the polynucleotide is about 10 to about 50 nucleotides in length. In some examples, the polynucleotide is about 10, 15, 18, 20, 22, 25, 30, 35, 40, 45, or 50 nucleotides in length. In some examples, the polynucleotide is about 10 to about 30, about 15 to about 30, about 18 to about 25, about 18 to about 24, about 19 to about 23, or about 20 to about 22 nucleotides in length.
[0094] In some embodiments, the polynucleotide is about 50 nucleotides in length. In some examples, the polynucleotide is about 45 nucleotides in length. In some examples, the polynucleotide is about 40 nucleotides in length. In some examples, the polynucleotide is about 35 nucleotides in length. In some examples, the polynucleotide is about 30 nucleotides in length. In some examples, the polynucleotide is about 25 nucleotides in length. In some examples, the polynucleotide is about 20 nucleotides in length. In some examples, the polynucleotide is about 19 nucleotides in length. In some examples, the polynucleotide is about 18 nucleotides in length. In some examples, the polynucleotide is about 17 nucleotides in length. In some examples, the polynucleotide is about 16 nucleotides in length. In some examples, the polynucleotide is about 15 nucleotides in length. In some examples, the polynucleotide is about 14 nucleotides in length. In some examples, the polynucleotide is about 13 nucleotides in length. In some examples, the polynucleotide is about 12 nucleotides in length. In some examples, the polynucleotide is about 11 nucleotides in length. In some examples, the polynucleotide is about 10 nucleotides in length. In some examples, the polynucleotide is about 8 nucleotides in length. In some examples, the polynucleotide is about 8 to about 50 nucleotides in length. In some examples, the polynucleotide is about 10 to about 50 nucleotides in length. In some examples, the polynucleotide is about 10 to about 45 nucleotides in length. In some examples, the polynucleotide is about 10 to about 40 nucleotides in length. In some examples, the polynucleotide is about 10 to about 35 nucleotides in length. In some examples, the polynucleotide is about 10 to about 30 nucleotides in length. In some examples, the polynucleotide is about 10 to about 25 nucleotides in length. In some examples, the polynucleotide is about 10 to about 20 nucleotides in length. In some examples, the polynucleotide is about 15 to about 25 nucleotides in length.In some examples, the polynucleotide is about 15 to about 30 nucleotides in length. In some examples, the polynucleotide is about 12 to about 30 nucleotides in length.
[0095] In some embodiments, the DNA or RNA is chemically modified. In some embodiments, the polynucleotide comprises natural or synthetic or artificial nucleotide analogs or bases. In some examples, the polynucleotide comprises a combination of DNA, RNA, and / or nucleotide analogs. The polynucleotide may be modified using LNA monomers. In some embodiments, the polynucleotide is modified using MOE, ANA, FANA, PS, or a combination thereof.
[0096] In some examples, the synthetic or artificial nucleotide analog or base comprises a modification at one or more of the ribose moiety, the phosphate moiety, the nucleoside moiety, or a combination thereof. In some embodiments, the nucleotide analog or artificial nucleotide base comprises a nucleic acid in which the 2' hydroxyl group of the ribose moiety is modified. In some examples, the modifications include H, OR, R, halo, SH, SR, NH 2 , N.H.R., N.R. 2or CN, where R is an alkyl moiety. Exemplary alkyl moieties include, but are not limited to, halogen, sulfur, thiol, thioether, thioester, amine (primary, secondary, or tertiary), amide, ether, ester, alcohol, and oxygen. In some examples, the alkyl moiety further comprises a modification. In some examples, the modification includes an azo group, a keto group, an aldehyde group, a carboxyl group, a nitro group, a nitroso group, a nitrile group, a heterocyclic (e.g., imidazole, hydrazino, or hydroxylamino) group, an isocyanate group or a cyanate group, or a sulfur-containing group (e.g., sulfoxide, sulfone, sulfide, and disulfide). In some examples, the alkyl moiety further comprises a heterosubstitution. In some examples, a carbon of a heterocyclic group is substituted with nitrogen, oxygen, or sulfur. In some examples, the heterocyclic substitution includes, but is not limited to, morpholino, imidazole, pyrrolidino.
[0097] Modified polynucleotides may also contain one or more substituted sugar moieties. In some embodiments, modified polynucleotides include one of the following at the 2' position: OH, F, O-, S-, or N-alkyl, O-, S-, or N-alkenyl, O-, S-, or N-alkynyl, or Oalkyl-O-alkyl (wherein alkyl, alkenyl, and alkynyl are substituted or unsubstituted C-CO alkyl or C-CO alkyl). 2 ~C 10 Particularly preferred is O(CH 2 ) n O m CH 3 , O(CH 2 ) n , O.C.H. 3 , O(CH 2 ) n NH 2 , O(CH 2 ) n CH 3 , O(CH 2 ) n O.N.H. 2 , and O(CH 2 ) n ON(CH3 ) 2 where n and m can be from 1 to about 10. In some embodiments, the modified polynucleotide comprises one of the following at the 2' position: C-CO, (lower alkyl, substituted lower alkyl, alkaryl, araalkyl, O-alkaryl or O-araalkyl, SH, SCH 3 , OCN, Cl, Br, CN, CF 3 , OCF 3 , SOCH 3 , S.O. 2 CH 3 , O.N.O. 2 , NO 2 , N 3 , N.H. 2 , heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving groups, reporter groups, intercalators, groups for improving the pharmacokinetic properties of a polynucleotide, or groups for improving the pharmacodynamic properties of a polynucleotide, and other substituents with similar properties. In some embodiments, the modification is 2'-methoxyethoxy (2'-O-CH 2 CH 2 OCH 3 , also known as 2'-O-(2-methoxyethyl) or 2'-MOE), i.e., an alkoxyalkoxy group. A further preferred modification is 2'-dimethylaminooxyethoxy, i.e., O(CH 2 ) 2 ON(CH 3 ) 2 groups (also known as 2'-DMAOE, as described in the Examples herein below), and 2'-dimethylaminoethoxyethoxy (also known in the art as 2'-O-dimethylaminoethoxyethyl or 2'-DMAEOE), i.e., 2'-O-CH 2 -O-CH 2 -N(CH 2 ) 2 Includes.
[0098] In some embodiments, the polynucleotide is one or more of the artificial nucleotide analogs described herein. In some examples, the polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 25, or more of the artificial nucleotide analogs described herein. In some embodiments, the artificial nucleotide analogs include 2'-O-methyl, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl, 2'-deoxy, T-deoxy-2'-fluoro, 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), TO-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-ON-methylacetamide (2'-O-NMA) modifications, LNA, ENA, PNA, HNA, morpholino, methylphosphonate nucleotides, thiolphosphonate nucleotides, 2'-fluoro N3-P5'-phosphoramidites, or combinations thereof. In some examples, the polynucleotide is selected from the group consisting of 2'-O-methyl, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl, 2'-deoxy, T-deoxy-2'-fluoro, 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), T-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE). The polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 25, or more of the artificial nucleotide analogs selected from 2'-O-N-methylacetamide (2'-O-NMA) modifications, LNA, ENA, PNA, HNA, morpholino, methylphosphonate nucleotides, thiolphosphonate nucleotides, 2'-fluoro N3-P5'-phosphoramidites, or combinations thereof. In some examples, the polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 25, or more 2'-O-methyl modified nucleotides.In some examples, the polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 25, or more 2'-O-methoxyethyl (2'-O-MOE) modified nucleotides. In some examples, the polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 25, or more thiol phosphonate nucleotides.
[0099] In some embodiments, the modification is 2'-methoxy (2'-OCH 3 ), 2'-aminopropoxy (2'-OCH 2 CH 2 CH 2 NH 2 ) and 2'-fluoro (2'-F). Similar modifications may also be made at other positions on a polynucleotide, particularly the 3' position of the sugar on the 3' terminal nucleotide or in 2'-5' linked polynucleotides and the 5' position of 5' terminal nucleotide. In some embodiments, polynucleotides include sugar mimetics such as cyclobutyl moieties in place of the pentofuranosyl sugar.
[0100] Polynucleotides may also contain nucleic acid base ("base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleotides include the purine bases adenine (A) and guanine (G) and the pyrimidine bases thymine (T), cytosine (C) and uracil (U). Modified nucleotides include 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and 5-halocytosine, 5-propynyluracil and 5-propynylcytosine, 6-azouracil, 6-azocytosine and 6-azothymine, 5-aminoadenine ... Other synthetic and natural nucleotides include uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine.
[0101] In some embodiments, the polynucleotide backbone is modified. In some embodiments, the polynucleotide backbone includes, but is not limited to, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methylphosphonates and other alkylphosphonates including 3' alkylenephosphonates and chiral phosphonates, phosphinates, phosphoramidates including 3'-aminophosphoramidates and aminoalkylphosphoramidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, and boranophosphates with normal 3'-5' linkages, their 2'-5' linkage analogs, and those with reverse polarity where adjacent pairs of nucleoside units are linked 3'-5' to 5'-3' or 2'-5' to 5'-2'. Also included are various salts, mixed salts, and free acid forms.
[0102] In some embodiments, modified polynucleotide backbones do not contain a phosphorus atom in the backbone and include backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages, including morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide backbones, sulfoxide backbones, and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, amide backbones, and other backbones with mixed N, O, S, and CH2 moieties.
[0103] In some embodiments, the polynucleotide is modified by chemically linking the polynucleotide to one or more moieties or conjugates. Exemplary moieties include, but are not limited to, cholesterol moieties, cholic acid, thioethers such as hexyl-S-tritylthiol, thiocholesterol, aliphatic chains such as dodecanediol or undecyl groups, phospholipids such as dihexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate, polyamine or polyethylene glycol chains, or lipid moieties such as adamantane acetic acid, palmityl moieties, or octadecylamine or hexylamino-carbonyl-toxycholesterol moieties.
[0104] Once the non-natural chemical linker is cleaved from the polynucleotide or polynucleotides, the remaining chemical moiety is referred to as a "scar." In some embodiments, the scarr is an olefinic or alkyne moiety. The methods described herein, in some embodiments, do not leave a scar. In some embodiments, no scarr remains after the bound phosphate is cleaved.
[0105] The methods of enzymatic polynucleotide synthesis disclosed herein can have a coupling efficiency of at least 95%, at least 95.5%, at least 96%, at least 96.5%, at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. In some embodiments, the methods can have a coupling efficiency of at least 99.5%. In some embodiments, the methods can have a coupling efficiency of at least 99.7%. In some embodiments, the methods can have a coupling efficiency of at least 99.9%.
[0106] The methods of enzymatic polynucleotide synthesis disclosed herein can have a coupling efficiency of about 95%, about 95.5%, about 96%, about 96.5%, about 97%, about 97.5%, about 98%, about 98.5%, about 99%, about 99.5%, about 99.6%, about 99.7%, about 99.8%, or about 99.9%. In some embodiments, the methods can have a coupling efficiency of about 99.5%. In some embodiments, the methods can have a coupling efficiency of about 99.7%. In some embodiments, the methods can have a coupling efficiency of about 99.9%.
[0107] The enzymatic polynucleotide synthesis methods described herein can have a total average error rate of less than about 1 in 100 bases, less than about 1 in 200 bases, less than about 1 in 300 bases, less than about 1 in 400 bases, less than about 1 in 500 bases, less than about 1 in 1000 bases, less than about 1 in 2000 bases, less than about 1 in 5000 bases, less than about 1 in 10,000 bases, less than about 1 in 15,000 bases, or less than about 1 in 20,000 bases. In some embodiments, the total average error rate is less than about 1 in 100. In some embodiments, the total average error rate is less than about 1 in 200. In some embodiments, the total average error rate is less than about 1 in 500. In some embodiments, the total average error rate is less than about 1 in 1000.
[0108] The methods of enzymatic polynucleotide synthesis described herein can have a total average error rate of less than about 95%, less than about 96%, less than about 97%, less than about 98%, less than about 99%, less than about 99.5%, less than about 99.6%, less than about 99.7%, less than about 99.8%, or less than about 99.9%. In some embodiments, the methods can have a total average error rate of less than about 99.5%. In some embodiments, the methods can have a total average error rate of less than about 99.7%. In some embodiments, the methods can have a total average error rate of less than about 99.9%.
[0109] The error rate of the methods disclosed herein is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, 99.5% or more of the polynucleotides synthesized. In some embodiments, the error rate is for at least 60% of the polynucleotides synthesized. In some embodiments, the error rate is for at least 80% of the polynucleotides synthesized. In some embodiments, the error rate is for at least 90% of the polynucleotides synthesized. In some embodiments, the error rate is for at least 99% of the polynucleotides synthesized. Individual types of error rate include mismatches, deletions, insertions, and / or substitutions to the polynucleotides synthesized on the substrate. The term "error rate" refers to a comparison of the total amount of biopolymers synthesized to a collection of pre-determined biopolymer sequences.
[0110] The methods of enzymatic polynucleotide synthesis disclosed herein can extend a primer by a single nucleotide in about 1 second (sec) to about 20 seconds. In some embodiments, the methods can extend a single nucleotide in about 1 second to about 5 seconds. In some embodiments, the methods can extend a single nucleotide in about 5 seconds to about 10 seconds. In some embodiments, the methods can extend a single nucleotide in about 10 seconds to about 15 seconds. In some embodiments, the methods can extend a single nucleotide in about 15 seconds to about 20 seconds. In some embodiments, the methods can extend a single nucleotide in about 10 seconds to about 20 seconds.
[0111] The methods of enzymatic polynucleotide synthesis disclosed herein can extend a primer by a single nucleotide in about 1 second (sec), about 2 seconds, about 3 seconds, about 4 seconds, about 5 seconds, about 6 seconds, about 7 seconds, about 8 seconds, about 9 seconds, about 10 seconds, about 11 seconds, about 12 seconds, about 13 seconds, about 14 seconds, about 15 seconds, about 16 seconds, about 17 seconds, about 18 seconds, about 19 seconds, or about 20 seconds. In some embodiments, the method can extend a single nucleotide in about 5 seconds. In some embodiments, the method can extend a single nucleotide in about 10 seconds. In some embodiments, the method can extend a single nucleotide in about 15 seconds. In some embodiments, the method can extend a single nucleotide in about 20 seconds.
[0112] The methods of enzymatic polynucleotide synthesis disclosed herein can extend a polynucleotide by at least about 10 nucleotides per hour, in some examples, the methods extend a polynucleotide by at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or 51 or more nucleotides per hour.
[0113] The synthetic polynucleotides of the present disclosure can be from about 50 bases to about 1000 bases. In some embodiments, the synthetic polynucleotides comprise at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 125, at least 150, at least 175, at least 200, at least 225, at least 250, at least 275, at least 300, at least 325, at least 350, at least 375, at least 400, at least 425, at least 450, at least 475, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, at least 1700, at least 1800, at least 1900, or at least 2000 bases. In some embodiments, the synthesized polynucleotides are about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, about 600, about 700, about 800, about 900, about 1000, about 1250, about 1500, about 1750, about 2000, about 2250, about 2500, about 2750, about 3000, about 3250, about 3500, about 3750, about 4000, about 4250, about 4500, about 4750, about 5000, about 6000, about 7000, about 8000, about 9000, about 10000, about 11000, about 1250, about 1500, about 1750, about 2000, about 2250, about 2500, about 2750, about 3000, about 3250, about 3500, about 3750, about 4000, about 4250, about 4500, about 4750, about 5000, about 6000, about 7000, about 8000, about 10000, about 11000, about 1250, about 1500, about 16 00, about 900, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, about 2100, about 2200, about 2300, about 2400, about 2500, about 2600, about 2700, about 2800, about 2900, about 3000, 4000, 5000 bases, or 5001 or more bases.
[0114] In some embodiments, the polymerase-nucleotide conjugate can include an additional moiety that stops the extension of a nucleic acid when the tethered nucleic acid is incorporated. In some embodiments, a 3'O-modified or base-modified reversible terminator deoxynucleoside triphosphate (RTdNTP) is tethered to the polymerase. In some embodiments, the reversible terminator may be attached to the oxygen atom of the 3-prime hydroxyl group of the nucleotide pentose (e.g., a 3'-O-protected reversible terminator). Alternatively or additionally, the reversible terminator may be attached to the nucleobase of the nucleotide (e.g., a 3'-unprotected reversible terminator). In some embodiments, the reversible terminator nucleotide is a chemically modified nucleoside triphosphate analog that stops the extension when incorporated into a nucleic acid molecule. When a conjugate comprising a polymerase and a RTdNTP is used to extend a nucleic acid, cleavage of the linker and deprotection of the RTdNTP may be required to allow the addition of additional nucleotides to the extended nucleic acid. The reversible terminator may comprise a detectable label. The reversible terminator may comprise an allyl, hydroxylamine, acetate, benzoate, phosphate, azidomethyl, or amide group. The reversible terminator may be removed by treatment with a reducing agent, an acid or base, an organic solvent, an ionic detergent, a photon (photolysis), or any combination thereof.
[0115] In the conjugate, the linker at least connects the alpha phosphate of the nucleotide to the C of the polymerase backbone. α In some embodiments, the polymerase and the nucleotide are covalently bonded such that the bond atom of the nucleotide and the C of the polymerase backbone are bonded together. α The distance between the bond atom of the nucleoside and the C atom of the polymerase backbone is about 4 Å to about 100 Å. α The distance between the bond atom of the nucleoside and the C atom of the polymerase backbone is about 5 Å to about 20 Å. αThe distance between the bond atom of the nucleoside and the C atom of the polymerase backbone is about 20 Å to about 50 Å. α In some embodiments, the distance between the bond atom of the nucleoside and the C atom of the polymerase backbone is about 50 Å to about 75 Å. α The distance between the atoms is about 75 Å to about 100 Å.
[0116] In some embodiments, the linker is linked to the base of the nucleotide at an atom that does not participate in base pairing. α It is the atom that connects the atom to the terminal phosphate group of a nucleotide.
[0117] The linker must be long enough to allow the nucleoside triphosphate access to the active site of the polymerase to which it is tethered, so that the conjugated polymerase can catalyze the addition of the nucleotide to which it is attached to the 3' end of the nucleic acid.
[0118] How to use
[0119] The compositions and methods described herein can be used in nucleic acid assembly. In some embodiments, the nucleic acid is DNA. In some embodiments, the nucleic acid is RNA. In some embodiments, the compositions and methods described herein can be used to assemble nucleic acids that are about 8 to about 100 nucleotides in length. In some embodiments, the compositions and methods described herein can be used to assemble nucleic acids that are about 8 to about 50 nucleotides in length. In some embodiments, the compositions and methods described herein can be used to assemble nucleic acids that are about 50 nucleic acids in length.
[0120] The compositions and methods described herein can be used in place of Gibson assembly. The compositions and methods described herein can be used to join multiple DNA fragments in a single isothermal reaction. In some embodiments, the compositions and methods described herein can be used to combine 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 DNA fragments based on sequence identity. In some embodiments, the compositions or methods described herein can be used to combine 10 DNA fragments. In some embodiments, the compositions or methods described herein can be used to combine 15 DNA fragments. In some embodiments, the compositions or methods described herein can be used to combine 20 DNA fragments. In some embodiments, the DNA fragments to be joined contain about 15, about 20, about 25, about 30, about 35, about 40, about 45, or about 50 base pairs of overlap with adjacent DNA fragments. In some embodiments, the DNA fragments to be joined using the methods described herein contain about 20 base pairs of overlap with adjacent DNA fragments. In some embodiments, the DNA fragments to be joined using the methods described herein contain about a 30 base pair overlap with the adjacent DNA fragments, hi some embodiments, the DNA fragments to be joined using the methods described herein contain about a 40 base pair overlap with the adjacent DNA fragments.
[0121] Described herein are compositions and methods for gene assembly to generate a gene library. A gene library can include a collection of genes. In some embodiments, the collection includes at least 100 different pre-selected synthetic genes, which may be at least 0.5 kb in length and have an error rate of less than 1 per 3000 bp compared to the predetermined sequence comprising the gene. The collection can include at least 100 pre-selected synthetic genes, each of which may be at least 0.5 kb in length. At least 90% of the pre-selected synthetic genes may have an error rate of less than 1 per 3000 bp compared to the predetermined sequence comprising the gene. The desired predetermined sequence may be provided in any manner, typically by a user, such as, for example, a user inputting data using a computerized system. In various embodiments, the synthesized nucleic acid is compared to these predetermined sequences, in some instances, for example, by sequencing at least a portion of the synthesized nucleic acid using next generation sequencing. In some embodiments related to any of the gene libraries described herein, at least 90% of the pre-selected synthetic genes have an error rate of less than 1 per 5000 bp compared to the predetermined sequence comprising the gene. In some embodiments, at least 0.05% of the preselected genes are error free. In some embodiments, at least 0.5% of the preselected genes are error free. In some embodiments, at least 90% of the preselected genes comprise an error rate of less than 1 per 3000 bp compared to a given sequence comprising the gene. In some embodiments, at least 90% of the preselected genes are error free or substantially error free. In some embodiments, the preselected genes comprise a deletion rate of less than 1 per 3000 bp compared to a given sequence comprising the gene. In some embodiments, the preselected genes comprise an insertion rate of less than 1 per 3000 bp compared to a given sequence comprising the gene. In some embodiments, the preselected genes comprise a substitution rate of less than 1 per 3000 bp compared to a given sequence comprising the gene.In some embodiments, the gene libraries described herein further comprise at least 10 copies of each gene. In some embodiments, the gene libraries described herein further comprise at least 100 copies of each gene. In some embodiments, the gene libraries described herein further comprise at least 1000 copies of each gene. In some embodiments, the gene libraries described herein further comprise at least 1,000,000 copies of each gene. In some embodiments, the collection of genes described herein comprises at least 500 genes. In some embodiments, the collection comprises at least 5,000 genes. In some embodiments, the collection comprises at least 10,000 genes. In some embodiments, the preselected genes are at least 1 kb. In some embodiments, the preselected genes are at least 2 kb. In some embodiments, the preselected genes are at least 3 kb. In some embodiments, the predetermined sequence comprises less than an additional 20 bp compared to the preselected genes. In some embodiments, the predetermined sequence comprises less than an additional 15 bp compared to the preselected genes. In some embodiments, at least one of the genes differs from any other gene by at least 0.1%. In some embodiments, each gene differs from any other gene by at least 0.1%. In some embodiments, at least one of the genes differs from any other gene by at least 10%. In some embodiments, each gene differs from any other gene by at least 10%. In some embodiments, at least one of the genes differs from any other gene by at least 2 base pairs. In some embodiments, each gene differs from any other gene by at least 2 base pairs. In some embodiments, the gene libraries described herein further comprise genes less than 2 kb in length that have an error rate of less than 1 per 20,000 bp compared to a preselected sequence of the genes. In some embodiments, a subset of the deliverable genes are covalently linked together.In some embodiments, a first subset of the collection of genes encodes components of a first metabolic pathway comprising one or more metabolic end products. In some embodiments, the gene library described herein further comprises selecting one or more metabolic end products, thereby constructing a collection of genes. In some embodiments, the one or more metabolic end products comprise a biofuel. In some embodiments, a second subset of the collection of genes encodes components of a second metabolic pathway comprising one or more metabolic end products. In some embodiments, the gene library comprises a 100m. 3 In some embodiments, the gene library is located within a space of less than 1 m 3 It is in a space of less than .
[0122] In some examples, a method for constructing a gene library is described herein. The method may include: before a first time point, inputting at least a first gene list and a second gene list into a computer-readable non-transitory medium, where the genes are at least 500bp, and when compiled into a consolidated list, the consolidated list includes at least 100 genes; before a second time point, synthesizing more than 90% of the genes in the consolidated list, thereby constructing a gene library with deliverable genes. In some embodiments, the second time point is less than one month away from the first time point.
[0123] When performing any of the methods of constructing a gene library provided herein, the methods described herein further include delivering at least one gene at a second time point. In some embodiments, at least one of the genes differs from any other gene in the gene library by at least 0.1%. In some embodiments, each gene differs from any other gene in the gene library by at least 0.1%. In some embodiments, at least one of the genes differs from any other gene in the gene library by at least 10%. In some embodiments, each gene differs from any other gene in the gene library by at least 10%. In some embodiments, at least one of the genes differs from any other gene in the gene library by at least 2 base pairs. In some embodiments, each gene differs from any other gene in the gene library by at least 2 base pairs. In some embodiments, at least 90% of the deliverable genes are error free. In some embodiments, the deliverable genes include an error rate of less than 1 / 3000 that results in the production of a sequence that deviates from the sequence of the genes in the integrated list of genes. In some embodiments, at least 90% of the deliverable genes comprise an error rate of less than 1 per 3000 bp that results in the production of sequences that deviate from the sequences of the genes in the integrated list of genes. In some embodiments, the genes in the subset of deliverable genes are covalently linked together. In some embodiments, a first subset of the integrated list of genes encodes a component of a first metabolic pathway that includes one or more metabolic end products. In some embodiments, any of the methods of constructing a gene library described herein further comprises selecting one or more metabolic end products, thereby constructing the first list, second list, or integrated list of genes. In some embodiments, the one or more metabolic end products include a biofuel. In some embodiments, the second subset of the integrated list of genes encodes a component of a second metabolic pathway that includes one or more metabolic end products. In some embodiments, the integrated list of genes comprises at least 500 genes.In some embodiments, the integrated list of genes includes at least 5000 genes. In some embodiments, the integrated list of genes includes at least 10000 genes. In some embodiments, the genes may be at least 1 kb. In some embodiments, the genes are at least 2 kb. In some embodiments, the genes are at least 3 kb. In some embodiments, the second time point is less than 25 days away from the first time point. In some embodiments, the second time point is less than 5 days away from the first time point. In some embodiments, the second time point is less than 2 days away from the first time point. It is noted that any of the embodiments described herein can be combined with any of the methods, devices, or systems provided in the present disclosure.
[0124] In another aspect, a method of constructing a gene library is provided herein. The method includes: inputting a list of genes into a computer-readable non-transitory medium at a first time point; synthesizing more than 90% of the list of genes to construct a gene library that includes deliverable genes; and delivering the deliverable genes at a second time point. In some embodiments, the list includes at least 100 genes, and the genes can be at least 500bp. In still some embodiments, the second time point is less than one month away from the first time point.
[0125] When performing any of the methods for constructing a gene library provided herein, in some embodiments, the methods described herein further include delivering at least one gene at a second time point. In some embodiments, at least one of the genes differs from any other gene in the gene library by at least 0.1%. In some embodiments, each gene differs from any other gene in the gene library by at least 0.1%. In some embodiments, at least one of the genes differs from any other gene in the gene library by at least 10%. In some embodiments, each gene differs from any other gene in the gene library by at least 10%. In some embodiments, at least one of the genes differs from any other gene in the gene library by at least 2 base pairs. In some embodiments, each gene differs from any other gene in the gene library by at least 2 base pairs. In some embodiments, at least 90% of the deliverable genes are error free. In some embodiments, the deliverable genes include an error rate of less than 1 in 3000 that results in the generation of a sequence that deviates from the sequence of the gene in the gene list. In some embodiments, at least 90% of the deliverable genes include an error rate of less than 1 per 3000 bp that results in the generation of a sequence that deviates from the sequence of the gene in the gene list. In some embodiments, the genes in the subset of deliverable genes are covalently linked together. In some embodiments, the first subset of the gene list encodes a component of a first metabolic pathway comprising one or more metabolic end products. In some embodiments, the method of constructing a gene library further comprises selecting one or more metabolic end products, thereby constructing a list of genes. In some embodiments, the one or more metabolic end products comprise biofuels. In some embodiments, the second subset of the gene list encodes a component of a second metabolic pathway comprising one or more metabolic end products. It should be noted that any of the embodiments described herein can be combined with any of the methods, devices, or systems provided in the present disclosure.
[0126] When performing any of the methods for constructing a gene library provided herein, in some embodiments, the list of genes comprises at least 500 genes. In some embodiments, the list comprises at least 5000 genes. In some embodiments, the list comprises at least 10000 genes. In some embodiments, the genes are at least 1 kb. In some embodiments, the genes are at least 2 kb. In some embodiments, the genes are at least 3 kb. In some embodiments, the second time point described in the method for constructing a gene library is less than 25 days away from the first time point. In some embodiments, the second time point is less than 5 days away from the first time point. In some embodiments, the second time point is less than 2 days away from the first time point. It should be noted that any of the embodiments described herein can be combined with any of the methods, devices, or systems provided in this disclosure.
[0127] The compositions and methods described herein can be used for DNA digital data storage. In some embodiments, the compositions and methods disclosed herein can be used to prepare DNA molecules for 4-bit information coding. An exemplary workflow is shown in FIG. 3. In a first step, a digital sequence (i.e., digital information in binary code for processing by a computer) encoding an information item is received 301. An encryption 302 scheme is applied to convert the digital sequence from binary code to a nucleic acid sequence 303. Design of surface material for nucleic acid extension, locations on the sequence for nucleic acid extension (aka placement spots), and reagents for nucleic acid synthesis are selected 304. The surface of the structure is prepared for nucleic acid synthesis 305. De novo polynucleotide synthesis is performed 306. The polynucleotide may be about 8-300 bases in length. In some examples, the polynucleotide is about 8, 10, 50, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, or 300 bases in length. In some examples, the polynucleotide is up to about 8, 10, 50, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, or 300 bases in length. In some examples, the polynucleotide is at least about 8, 10, 50, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, or 300 bases in length. In some examples, the polynucleotides are about 10-100, 10-150, 10-200, 50-100, 50-150, 50-200, 100-150, 100-200, 100-300, 150-200, 150-250, 150-300, or 200-300 bases in length. The synthesized polynucleotides are stored 307 and available for subsequent release 308 in whole or in part. For example, selected polynucleotides may be independently cleaved and released from the surface. In some examples, the polynucleotides are stored on the surface on which they were synthesized. However, in alternative cases, the polynucleotides are released from the synthesis surface and stored in an alternative environment (e.g., a storage vessel).Once released, the polynucleotides are sequenced in whole or in part 309 and decoded 310, and the nucleic acid sequences are converted back into digital sequences 311. The digital sequences are then assembled 311 to obtain an alignment that encodes the original information items.
[0128] Nucleic acid-based information storage Provided herein are devices, compositions, systems, and methods for nucleic acid-based information (data) storage. In some examples, biomolecules synthesized and / or extracted from a substrate using the methods and compositions described herein may encode information for DNA data storage. Biomolecules such as DNA molecules provide suitable hosts for storing information, such as digital information, due in part to their stability over time and enhanced information coding capabilities, as opposed to traditional binary information coding. Furthermore, biomolecules such as DNA molecules can provide high volumetric storage density. In a first step, a digital sequence is received that encodes an information item (e.g., digital information in binary code for processing by a computer). The digital sequence may include a first plurality of symbols, such as binary data, octal data, decimal data, or hexadecimal data. An encryption scheme is applied to convert the digital sequence from a first string of symbols to a second string of symbols. The second string of symbols may include alternate representations of the first string of symbols. In some examples, the second string of symbols includes a nucleic acid sequence.
[0129] Once the information items have been converted into nucleic acid sequences, the nucleic acids can be synthesized. The design of the surface material for nucleic acid extension, the locations on the sequence for nucleic acid extension (also known as placement spots), and the reagents for nucleic acid synthesis are selected. The surface of the structure is prepared for nucleic acid synthesis. De novo polynucleotide synthesis is then performed. The synthesized polynucleotides can be extracted in whole or in part using the systems, devices, methods, or platforms provided herein. The synthesized polynucleotides are stored in the structure and, in some instances, are available for later release in whole or in part. The synthesized polynucleotides may be stored in a structure suitable for long-term storage (e.g., weeks, months, years, etc.). The structures suitable for long-term storage may be identifiable and / or cataloguable, for example, by using tags (e.g., barcodes or tags, etc.). Once released, the polynucleotides are sequenced in whole or in part and decoded to convert the nucleic acid sequence back into a digital sequence. The digital sequence is then assembled to obtain an alignment that codes for the original information items.
[0130] Information item Optionally, an initial step of the data storage process disclosed herein includes obtaining or receiving one or more information items in the form of an initial code. In some examples, the information items are encoded as a plurality of polynucleotides extracted from a substrate using a system, method, platform, or device provided herein. Information items (e.g., digital information) include, but are not limited to, text, audio, and visual information. Exemplary sources of information items include, but are not limited to, books, periodicals, electronic databases, medical records, letters, forms, audio recordings, animal registries, biological profiles, broadcasts, movies, short videos, emails, bookkeeping phone logs, Internet activity logs, drawings, paintings, printouts, photographs, pixelated graphics, software code, and the like. Exemplary biological profile sources of information items include, but are not limited to, gene libraries, genomes, gene expression data, and protein activity data. Exemplary formats for an item of information include, but are not limited to, .txt, .PDF, .doc, .docx, .ppt, .pptx, .xls, .xlsx, .rtf, .jpg, .gif, .psd, .bmp, .tiff, .png, and .mpeg. The amount of individual file sizes encoding the items of information in digital form, or the amount of multiple files encoding the items of information, may include, but are not limited to, up to 1024 bytes (equivalent to 1 KB), 1024 KB (equivalent to 1 MB), 1024 MB (equivalent to 1 GB), 1024 GB (equivalent to 1 TB), 1024 TB (equivalent to 1 PB), 1 exabyte, 1 zettabyte, 1 yottabyte, 1 xenottabyte, or more. In some examples, the amount of digital information is at least 1 gigabyte (GB). In some examples, the amount of digital information is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 gigabytes, or more than 1000 gigabytes. In some examples, the amount of digital information is at least 1 terabyte (TB).In some examples, the amount of digital information is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 terabytes, or more than 1000 terabytes. In some examples, the amount of digital information is at least 1 petabyte (PB). In some examples, the amount of digital information is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 petabytes, or more than 1000 petabytes. In some examples, the digital information does not contain genomic data obtained from an organism. The information item is, in some examples, coded. Non-limiting examples of encoding methods include 1 bit / base, 2 bits / base, 4 bits / base, or other encoding methods.
[0131] Sequencing
[0132] Polynucleotides are extracted and / or amplified from the surface where they are synthesized or stored. After extracting and / or amplifying the polynucleotides from the surface of the structure, the polynucleotides may be sequenced using a suitable sequencing technique. In some examples, the DNA sequence is read on the substrate or within the features of the structure. In some examples, polynucleotides stored on the substrate are extracted, optionally assembled into longer polynucleotides, and then sequenced. Polynucleotides may be extracted from the substrate using the systems and methods described herein.
[0133] Polynucleotides synthesized and stored on the structures described herein encode data that can be interpreted by reading the sequence of the synthesized polynucleotide and converting the sequence into computer readable binary code. In some instances, the sequence requires assembly, and assembly steps may need to be performed at the nucleic acid sequence stage or at the digital sequence stage.
[0134] Provided herein is a detection system that includes a device capable of sequencing stored polynucleotides either directly on the synthesis structure and / or after removal from a host structure (e.g., synthesis structure, storage structure, etc.). When the synthesis structure is a reel-to-reel tape of flexible material, the detection system includes a device that holds and advances the structure through a detection location, and a detector positioned near the detection location to detect a signal emanating from a section of the tape when that section is at the detection location. In some examples, the signal indicates the presence of the polynucleotide. In some examples, the signal indicates the sequence of the polynucleotide (e.g., a fluorescent signal). In some examples, the information encoded within the polynucleotide on the continuous tape is read by a computer as the tape is continuously transported past a detector operably connected to the computer. In some examples, the detection system comprises a computer system including a polynucleotide sequencing device, a database for storage and retrieval of data related to the polynucleotide sequence, software for converting the DNA code of the polynucleotide sequence to binary code, a computer for reading the binary code, or any combination thereof.
[0135] Provided herein is a sequencing system that can be integrated into the devices described herein. Various methods of sequencing are well known in the art, including "base calling," in which the identity of a base in a target polynucleotide is identified. In some examples, polynucleotides synthesized using the methods, devices, compositions, and systems described herein are sequenced after cleavage from the synthesis surface. In some examples, sequencing is performed during or simultaneously with polynucleotide synthesis, and base calling is performed immediately after or before the extension of nucleoside monomers into a growing polynucleotide chain. Methods for base calling include measuring the current / voltage generated by the polymerase-catalyzed addition of a base to a template strand. In some examples, the synthesis surface includes an enzyme, such as a polymerase. In some examples, such an enzyme is tethered to an electrode or synthesis surface. In some examples, the enzyme includes a terminal deoxynucleotidyl transferase or a variant thereof.
[0136] In some examples, the polynucleotides cleaved from the substrate surface or the amplified polynucleotides can be processed by techniques such as conventional sequencing or massively parallel sequencing. Sequencing can be performed by a variety of methods available in the art, including, for example, methods that include the incorporation of one or more end-of-read nucleotides, such as Sanger sequencing, which can be performed by Applied Biosystems' SeqStudio® Genetic Analyzer. In other embodiments, sequencing can include performing next-generation sequencing (NGS) methods, such as primer extension followed by semiconductor-based detection (e.g., Thermo Fisher Scientific's Ion Torrent™ system) or fluorescence detection (e.g., Illumina system).
[0137] Computer Systems
[0138] Any of the systems described herein may be operably linked to a computer and may be automated via a computer, either locally or remotely. In various cases, the methods and systems of the present disclosure may further include software programs on a computer system and their use. Thus, computer control for synchronization of dispense / vacuum / replenish functions, such as coordinating and synchronizing the movement, dispense operation, and vacuum operation of the material deposition device, is within the scope of the present disclosure. The computer system may be programmed to interface between the user-specified base sequence and the location of the material deposition device to deliver the correct reagent to the specified area of the substrate. The computer system may also be programmed to independently address one or more areas of a solid support as provided herein.
[0139] The computer system 400 shown in FIG. 4 may be understood as a logical device that can read instructions from a medium 411 and / or a network port 405 and can be connected to a server 409, optionally having a fixed medium 412. A system as shown in FIG. 4 may include a CPU 401, a disk drive 403, optional input devices such as a keyboard 415 and / or a mouse 416, and an optional monitor 407. Data communication may be achieved to a server at a local or remote location over a designated communication medium. A communication medium may include any means of transmitting and / or receiving data. For example, a communication medium may be a network connection, a wireless connection, or an Internet connection. Such a connection may provide communication over the World Wide Web. It is envisioned that data related to the present disclosure may be transmitted over such a network or connection for receipt and / or review by a party 422, as shown in FIG. 4.
[0140] FIG. 5 is a block diagram illustrating a first exemplary architecture of a computer system 500 that can be used in connection with an exemplary embodiment of the present disclosure. As shown in FIG. 5, the exemplary computer system can include a processor 502 for processing instructions. Non-limiting examples of processors include Intel Xeon™ processors, AMD Opteron™ processors, Samsung 32-bit RISC ARM 1176JZX(F)-S v1.0™ processors, ARM Cortex-A8 Samsung S5PC100™ processors, ARM Cortex-A8 Apple A4™ processors, Marvell PXA 930™ processors, or functionally equivalent processors. Multiple execution threads can be used for parallel processing. In some examples, multiple processors or processors with multiple cores can also be used, either within a single computer system, within a cluster, or distributed among systems on a network, including multiple computers, mobile phones, and / or personal data assistant devices.
[0141] As shown in FIG. 5, a high speed cache 504 may be connected to or incorporated into the processor 502 to provide high speed memory for instructions or data recently or frequently used by the processor 502. The processor 502 is connected to a north bridge 506 by a processor bus 508. The north bridge 506 is connected to a random access memory (RAM) 510 by a memory bus 512 and manages access to the RAM 510 by the processor 502. The north bridge 506 is also connected to a south bridge 514 by a chipset bus 516. The south bridge 514 is in turn connected to a peripheral bus 518. The peripheral bus may be, for example, a PCI, PCI-X, PCI Express, or other peripheral bus. The north bridge and south bridge are often referred to as the processor chipset and manage data transfers between the processor, the RAM, and peripheral components on the peripheral bus 518. In some alternative architectures, instead of using a separate north bridge chip, the functionality of the north bridge may be incorporated into the processor. In some examples, the system 500 may include an accelerator card 522 attached to the peripheral bus 518. An accelerator may include a field programmable gate array (FPGA) or other hardware to speed up a particular process. For example, an accelerator may be used for adaptive data restructuring or to evaluate algebraic expressions used in extended set processing.
[0142] Software and data may be stored on external storage 524 and loaded into RAM 510 and / or cache 504 for use by the processor. System 500 includes an operating system for managing system resources, non-limiting examples of which include Linux, Windows™, MACOS™, Blackberry OS™, iOS™, and other functionally equivalent operating systems, as well as application software that runs on the operating system to manage data storage and optimization according to exemplary embodiments of the present disclosure. In this example, system 500 also includes network interface cards (NICs) 520 and 521 connected to the peripheral bus to provide a network interface to external storage, such as network attached storage (NAS) and other computer systems that can be used for distributed parallel processing.
[0143] FIG. 6 illustrates a network 600 with multiple computer systems 602a and 602b, multiple mobile phones and personal data assistants 602c, and network attached storage (NAS) 604a and 604b. In an example, the systems 602a, 602b, and 602c can manage data storage and optimize data access for data stored in the network attached storage (NAS) 604a and 604b. Mathematical models can be used on the data and evaluated using distributed parallel processing across the computer systems 602a, 602b, and the mobile phone and personal data assistant system 602c. The computer systems 602a and 602b, and the mobile phone and personal data assistant system 602c can also provide parallel processing for adaptive data reconstruction of data stored in the network attached storage (NAS) 604a and 604b. FIG. 6 illustrates only one example, and a variety of other computer architectures and systems can be used in combination with various aspects of the present disclosure. For example, blade servers can be used to provide parallel processing. The processor blades can be connected through a backplane to provide parallel processing. Storage can also be connected to the backplane or through another network interface as network attached storage (NAS). In some cases, the processors can maintain separate memory spaces and send data through a network interface, backplane, or other connector for parallel processing by other processors. In other cases, some or all of the processors can use a shared virtual address memory space.
[0144] FIG. 7 is a block diagram of a multiprocessor computer system using a shared virtual address memory space according to one example. The system includes multiple processors 702a-f that can access a shared memory subsystem 704. The system incorporates multiple programmable hardware memory algorithm processors (MAPs) 706a-f in the memory subsystem 704. Each MAP 706a-f can include a memory 708a-f and one or more field programmable gate arrays (FPGAs) 710a-f. The MAPs provide configurable functional units and can provide specific algorithms or parts of algorithms to the FPGAs 710a-f for processing in close cooperation with the respective processors. For example, the MAPs can be used to evaluate algebraic expressions related to a data model and, in the exemplary case, perform adaptive data restructuring. In this example, each MAP is globally accessible to all processors for these purposes. In one configuration, each MAP can access the associated memory 708a-f using direct memory access (DMA), allowing it to perform tasks asynchronously and independently of the respective microprocessors 702a-f. In this configuration, a MAP can feed results directly to another MAP for pipelined and parallel execution of algorithms.
[0145] The above computer architectures and systems are merely examples, and a variety of other computer, cell phone, and personal data assistant architectures and systems can be used in conjunction with the exemplary embodiments, including systems using any combination of general purpose processors, co-processors, FPGAs and other programmable logic devices, systems on a chip (SOC), application specific integrated circuits (ASICs), and other processing and logic elements. In some examples, the computer system can be implemented in whole or in part in software or hardware. Any of a variety of data storage media can be used in conjunction with the exemplary embodiments, including random access memory, hard drives, flash memory, tape drives, disk arrays, network attached storage (NAS), and other local or distributed data storage devices and systems.
[0146] In an exemplary form, the computer system may be implemented using software modules executing on any of the above or other computer architectures and systems. In other examples, the functionality of the system may be implemented partially or fully in firmware, programmable logic devices such as a field programmable gate array (FPGA) as referenced in FIG. 5, a system on a chip (SOC), an application specific integrated circuit (ASIC), or other processing and logic elements. For example, the set processor and optimizer may be implemented with hardware acceleration through the use of a hardware accelerator card such as accelerator card 522 shown in FIG. 5. EXAMPLES
[0147] The following examples are presented to more clearly illustrate to one of ordinary skill in the art the principles and practice of the embodiments disclosed herein, and should not be construed as limiting the scope of the claimed embodiments. Unless otherwise specified, all parts and percentages are by weight.
[0148] Example 1: Single-stranded chain extension using dN6P substrate and TdT enzyme TdT was used for single-strand elongation. dNTP-TdT conjugates were constructed with modifications according to the general method of Palluk, et al., 2018, "De novo DNA synthesis using polymerase-nucleotide conjugates", Nat. Biotechnol. 36, 645-650. A scarless linker was incorporated.
[0149] Briefly, TdT was incubated with single-stranded DNA, manganese, and the dA6P (deoxyadenosine hexaphosphate) substrate. Because no blocking group was used at the 3' end, multiple dAs were added.
[0150] The TdT cysteine variant NTT-1 was also used for single strand extension. Using such NTT-TIDES conjugates, NTT-1 was found to exhibit extension activity. Enzymatic synthesis was then performed on the surface. Briefly, reverse phosphoramidites (phosphoramidites on the 5' hydroxyl monomer) were used as well as diethylamine to gently remove the cyanoethyl group, leaving the linker bond in the appropriate position. dT was also used for successful extension. Single strand extension was also performed using dATP and dA6P.
[0151] Example 2: Uracil-mediated substrate cleavage Polynucleotide synthesis is performed on the surface. In enzymatic DNA synthesis, extension is usually performed from 5' to 3', and synthesis begins with a natural or natural-like nucleic acid strand as a substrate for a terminal transferase (e.g., TdT). The generation of this strand on the surface is in some instances performed by chemical synthesis using a reverse thymidine phosphoramidite in the 5' to 3' direction. The strand is in some instances treated with a base such as diethylamine or other substituted amine to remove the cyanoethyl protecting group, leaving a tethered natural DNA strand. The strand may also be prepared with a 5' modification that can then be reacted with the surface. This conjugation may be thiol / maleimide, NHS ester / amine, copper-assisted or copper-free Huisgen cycloaddition, TCO / tetrazine. The strand can then be acted upon by a terminal transferase, which in some instances results in cleavage of the entire strand, leaving an oligothymidine "pillar." This example describes a method for cleaving an enzymatically derived oligonucleotide from a chemically synthesized "pillar."
[0152] If deoxyuracil is chemically synthesized as the last nucleotide at the 3' end of the strut, enzymatic synthesis may be initiated upon TdT recognition of this nucleotide to extend the chain (Figure 2A). After the desired sequence is enzymatically synthesized, the base is excised by treatment with uracil DNA glycosylase, leaving an aldehyde anomeric carbon. This sugar can then be treated with a weak base to cleave the strand, leaving a 5' phosphate and a 3' phosphate strand. Alternatively, after base excision, the strand is cleaved by treatment with an apurinic-apyrimidinic site (AP) endonuclease. AP classes I-IV may be used to phosphorylate or dephosphorylate alternating 3' and 5' ends of the cleaved strand.
[0153] Base excision repair (BER) enzymes may be used for different endogenous targets. These targets are "damaged" bases such as 3-methyladenine, 8-oxoguanine, 2,6-diamino-4-hydroxy-5-formamidopyrimidine (FapyG), 4,6-diamino-5-formamidopyrimidine (FapyA), 5-hydroxyuracil, 5-hydroxymethyluracil, 5-formyluracil. These bases may be incorporated using phosphoramidite chemistry with phosphoramidites containing a labile base protecting group that may be cleaved before the start of enzymatic synthesis. Alkylpurines may be further excised by alkylpurine glycosylases C and D (AlkC, AlkD). Bifunctional DNA glycosylases such as OGG1, NTH1, NEIL1-3, and their homologs may also be used so that secondary enzyme treatment is not necessary. Endonuclease V may be used to cleave the inserted inosine. In some examples, the site where cleavage occurs is further away from the start position of enzymatic synthesis.
[0154] Example 3: Enzymatic substrate cleavage via uracil Following the general procedure of Example 2, a polynucleotide having deoxyuracil (A in Figure 8A) was synthesized. After the desired sequence was synthesized, treatment with uracil deglycosylase excised the base, leaving the aldehyde anomeric carbon (B in Figure 8A). After base excision, the strand was cleaved by treatment with endonuclease VIII (C in Figure 8A). Analysis of the reaction results by LCMS showed both the intermediate product B and the cleavage product C (upper in Figure 8B).
[0155] Subsequent treatment was carried out for 1 hour using an aqueous base (NH 3 / CH 3 NH 2 ) and heating (65 degrees Celsius). Analysis of the reaction results showed an increase in the yield of the cleavage product (lower in Figure 8B).
[0156] Example 4: Substrate cleavage via ribonucleotides The general procedure of Example 2 may be carried out with modifications to also incorporate RNA nucleotides at the 3' end of the strut (Figure 2B). Treatment of this DNA / RNA hybrid with basic conditions generates a 3'-cyclic phosphate on the strut and a 5'-OH on the enzymatically synthesized strand. In many of these embodiments, many of these enzymatic cleavage pathways require a complementary strand to the region surrounding the excision site. Mismatches can also be introduced in this way, providing T:G mismatches that are excised by thymidine DNA glycosylase (TDG) and / or methylated CpG binding domain protein 4 (MBD4).
[0157] In some embodiments, multiple RNA bases may be added to the ends of the struts. Addition of a DNA complement to the RNA region in the presence of RNase H cleaves the synthesized nucleic acid from the surface. Any uncleaved RNA still present can be subsequently removed enzymatically or by incubation under basic conditions. Hybridization of a DNA complement to the strut region allows selective cleavage of specific enzymatically synthesized sequences using restriction endonucleases such as BamHI, EcoRI, EcoRV, HindIII, HaeIII, etc.
[0158] Example 5: Substrate cleavage via electrochemically generated acid or base The general procedure of Example 2 is carried out with a modification, where cleavage of the polynucleotide from the surface is achieved through the use of an acid- or base-sensitive linker connecting the polynucleotide to the surface.
[0159] In one embodiment, the acid is generated by applying a potential to a solution containing a mixture of benzoquinone and hydroquinone. The acid-labile linker may include aldol or tetrahydrofuran based linkers, trityl or variously substituted trityl based linkers.
[0160] In another embodiment, the base is generated in a solution containing unsubstituted or 1,6 or 2,7 disubstituted phenazine or tetrasubstituted phenazine with their respective corresponding hydrophenazine compounds. The protic solvent in the solution is a primary alcohol, a secondary alcohol, or a tertiary alcohol. Deprotonation of these compounds generates species that can initiate cleavage from the surface. These molecules may be phenolic, cresol, or catechol. The molecules may also be amine-based, in which case the pKa of the amino proton can be manipulated by various substituents, including but not limited to trifluoromethylsulfonyl, hexafluoropropyl, trifluoromethyl, pentafluorophenyl, or nitrophenyl, and optionally contain various numbers of halogens to manipulate the pKa of each compound.
[0161] Example 6: Substrate cleavage using a redox-active linker The general procedure of Example 5 is carried out with modifications, where the linker contains a redox-active chemical group. The linker may be cleaved by a (3) elimination reaction in a manner similar to decyanoethylation of the phosphate backbone in standard phosphoramidite chemistry. The linker may contain electron-withdrawing functional groups, including but not limited to sulfone, fluorine(s), nitro, sulfonyl, or cyano. The linker may have a "bite" on itself. The linker may be cleaved by unmasking the internal nucleophile "back" resulting in the release of the biological on non-biological molecule of interest. The linker may include a levulinyl fragment or moiety. The linker may be an ester derivative of hydroquinone-O,O-diacetic acid (Q-linker). The linker may be a variety of alkyl substituted silanes that may be cleaved by electrochemical generation of an alkoxide. The linker may be cleaved by an active metal center that may be generated by oxidation or reduction of the metal center. The metal may belong to, but is not limited to, groups 8-10 of the periodic table. The linker may contain an organoborane that may be cleaved by a mechanism of oxidative elimination followed by reductive elimination (considering Suzuki coupling, etc.). The linker may be oxidatively added to an electrochemically generated metal center. The linker may consist of a suitable aryl or alkyl sulfonate. The linker itself may contain a transition metal complex that undergoes a conformational change upon oxidation or reduction to release the ligand-modified biomolecule. The linker may contain one or more embedded or pendant redox-active molecules, such as quinones, imides, carbazole viologens, organosulfur compounds, triphenylamines, ferrocene, or radical compounds, such as nitroxyl, phenoxyl, verdazyl groups, with stable charge / discharge voltages and high reactivity. The biomolecule may be tethered to the surface by a ligation reaction that can be competed by deprotonation or otherwise unmasking of a ligand with a lower kD to the metal center. These metal complexes can be immobilized on the surface of the device or can be freely floating in solution.
[0162] Example 7: Substrate cleavage using a photocleavable linker The general procedure of Example 2 is modified and polynucleotide synthesis is carried out on the surface. The chain can also be prepared with a 5' modification, which can then be reacted with a suitably modified surface. The conjugation can be thiol / maleimide, NHS ester / amine, copper-assisted or copper-free Huisgen cycloaddition, TCO / tetrazine. The carrying linker contains one or more photocleavable units. In some embodiments, the photocleavable linker is an ortho-nitrobenzyl-based linker, a phenacyl linker, an alkoxybenzoin linker, a chromarene complex linker, an NpSSMpact linker, or a pivaloyl glycol linker. In some embodiments, the photocleavable linker can be cleaved by irradiating the linker at about 312 nm, 365 nm, or about 405 nm (e.g., FIG. 9A). In enzymatic DNA synthesis, extension is usually performed from 5' to 3', and synthesis begins with a natural or natural-like nucleic acid strand as a substrate for terminal transferase (e.g., TdT). Generation of this strand on a surface is in some instances achieved by chemical synthesis using a reverse thymidine phosphoramidite in the 5' to 3' direction. This strand is in some instances treated with a base such as diethylamine or other substituted amine to remove the cyanoethyl protecting group, leaving a tethered natural DNA strand. This strand can then be acted upon by terminal transferase, which in some instances results in cleavage of the entire strand, leaving an oligothymidine "pillar." Additionally, methods are described herein for cleaving enzymatically derived oligonucleotides from chemically synthesized "pillars" at photocleavable sites introduced into the support linker.
[0163] Example 8: Substrate cleavage using an ortho-nitrobenzyl-based photolabile linker Following the general procedure of Example 7, an ortho-nitrobenzyl-based linker was used as the supported linker (Figure 9A). The sample contained 1 uM of polynucleotide (A) with a photocleavable linker in 100 uL of pH 7.0 buffer. The sample was exposed to a wavelength of 365 nm to cleave the linker (B) and analyzed by LCMS.
[0164] Samples were irradiated for 3 min (Figure 9B top), 5 min (Figure 9B bottom), 10 min (Figure 9C top), and 15 min (Figure 9C bottom). As shown in the LCMS chromatograms, with increasing exposure time, the peak corresponding to the cleavage product (B) increased and the peak corresponding to the uncleaved polynucleotide (A) decreased. After about 10 min of exposure time, about 95% cleavage of the ortho-nitrobenzyl-based linker was achieved (Figure 9C top).
[0165] While preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed in the practice of the disclosure. The following claims define the scope of the disclosure, and it is intended that methods and structures within the scope of these claims, and their equivalents, be covered thereby.
Claims
1. A method for cleaving polynucleotides, (a) Synthesizing multiple polynucleotides, each containing one or more bases that are easily cleaved enzymatically, (b) Exposing the plurality of polynucleotides to one or more enzymes, (c) A method comprising treating the plurality of polynucleotides in an aqueous base at a temperature of about 55 to 75 degrees Celsius.
2. The method according to claim 1, wherein exposing the plurality of polynucleotides to one or more enzymes includes exposing the plurality of polynucleotides to a first enzyme among the one or more enzymes.
3. The method according to claim 2, wherein exposing the plurality of polynucleotides to one or more enzymes further comprises exposing the plurality of polynucleotides to a second enzyme among the one or more enzymes.
4. The method according to claim 3, wherein the first enzyme and the second enzyme are different enzymes.
5. The method according to claim 1, wherein synthesis includes enzymatic synthesis or chemical synthesis.
6. The method according to claim 1, wherein synthesis includes synthesizing the plurality of polynucleotides on a solid support.
7. The method according to claim 6, wherein the plurality of polynucleotides are attached to the surface of the solid support via a supported linker.
8. The method according to claim 7, wherein the supporting linker includes a support column.
9. The support column is the method according to claim 8, wherein the support column contains thymidine.
10. The method according to claim 1, wherein the one or more bases include deoxyuracil.
11. The method according to claim 1, wherein the one or more enzymes include one or more of uracil DNA glycosylase, depurine-depyrimidine site (AP) endonuclease, alkylpurine glycosylases C and D, OGG1, NTH1, NEIL1-3, endonuclease V, or endonuclease VII.
12. The method according to claim 1, wherein the plurality of polynucleotides are treated in the aqueous base for about 1 hour.
13. The method according to claim 1, wherein the temperature is approximately 65 degrees Celsius.
14. The method according to claim 1, wherein the aqueous base comprises NH3 / CH3NH2.
15. The method according to claim 6, further comprising cleaving the plurality of polynucleotides from the solid support.
16. The method according to claim 7, wherein the supported linker comprises one or more of uracil, 3-methyladenine, 8-oxoguanine, oxoinosine, 2,6-diamino-4-hydroxy-5-formamidopyrimidine (FapyG), 4,6-diamino-5-formamidopyrimidine (FapyA), 5-hydroxyuracil, 5-hydroxymethyluracil, or 5-formyluracil.
17. Hybridizing a complementary polynucleotide to the supported linker, Using a restriction endonuclease, one or more polynucleotides from the plurality of polynucleotides are cleaved from the surface of the solid support. The method according to claim 7, further comprising:
18. The method according to claim 7, wherein the supported linker comprises one or more ribonucleosides having a protecting group at one or both of the 2'OH and 3'OH positions, and the one or more enzymes comprises RNase H.
19. The method according to claim 6, wherein the solid support comprises an addressable gene locus, and the cleavage of polynucleotides from the solid support is independently addressable.
20. The method according to claim 1, wherein the synthesis includes enzymatic synthesis, and the cleavage site of the synthesized polynucleotide is 1 to 30 bases from the start site of enzymatic synthesis.