Compositions and methods for solution-phase, phosphoramidite-free synthesis of nucleic acids
A solution-phase, phosphoramidite-free method using addamers and restriction enzymes enables efficient and error-reduced synthesis of long DNA sequences, addressing the limitations of existing technologies while promoting environmental sustainability.
Patent Information
- Application Number
- PCT/IB2024/000703
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-12
AI Technical Summary
Existing methods for synthesizing nucleic acids often rely on solid supports and phosphoramidite chemistry, which can be costly, error-prone, and environmentally unfriendly, limiting the ability to generate long sequences efficiently.
A solution-phase, phosphoramidite-free method using addamers, which are short, exonuclease-resistant, double-stranded double-hairpin structures, to assemble DNA sequences in solution, employing Type II S restriction endonucleases and nickases to manipulate and ligate DNA fragments.
This method allows for the efficient and cost-effective synthesis of long DNA sequences with reduced error rates, eliminating the need for solid supports and phosphoramidite reagents, thereby enhancing environmental sustainability.
Smart Images

Figure IB2024000703_12062025_PF_FP_ABST
Abstract
Description
COMPOSITIONS AND METHODS FOR SOLUTION-PHASE, PHOSPHORAMIDITE- FREE SYNTHESIS OF NUCLEIC ACIDSCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 608,002, filed December 8, 2023, which is hereby incorporated by reference in its entirety.BACKGROUND
[0002] The invention described herein describes a gene assembly method, in particular double stranded DNA assembly methods referred to as geometric synthesis (gSynth) methods. Environmentally-friendly (e.g., “green”) double-stranded DNA synthesis methods are described herein using double-stranded DNA constructs referred to as “addamers”. Addamers are short, exonuclease resistant, double-stranded double-hairpin structures that carry a DNA payload as well as a variety of control elements that allow the manipulation of the sequence. The methods have previously relied on solid supports to generate new sequences in a phosphoramidite-free way. The invention described herein discloses a cost effective, efficient, and environmentally conscious process for phosphoramidite synthesis. The invention describes herein allows for longer sequence with reduced risk of error. The compositions and methods allow for the generation of sequences of DNA at any arbitrary’ length in solution. Furthermore, described herein are addamer reagents generated from bacterial plasmids as well as bacterially expressed, purified enzymes used to generate a sequence.SUMMARY
[0003] Described herein, in some aspects, is a composition comprising: (a) a double stranded DNA sequence comprising: (i) an N-mer sequence; (ii) at least one offset-cutting nickase DNA binding sequence configured to bind a nickase that cleaves a first strand or a second strand of the double stranded DNA sequence; and (iii) at least one Type II S restrictionendonuclease (IISRE) DNA binding sequence configured to bind a IISRE that cleaves the first stand and the second strand of the N-mer sequence, wherein the nickase DNA binding sequence is located between the N-mer sequence and the IISRE DNA binding sequence.
[0004] In some aspects, a nickase.
[0005] In some aspects, the nickase and the double stranded DNA sequence is in a single volume.
[0006] In some aspects, the double stranded DNA sequence is not coupled to a solid support.
[0007] In some aspects, a Type II S restriction endonuclease (IISRE).
[0008] In some aspects, the nickase cleaved site is within or adjacent to the N-mer sequence.
[0009] In some aspects, the nickase DNA binding sequence comprises enzymes: BbvCI,BsmI, BsrDI, BssSI, BtsI, Alwl, BbvCI, BsmAI, BspQI, or BstNBI.
[0010] In some aspects, the IISRE DNA binding sequence comprises: Mlyl, NgoAVII, SspD5I, Alwl, Ajul, Alol, BccI, Bcefl, Piel, BceAI, BceSIV, BscAI, BspD6I, Faul, Earl, BspQI, BfuAI, PaqCI, Esp3I, BbsI, Bbvl, BtgZI, FokI, BsmFI, Bsal, BcoDI, Hgal, or SfaNI.
[0011] In some aspects, the IISRE DNA binding sequence comprises Earl.
[0012] In some aspects, at least one hairpin structure.
[0013] In some aspects, at least one hairpin structure comprises an aptamer sequence.
[0014] In some aspects, a second IISRE DNA binding sequence.
[0015] In some aspects, the second IISRE DNA binding sequence comprises a BspQI DNA binding sequence.
[0016] In some aspects, a core melting temperature of the double stranded DNA sequence is at least 60°C. In some aspects, a core melting temperature of the double stranded DNA sequence is between 50°C to 70°C. In some aspects, a core melting temperature of the double stranded DNA sequence is about 50°C. 51 °C. 52°C, 53°C, 54°C. 55°C, 56°C, 57°C. 58°C, 59°C, 60°C. 61°C, 62°C, 63°C. 64°C, 65°C, 66°C, 67°C. 68°C, 69°C, or 70°C.
[0017] Described herein, in some aspects, is a method of synthesizing a target sequence in solution comprising: (a) providing a plurality of DNA sequences in solution, wherein a DNA sequence of the plurality of DNA sequences comprises aN-mer, a IISRE sequence, and at least one hairpin structures; (b) providing at least one IISRE to the solution, wherein the at least IISRE cleaves the IISRE sequence, thereby exposing an M-base overhang; (c) ligating at least two DNA sequences of the plurality of the DNA sequences having the M-base overhang to generate one or more ligated DNA sequences, wherein a ligated DNA sequence comprises a sequence having a length of 2N-M; and (d) repeating (a)-(c) using the one or more ligated DNA sequences to generate a final DNA sequence comprising a target sequence.
[0018] In some aspects, M is 3. In some aspects, M is at least 3.
[0019] In some aspects, the N-mer is at least a 4-mer. In some aspects, the N-mer is no more than the 4-mer.
[0020] In some aspects, the target sequence comprises a length of about 50 bases to 5000 bases. In some aspects, the target sequence comprises the length of about 100 bases to 5000 bases. In some aspects, the length is about 50 bases, 60 bases, 70 bases, 80 bases, 90 bases, 100 bases, 200 bases, 300 bases, 400 bases, 500 bases, 600 bases, 700 bases, 800 bases, 900 bases, 1000 bases, 1100 bases, 1200 bases, 1300 bases, 1400 bases, 1500 bases, 1600 bases, 1700 bases, 1800 bases, 1900 bases, 2000 bases, 2100 bases, 2200 bases, 2300 bases, 2400 bases, 2500 bases, 2600 bases, 2700 bases, 2800 bases, 2900 bases, 3000 bases, 3100 bases, 3200 bases, 3300 bases, 3400 bases, 3500 bases, 3600 bases, 3700 bases, 3800 bases, 3900 bases, 4000 bases, 4100 bases, 4200 bases, 4300 bases, 4400 bases, 4500 bases, 4600 bases, 4700 bases, 4800 bases, 4900 bases, or 5000 bases.
[0021] In some aspects, the target sequence comprises a length of about 50 bases to 300 bases. In some aspects, the length is about 50 bases. 60 bases. 70 bases. 80 bases, 90 bases, 100 bases, 110 bases. 120 bases, 130 bases, 140 bases, 150 bases. 160 bases, 170 bases, 180 bases.190 bases, 200 bases, 210 bases, 220 bases, 230 bases, 240 bases. 250 bases, 260 bases, 270 bases, 280 bases, 290 bases, or 300 bases.
[0022] In some aspects, (d) providing an exonuclease to the solution to remove any DNA sequence of the plurality of DNA sequences having the M-base overhang.
[0023] In some aspects, the DNA sequence further comprises a nickase DNA binding sequence.
[0024] In some aspects, providing at least one nickase to the solution prior to providing the at least one IISRE to the solution, wherein the at least one nickase cleaves the nickase sequence.
[0025] In some aspects, the target sequence is generated w ith an error rate of less than 1 in 100,000.
[0026] In some aspects, the solution comprises: (a) 50 mM Potassium acetate, 20 mM Trisacetate, 10 mM Magnesium acetate, 100 pg / mL rAlbumin, pH 7.9 at 25°C, (b) 100 mM NaCl, 50 mM Tris-HCL 10 mM MgCh, 100 pg / mL rAlbumin, pH 7.9 at 25°C, or (c) 66 mM Potassium acetate, 33 mM Tris-acetate, 10 mM Magnesium acetate, 100 pg / mL Bovine Serum Albumin, pH 7.9 at 37°C.
[0027] In some aspects, the solution further comprises 10 mM DTT and either 1 mM ATP or 1 pM ATP. In some aspects, DTT is about 5mm, 6mm, 7mm, 8mm, 9mm, 10 mm, 11 mm, 12 mm, 13 mm, 14 mm, or 15 mm. In some aspects, ATP is about 0.5 mM, 06 mM, 0.7 mM, 0.8 mM, 0.9 mM, 1.0 mM, 1.1 mM, 1.2 mM, 1.3 mM, 1.4 mM, or 1.5 mM. In some aspects, ATP is about 0.5 pM, 06 pM, 0.7 pM, 0.8 pM, 0.9 pM, 1.0 pM, 1.1 pM, 1.2 pM, 1.3 pM, 1.4 pM, or 1.5 pM.
[0028] In some aspects, (a), (b), or (c) is performed at a temperature of at least 10°C. In some aspects, the temperature is about 10°C, 16°C, 25°C, 37°C, 50°C, 55°C, 60°C or 65°C. In some aspects, the temperature is about 10°C, 11°C, 12°C. 13°C, 14°C, 15°C. 16°C, 17°C . 18°C, 19°C, 20°C, 21°C. 22°C, 23°C, 24°C, 25°C. 26°C, 27°C, 28°C. 29°C, 30°C, 31°C. 32°C, 33°C,34°C, 35°C, 36°C, 37°C. 38°C, 39°C, 40°C. 41°C, 42°C, 43°C, 44°C, 45°C, 46°C, 47°C, 48°C, 49°C, 50°C, 51°C, 52°C, 53°C, 54°C, 55°C. 56°C, 57°C, 58°C, 59°C, 60°C, 61°C, 62°C, 63°C, 64°C or 65°C.
[0029] In some aspect, (a), (b), or (c) is performed cycling between one or more temperatures: about 10°C, 16°C, 25°C, 37°C, 50°C, 55°C, 60°C or 65°C. In some aspects, the temperature is about 5°C, 6°C, 7°C, 8°C, 9°C. 10°C, 11°C, 12°C, 13°C, 14°C, 15°C, 16°C, 17°C, 18°C, 19°C, 20°C, 21°C, 22°C, 23°C, 24°C, 25°C, 26°C, 27°C, 28°C, 29°C, 30°C, 31°C,32°C, 33°C, 34°C, 35°C, 36°C, 37°C, 38°C, 39°C, 40°C, 41°C, 42°C, 43°C, 44°C, 45°C, 46°C,47°C, 48°C, 49°C, 50°C, 51°C, 52°C, 53°C, 54°C, 55°C, 56°C, 57°C, 58°C, 59°C, 60°C, 61°C,62°C, 63°C, 64°C 65°C, 66°C, 67°C, 68°C, 69°C, 70°C, 71°C, 72°C, 73°C, 74°C, or 75°C.
[0030] In some aspects, the DNA sequence is formed using a ligase enzyme.
[0031] In some aspects, the ligase enzyme is T4 DNA ligase, T7 DNA ligase, T3 DNA ligase, or human DNA ligase III.
[0032] In some aspects, the synthesized target sequence has a purity of at least 80%. In some aspects, the synthesized target sequence has the purity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%. 82%, 83%. 84%. 85%. 86%. 87%. 88%. 89%. or 90.
[0033] In some aspects, the DNA sequence further undergoes heat inactivation of the enzymes directly or through protease treatment, wherein the protease treatment is inactivated by heat.
[0034] Described herein, in some aspects, a composition comprising: (a) a double stranded DNA sequence, wherein the double stranded DNA sequence comprises: (i) an N-mer sequence; (ii) a nickase sequence, wherein the nickase binding sequence is at most 4 bases from a nick site; and (iii) a first Type II S restriction endonuclease (IISRE) binding sequence configured to bind a IISRE that cleave the N-mer sequence, wherein the nickase binding sequence is located between the N-mer sequence and the IISRE binding sequence.
[0035] In some aspects, a nickase.
[0036] In some aspects, the nickase and the double stranded DNA sequence is in a single volume.
[0037] In some aspects, the double stranded DNA sequence is not coupled to a solid support.
[0038] In some aspects, a Type II S restriction endonuclease (IISRE).
[0039] In some aspects, the nickase cleaved site is within or adjacent to the N-mer sequence.
[0040] In some aspects, the nickase binding sequence comprises enzymes: BbvCI, Bsml,BsrDI, BssSI, BtsI, Alwl, BbvCI, BsmAI, BspQI, or BstNBI.
[0041] In some aspects, the IISRE binding sequence comprises: Mlyl, NgoAVII, SspD5I, Alwl, Ajul, Alol, BccI, Bcefl, Piel, BceAI, BceSIV, BscAI, BspD6I, Faul, Earl, BspQI, BfuAI, PaqCI, Esp3I, BbsI, Bbvl, BtgZI, FokI, BsmFI, Bsal, BcoDI, Hgal, or SfaNI.
[0042] In some aspects, the IISRE sequence comprises Earl.
[0043] In some aspects, at least one hairpin structure.
[0044] In some aspects, at least one hairpin structure comprises an aptamer sequence.
[0045] In some aspects, a second IISRE binding sequence.
[0046] In some aspects, the second IISRE binding sequence comprises a BspQI sequence.
[0047] In some aspects, a core melting temperature of the double stranded DNA sequence is at least 60°C. In some aspects, a core melting temperature of the double stranded DNA sequence is between 50°C to 70°C. In some aspects, a core melting temperature of the double stranded DNA sequence is about 50°C, 51°C, 52°C, 53°C, 54°C, 55°C, 56°C, 57°C, 58°C, 59°C, 60°C, 61°C, 62°C, 63°C, 64°C, 65°C, 66°C, 67°C, 68°C, 69°C, or 70°C.
[0048] Described herein, in some aspects, a system for synthesizing a target sequence in solution comprising: (a) a plurality of DNA sequences in solution, wherein a DNA sequence of the plurality of DNA sequences comprises a N-mer, a IISRE sequence, and at least one hairpin structures; (b) at least one IISRE in the solution, wherein the at least IISRE is configured to cleave the IISRE sequence, thereby exposing an M-base overhang; (c) at least two DNAsequences of the plurality of the DNA sequences configured to have the M-base overhang are configured to generate one or more ligated DNA sequences, wherein a ligated DNA sequence comprises a sequence having a length of 2N-M; and (d) one or more ligated DNA sequences are configured to generate a final DNA sequences comprising the target sequence.
[0049] In some aspects, M is 3. In some aspects, M is at least 3.
[0050] In some aspects, the N-mer is at least a 4-mer. In some aspects, the N-mer is no more than the 4-mer.
[0051] In some aspects, the target sequence comprises a length of about 50 bases to 5000 bases. In some aspects, the target sequence comprises the length of about 100 bases to 5000 bases. In some aspects, the length is about 50 bases, 60 bases, 70 bases, 80 bases, 90 bases, 100 bases, 200 bases, 300 bases, 400 bases, 500 bases, 600 bases, 700 bases, 800 bases, 900 bases, 1000 bases, 1100 bases, 1200 bases, 1300 bases, 1400 bases, 1500 bases, 1600 bases, 1700 bases, 1800 bases, 1900 bases, 2000 bases, 2100 bases, 2200 bases, 2300 bases, 2400 bases, 2500 bases, 2600 bases, 2700 bases, 2800 bases, 2900 bases, 3000 bases, 3100 bases, 3200 bases, 3300 bases, 3400 bases, 3500 bases, 3600 bases, 3700 bases, 3800 bases, 3900 bases, 4000 bases, 4100 bases, 4200 bases, 4300 bases, 4400 bases, 4500 bases, 4600 bases, 4700 bases, 4800 bases, 4900 bases, or 5000 bases.
[0052] In some aspects, the target sequence comprises a length of about 50 bases to 300 bases. In some aspects, the length is about 50 bases, 60 bases, 70 bases, 80 bases, 90 bases, 100 bases, 110 bases, 120 bases, 130 bases, 140 bases, 150 bases, 160 bases, 170 bases, 180 bases, 190 bases, 200 bases, 210 bases, 220 bases, 230 bases, 240 bases, 250 bases, 260 bases, 270 bases, 280 bases, 290 bases, or 300 bases.
[0053] In some aspects, (d) an exonuclease in the solution configured to remove any DNA sequence of the plurality of DNA sequences having the M-base overhang.
[0054] In some aspects, the DNA sequence further comprises a nickase sequence.
[0055] In some aspects, at least one nickase in the solution, wherein the at least one nickase cleaves the nickase sequence.
[0056] In some aspects, the target sequence is generated with an error rate of less than 1 in 100,000.
[0057] In some aspects, the solution comprises: (a) 50 mM Potassium acetate, 20 mM Trisacetate, 10 mM Magnesium acetate, 100 pg / mL rAlbumin, pH 7.9 at 25°C, (b) 100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCh, 100 pg / mL rAlbumin, pH 7.9 at 25°C, or (c) 66 mM Potassium acetate, 33 mM Tris-acetate, 10 mM Magnesium acetate, 100 pg / mL Bovine Serum Albumin, pH 7.9 at 37°C.
[0058] In some aspects, the solution further comprises 10 mM DTT and either 1 mM ATP or 1 pM ATP. In some aspects, DTT is about 5mm, 6mm, 7mm, 8mm, 9mm, 10 mm, 11 mm, 12 mm, 13 mm, 14 mm, or 15 mm. In some aspects, ATP is about 0.5 mM, 06 mM, 0.7 mM, 0.8 mM, 0.9 mM, 1.0 mM, 1.1 mM, 1.2 mM, 1.3 mM, 1.4 mM, or 1.5 mM. In some aspects, ATP is about 0.5 pM, 06 pM, 0.7 pM, 0.8 pM, 0.9 pM, 1.0 pM, 1.1 pM, 1.2 pM, 1.3 pM, 1.4 pM, or 1.5 pM.
[0059] In some aspects, a temperature of at least 10°C. In some aspects, the temperature is about 10°C, 16°C, 25°C, 37°C, 50°C, 55°C, 60°C or 65°C. In some aspects, the temperature is about 10°C, 1 1 °C, 12°C, 13°C, 14°C, 15°C, 16°C, 17°C , 18°C, 19°C, 20°C, 21 °C, 22°C, 23°C, 24°C, 25°C, 26°C, 27°C, 28°C, 29°C, 30°C, 31 °C, 32°C, 33°C, 34°C, 35°C, 36°C, 37°C, 38°C, 39°C, 40°C, 41°C, 42°C, 43°C, 44°C, 45°C, 46°C, 47°C, 48°C, 49°C, 50°C, 51°C, 52°C, 53°C, 54°C, 55°C, 56°C, 57°C, 58°C, 59°C, 60°C, 61°C, 62°C, 63°C, 64°C or 65°C
[0060] In some aspects, the DNA sequence is formed using a ligase enzyme.
[0061] In some aspects, the ligase enzyme is T4 DNA ligase, T7 DNA ligase, T3 DNA ligase, or human DNA ligase 111.
[0062] In some aspects, the synthesized target sequence has a purity of at least 80%. In some aspects, the synthesized target sequence has the purity of at least 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 90.
[0063] In some aspects, the DNA sequence further undergoes heat inactivation of the enzymes directly or protease treatment, wherein the protease treatment is inactivated by heat.
[0064] Described herein, in some aspects, a composition comprising: (a) a double stranded DNA sequence, wherein the double stranded DNA sequence comprises: (i) an N-mer sequence; (ii) a first enzyme binding sequence, wherein the first enzyme binding sequence is at most 4 bases from a nick site; and (iii) a second binding sequence configured to bind a IISRE that cleave the N-mer sequence, wherein the nickase binding sequence is located between the N- mer sequence and the IISRE binding sequence.
[0065] In some aspects, the first enzyme binding sequence is a nickase binding sequence.
[0066] In some aspects, the second binding sequence is a first Type II S restriction endonuclease (IISRE) binding sequence.
[0067] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0068] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0070] FIG. 1 depicts an exemplary schematic of a pay load addamer.
[0071] FIG. 2 depicts an exemplary schematic of control addamers including 3-base overhang systems using a TTCG linker.
[0072] FIG. 3 depicts an exemplary schematic of control addamers including 3-base overhang systems using a TCG or GTG linker.
[0073] FIG. 4 depicts an exemplary schematic of control addamers including 4-base overhang systems using a TTCG linker.
[0074] FIG. 5 depicts an exemplary' schematic of control addamers including Bsal and BsmBI sites using a TCG linker.
[0075] FIG. 6 depicts an exemplary' schematic of generic libraries of control payload addamer sets.
[0076] FIG. 7 depicts an exemplary' schematic of a method of addamer excision.
[0077] FIG. 8 depicts an exemplary' schematic of addamer excision systems.
[0078] FIGs. 9A-C depicts an exemplary schematic of building addamers.
[0079] FIG. 10 depicts an exemplary schematic using gSynth to assemble addamer arrays.
[0080] FIG. 11 depicts an exemplary schematic of combining a pay load addamer with a control addamer.
[0081] FIG. 12 depicts an exemplary schematic combining two control payload addamers to make a 5 bp payload.
[0082] FIG. 13 depicts an exemplary7schematic combining two 5 bp payload addamers to make a 7 bp payload.
[0083] FIG. 14 depicts an exemplary7schematic combining two 7 bp payload addamers to make an 11 bp payload.
[0084] FIG. 15 depicts an exemplary schematic combining two 11 bp payload addamers to make a 19 bp payload.
[0085] FIG. 16 depicts an exemplary schematic illustrating the generation of a Level 0 fragment assembly using a mixture of 3- and 4-base overhangs.
[0086] FIG. 17 depicts an exemplary schematic illustrating the assembly from reagents to Level £ through Level 0.
[0087] FIG. 18 depicts an exemplary schematic of GreenSynth Assembly using only 3-base overhangs.
[0088] FIG. 19 depicts an exemplary embodiment of the assembly pathways for Fragment 1.
[0089] FIG. 20 depicts an exemplary embodiment of the assembly pathways for Fragment 2.
[0090] FIG. 21 depicts an exemplary embodiment of the assembly pathway for Fragment 3.
[0091] FIG. 22 depicts an exemplary embodiment of the assembly pathway for Fragment 4.
[0092] FIG. 23 depicts an exemplary embodiment of the final pooled assembly to Level 0.
[0093] FIG. 24 depicts an example of the addamers that earn74 bp pay loads.
[0094] FIG. 25 depicts an example of the pay load-free reagent addamer.
[0095] FIG. 26 depicts a partial assembly path diagram for generation of a Level 0 addamer.
[0096] FIG. 27 depicts in a diagram the sequence verification of the Level 0 pay load.
[0097] FIG. 28 depicts an exemplary schematic for the plasmid insert design and recovery method.
[0098] FIG. 29 depicts the stages of plasmid digestion and treatment to recover the addamer.
[0099] FIG. 30 depicts an exemplar}7schematic two enzy me-system for generating addamers for 3’ base overhangs.
[0100] FIG. 31 depicts an exemplar}7schematic two enzy me-system for generating addamers for 3’ base overhangs.
[0101] FIGs. 32A-B shows an exemplary7schematic of the tests performed to validate the two-enzy me systems.
[0102] FIG. 33 depicts exemplary7schematics of blunt cutting two-enzyme system addamers.
[0103] FIG. 34 depicts exemplary7schematics of bi-active addamer generic library designs.DETAILED DESCRIPTION
[0104] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0105] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least” or “greater than” applies to each one of the numerical values in that series of numerical values.Whenever the term “no more than,” “less than,’" or “less than or equal to'’ precedes the first numerical value in a series of two or more numerical values, the term “no more than"’ or “less than’" applies to each one of the numerical values in that series of numerical values.
[0106] The term “about’' or “nearly” as used herein generally refers to within (plus or minus) 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of a designated value.
[0107] As used herein, the singular forms “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise.
[0108] Solution phase GreenSynth is a non-templated DNA synthesis system capable of generating any arbitrary DNA sequence. In some embodiments, small DNA fragments carrying 4 bp payloads are pieced together pairwise by a process of cutting with a Type II S restriction endonuclease (IISRE), combining the pair of cut Addamers, ligating them together, then digesting away unused reactants using an exonuclease, which degrades any DNA with a nick or free end. In this process geometric pairwise combination is employed to generate longer sequences of DNA. In some embodiments, once a certain threshold length is reached, a small group of fragments, for example five fragments each with payload just over 18 bp long, are pooled and combined to generate a Level 0 Addamer of 50-100 bp in length. By using a combination of 3- and 4-base overhang cutting IISRE systems, the system may generate accurate, specific lengths of Level 0 Addamer sequence for further assembly using gSynth. Addamer-based methods and compositions of the present disclosure
[0109] Addamers
[0110] The key element of GreenSynth is the Addamer. In some embodiments, an Addamer (FIG. 1) comprises a double stranded DNA capped at both ends with a hairpin, which is a stemloop structure, like the sequence 5’-GCGCAGC-3’ that forms a small stem-loop with a 3-base loop. An Addamer also carries a payload, which is an original small insert payload (such as NNNN), a linker sequence, or a target sequence in mid-synthesis. Addamers have DNA bindingsite sequences for offset cutting IISREs systems that flank the payload, such as BtgZI / Nt.BstNBI and BspQI as shown in FIG. 1. Additionally, an Addamer may also contain other elements such as primer sequences for amplification (e.g., Payload Primer) and the remnants of a plasmid excision system, such as PacI / Nb.BssSI adjacent to the hairpins in FIG. 1.
[0111] In some embodiments, FIG. 1 shows three different payload Addamer libraries. In some embodiments, any element of these libraries can be combined with a Control Addamer to generate a new Addamer with a new Control element on the left side and BspQI on the right. For the first library' NNNN (pl-p256) a, the combination of two enzy mes BtgZI and Nt.BstNBI may be used to reveal the 3 base overhang linker TCG (See FIG. 3 and 5 for control Addamers), which will allow the new Control element to have access to the payload sequence. In some embodiments, for the second library' NNNN (pl-p256) b, BbsI digestion reveals the 4 base overhang linker TTCG (See FIG. 2 and 4 for control Addamers). In some embodiments, NNNN (pl-p256) c, the combination of two enzymes BtgZI and Nt.BstNBI is used to reveal the 3 base overhang linker GTG (See FIG. 3 and 5 for control Addamers).
[0112] Two Enzvme Systems
[0113] Of the known IISREs, there are onty three publicly available IISRE enzyme classes that yield 3-base overhangs. These are exemplified by BspQI [GCTCTTC(l / 4)], Earl [CTCTTC(l / 4)] and BsaXI [(9 / 12)ACNNNNNCTCC(10 / 7)]; however, the DNA binding site for Earl is fully a part of the DNA binding domain of BspQI which means it is not useful for generating 3-base overhangs. However, performing solution synthesis requires at least 3 enzymes that have unrelated DNA binding specificities. Therefore, there is a need to identify new means of generating 3-base overhangs.
[0114] In some embodiments, combining another class of restriction endonucleases, with other IISREs can generate 3-base overhangs in 6 additional ways (FIG. 2 and 3). In some embodiments, the other class of restriction endonucleases comprises nickases. In someembodiments, the nickases comprise nickases that are offset cutting. In some embodiments, by first nicking one strand at a specific position then full double-stranded digestion of the Addamer with the IISRE, a 3 -base overhang is generated on the Addamer half that is ligated in a subsequent reaction. Additionally, this system has the advantage of making the remaining half of the cut Addamer unable to be ligated and sensitive to exonuclease activity', which may be used to purify one or more final reaction products.
[0115] Nickases are restriction endonucleases that cut only one strand. The two major types are Top strand cutters (Nt.) and Bottom strand cutters (Nb.) as described herein in Table 1. Additionally, there are nickases that cut internally to their respective DNA binding site (recognition sequence). The DNA binding sequences are oriented 5’-3’. Consider Nb.BbvCI and its recognition sequence CCTCAGC (none / -2). Here ‘none’, means that the enzy me does not cut the top strand and ‘-2’ means that the enzy me will cut 2 bases before the 3’ most base of the recognition sequence, i.e., between the A and the G. Thus, for Nt.AlwI, its recognition sequence code GGATC (4 / none) indicates that it will cut the top strand 4 bases downstream and not cut the bottom strand. Simply put, an offset cutting nickase will cut a single strand of DNA outside its DNA binding sequence. In some aspects, an Addamer can comprise, consists essentially of, or consist of DNA.
[0116] Table 1
[0117] Addamer Sets
[0118] The Addamers considered in this disclosure are detailed in FIGs. 1 - 6. The Addamers are grouped into several categories, payload Addamers (FIG. 1), 3-base overhang control Addamers (FIGs. 2 and 3), 4-base overhang control Addamers (FIG. 4 and 5) and generic control: :payload Addamer sets (FIG 6). There are also categories according to the appropriate linker sequence to use when combining a payload with a specific control Addamer. The linker sequences are TTCG, TCG, CGTG or GTG.
[0119] Assembly Process
[0120] The assembly process is arbitrarily divided into Levels 0, 1, 2 and 3. Originally Level 0 was designated as Addamers, generally with payloads under 100 bp, which are generated by hybridization and ligation of phosphoramidite synthesized oligonucleotides. For Level 1 assembly, a group of Level 0 Addamers are assembled into a longer Addamer. up to about 500 bp. through a cycling assembly using a IISRE and DNA ligase, followed by exonuclease digestion. Level 1 Addamers can be cloned into plasmids, in which case Levels 2 and 3 are assembled by a Golden-Gate process into destination vectors.
[0121] To replace phosphoramidite generated oligonucleotides, assemble DNA is assembled from DNA fragments with generic DNA reagents generated from plasmids harvested from bacteria. It is sufficient to have one library of 256 4-base long payload Addamers that can be combined with at least 16 different control Addamers as needed. In this way, any DNA sequence can be built from generic reagent Addamers.
[0122] In one preferred embodiment, the first step is to put together specific 4 bp payload Addamers each together with prescribed Control Addamers (FIG. 12). For a given assembly, the combination of payloads and controls is determined beforehand by an algorithm that identifies a set of Level 0 gene fragments, then specifies a set of overlapping 5 bp fragments, for each Level 0 gene fragment. The overlaps between the 5 bp fragments can be 3- or 4-bases.
[0123] As the original payloads are 4 bp long, the first extension reaction is carried out. A 3- base overhang cutting enzyme BspQI exposes the proximal 3 bases each of a left and right Addamer. In the same reaction mixture, ligase is added after digestion is complete and the exposed 3-base overhangs hybridize and are ligated together. After ligation, the entire reaction is further processed with an exonuclease that degrades all DNA except the desired reaction product, which is a new Addamer with a 5 bp payload with distinct left and right side Control elements, including specific left and right restriction enzyme system DNA binding sites and specific left and right side primer binding sites. Finally, all enzymes are destroyed using a heat labile protease followed by heat inactivation (FIG. 12). This describes an extension reaction cycle. In subsequent cycles, for example using all 3-base overhang cutting systems, two 5 bp carrying Addamers are combined to produce a new Addamer carrying a 7 bp payload. Two 7 bp carrying Addamers are combined to form a new 1 1 bp carrying Addamer, then two 1 1 bp payload carrying Addamers are combined to form a 19 bp carrying Addamer.
[0124] Plasmid Construction and Excision
[0125] Provided herein is a system for manufacture at scale of any reagent Addamer. Addamers are cloned as a pre- Addamer construct, either as a single Addamer or an array of identical Addamers. In three preferred embodiments, using a two-enzyme system, where by first nicking one strand at a specific position then full double-stranded digestion of the Addamer with an IISRE, a long single stranded overhang is generated that encodes a strong stem-loop structure (FIG. 7). Following a brief exposure to a higher temperature to melt away the complementary single-stranded oligonucleotide and an annealing step, the nicks in the excised Addamer are healed with a brief DNA ligation step. After ligation the remaining non-Addamer DNA is degraded by exonuclease treatment. In some embodiments, target Addamers may be excised from larger Addamers using three different combinations of Nickase’s and IISRE’s In some embodiments, Addamers may also be excised from plasmids.
[0126] The sequence of events in Addamer excision is depicted in FIG. 7. In this case a larger Addamer that contains the target Addamer is generated. This larger Addamer is cloned into a plasmid as a single copy insert. The larger Addamer is digested with Nb. BssSI then digested with BseRI and held at a temperature below that of the Core melting temperature and above the melting temperature of the long overhang generated by the restriction enzy mes. The melting temperature of the formed hairpins is well above the melting temperature of the overhang. A quick ligation at an elevated temperature favors the healing of the nicks in the newly formed addamer. The other reaction products are then degraded by exonuclease leaving the excised target Addamer.
[0127] Three different two enzyme systems were tested for excision activity. Depicted in FIG. 8 are three test constructs, using a common core Addamer (a version of Addamer 3a AAAA), and the same external IISREs, but differing in the hairpin forming enzyme systems. In some implementations, ExA uses the nickase Nt.BbvCl and the IISRE Btsl to generate a long overhang that will fold to be the hairpin of the target Addamer. In some implementations, forExB the enzyme combination Nb.BssSI and BseRI is used to generate the long overhang. In some implementations, the combination of Nb.BsmI and Acul is used for ExC.
[0128] In some implementations, plasmids may have a single pre-Addamers. In some implementations, to maximize Addamer production it is desirable to have as many copies of the same pre- Addamer as possible in one plasmid. By making three different large Addamer constructs with three different IISREs and three different linker sequences, very long arrays may be generated with gSynth (FIGs. 9A-C). FIG. 10 depicts the assembly of an array of 4 copies of the same pre- Addamer, which can be inserted into a plasmid backbone and propagate in bacteria. The copy number can be doubled with each new round of gSynth.
[0129] Payload addamer libraries
[0130] In some embodiments, FIG. 2 shows schematics for three 3-base overhang generating Control Addamers. In some embodiments, the digestion with BbsI leads to the generation of a TTCG overhang linker which will allow ligation with predigested Generic Payload Addamers NNNN (pl-p256) b (FIG. 1). The Controls elements are: 3a (Earl), which confers an Earl IISRE control to the specific Payload Addamer; 3b (BsaXI), which confers a BsaXI IISRE to the payload Addamer, to generate 3’ 3-base overhangs; and 3c (BbvI / Nt.BstNBI), which confers a combination of Bbvl and Nt.BstNBI control to the Payload Addamer to generate a 5’ 3-base overhang using a two-enzyme system.
[0131] In some embodiments, FIG. 3 shows five Control Addamers that can each generate 3-base overhangs for their respective payloads. In some embodiments, Earl is used to generate either a TCG overhang linker for combination with the NNNN (pl-p256) a Pay load Addamer (FIG. 1) or to generate a GTG overhang linker for combination with the NNNN (pl-p256) c Payload Addamer (FIG. 1).
[0132] FIG. 4. in some embodiments, shows three 4-base overhang generating Control Addamers. In some embodiments, digestion with BbsI leads to the generation of a TTCGoverhang linker which will allow7ligation with predigested Generic Payload Addamers NNNN (pl-p256) b (FIG. 1). Each of the five Addamers confers a different control element, which include a IISRE and a control-specific primer, to the specific payload. In some embodiments, FIG. 5 shows 3 Control Addamers that can each generate 4-base overhangs for their respective payloads. In In some embodiments, BspQI is used to generate a TCG overhang linker for combination with the NNNN (pl-p256) a Pay load Addamer (FIG. 1). In some embodiments, the BsmBI Control Addamers are of two flavors, with distinct primer sequences so that synthesized Addamers with the distinct T3 or T7 primer sites can be amplified by PCR or an equivalent in vitro amplification system.
[0133] In some embodiments, FIG. 6 show s a pre-configured control payload Addamer libraries. In some embodiments, each of these libraries carries a distinct IISRE and primer sequence. Table 2 disclosed herein shows a list of addamer designs.
[0135] In some embodiments, FIG. 7 shows the process of Addamer excision from a construct that can be recovered from a plasmid. The Addamer 3a AAAA may be contained within a construct that possesses a IISRE and Nickase site on either side of the Addamer sequence. The Nickase, Nb.BssSI, may be first used to cut the bottom strand on the left side and the top strand on the right. After Nickase digestion, the IISRE, BseRI, may be used to generatedouble stranded cuts that leave the hairpin sequences to properly fold under the appropriate temperature conditions. The resulting nicks in the reformed Addamer may be sealed by T4 DNA ligase treatment. In some embodiments, the exonuclease is used to remove all non-Addamer DNA.
[0136] In some embodiments, FIG. 8 shows three different Addamer excision systems. These systems can be used to generate plasmid inserts for production scale Addamer formation.
[0137] In some embodiments, FIGs. 9A-C shows addamer arrays. These addamer arrays can be built from the use of three different Building Addamers. Each of these Addamer may have a pair of 4-base overhang sites, which can be exposed by BsmFI, PaqCI or Bbsl. The arrays of reagent addamers can be generated using a Golden-Gate like assembly.
[0138] In some embodiments, described herein is a process for generating a four element array of Addamer using the constructs described in FIG. 10. A similar process can be used to generate arbitrarily long (i.e., > 4) Addamer arrays.
[0139] FIG. 11 may depict combining a payload Addamer pl 5a (AATG) with a Control Addamer 3a (Earl). The final product can be used in subsequent steps to assemble a de novo DNA sequence.
[0140] In the first extension step, two loaded Addamers may be combined (FIG. 12). In some embodiments, in a pooled reaction, BspQI is used to cut the two 4 bp payload Addamer to form a 5 bp payload Addamer mediated by a 3-base overhang. The Addamers 3a_AATG and 3d_CCAT (whose reverse complement is ATGG) may be combined to form 3a_AATGG_3d.
[0141] In the second extension step, two loaded Addamers may be combined (FIG. 13). In a pooled reaction, Nt.BsmAI and BsmFI may be used to cut the two 5 bp payload Addamer to form a 7 bp payload Addamer mediated by a 3-base overhang. The Addamers 3a_AATGG_3d (FIG. 13) and 3d_TGGAC_3c may be combined to form 3a_AATGGAC_3c.
[0142] In the third extension step, two loaded Addamers may be combined (FIG. 14). In a pooled reaction, Nt.BstNBI and Bbvl may be used to cut the two 7 bp payload Addamer to form an 11 bp payload Addamer mediated by a 3-base overhang. The Addamers 3a_AATGGAC_3c (FIG. 14) and 3c_GACCGTG_3d may be combined to form 3a_AATGGACCGTG_3d.
[0143] In the fourth extension step, two loaded Addamers may be combined (FIG. 15). In a pooled reaction, Nt.BsmAI and BsmFI may be used to cut the two 11 bp payload Addamer to form a 19 bp pay load Addamer mediated by a 3-base overhang. The Addamers 3a_AATGGACCGTG_3d (FIG. 15) and 3d_GTGACCGCTAC_3c may be combined to form 3a AATGGACCGTGACCGCTAC 3c.
[0144] A 65 bp Level 0 Addamer can be assembled using a mixture of 3- and 4-base overhangs (FIG. 16). The sequence may be divided into four fragments with 3-base overlaps. In some embodiments, each of the four fragments are individually assembled then ultimately mixed and assembled together in a Golden-Gate like reaction. In some embodiments, 5 bp payloads contribute to each of the four fragments.
[0145] The overall assembly strategy may be depicted in FIG. 17 in Fragment 1 assembly at sub-Level 0. The four fragments going into the final step to Level 0 may be all Level a fragments.
[0146] In some embodiments, are 31 5 bp payload Addamers that are combined to form the final Level 0 fragment (FIG. 18). The 31 5 bp pay load Addamer may be formed from 62 original reagent Addamers or 124 reactions combining the payload and control Addamers. In some embodiments, the assembly of Fragment 1 (FIG. 19), Fragment 2 (FIG. 20), Fragment 3 (FIG. 21), and Fragment 4 (FIG. 22) may have intermediate Addamers.
[0147] In a final Golden-Gate like reaction, using Earl, the four fragments may be combined to form the 65 bp payload Level 0 Addamer (FIG. 23). For this final reaction, theNEB ligation fidelity calculator can predict the ligation fidelity to be 97% after 1 hour at 25°C with T4 DNA ligase.
[0148] FIG. 24 may depict addamers that cany 4 bp payloads, wherein each has a BspQI control element that allows combinations using 3 base overhangs in initial assembly reactions. In some embodiments, each class of addamer carries a unique binding site for the IISRE control and a unique primer sequence for amplification. Additionally, each Addamer may carry a nickase DNA binding site and a binding site for the restriction endonuclease Pad.
[0149] FIG. 25 may show two examples of this type of payload-free reagent Addamer. In some embodiments, they carry' a TCT sequence that allows them to be combined with any other reagent Addamer with a payload sequence of the form NAGA. In some embodiments, the combined use of the reagent Addamer and the adapter Addamer will only contribute 1 bp to the final assembly product. FIG. 26 may depict a diagram of two of the eight way pools are highlighted. In some embodiments, in the assembly a total of 80 reagent Addamers and adapter Addamers were used. The assembly may level Epsilon through Level zero, corresponding to the maximum possible payload length for each column, i.e., step, in the assembly process. FIG. 27 may depict multiple clones of the cloned final product. In some embodiments, the clones were sequenced using capillary Sanger sequencing. The results may verify that the expected sequence was generated. FIG. 28 may show the plasmid insert design and the recovery' method. FIG. 29 shows a schematic shows stages of plasmid digestion and treatment to recover the addamer from a mini-prep. FIG. 30 show four of the two enzyme systems. In some embodiments, the IISRE cut site are denoted by the colored line and the Nickase cut site by a colored arrowhead.
[0150] FIG. 31 may show four of the two enzyme systems. In some embodiments, the IISRE cut site are denoted by the colored line and the Nickase cut site by a colored arrowhead.
[0151] FIGs. 32A-B may show the two Addamer for each of the two enzyme systems. Each pair may combine the payload TCGA and CGAG to generate the BspQI CTCGA BspQIAddamer product. FIG. 32A may show are a schematic and a before and after analytical gel of the results of cutting, ligation and exonuclease treatments for each of the three enzy me systems. In some embodiments, the 3d system both BsmFI and its isoschizomer FaqI were tested. FIG. 32B may show before and after analytical gel of the results of cutting, ligation and exonuclease treatments for each of the three enzyme systems. Table 3 described herein provides a list of Addamer designs describing Addamers in terms of their constituent parts.
[0153] FIG. 33 may show Addamer designs in which the payload can be exposed with a blunt end. In some embodiments, each of the Addamer designs comprise a different pair of nickase and IISREs are used to generate the blunt ends.
[0154] FIG. 34 may show the Bi-active Addamer Generic Library Designs. In some embodiments, the two enzyme systems use the same pair of enzymes. In some embodiments, BtgZI and Nt.AlwI are used to generate both 3’ overhangs (3k) and 5’ overhangs (31). In someembodiments, Bbvl and Nt.BstNBI are used to generate 3’ (3m) or 5‘ (3n) overhangs. In some embodiments, the methods described herein produce deeper pooling of Addamers at the start of Addamer synthesis.
[0155] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.EXAMPLES
[0156] Example 1. Synthesis of 19 bp Sequence
[0157] Described herein is an example of GreenSynth (gSynth) reactions. In some instances, the gSynth reaction generates a 19 bp pay load. Described herein the 19 bp pay load is generated using eight, 3 -base overlapping 5 bp fragments. Each of the eight, 3 -base overlapping 5 bp fragment are combined using the Control and Payload Addamers as described herein. In some instances, a reaction is depicted as in FIG. 11. As shown in FIG. 11. the payload Addamerpl 5 (AATG) is combined with the control Addamer 3a (Earl). The enzyme BtgZI is further used for digestion to expose the linker sequence TTCG on both the control and payload Addamers described herein. In some instances, BspQI acts as the payload primer. The digestion products are further recombined by annealing and DNA ligation. The reaction product is designated as 3a_AATG (Earl). 3a_AATG (Earl) is further purified using exonuclease treatment and Proteinase K treatment.
[0158] In the next step, two control: payload Addamers are combined (FIG. 12) as described herein. The two control: pay load Addamers are combined to generate a 5 bp Addamer described herein. The Addamer 3a_AATG (Earl) is combined with 3d_CCAT (BsmFI / Nt.BsmAI), which is equivalent to ATGG_3d (FIG. 12). In some instances, the nickase DNA binding sequence is an enzyme. The enzyme described herein is used to expose the overlapping 3 base ATG for each of the Addamers. In some instances, the enzyme is BspQI. In some instances, the enzy me is BbvCI. In some instances, the enzyme is BsmI. In some instances, the enzyme is BsrDI. In some instances, the enzyme is BssSI. In some instances, the enzyme is Btsl. In some instances, the enzyme is AlwI. In some instances, the enzy me is BbvCI. In some instances, the enzyme is BsmAI. In some instances, the enzyme is BstNBI.
[0159] The Addamers are further recombined by annealing. The Addamers further undergo DNA ligation. The Addamers are further purified to produce the product 3a_AATGG_3d. The two 5 bp Addamers are cut with a serial combination of Nt. BsmAI followed by BsmFI. The two 5 bp Addamers are further recombined to form a 7 bp Addamer. The 7 bp Addamer are designated as 3a_AATGGAC_3c (FIG. 13). In some instances, the 7 bp Addamers are 3a_AATGGAC_3c and 3c_GACCGTG_3d. The 7 bp Addamers, 3a_AATGGAC_3c and 3c_GACCGTG_3d, are cut with a nickase enzyme and a IISRE enzyme. In some instances, the nickase enzyme is Nt.BstNBl and the IISRE enzyme is Bbvl. In some instances, the nickase enzyme is BspQI. In some instances, the nickase enzyme is BbvCI. Insome instances, the nickase enzyme is BsmI. In some instances, the nickase enzyme is BsrDI. In some instances, the nickase enzyme is BssSI. In some instances, the nickase enzyme is Btsl. In some instances, the nickase enzyme is AlwI. In some instances, the nickase enzyme is BbvCI. In some instances, the nickase enzyme is BsmAI. In some instances, the nickase enzyme is BstNBI. In some instances, the IISRE sequence is Mlyl. In some instances, the IISRE sequence is NgoAVII. In some instances, the IISRE sequence is SspD5I. In some instances, the IISRE sequence is AlwI. In some instances, the IISRE sequence is Ajul. In some instances, the IISRE sequence is Alol. In some instances, the IISRE sequence is Bed. In some instances, the IISRE sequence is Bcefl. In some instances, the IISRE sequence is Piel. In some instances, the IISRE sequence is BceAI. In some instances, the IISRE sequence is BceSIV. In some instances, the IISRE sequence is BscAI. In some instances, the IISRE sequence is BspD6I. In some instances, the IISRE sequence is Faul. In some instances, the IISRE sequence is Earl. In some instances, the IISRE sequence is BspQI. In some instances, the IISRE sequence is BfuAI. In some instances, the IISRE sequence is PaqCI. In some instances, the IISRE sequence is Esp3I. In some instances, the IISRE sequence is Bbsl. In some instances, the IISRE sequence is BtgZI. In some instances, the IISRE sequence is Fokl. In some instances, the IISRE sequence is BsrnFI. In some instances, the IISRE sequence is Bsal. In some instances, the IISRE sequence is BcoDI. In some instances, the IISRE sequence is Hgal. In some instances, the IISRE sequence is SfaNI.
[0160] The 7 bp Addamers are further recombined to generate the 11 bp Addamer. In some instances, the 11 bp Addamer is 3a_AATGGACCGTG_3d (FIG. 14). In the final step, two 11 bp Addamers are recombined to form the final 19 bp product. In some instances, the final 19 bp product is 3a_AATGGACCGTGACCGCTAC_3c (FIG. 15).
[0161] Example 2. GreenSynth Assembly: Generation of a Level 0 Sequence Fragment
[0162] The known Level 0 gene is generated using a mix of 4-base and 3-base overhang cutting (FIG. 16). There are three types of Level 0 gene fragments. The generation of Level 0Addamers from Green Synth uses adapter Addamers. In some instances, the adaptor Addamers attach Bsal or BsmBI site to the ends of the Level 0 Addamer. In the cloning system, the leftmost Level 0 fragment of a Level 1 gene carries a left-hand BsmBI site and a left-hand Bsal site [4f (Bsal)]. The left-hand BsmBI site has a T7 primer sequence [4gT7 (BsmBI)]. The central Level 0 fragments of a Level 1 gene carry Bsal sites. The Bsal sites exist on either side of the pay load. The rightmost Level 0 fragment carries a left-hand Bsal site and a right-hand BsmBI site with a T3 primer sequence [4gT3 (BsmBI)]. In some instances, a central Level 0 gene fragment is assembled. The payload needs the linker sequence CTCG added at either end for addition of Bsal (4f) caps.
[0163] The gene is divided into 4 fragments at the 4-base overhangs. The 4-base overhangs are CTCC, TCAT, and CGTC as described herein. The 4-base overhangs are used for Level 0 assembly. The ligation fidelity for the mixed ligation reaction of these three, 4-base overhangs appears to be about 100%. The sequence contains two 4-base self-complementary' sub-sequences. In some instances, the two 4-base self-complementary sub-sequences are TCGA and CCGG. The two 4-base self-complementary' sub-sequences need to be avoided in this assembly' as 4-base overhangs are being used (FIG. 16).
[0164] The Levels of synthesis for this assembly are depicted in FIG. 17. In some instances, the Levels of synthesis are Level s through Level a. The Addamers assembly proceeds in a pairwise fashion. In some instances, the Level a to Level 0 assembly is carried out by digestion with Bsal. The Addamer further undergoes ligation to generate the Level 0 Addamer.
[0165] Example 3. GreenSynth Assembly Using Only 3-base Overhangs
[0166] GreenSynth assembly with 4-base overhangs provides improved ligation efficiency, but for the stepwise growth of Addamers 3-base overhangs will lead to fewer needed steps and fewer generic 4 bp payload Addamers at the beginning of the assembly. FIG. 18 shows the list of 5 bp Addamers and their control elements with overlaps of 3-bases. A total of 31 5 bpAddamers are needed to generate the Level 0 sequence of 65 bases. FIGs. 19-22 show the assembly pathways for each of the 4 sub-Level 0 Addamers that are pooled (FIG. 23) and assembled to the Level 0 addamer using the two enzyme system Nt.BstNBI / Bbvl to generate the 3-base overhangs. Thereafter, ligation is performed.
[0167] Example 4. Synthesis of a Level 0 sequence fragment from generic reagent Addamers
[0168] In the next step, the generic reagent addamers are used to synthesize an arbitrary DNA sequence. Described herein the first 56 bp of a modified PhiX174 genome are generated as an addamer payload. The genomic sequence is divided into fragments to prepare a full assembly of the PhiX174 genome. In the first step, each fragment possesses compatible 4-base sites at each end. Each level has a different average payload length. The base level for geometric synthesis is designated Level 0. In some instances, this level of Addamers carry payloads of 50- 100 bp.
[0169] For this assembly, five different libraries of reagent Addamers and two different reagent Addamer adapters were used (FIGs. 1 and 2). Specific reagent Addamers or adapters were combined eight in a pool.
[0170] Step 1 described herein outlines BspQI digestion and Ligation Within Pools. Reagent Addamers are joined in a pairwise process with a 3 base overhang, generated by cleavage by BspQI. In the first block of eight the Addamers and adaptors are described herein in Table IX:
[0172] The 4c TCT Addamer is an adaptor and the TCT sequence is used to connect the two reagent Addamers. In some instances, the two reagent Addamers are 4c TCT and 3a CAGA. In some instances, the 3a CAGA Addamer is equivalent to the TCTG 3a Addamer. The next step further involves pairwise hybridization and ligation. The Addamer further undergoes exonuclease treatment. As a result, the process yields the following 5 bp payload Addamers as described herein in Table 2X:
[0173] Table 2X
[0174] Described herein Step 2 outlines Earl (3a) digestion and ligation. In some instances, to extend from 5 bp to 7 bp, the pool of four 5 bp Addamers is cut by the IISRE. In some instances, IISRE sequence is Earl. Earl generates 3 base overhangs as described herein. To generate the following < 7 bp payload Addamers see Table 3X described herein (FIG. 3).
[0175] Table 3X
[0176] Described herein Step 3 outlines BsmBI (4g) digestion and ligation. As a pool, the final reaction cycle uses the IISRE sequence to cut the two < 7 bp payload Addamers. In someinstances, the IISRE sequence is the BsmBI sequence. The 4 base overhangs can hybridize and ligate to generate the following < 10 bp Addamer (FIG. 3) as described herein in Table 4X:
[0177] Table 4X
[0178] Described herein steps 4 and 5 outline digestion and litigation of the first enzyme followed by the second enzyme. In some instances, the enzyme is BbsI (4h). In some instances, the enzyme is BsmFI (4d). Pairs of < 10 bp Addamers are combined using BbsI to generate 4base overhangs. The Addamers are used to form < 16 bp Addamers. The pairs of < 16 bpAddamers are further combined using BsmFI to generate 4 base overhangs. The final product will form < 28 bp Addamers (FIG. 3) as described herein in Table 5X.
[0179] Table 5X:
[0180] Step 6 involve combining pool of < 28 bp Addamers using Bsal (4f). In the final step, a pool of arbitrary size, in this case three fragment Addamers, are combined in a Golden- gate like reaction using the IISRE Bsal to generate the Level 0 Addamer (FIG. 3). By molecular cloning into a plasmid, followed by capillary Sanger sequencing of the inserts the target sequence was confirmed in the final Level 0 Addamer (FIG. 4).
[0181] Plasmid Cloning and Excision of a reagent Addamer
[0182] The method described herein provides a cost effective method for producing reagent Addamers in bulk. In some aspects, the reagent Addamers can be produced through fermentation or in vitro amplification methods. Given that these methods do not involve organic-phase chemical synthesis, the whole gene synthesis pathway using reagent Addamer gSynth is an environmentally friendly process producing little in non-compostable waste. An example reagent Addamer 3a AAAA was assembled into a special construct as described in FIG. 5 (e.g., a pUC19 based plasmid vector). The construct was designed so that a Nickase, a Nb.BssSI and a IISRE BseRI are used together. The Addamer is released from the plasmid backbone. The Addamer can refold the hairpins at either end. The hairpins can be refold by a temperature change. In some instances, cleanup inactivates the enzymes. The remaining nicks are further a refolded Addamer. The refolded Addamer are healed by ligation using a ligase enzyme (e.g., T4DNA ligase). This is followed by exonuclease treatment. The final Addamer is recovered (FIG. 6).
[0183] Generation of a 3-base overhangs using a two-enzyme system
[0184] Described herein is a generation of a 3 ’base overhangs using a two-enzyme system. In some instances, Addamer usage efficiency is increased as the length of the overhangs decreases while 4-base overhangs. The 4-base overhangs are generated by some IISREs. The IISREs are robust and efficient for ligation. In some instances, they have some technical difficulties. For example, there is a huge reduction in ligation efficiency for those 4 base overhangs that are palindromic. In some instances, sequences such as CGCG or ACGT can pose a challenge during assembly. Palindromes force the assembly path to use adapter Addamers to shift the sequence phase to avoid palindromes.
[0185] There is thus a need in the art for new offset cutting enzymes or enzyme systems that can result in unproblematic overhangs. There are a few 3 base overhang cutting IISREs but they are very limited and, for example, Earl and BspQI have the same core DNA recognition sequence but BspQI has one extra base at the end. Hence, if one wants to use both enzymes, BspQI must be used first.
[0186] To generate new overhang lengths, in this disclosure, at least eight enzy me systems have been validated as generating 3 base overhangs that are specifically ligatable. Each of the enzyme systems uses an offset cutting Nickase, which is deployed first, followed by a double strand cutting IISRE (FIG. 7).
[0187] In a first test, each of the eight Addamers were cut then ligated to produce a new payload, which was verified by NGS sequencing. In a second test (FIG. 8) each of six Addamer pairs were combined to generate BspQI-CTCGA-BspQI payload Addamers (FIGs. 9 and 10).
[0188] The general notion of using these two enzyme systems was shown here but also shown in the previous Additional Example to generate a flap that folds into a hairpin allowingAddamer recovery. These systems can be used to generate essentially any overhang length needed.
[0189] Generation of blunt ends using a two-enzyme system
[0190] Described herein is the process of generating blunt ends using the two-enzyme system. Blunt ends are needed for DNA synthesis. In anticipation of DNA synthesis, the two- enzyme system is used to generate blunt ends. The two-enzyme system relies on two enzymes to generate blunt ends. In some instances, DNA binding sites for distinct pairs of nickase and IISREs w ere designed. Described herein are unique combinations of nickase sequences and IISRE sequences (FIG. 33). In some instances, the nickase enzyme is BspQI. In some instances, the nickase enzyme is BbvCI. In some instances, the nickase enzy me is BsmI. In some instances, the nickase enzyme is BsrDI. In some instances, the nickase enzy me is BssSI. In some instances, the nickase enzyme is Btsl. In some instances, the nickase enzy me is AlwI. In some instances, the nickase enzyme is BbvCI. In some instances, the nickase enzy me is BsmAI. In some instances, the nickase enzy me is BstNBI. In some instances, the IISRE sequence is Mlyl. In some instances, the IISRE sequence is NgoAVII. In some instances, the IISRE sequence is SspD5I. In some instances, the IISRE sequence is AlwI. In some instances, the IISRE sequence is Ajul. In some instances, the IISRE sequence is Alol. In some instances, the IISRE sequence is Bccl. In some instances, the IISRE sequence is Bcefl. In some instances, the IISRE sequence is Piel. In some instances, the IISRE sequence is BceAI. In some instances, the IISRE sequence is BceSIV. In some instances, the IISRE sequence is BscAI. In some instances, the IISRE sequence is BspD6I. In some instances, the IISRE sequence is Faul. In some instances, the IISRE sequence is Earl. In some instances, the IISRE sequence is BspQI. In some instances, the IISRE sequence is BfuAI. In some instances, the IISRE sequence is PaqCI. In some instances, the IISRE sequence is Esp31. in some instances, the IISRE sequence is Bbsl. In some instances, the IISRE sequence is BtgZI. In some instances, the IISRE sequence is Fokl. In some instances,the IISRE sequence is BsmFI. In some instances, the IISRE sequence is Bsal. In some instances, the IISRE sequence is BcoDI. In some instances, the IISRE sequence is Hgal. In some instances, the IISRE sequence is SfaNI. In some instances, the IISRE enzy me is BbvI.
[0191] Biactive Addamer Designs
[0192] The two enzyme system can further be generated with a 3’ overhang or a 5‘ overhang. The two enzyme system generated with a 5’ overhang can rely on a variety' of pairs of enzy mes (FIG. 34) as described herein.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A composition comprising: a. a double stranded DNA sequence comprising: i. an N-mer sequence; ii. at least one offset-cutting nickase DNA binding sequence configured to bind a nickase that cleaves a first strand or a second strand of the double stranded DNA sequence; and iii. at least one Type II S restriction endonuclease (1ISRE) DNA binding sequence configured to bind a IISRE that cleaves the first stand and the second strand of the N-mer sequence, wherein the nickase DNA binding sequence is located between the N-mer sequence and the IISRE DNA binding sequence.
2. The composition of claim 1 , further comprising a nickase.
3. The composition of claim 2, wherein the nickase and the double stranded DNA sequence is in a single volume.
4. The composition of any one of claims 1-3, wherein the double stranded DNA sequence is not coupled to a solid support.
5. The composition of claim 4, further comprising a Type II S restriction endonuclease (IISRE).
6. The composition of any one of claims 1-5, wherein the nickase cleaved site is within or adjacent to the N-mer sequence.
7. The composition of any one of claims 1 -6, wherein the nickase DNA binding sequence comprises enzymes: BbvCI, BsmI, BsrDI, BssSI, BtsI, Alwl, BbvCI, BsmAI, BspQI, or BstNBI.
8. The composition of any one of claims 1-7, wherein the IISRE DNA binding sequence comprises: Mlyl, NgoAVII, SspD5I, Alwl, Ajul, Alol, BccI, Bcefl, Piel, BceAI, BceSIV, BscAI, BspD6I, Faul, Earl, BspQI, BfuAI, PaqCI, Esp3I, BbsI, Bbvl, BtgZI, FokI, BsmFI, Bsal, BcoDI, Hgal, or SfaNI.
9. The composition of any one of claims 1-8, wherein the IISRE DNA binding sequence comprises Earl.
10. The composition of any one of claims 1-9, further comprising at least one hairpin structure.
11. The composition of any one of claims 1-1044, wherein at least one hairpin structure comprises an aptamer sequence.
12. The composition of any one of claims 1-11, further comprising a second IISRE DNA binding sequence.
13. The composition of claim 12, wherein the second IISRE DNA binding sequence comprises a BspQI DNA binding sequence.
14. The composition of any one of claims 1-13, wherein a core melting temperature of the double stranded DNA sequence is at least 60°C.
15. A method of synthesizing a target sequence in solution comprising: a. providing a plurality of DNA sequences in solution, wherein a DNA sequence of the plurality' of DNA sequences comprises aN-mer, a IISRE sequence, and at least one hairpin structures; b. providing at least one IISRE to the solution, wherein the at least IISRE cleaves the IISRE sequence, thereby exposing an M-base overhang; c. ligating at least two DNA sequences of the plurality of the DNA sequences having the M-base overhang to generate one or more ligated DNA sequences, wherein a ligated DNA sequence comprises a sequence having a length of 2N-M; and d. repeating (a)-(c) using the one or more ligated DNA sequences to generate a final DNA sequence comprising a target sequence.
16. The method of claim 15, wherein M is at least 3.
17. The method of claim 15, wherein M is 3.
18. The method of any one of claims 15-17, wherein the N-mer is at least a 4-mer.
19. The method of any one of claims 15-17, wherein the N-mer is no more than the 4-mer.
20. The method of any one of claims 15-19, wherein the target sequence comprises a length of about 50 bases to 5000 bases.
21. The method of any one of claims 15-19, wherein the target sequence comprises the length of about 100 bases to 5000 bases.
22. The method of any one of claims 15-19, wherein the target sequence comprises the length of about 50 bases to 300 bases.
23. The method of any one of claims 15-22, further comprising (d) providing an exonuclease to the solution to remove any DNA sequence of the plurality of DNA sequences having the M-base overhang.
24. The method of any one of claims 15-23, wherein the DNA sequence further comprises a nickase DNA binding sequence.
25. The method of any one of claims 15-24, further comprising providing at least one nickase to the solution prior to providing the at least one IISRE to the solution, wherein the at least one nickase cleaves the nickase sequence.
26. The method of any one of claims 15-25, wherein the target sequence is generated with an error rate of less than 1 in 100,000.
27. The method of any one of claims 15-26, wherein the solution comprises: a. 50 mM Potassium acetate, 20 mM Tris-acetate, 10 mM Magnesium acetate, 100 pg / mL rAlbumin, pH 7.9 at 25°C, b. 100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCh, 100 pg / mL rAlbumin. pH 7.9 at 25°C. or c. 66 mM Potassium acetate, 33 mM Tris-acetate, 10 mM Magnesium acetate, 100 pg / mL Bovine Serum Albumin, pH 7.9 at 37°C.
28. The method of claim 27, wherein the solution further comprises 10 mM DTT and either 1 mM ATP or 1 pM ATP.
29. The method of any one of claims 15-28, wherein (a), (b). or (c) is performed at a temperature of at least 10°C.
30. The method of any one of claims 15-29, wherein (a), (b), or (c) is performed cycling between one or more temperatures: about 10°C, 16°C, 25°C, 37°C, 50°C, 55°C, 60°C or 65°C.
31. The method of any one of claims 15-30, wherein the DNA sequence is formed using a ligase enzyme.
32. The method of any one of claims 15-31, wherein the ligase enzyme is T4 DNA ligase, T7 DNA ligase, T3 DNA ligase, or human DNA ligase III.
33. The method of any one of claims 15-32, wherein the synthesized target sequence has a purity of at least 80%.
34. The method of any one of claims 15-33, wherein the DNA sequence further undergoes heat inactivation of the enzymes directly or through protease treatment, wherein the protease treatment is inactivated by heat.
35. A composition comprising:(a) a double stranded DNA sequence, wherein the double stranded DNA sequence comprises: i. an N-mer sequence; ii. a nickase binding sequence, wherein the nickase binding sequence is at most 4 bases from a nick site; and iii. a first Type II S restriction endonuclease (IISRE) binding sequence configured to bind a IISRE that cleave the N-mer sequence, wherein the nickase binding sequence is located between the N-mer sequence and the IISRE binding sequence.
36. The composition of claim 35, further comprising a nickase.
37. The composition of claim 36, wherein the nickase and the double stranded DNA sequence is in a single volume.
38. The composition of any one of claims 35-37, wherein the double stranded DNA sequence is not coupled to a solid support.
39. The composition of claim 38, further comprising a Type II S restriction endonuclease (IISRE).
40. The composition any one of claims 35-39, wherein the nickase cleaved site is within or adjacent to the N-mer sequence.
41. The composition of any one of claims 35-40, wherein the nickase binding sequence comprises enzy mes: BbvCI, BsmI, BsrDI, BssSI, BtsI, Alwl, BbvCI, BsmAI, BspQI, or BstNBI.
42. The composition of any one of claims 35-41, wherein the IISRE binding sequence comprises: Mlyl. NgoAVII, SspD5I. Alwl, Ajul, Alol, BccI, Bcefl. Piel, BceAI, BceSIV, BscAI, BspD6I, Faul, Earl, BspQI, BfuAI, PaqCI, Esp3I, BbsI, Bbvl, BtgZI, FokI, BsmFI, Bsal, BcoDI, Hgal, or SfaNI.
43. The composition of any one of claims 35-42, wherein the IISRE binding sequence comprises Earl.
44. The composition of any one of claims 35-43, further comprising at least one hairpin structure.
45. The composition of any one of claims 35-44, wherein at least one hairpin structure comprises an aptamer sequence.
46. The composition of any one of claims 35-45, further comprising a second IISRE binding sequence.
47. The composition of claim 46, wherein the second IISRE binding sequence comprises a BspQI sequence.
48. The composition of any one of claims 35-47, wherein a core melting temperature of the double stranded DNA sequence is at least 60°C.
49. A system for synthesizing a target sequence in solution comprising: a. a plurality of DNA sequences in solution, wherein a DNA sequence of the plurality' of DNA sequences comprises aN-mer, a IISRE sequence, and at least one hairpin structures; b. at least one IISRE in the solution, wherein the at least IISRE is configured to cleave the IISRE sequence, thereby exposing an M-base overhang; c. at least two DNA sequences of the plurality of the DNA sequences configured to have the M-base overhang are configured to generate one or more ligated DNA sequences, wherein a ligated DNA sequence comprises a sequence having a length of 2N-M; and d. one or more ligated DNA sequences are configured to generate a final DNA sequences comprising the target sequence.
50. The system of claim 49, wherein M is at least 3.
51. The system of claim 49, wherein M is 3.
52. The system of any one of claims 49-51, wherein the N-mer is at least a 4-mer.
53. The system of any one of claims 49-51, wherein the N-mer is at most the 4-mer.
54. The system of any one of claims 49-53, wherein the target sequence comprises a length of about 50 bases to 5000 bases.
55. The system of any one of claims 49-53, wherein the target sequence comprises the length of about 100 bases to 5000 bases.
56. The system of any one of claims 49-53, wherein the target sequence comprises the length of about 50 bases to 300 bases.
57. The system of any one of claims 49-56, further comprising (d) an exonuclease in the solution configured to remove any DNA sequence of the plurality of DNA sequences having the M-base overhang.
58. The system of any one of claims 49-57, wherein the DNA sequence further comprises a nickase sequence.
59. The system of any one of claims 49-58, further comprising at least one nickase in the solution, wherein the at least one nickase cleaves the nickase sequence.
60. The system of any one of claims 49-59, wherein the target sequence is generated with an error rate of less than 1 in 100,000.
61. The system of any one of claims 49-60, wherein the solution comprises: a. 50 mM Potassium acetate, 20 mM Tris-acetate, 10 mM Magnesium acetate, 100 pg / mL rAlbumin, pH 7.9 at 25°C, b. 100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCty, 100 pg / mL rAlbumin, pH 7.9 at 25°C, or c. 66 mM Potassium acetate, 33 mM Tris-acetate, 10 mM Magnesium acetate, 100 pg / mL Bovine Serum Albumin, pH 7.9 at 37°C.
62. The system of any one of claims 49-61, wherein the solution further comprises 10 mM DTT and either 1 mM ATP or 1 pM ATP.
63. The system of any one of claims 49-62, further comprising a temperature of at least 10°C.
64. The system of any one of claims 49-63, wherein the DNA sequence is formed using a ligase enzyme.
65. The system of any one of claims 49-64, wherein the ligase enzyme is T4 DNA ligase, T7 DNA ligase, T3 DNA ligase, or human DNA ligase III.
66. The system of any one of claims 49-65, wherein the synthesized target sequence has a purity of at least 80%.
67. The system of any one of claims 49-66, wherein the DNA sequence further undergoes heat inactivation of the enzymes directly or protease treatment, wherein the protease treatment is inactivated by heat.
Citation Information
Patent Citations
Linear DNA with enhanced resistance against exonucleases and methods for the production thereof
CA3224561A1
Compositions and methods for template-free double stranded geometric enzymatic nucleic acid synthesis
WO2021055962A1
Linked-read sequencing library preparation
WO2022081940A1
Compositions and methods for phosphoramedite-free enzymatic synthesis of nucleic acids
WO2022232134A1
Gene assembly from oligonucleotide pools
WO2022240991A1