Gene assembly from oligonucleotide pools
Patent Information
- Application Number
- JP2023569973
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-11
- Filing Date
- 2022-05-11
- Publication Date
- 2025-05-20
AI Technical Summary
Existing gene assembly methods, such as the gSynth method, face challenges in achieving high-fidelity and high-throughput production of double-stranded DNA sequences, with a need for improved efficiency and throughput in oligonucleotide synthesis.
The use of compositions comprising multiple nucleic acid molecules with different sequences that hybridize to form hybridized complexes, combined with adamer technology, which involves hybridizing single-stranded nucleic acid molecules to form double-stranded hairpin structures capped by ligase enzymes, followed by exonuclease treatment to purify the adamers.
This approach enhances the production of high-fidelity gene assemblies with increased throughput by utilizing high-throughput oligonucleotide synthesis, improving the efficiency and purity of DNA synthesis while reducing the need for error-prone DNA polymerases and large amounts of purified nucleotides.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 186,871, filed May 11, 2021, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0002] Sequence Listing This application has been submitted in ASCII format via EFS-Web and contains a Sequence Listing, which is incorporated herein by reference in its entirety. The ASCII copy, created on May 11, 2022, is named "DNWR-010_001WO_SeqList.txt" and is approximately 23,462 bytes in size. [Background technology]
[0003] There is a need in the art for gene assembly that efficiently produces arbitrary double-stranded DNA sequences with high fidelity and high throughput. Recent advances in gene assembly include a double-stranded DNA assembly method called the gSynth method, which is described in detail in PCT Application No. PCT / US2020 / 051838 and published as International Publication No. WO2021055962A1. Although the gSynth method produces high fidelity DNA sequence assembly, the use of high throughput DNA synthesis techniques such as array-based oligonucleotide synthesis (see Lipshutz, RJ et al., High density synthetic oligonucleotide arrays. Nature Genetics volume 21, pages 20-24, 1999) to produce the materials and double-stranded elements required for the gSynth method requires intermediate amplification (Saiki RK et al., Enzymatic amplification of beta-globin genomic sequences and restriction site analysis for diagnosis of sickle cell anemia. Science 20 Dec 1985: Vol. 230, Issue 4732, pp. 1350-1354), thus requiring an increase in the throughput of the gSynth method.
[0004] In addition to the gSynth method, an environmentally friendly (e.g., "green") double-stranded DNA assembly / synthesis method has recently been developed that uses as a basic building block double-stranded DNA construct called an "adamer." An adamer is a double-stranded duplex hairpin structure that carries the DNA payload as well as various regulatory elements that allow for sequence manipulation. Adamer elements are composed of a variety of binding sites, including binding sites for restriction endonucleases (RE), binding sites for type II S restriction endonucleases (IISRE), payload DNA sequences, and a wide variety of nucleotide sequences ranging from simple GNA motifs (Yoshizawa S et al., GNA Trinucleotide Loop Sequences Producing Extraordinarily Stable DNA Minihairpins. Biochemistry 1997,36,16,4761-4767) to complex three-dimensional aptamers with high affinity ligand binding (Ellington. AD and Szostak, JW In vitro selection of RNA molecules that bind specific ligands. Nature volume 346,pages818-822.1990; Tuerk C et al., Systematic evolution of ligands by exponential enrichment: RNA ligands to bacteriophage T4 DNA polymerase. Science 03 Aug 1990:Vol.249,Issue 4968, pp. 505-510).
[0005] Disclosed herein are compositions and methods that utilize high-throughput oligonucleotide synthesis methods, including array-based oligonucleotide synthesis, to generate the materials required for the gSynth and Adamer-based DNA synthesis methods described above and herein. Thus, disclosed herein are compositions and methods that combine gSynth and Adamer technologies with high-throughput oligonucleotide synthesis to generate conventional high-fidelity gene (target nucleic acid) synthesis. Summary of the Invention
[0006] overview The disclosure provides compositions comprising two or more pluralities of nucleic acid molecules, each of the pluralities of nucleic acid molecules comprising two or more species of nucleic acid molecules, wherein the different species of nucleic acid molecules comprise different nucleic acid sequences, and wherein at least one set of corresponding pluralities is present within the two or more pluralities of nucleic acid molecules such that when at least one set of corresponding pluralities are combined in a single reaction volume, nucleic acid molecules from the different pluralities within the set hybridize together to form at least one hybridized complex.
[0007] In some embodiments, a composition of the disclosure comprises a plurality of about a) 6, b) 10, c) 15, d) 20, or e) 50 nucleic acid molecules.
[0008] In some embodiments of the disclosed compositions, each plurality of nucleic acid molecules comprises at least about 25 different species of nucleic acid molecules.
[0009] In some aspects of compositions of the present disclosure, each set of corresponding pluralities includes the same number of pluralities.
[0010] In some aspects of the compositions of the present disclosure, at least one species of hybridized complex comprises one nucleic acid species from each of the plurality within the corresponding plurality of sets.
[0011] In some aspects of the disclosed compositions, the two or more plurality of nucleic acid molecules are present in separate volumes.
[0012] In some embodiments of compositions of the present disclosure, at least one set of corresponding plurality includes a plurality of at least about a) two, b) three, c) four, or d) five nucleic acid molecules.
[0013] In some aspects of the disclosed compositions, the number of corresponding plurality of sets is:
number
[0014] In some embodiments of the compositions of the present disclosure, a) at least one set of corresponding pluralities comprises a plurality of two, and at least one hybridized complex comprises two nucleic acid molecules; b) at least one set of corresponding pluralities comprises a plurality of three, and at least one hybridized complex comprises three nucleic acid molecules; or c) at least one set of corresponding pluralities comprises a plurality of four, and at least one hybridized complex comprises four nucleic acid molecules.
[0015] In some aspects of the disclosed compositions, when corresponding sets are combined in a single reaction, at least about five different hybridized complex species are formed.
[0016] In some aspects of the disclosed compositions, nucleic acid molecules of different species within a single plurality of nucleic acid molecules are not complementary to one another.
[0017] The present disclosure provides a method for producing at least one adamer, the method comprising: a) providing a composition of the present disclosure; b) combining at least one set of corresponding pluralities of nucleic acid molecules in a single reaction volume such that at least one hybridized complex is formed; and c) contacting at least one of the hybridized complexes with at least one ligase enzyme to form the at least one adamer capped at both ends by a hairpin.
[0018] In some embodiments, the methods of the disclosure further comprise treating the product of step (c) with an exonuclease, thereby purifying the properly ligated adamers.
[0019] In some embodiments, the methods of the disclosure further comprise contacting at least one hybridized complex with a MutS enzyme.
[0020] In some embodiments, the adamer comprises a) a first type II S restriction endonuclease (IISRE) sequence, b) a payload sequence, and c) at least a second IISRE sequence, and at least one end of the adamer comprises a hairpin structure.
[0021] In some embodiments, the adamer comprises a hairpin structure at both ends of the adamer.
[0022] In some embodiments, the adamer comprises a) a first IISRE sequence, b) a second IISRE sequence, c) a payload sequence, and d) at least a third IISRE sequence.
[0023] In some embodiments, the adamer comprises a) a first IISRE sequence, b) a second IISRE sequence, c) a payload sequence, d) a third IISRE sequence, and e) at least a fourth IISRE sequence.
[0024] In some embodiments, the adamer further comprises a multiple cloning site (MCS) sequence, which comprises one or more restriction endonuclease sequences.
[0025] In some embodiments, at least one of the IISRE sequences is selected from the group consisting of MlyI, NgoAVII, SspD5I, AlwI, BccI, BcefI, PleI, BceAI, BceSIV, BscAI, BspD6I, FauI, EarI, BspQI, BfuAI, PaqCI, Esp3I, BbsI, BbvI, BtgZI, FokI, BsmFI, BsaI, BcoDI, and HgaI sequences.
[0026] In some embodiments, at least one hairpin structure comprises an aptamer sequence, hi some embodiments, the aptamer sequence is selected from a pL1 aptamer sequence, a thrombin 29-mer aptamer sequence, a S2.2 aptamer sequence, an ART1172 aptamer sequence, a R12.45 aptamer sequence, a Rb008 aptamer sequence, and a 38NT SELEX aptamer sequence.
[0027] The present disclosure provides a method for synthesizing a nucleic acid molecule comprising a target nucleic acid sequence, the method comprising: a) providing a composition of any one of the preceding claims, the composition comprising corresponding sets of nucleic acid molecules such that when the corresponding sets are combined in a single reaction volume, nucleic acid molecules from different sets within the sets hybridize together to form two or more hybridized complexes, the two or more hybridized complexes comprising fragments of the target nucleic acid sequence; b) combining the corresponding sets of nucleic acid molecules in a single reaction volume, such that two or more hybridized complexes are formed; c) contacting the two or more hybridized complexes with at least one ligase enzyme to form two or more adamers capped at both ends by hairpins; and d) assembling the two or more adamers to synthesize a nucleic acid molecule comprising the target nucleic acid sequence.
[0028] In some embodiments of the methods disclosed herein, assembling two or more adumers comprises treating the adumers with a) one or more restriction enzymes and b) one or more ligases, either simultaneously or sequentially, thereby assembling a nucleic acid molecule that includes a target nucleic acid sequence.
[0029] In some embodiments of the disclosed methods, assembling two or more adamers comprises treating the adamers with a) an engineered Cas9 exhibiting nickase activity in combination with at least one guide RNA, and b) one or more ligases, either simultaneously or sequentially, thereby assembling a nucleic acid molecule comprising a target nucleic acid sequence.
[0030] Any of the aspects and / or embodiments described above and herein may be combined with any other aspect and / or embodiment described above and herein.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs. In this specification, the singular also includes the plural unless the context clearly indicates otherwise, and for example, the terms "a", "an" and "the" are understood to be singular or plural, and the term "or" is understood to be inclusive. As an example, "element" means one or more elements. Throughout this specification, the term "comprises" or variations such as "comprises" or "comprises" will be understood to mean the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps. About can be understood to be within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Throughout this specification, recitations of ranges of values stated as "x to y" or "x to y" (where x and y are two values) are understood to include x and y. Unless otherwise clear from the context, all numerical values provided herein are modified by the term "about."
[0032] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below. All publications, patent applications, patents, and other references described herein are incorporated by reference in their entirety for all purposes. References cited herein are not admitted to be prior art to the claimed invention. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting. Other features and advantages of the present disclosure will be apparent from the following detailed description and claims. [Brief description of the drawings]
[0033] The above and further features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings.
[0034] [Figure 1] FIG. 1 shows a schematic diagram illustrating the generation of an adamer from a chemically synthesized oligonucleotide pair. FIG. 1 shows the construction of a DNA adamer from two single-stranded DNA molecules. The top and bottom strands, the DUPLEX INSERT and the two nicking sites are as shown. The nucleotide sequences of the 10 or more nucleotides presented in FIG. 1 are set forth in SEQ ID NOs: 1-7. [Diagram 2] FIG. 2 shows a double-stranded system with pairwise combination pools to generate hadamar sets for gene assembly. The left panel shows the top and bottom strand pool positions for hadamar generation. Each combination of six pools provides a unique combination of complementary top and bottom strands. There are 15 possible pairwise combinations from the six pools, which correspond to genes A through O. The right panel shows the relationship between the number of pools and the number of possible gene assemblies, the number of hadamar gene sets per pool, and the total number of different oligonucleotides per pool. [Diagram 3]Figure 3 shows the generation of adamers from multiple oligonucleotide array pools, two-stranded and three-stranded designs for adamer generation. For the 300 base strand, the overlapping complementary region is long, 200 base pairs for the double-stranded system and 200 base pairs between the left and middle strands and 100 base pairs for the middle and right strands for the triple-stranded system. For the double-stranded system, there are two nicks resulting from successful hybridization and three nicks for the triple-stranded system. These nicks are rapidly and efficiently resolved by T4 DNA ligase. [Figure 4] Figure 4 shows three standard systems with ternary combinations for generating hadamar sets for gene assembly. The left panel shows the pool positions of the middle and right strands for hadamar generation. Each combination of six pools provides a unique combination of complementary left, middle and right strands. There are 20 possible pairwise combinations from the six pools, corresponding to genes A through T. The right panel shows the relationship between the number of pools and the number of possible gene assemblies, the number of hadamar gene sets per pool, and the total number of different oligonucleotides per pool. [Diagram 5] Figure 5 shows a three stranded system with six pool array layout. On the left, Figure 5 shows a combination map of the three stranded system across six oligonucleotide array pools, and on the right, an exemplary array layout. Each combination of the three pools results in a single gene set of adamers. Here, gene S (light purple) is constructed from five adamers encoded by five separate oligonucleotides (A, B, C, D, and E) on each of pools 1, 4, and 6. [Figure 6] Figure 6 shows a schematic illustrating the use of a pooled approach to combine adamers to generate high-fidelity gene sequences. Internal IISRE sites are indicated by black speckles. gSynth sites are indicated by numbered boxes. [Figure 7]FIG. 7 is an exemplary schematic diagram of an adamer, a double-stranded nucleic acid molecule containing hairpins at both ends. [Figure 8] 8 shows an exemplary schematic of various adamer designs of the present disclosure. Several examples of adamer designs are shown, including means for attachment to solid supports and double hairpins to promote exonuclease resistance, with each adamer type having a payload with flanking IISRE sites. The legend shows several possible IISRE and RE sites, as well as the thrombin aptamer hairpin. [Figure 9] 9 is an exemplary schematic diagram of a method of synthesizing nucleic acids of the present disclosure, including the use of the disclosed adamers. First, a first adamer and a second adamer are ligated to a binding stud with an MCS already loaded on a solid support using DNA ligation. A donor construct and an acceptor construct are generated in separate volumes. The donor construct and the acceptor construct are treated with separate IISREs to generate ligatable ends. In this figure, the acceptor is generated by digestion with the R1 IISRE, and the released ends and enzyme are discarded by rinsing. A donor construct is generated by digestion with the purple L2 enzyme. The donor construct solution (with L2 enzyme) is transferred to the acceptor well and ligated using T4 DNA ligase, which has a high efficiency of over 80% for 2-, 3-, or 4-base sticky end ligation. The well is treated with exonuclease and rinsed. The resulting adamer construct is then ready for the subsequent extension cycle. [Figure 10A]10A, 10B, 10C, 10D, 10E, 10F, 10G, and 10H are exemplary schematic diagrams of a method of synthesizing a nucleic acid of the present disclosure, including the use of an adamer of the present disclosure to synthesize a target nucleic acid molecule of 27 nucleotides in length. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10A corresponds to that set forth in SEQ ID NO:8. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10E corresponds to that set forth in SEQ ID NO:9-10. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10G corresponds to that set forth in SEQ ID NO:11-22. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10H corresponds to that set forth in SEQ ID NO:23-33. [Figure 10B] 10A, 10B, 10C, 10D, 10E, 10F, 10G, and 10H are exemplary schematic diagrams of a method of synthesizing a nucleic acid of the present disclosure, including the use of an adamer of the present disclosure to synthesize a target nucleic acid molecule of 27 nucleotides in length. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10A corresponds to that set forth in SEQ ID NO:8. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10E corresponds to that set forth in SEQ ID NO:9-10. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10G corresponds to that set forth in SEQ ID NO:11-22. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10H corresponds to that set forth in SEQ ID NO:23-33. [Figure 10C]10A, 10B, 10C, 10D, 10E, 10F, 10G, and 10H are exemplary schematic diagrams of a method of synthesizing a nucleic acid of the present disclosure, including the use of an adamer of the present disclosure to synthesize a target nucleic acid molecule of 27 nucleotides in length. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10A corresponds to that set forth in SEQ ID NO:8. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10E corresponds to that set forth in SEQ ID NO:9-10. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10G corresponds to that set forth in SEQ ID NO:11-22. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10H corresponds to that set forth in SEQ ID NO:23-33. [Figure 10D] 10A, 10B, 10C, 10D, 10E, 10F, 10G, and 10H are exemplary schematic diagrams of a method of synthesizing a nucleic acid of the present disclosure, including the use of an adamer of the present disclosure to synthesize a target nucleic acid molecule of 27 nucleotides in length. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10A corresponds to that set forth in SEQ ID NO:8. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10E corresponds to that set forth in SEQ ID NO:9-10. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10G corresponds to that set forth in SEQ ID NO:11-22. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10H corresponds to that set forth in SEQ ID NO:23-33. [Figure 10E]10A, 10B, 10C, 10D, 10E, 10F, 10G, and 10H are exemplary schematic diagrams of a method of synthesizing a nucleic acid of the present disclosure, including the use of an adamer of the present disclosure to synthesize a target nucleic acid molecule of 27 nucleotides in length. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10A corresponds to that set forth in SEQ ID NO:8. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10E corresponds to that set forth in SEQ ID NO:9-10. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10G corresponds to that set forth in SEQ ID NO:11-22. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10H corresponds to that set forth in SEQ ID NO:23-33. [Figure 10F] 10A, 10B, 10C, 10D, 10E, 10F, 10G, and 10H are exemplary schematic diagrams of a method of synthesizing a nucleic acid of the present disclosure, including the use of an adamer of the present disclosure to synthesize a target nucleic acid molecule of 27 nucleotides in length. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10A corresponds to that set forth in SEQ ID NO:8. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10E corresponds to that set forth in SEQ ID NO:9-10. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10G corresponds to that set forth in SEQ ID NO:11-22. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10H corresponds to that set forth in SEQ ID NO:23-33. [Figure 10G]10A, 10B, 10C, 10D, 10E, 10F, 10G, and 10H are exemplary schematic diagrams of a method of synthesizing a nucleic acid of the present disclosure, including the use of an adamer of the present disclosure to synthesize a target nucleic acid molecule of 27 nucleotides in length. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10A corresponds to that set forth in SEQ ID NO:8. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10E corresponds to that set forth in SEQ ID NO:9-10. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10G corresponds to that set forth in SEQ ID NO:11-22. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10H corresponds to that set forth in SEQ ID NO:23-33. [Figure 10H] 10A, 10B, 10C, 10D, 10E, 10F, 10G, and 10H are exemplary schematic diagrams of a method of synthesizing a nucleic acid of the present disclosure, including the use of an adamer of the present disclosure to synthesize a target nucleic acid molecule of 27 nucleotides in length. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10A corresponds to that set forth in SEQ ID NO:8. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10E corresponds to that set forth in SEQ ID NO:9-10. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10G corresponds to that set forth in SEQ ID NO:11-22. The nucleotide sequence of 10 or more nucleotides presented in FIG. 10H corresponds to that set forth in SEQ ID NO:23-33. [Figure 11A] 11A-11F are schematic diagrams of the double-stranded geometric synthesis (gSynth) of the present disclosure. FIG. 11A is the sequence to be synthesized using the double-stranded gSynth method of the present disclosure. The bold and underlined portions of the sequence correspond to the selected 4-mer overhangs and thus define the fragments used to synthesize the entire sequence. The nucleotide sequence of 10 or more nucleotides presented in FIG. 11A corresponds to that set forth in SEQ ID NO:34. [Figure 11B]Figure 11B shows the individual double-stranded nucleic acid fragments of the sequence shown in Figure 11A that are used in the double-stranded gSynth method of the present disclosure to construct the sequence shown in Figure 11A. These fragments are chosen based on the sites selected in Figure 11A. The nucleotide sequences of 10 or more nucleotides presented in Figure 11B correspond to those set forth in SEQ ID NOs: 35-62. [Figure 11C] FIG. 11C is a schematic diagram of a binary tree showing the order in which the fragments in FIG. 11B should be assembled to produce the sequence shown in FIG. 11A. [Figure 11D] Figure 11D is a schematic diagram of the first round of ligation in the double-stranded gSynth method for synthesizing the sequence shown in Figure 11A. In the first ligation round, fragments 1 and 2, fragments 3 and 4, fragments 5 and 6, fragments 7 and 8, fragments 9 and 10, fragments 11 and 12, and fragments 13 and 14 are hybridized via their complementary 5' overhangs and then ligated together to generate fragments 1+2, fragments 3+4, fragments 5+6, fragments 7+8, fragments 9+10, fragments 11+12, and fragments 13+14. The nucleotide sequences of the 10 or more nucleotides presented in Figure 11D correspond to those set forth in SEQ ID NOs: 35-62. [Figure 11E] Figure 11E is a schematic diagram of the second round of ligation in the double-stranded gSynth method for synthesizing the sequence shown in Figure 11A. In the second ligation round, fragments 1+2 and 3+4, fragments 5+6 and 7+8, and fragments 11+12 and 13+14 are hybridized via their complementary 5' overhangs and then ligated together to generate fragments 1+2+3+4, fragments 5+6+7+8, and fragments 11+12+13+14. The nucleotide sequences of the 10 or more nucleotides presented in Figure 11E correspond to those set forth in SEQ ID NOs: 63-76. [Figure 11F]Figure 11F is a schematic diagram of the third round of ligation in the double-stranded gSynth method for synthesizing the sequence shown in Figure 11A. In the third ligation round, fragments 1+2+3+4 and 5+6+7+8, and fragments 9+10 and 11+12+13+14 are hybridized via their complementary 5' overhangs and then ligated together to generate fragments 1+2+3+4+5+6+7+8 and fragments 9+10+11+12+13+14. The nucleotide sequences of the 10 or more nucleotides presented in Figure 11F correspond to those set forth in SEQ ID NOs: 77-84. [Figure 11G] Figure 11G is a schematic diagram of the fourth and final round of ligation in the double-stranded gSynth method for synthesizing the sequence shown in Figure 11A. In the fourth ligation round, fragments 1+2+3+4+5+6+7+8 and 9+10+11+12+13+14 are hybridized via their complementary 5' overhangs and ligated together, thereby generating the sequence shown in Figure 1A. The nucleotide sequences of the 10 or more nucleotides presented in Figure 11G correspond to those set forth in SEQ ID NOs: 85-88. [Figure 12] FIG. 12 shows an exemplary processing and analysis of a target nucleic acid sequence to be synthesized using the disclosed method. The full-length sequence of 431 bp was divided into five variable-sized fragments (F1-F5). These fragments are the payload of the adumers used to generate the final sequence. A computer program was used to reliably predict well-spaced 4-base overhang sites compatible with the ligation reaction, as well as other features such as optimal GC content. The entire sequence was analyzed for the presence of IISRE sites (BsmFI, FokI, BtgZI, SfaNI). Sites present in the target sequence exclude certain IISREs for assembly purposes. [Figure 13]FIG. 13 shows the adamers corresponding to the fragments identified in FIG. 12. The payload of the adamers, which are predicted fragments (blue arrows indicate a 5'->3' orientation), is flanked by IISRE sites (e.g., BsaI for the internal fragments). The 4-base overhangs are indicated by light orange boxes (top strand) and light green boxes (bottom strand). For terminal fragments F1 and F5, the external IISRE site is different (BbsI) to allow for subsequent assembly after cloning of the sequences. Fragments F1 and F5 also have forward and reverse primer sites for amplification (extended T7 and extended T3) and RE sites for cloning (XbaI and EcoRI). The nucleotide sequences of 10 or more nucleotides presented in FIG. 11G correspond to those set forth in SEQ ID NOs: 89-108. [Figure 14] Figure 14 shows the validation of the target nucleic acid sequence assembled using the adamer presented in Figure 13 by sequencing. Clones were picked and the plasmids sequenced using capillary sequencing. The reverse and forward sequences were perfectly aligned, thus validating the accuracy of the adamer-based assembly from the pool of oligonucleotides. [Figure 15-1] FIG. 15 shows the design of an assembly strategy for additional target nucleic acid sequences assembled using the methods of the present disclosure. [Figure 15-2] FIG. 15 shows the design of an assembly strategy for additional target nucleic acid sequences assembled using the methods of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0035] The present disclosure relates to compositions and methods that utilize pooled nucleic acid molecules to generate double-stranded or partially double-stranded nucleic acid molecules, including but not limited to adumers, for use in double-stranded DNA assembly / synthesis methods, including but not limited to gSynth-based and adumers-based methods. That is, the compositions and methods described herein improve upon existing gSynth-based and adumers-based DNA assembly / synthesis methods by utilizing the very high oligonucleotide production potential of pooled nucleic acid molecules without the need for double-stranded mediated amplification, which requires large amounts of purified nucleotides and error-prone DNA polymerase. In some embodiments, these pooled nucleic acid molecules can be generated using array-based oligonucleotide synthesis methods.
[0036] In one embodiment, multiple oligonucleotide pools can be generated and then combined to construct individual genes uniquely from the adumers generated in these combined pools.For example, the gSynth algorithm can be used to break down any gene into a relatively small number of fragments with unique high-fidelity overhang sequences at each junction, allowing for straight forward sticky end ligation to generate a final product similar to the Golden Gate assembly approach (Pryor JMet al., Enabling one-pot Golden Gate assemblies of unprecedented complexity using data-optimized assembly design.PLoS ONE 15(9):e0238592.2020).It is important to note that Golden Gate assembly works best with cloned sequences that are uniformly digested and selected for high fidelity overlap. If the double stranded sequence is 500 base pairs (assuming two 300 base oligos are combined to create an adamer with 100 bases for the regulatory elements, a 500 base or 250 base pair duplex remains), then ligation of five consecutive sequences will generate a 2.5 kb gene. In some embodiments, the user of the disclosed method can vary the length of any individual gene fragment widely to best suit assembly, IISRE site potential and secondary structure parameters.
[0037] In some embodiments, an adamer (whose structure is described in more detail herein) can be generated by a unique combination of a pair of complementary DNA strands that hybridize with each other to form a double-stranded double hairpin structure with a pair of unresolved niches. These nicks can be repaired by application of a DNA ligase, such as T4 DNA ligase. Once the nicks are resolved, possible mismatches can be removed by His-tagged MutS protein in combination with an affinity reagent (Wang J et al., Directly fishing out subtle mutations in genomic DNA with histidine-tagged Thermus thermophilus MutS. Volume 547, Issues 1-2, 22 March 2004, Pages 41-47). In some embodiments, all non-adamer DNA is removed by application of T7 exonuclease.
[0038] DNA adamers can be constructed from two single-stranded DNA molecules. DUPLEX INSERTs can be large, on the order of 250 base pairs (which allows for 100 bases of regulatory sequence from the estimated single-stranded oligonucleotide size of 300 bases). Because the total sequence amount can be high, the fidelity of hybridization is also high in terms of both selectivity and double-stranded recovery. Annealing can generate almost completely double-stranded sequences with stable hairpins. The two "nicks" can be efficiently and quickly repaired by simple treatment with T4 DNA ligase. Combining limited ligation with high-fidelity hybridization, only T7 exonuclease-resistant material can be the desired adamer product. In addition, treatment can be performed with His-tagged Taq MutS protein to bind potential mismatched base pairs, followed by contact with Ni-NTA (nickel nitriloacetic acid) agarose affinity resin to remove MutS-bound mismatches containing DNA, which allows for a substantially enriched and pure desired product to be obtained (see Figure 1).
[0039] That is, in some embodiments, the adamer of the present disclosure can be generated by hybridizing a first single-stranded nucleic acid molecule and a second single-stranded nucleic acid molecule, as shown in the top panel of FIG. 3, where the first single-stranded nucleic acid molecule comprises a first region that is complementary to a second region on the second single-stranded nucleic acid molecule and a second region that is self-complementary, and the second single-stranded nucleic acid molecule comprises a first region that is self-complementary and a second region that is complementary to a first region on the first single-stranded nucleic acid molecule. The first single-stranded nucleic acid molecule and the second single-stranded nucleic acid molecule are hybridized together to generate a hybridized complex. The hybridized complex can then be optionally contacted with the enzyme MutS, which binds to mismatched bases and exposes the DNA to exonuclease digestion. The hybridized complex can then be contacted with a ligase enzyme to form a double-stranded adamer structure capped at both ends by hairpins. After contact with the ligase enzyme, the product can be contacted with T7 exonuclease to purify and concentrate the properly formed adamers, this method being referred to herein as the "double-stranded adammer assembly method."
[0040] Accordingly, the disclosure provides a method of generating an adamer as described herein, the method comprising: a) providing a first single-stranded nucleic acid molecule and a second single-stranded nucleic acid molecule, the sequences of the first single-stranded nucleic acid molecule and the second single-stranded nucleic acid molecule comprising a portion of the adamer to be generated, the first single-stranded nucleic acid molecule comprising a first region that is complementary to a second region on the second single-stranded nucleic acid molecule and a second region that is self-complementary, and the second single-stranded nucleic acid molecule comprising a first region that is self-complementary and a second region that is complementary to the first region on the first single-stranded nucleic acid molecule; b) hybridizing the first single-stranded nucleic acid molecule and the second single-stranded nucleic acid molecule; and c) contacting the partial double-stranded nucleic acid molecule with a ligase enzyme to form a double-stranded adamer structure capped at both ends by a hairpin.
[0041] In some embodiments, the above method may further comprise treating the product of step (c) with an exonuclease, thereby purifying the properly ligated adamers.
[0042] In some embodiments, the above methods can further comprise contacting the partially double-stranded nucleic acid molecule with a MutS enzyme after step (b) and before step (c).
[0043] In some embodiments, an adamer can be generated by hybridizing a first single-stranded nucleic acid molecule, a second single-stranded nucleic acid molecule, and a third single-stranded nucleic acid molecule, as shown in the middle panel of Figure 3, where the first single-stranded nucleic acid molecule comprises a first region that is complementary to a first region on the second single-stranded nucleic acid molecule and a second region that is self-complementary, the second single-stranded nucleic acid molecule comprises a first region that is complementary to a first region of the first single-stranded nucleic acid molecule and a second region that is complementary to a first region of the third single-stranded nucleic acid molecule, and the third single-stranded nucleic acid molecule comprises a first region that is complementary to a second region of the second single-stranded nucleic acid molecule and a second region that is self-complementary. In Figure 3, the first single-stranded nucleic acid molecule is referred to as the left strand, the second single-stranded nucleic acid molecule is referred to as the middle strand, and the third single-stranded nucleic acid molecule is referred to as the right strand. The first single-stranded nucleic acid molecule, the second single-stranded nucleic acid molecule, and the third single-stranded nucleic acid molecule are hybridized together to generate a hybridized complex. The hybridized complex can then be optionally contacted with the enzyme MutS, which binds to mismatched bases and exposes the DNA to exonuclease digestion. The hybridized complex can then be contacted with a ligase enzyme to form a double-stranded adamer structure capped at both ends by a hairpin. After contact with the ligase enzyme, the product can be contacted with T7 exonuclease to purify and concentrate the properly formed adamers. In a triple-stranded system, the position of the middle strand relative to the left and right strands can be adjusted to balance the degree of hybridization between the strands. However, there is a trade-off between the amount of self-hybridization to form a hairpin and the amount of cross-hybridization to form a complete adamer structure. This method is referred to herein as the "triple-stranded adamer assembly method."
[0044] Accordingly, the disclosure provides a method of generating an adamer as described herein, the method comprising: a) providing a first single-stranded nucleic acid molecule, a second single-stranded nucleic acid molecule, and a third single-stranded nucleic acid molecule, the sequences of the first single-stranded nucleic acid molecule, the second single-stranded nucleic acid molecule, and the third single-stranded nucleic acid molecule comprising a portion of the adamer to be generated, the first single-stranded nucleic acid molecule comprising a first region that is complementary to a first region on the second single-stranded nucleic acid molecule and a second region that is self-complementary, the second ... first single-stranded nucleic acid molecule and a second region that is self-complementary, a) hybridizing the first single-stranded nucleic acid molecule, the second single-stranded nucleic acid molecule, and the third single-stranded nucleic acid molecule; and b) contacting the partially double-stranded nucleic acid molecule with a ligase enzyme to form a double-stranded adamer structure capped at both ends by a hairpin.
[0045] In some embodiments, the above method may further comprise treating the product of step (c) with an exonuclease, thereby purifying the properly ligated adamers.
[0046] In some embodiments, the above methods can further comprise contacting the partially double-stranded nucleic acid molecule with a MutS enzyme after step (b) and before step (c).
[0047] In some embodiments, as shown in the bottom panel of FIG. 3, an adamer can be generated by hybridizing a first single-stranded nucleic acid molecule, a second single-stranded nucleic acid molecule, a third single-stranded nucleic acid molecule, and a fourth single-stranded nucleic acid molecule, where the first single-stranded nucleic acid molecule comprises a first region that is complementary to a first region on the second single-stranded nucleic acid molecule and a second region that is self-complementary, the second single-stranded nucleic acid molecule comprises a first region that is complementary to a first region of the first single-stranded nucleic acid molecule and a second region that is complementary to a first region of the third single-stranded nucleic acid molecule, the third single-stranded nucleic acid molecule comprises a first region that is complementary to a second region of the second single-stranded nucleic acid molecule and a second region that is complementary to a first region on the fourth single-stranded nucleic acid molecule, and the fourth single-stranded nucleic acid molecule comprises a first region that is complementary to a second region of the third single-stranded nucleic acid molecule and a second region that is self-complementary. The first single-stranded nucleic acid molecule, the second single-stranded nucleic acid molecule, the third single-stranded nucleic acid molecule, and the fourth single-stranded nucleic acid molecule are hybridized together to generate a hybridized complex. The hybridized complex can then be optionally contacted with the enzyme MutS, which binds to mismatched bases and exposes the DNA to exonuclease digestion. The hybridized complex can then be contacted with a ligase enzyme to form a double-stranded adamer structure capped at both ends by a hairpin. After contact with the ligase enzyme, the product can be contacted with T7 exonuclease to purify and concentrate the properly formed adamers. In a triple-stranded system, the position of the middle strand relative to the left and right strands can be adjusted to balance the degree of hybridization between the strands. However, there is a trade-off between the amount of self-hybridization to form a hairpin and the amount of cross-hybridization to form a complete adamer structure. This method is referred to herein as the "four-stranded adamer assembly method."
[0048] Accordingly, the disclosure provides a method of generating an adamer as described herein, the method comprising: a) providing a first single-stranded nucleic acid molecule, a second single-stranded nucleic acid molecule, a third single-stranded nucleic acid molecule, and a fourth single-stranded nucleic acid molecule, wherein the sequences of the first single-stranded nucleic acid molecule, the second single-stranded nucleic acid molecule, the third single-stranded nucleic acid molecule, and the fourth single-stranded nucleic acid molecule comprise a portion of the adamer to be generated, the first single-stranded nucleic acid molecule comprises a first region that is complementary to a first region on the second single-stranded nucleic acid molecule and a second region that is self-complementary, the second single-stranded nucleic acid molecule comprises a first region that is complementary to a first region on the first single-stranded nucleic acid molecule and a second region that is self-complementary, and the second single-stranded nucleic acid molecule comprises a first region that is complementary to a first region on the first single-stranded nucleic acid molecule and a second region that is self-complementary. a) hybridizing the first single-stranded nucleic acid molecule, the second single-stranded nucleic acid molecule, and the third single-stranded nucleic acid molecule; a) hybridizing the first single-stranded nucleic acid molecule, the second single-stranded nucleic acid molecule, and the third single-stranded nucleic acid molecule; and c) contacting the partially double-stranded nucleic acid molecule with a ligase enzyme to form a double-stranded adamer structure capped at both ends by a hairpin.
[0049] In some embodiments, the above method may further comprise treating the product of step (c) with an exonuclease, thereby purifying the properly ligated adamers.
[0050] In some embodiments, the above methods can further comprise contacting the partially double-stranded nucleic acid molecule with a MutS enzyme after step (b) and before step (c).
[0051] In the context of the double-stranded, triple-stranded and / or quadruple-stranded adumer assembly methods, the individual single-stranded nucleic acid molecules used in each method can be provided individually in a separate plurality of nucleic acid molecules (also referred to herein as a "pool"), the separate plurality of nucleic acid molecules being generated using a method such as array-based oligonucleotide synthesis. As will be appreciated by those skilled in the art, array-based oligonucleotide synthesis methods include, but are not limited to, electrochemical methods, light-based chemical methods, inkjet printing methods or any combination thereof.
[0052] That is, gene fragment sets can be distributed across several oligonucleotide pools, such that a unique pair of pools is combined to generate an adamer that represents each of the fragments of the final gene (target nucleic acid) construct to be synthesized. As an example, for six oligonucleotide array pools, where each gene is constructed from five fragments, a total of 15 gene constructs are possible, with each pool containing the top or bottom five fragments of five different genes. Thus, six pools of 25 oligonucleotides each, where the oligonucleotides are about 300 bases long, would be sufficient to construct 15 genes of 2.5 kb in length.
[0053] It is clear that pairwise combinations are underutilizing the capacity of current array pools. Even with 50 pools into which 1225 genes can be assembled, the total number of distinct oligonucleotides per pool is only 245 (see FIG. 2). A low number of distinct oligonucleotides per pool increases the amount of each specific oligonucleotide, but limits the total amount of each oligonucleotide required to effectively assemble a given gene. By increasing the number of strands used to generate the adamer, the total number of genes and distinct oligonucleotides per pool increases dramatically, especially with a large number of pools.
[0054] In another embodiment, a triplex, rather than a doublex, adamer assembly method can be utilized (see FIG. 3). In a triplex system, the number of possible combinations of pools increases dramatically. For both double and triplex designs, there is a strong selective advantage for the formation of specific adamers rather than random combinations. First, by using long oligos from the array pool, e.g., 300 bases per strand, the overlapping complementary regions are very long, which ensures robust hybridization under normal conditions. Second, the process includes T7 exonuclease treatment to remove non-adamer DNA, so ligation of the nicks is a selective event. Thus, if hybridization is incorrect, the adamers will not ligate together to form a double hairpin structure, but will be degraded by T7 exonuclease.
[0055] In another preferred embodiment, a four-stranded system can be used. With four strands and four pools per gene, the number of genes increases when the number of pools is 10 or more. For example, with 15 pools, a total of 1365 genes can be generated (see Table 1). The increased complexity of the pools leads to more efficient use of oligonucleotide arrays.
[0056] [Table 1]
[0057] Some commercially available oligonucleotide array pools may be generated with any number of distinct oligonucleotides. In this case, for a fixed array surface, a smaller number of distinct oligonucleotides results in a larger total mass for each oligonucleotide sequence. Considering a triplex system with six pools, each pool has 50 different oligonucleotide species, and thus the three pools combined will identify 150 distinct oligonucleotides (see Figure 5). The combined pools can be used to generate any number of adamars, but it is best to keep the number of adamars in the assembly to a minimum to avoid inappropriate ligation of overhangs.
[0058] Thus, the disclosure provides compositions comprising two or more pluralities of nucleic acid molecules, each of the pluralities comprising two or more species of nucleic acid molecules, wherein the different species of nucleic acid molecules comprise different nucleic acid sequences, and within the two or more pluralities of nucleic acid molecules, there is at least one set of a corresponding plurality such that when the corresponding sets of the plurality are combined in a single reaction volume, at least one nucleic acid from at least one species in the plurality in the set hybridizes to at least one nucleic acid molecule from at least one other species in the plurality to form at least one hybridized complex.
[0059] Thus, the disclosure provides compositions comprising two or more pluralities of nucleic acid molecules, each of the pluralities comprising two or more species of nucleic acid molecules, wherein the different species of nucleic acid molecules comprise different nucleic acid sequences, and within the two or more pluralities of nucleic acid molecules, there is at least one set of corresponding pluralities such that when the sets of corresponding pluralities are combined in a single reaction volume, nucleic acid molecules from the different pluralities within the sets hybridize together to form at least one hybridized complex.
[0060] Thus, the disclosure provides compositions comprising two or more pluralities of nucleic acid molecules, each of the pluralities comprising two or more species of nucleic acid molecules, wherein the different species of nucleic acid molecules comprise different nucleic acid sequences, and within the two or more pluralities of nucleic acid molecules, there is at least one set of corresponding pluralities such that when the sets of corresponding pluralities are combined in a single reaction volume, the nucleic acid molecules from each different plurality within the set hybridize together to form at least one hybridized complex.
[0061] In some embodiments of the aforementioned compositions, the hybridized complex can be any of the hybridized complexes shown in FIG.
[0062] In some embodiments of the aforementioned compositions, within each plurality of nucleic acid molecules, a single species of nucleic acid molecule is present in a plurality (i.e., there is more than one copy of that species of nucleic acid molecule present in the plurality).
[0063] In some embodiments, the compositions include at least about one, or at least about two, or at least about three, or at least about four, or at least about five, or at least about six, or at least about seven, or at least about eight, or at least about nine, or at least about 10, or at least about 11, or at least about 12, or at least about 13, or at least about 14, or at least about 15, or at least about 16, or at least about 17, or at least about 18, or at least about 19, or at least about 20, or at least about 25, or at least about 30, or at least about 35, or at least about 40, or at least about 45, or at least about 50, or at least about 55, or at least about 60, or at least about 65, or at least about 70, or at least about 75, or at least about 80, or at least about 85, or at least about 90, or at least about 95, or at least about 100, or at least about 150, or at least about 20 0, or at least about 250, or at least about 300, or at least about 350, or at least about 400, or at least about 450, or at least about 500, or at least about 550, or at least about 600, or at least about 650, or at least about 700, or at least about 750, or at least about 800, or at least about 850, or at least about 900, or at least about 950, or at least about 1000, or at least about 1500, or at least about 2000 , or at least about 2500, or at least about 3000, or at least about 3500, or at least about 4000, or at least about 4500, or at least about 5000, or at least about 5500, or at least about 6000, or at least about 6500, or at least about 7000, or at least about 7500, or at least 8000, or at least about 8500, or at least about 9000, or at least about 9500, or at least about 10000 species.
[0064] In some embodiments, the compositions contain about one, or about two, or about three, or about four, or about five, or about six, or about seven, or about eight, or about nine, or about ten, or about eleven, or about twelve, or about thirteen, or about fourteen, or about fifteen, or about sixteen, or about seventeen, or about eighteen, or about nineteen, or about twenty, or about twenty-five, or about thirty, or about thirty-five, or about forty, or about forty-five, or about fifty, or about fifty, or about sixty, or about sixty-five, or about seventy-five, or about eighty, or about eighty-five, or about ninety, or about one hundred, or about fifty, or about twenty-five ... or about 300, or about 350, or about 400, or about 450, or about 500, or about 550, or about 600, or about 650, or about 700, or about 750, or about 800, or about 850, or about 900, or about 950, or about 1000, or about 1500, or about 2000, or about 2500, or about 3000, or about 3500, or about 4000, or about 4500, or about 5000, or about 5500, or about 6000, or about 6500, or about 7000, or about 7500, or about 8000, or about 8500, or about 9000, or about 9500, or about 10000.
[0065] In some embodiments of the aforementioned compositions, each of the plurality of nucleic acid molecules is at least about one, or at least about two, or at least about three, or at least about four, or at least about five, or at least about six, or at least about seven, or at least about eight, or at least about nine, or at least about ten, or at least about eleven, or at least about twelve, or at least about thirteen, or at least about fourteen, or at least about fifteen, or at least about sixteen, or at least about seventeen, or at least about eighteen, or less. or at least about 19, or at least about 20, or at least about 25, or at least about 30, or at least about 35, or at least about 40, or at least about 45, or at least about 50, or at least about 55, or at least about 60, or at least about 65, or at least about 70, or at least about 75, or at least about 80, or at least about 85, or at least about 90, or at least about 95, or at least about 100, or at least about 150, or at least about 200, or at least about 2 50, or at least about 300, or at least about 350, or at least about 400, or at least about 450, or at least about 500, or at least about 550, or at least about 600, or at least about 650, or at least about 700, or at least about 750, or at least about 800, or at least about 850, or at least about 900, or at least about 950, or at least about 1000, or at least about 1500, or at least about 2000, or at least about 2500, or at least or at least about 3000, or at least about 3500, or at least about 4000, or at least about 4500, or at least about 5000, or at least about 5500, or at least about 6000, or at least about 6500, or at least about 7000, or at least about 7500, or at least about 8000, or at least about 8500, or at least about 9000, or at least about 9500, or at least about 10000, or at least about 20000, or at least about 30000, or at least about 40000,Or at least about 50,000, or at least about 60,000, or at least about 70,000, or at least about 80,000, or at least about 90,000, or at least about 100,000 different species of nucleic acid molecules.
[0066] In some embodiments of the aforementioned compositions, each of the plurality of nucleic acid molecules is about one, or about two, or about three, or about four, or about five, or about six, or about seven, or about eight, or about nine, or about ten, or about eleven, or about twelve, or about thirteen, or about fourteen, or about fifteen, or about sixteen, or about seventeen, or about eighteen, or about nine, or about twenty, or about twenty-five, or about thirty, or about thirty-five, or about forty, or about forty-five, or about fifty, or about sixty, or about sixty-five, or about seventy-five, or about eighty, or about eighty-five, or about ninety-five, or about one hundred, or about fifty, or about two hundred, or about twenty-five ... 550, or about 600, or about 650, or about 700, or about 750, or about 800, or about 850, or about 900, or about 950, or about 1000, or about 1500, or about 2000, or about 2500, or about 3000, or about 3500, or about 4000, or about 4500, or about 5000, or about 5500, or about 6000, or about 6500 or about 7000, or about 7500, or about 8000, or about 8500, or about 9000, or about 9500, or about 10000, or about 20000, or about 30000, or about 40000, or about 50000, or about 60000, or about 70000, or about 80000, or about 90000, or about 100000 different species of nucleic acid molecules.
[0067] In some aspects of the aforementioned compositions, the nucleic acid molecules of different species within a single plurality of nucleic acid molecules are not complementary to one another.
[0068] In some aspects of the foregoing compositions, each set of corresponding plurality comprises the same number of plurality. In some aspects, at least one hybridized complex comprises one nucleic acid species from each of the plurality in the corresponding set of plurality.
[0069] In some embodiments of the aforementioned compositions, the two or more pluralities of nucleic acid molecules are present in separate volumes (i.e., they are physically separate from one another, such as in different containers or physically separate portions of an array).
[0070] In some embodiments of the aforementioned compositions, the corresponding set of pluralities can include at least about two pluralities, at least about three pluralities, at least about four pluralities, at least about five pluralities, at least about six pluralities, at least about seven pluralities, at least about eight pluralities, at least about nine pluralities, or at least about ten pluralities.
[0071] In some embodiments of the aforementioned compositions, the corresponding set of pluralities can include a plurality of about two, a plurality of about three, a plurality of about four, a plurality of about five, a plurality of about six, a plurality of seven, a plurality of about eight, a plurality of about nine, or a plurality of about ten.
[0072] In some embodiments of the aforementioned compositions, the corresponding set of pluralities can include a plural of two, a plural of three, a plural of four, a plural of five, a plural of six, a plural of seven, a plural of eight, a plural of nine, or a plural of ten.
[0073] In some embodiments of the aforementioned composition, the number of corresponding plurality of sets is:
number
[0074] In a non-limiting example, if the corresponding set of pluralities includes two pluralities, then the hybridized complex includes two nucleic acid molecules, one from each of the two pluralities. In a non-limiting example, if the corresponding set of pluralities includes three pluralities, then the hybridized complex includes three nucleic acid molecules, one from each of the three pluralities. In a non-limiting example, if the corresponding set of pluralities includes four pluralities, then the hybridized complex includes four nucleic acid molecules, one from each of the four pluralities.
[0075] In some embodiments of the aforementioned compositions, the corresponding sets, when combined in a single reaction, contain at least about one species, or at least about two species, or at least about three species, or at least about four species, or at least about five species, or at least about six species, or at least about seven species, or at least about eight species, or at least about nine species, or at least about ten species, or at least about eleven species, or at least about twelve species, or at least about thirteen species, or at least about fourteen species, or at least about fifteen species, or at least about sixteen species, or at least about seventeen species, or at least about seventeen species, or at least about eight species, or at least about nine species, or at least about ten species, or at least about eleven species, or at least about twelve species, or at least about thirteen ... At least about 18, or at least about 19, or at least about 20, or at least about 25, or at least about 30, or at least about 35, or at least about 40, or at least about 45, or at least about 50, or at least about 55, or at least about 60, or at least about 65, or at least about 70, or at least about 75, or at least about 80, or at least about 85, or at least about 90, or at least about 95, or at least about 100 different hybridized complex species.
[0076] In some embodiments of the foregoing compositions, corresponding sets, when combined in a single reaction, can form about one, or about two, or about three, or about four, or about five, or about six, or about seven, or about eight, or about nine, or about ten, or about eleven, or about twelve, or about thirteen, or about fourteen, or about fifteen, or about sixteen, or about seventeen, or about eighteen, or about nineteen, or about twenty, or about twenty-five, or about thirty, or about thirty-five, or about forty, or about forty-five, or about fifty, or about fifty, or about sixty-five, or about sixty-five, or about seventy, or about seventy-five, or about eighty, or about eighty-five, or about ninety, or about ninety-five, or about one hundred different hybridized complex species.
[0077] In some aspects of the aforementioned compositions, the at least one hybridized complex formed corresponds to an adammer of the present disclosure, such that when the at least one hybridized complex is contacted with a suitable ligase and, optionally, a MutS enzyme, then the adammer of the present disclosure forms a hybridized complex.
[0078] Thus, the present disclosure provides a method of producing at least one adamer of the present disclosure, the method comprising: a) providing a composition of the present disclosure; b) combining at least one set of a corresponding plurality of nucleic acid molecules in a single reaction volume such that at least one hybridized complex is formed; and c) contacting at least one hybridized complex with a ligase enzyme to form a double-stranded adamer structure capped at both ends by a hairpin. In some embodiments, at least one of the hybridized complexes comprises two single-stranded nucleic acid molecules, as described in the double-stranded adamer assembly methods described herein. In some embodiments, at least one of the hybridized complexes comprises three single-stranded nucleic acid molecules, as described in the triple-stranded adamer assembly methods described herein. In some embodiments, at least one of the hybridized complexes comprises four single-stranded nucleic acid molecules, as described in the quadruple-stranded adamer assembly methods described herein.
[0079] In some embodiments, the above method may further comprise treating the product of step (c) with an exonuclease, thereby purifying the properly ligated adamers.
[0080] In some embodiments, the above methods can further comprise contacting the partially double-stranded nucleic acid molecule with a MutS enzyme after step (b) and before step (c).
[0081] In addition, the present disclosure provides a method of purifying at least one double-stranded fragment of the present disclosure, the method comprising: a) providing a composition of the present disclosure; b) combining at least one set of a corresponding plurality of nucleic acid molecules in a single reaction volume such that at least one hybridized complex is formed, the at least one hybridized complex comprising at least one double-stranded fragment. In some embodiments, the at least one double-stranded fragment is a fragment for use in a gSynth synthesis method, as described herein.
[0082] Hadamer-Based Methods and Compositions of the Disclosure
[0083] Adamar
[0084] The present disclosure provides a composition comprising at least one adamer. As used herein, the term adamer is used to describe a double-stranded nucleic acid molecule that comprises a hairpin structure at both ends. In some embodiments in which the adamer is immobilized on a solid surface, the adamer may comprise a single hairpin located at the end of the molecule that is not bound to the solid surface. An exemplary schematic diagram of an adamer and two adamers immobilized on a solid surface is shown in FIG. 7. The adamer may comprise one or more features described herein.
[0085] In some embodiments, an adamer can comprise, consist essentially of, or consist of DNA.
[0086] In some embodiments, an adamer can include one or more multiple cloning site (MCS) sequences. In some embodiments, the MCS sequences can include one or more restriction endonuclease (RE) sequences that can be cleaved with a corresponding restriction endonuclease to generate a 3' overhang, a 5' overhang, or a blunt end.
[0087] As will be understood by one of skill in the art, "blunt end" is used to describe the end of a DNA fragment that is free of unpaired nucleotides.
[0088] As will be understood by one of skill in the art, the term 5' overhang is used to refer to a single-stranded portion of a partially double-stranded nucleic acid molecule that is located at the 5' end of one of the strands.
[0089] As will be understood by one of skill in the art, the term 3' overhang is used to refer to a single-stranded portion of a partially double-stranded nucleic acid molecule located at the 3' end of one of the strands.
[0090] In some embodiments, an adamer can include at least one offset-cleaving type II S restriction endonuclease (IISRE) sequence (hereinafter IISRE sequence) that can be cleaved with a corresponding type II S restriction endonuclease (hereinafter IISRE). In some embodiments, an adamer can include at least one IISRE sequence. In some embodiments, an adamer can include at least three IISRE sequences. In some embodiments, an adamer can include at least four IISRE sequences.
[0091] In some embodiments, the IISRE sequence is such that cleavage at the corresponding IISRE results in the creation of a "blunt end."
[0092] In some embodiments, the IISRE sequence is such that cleavage with the corresponding IISRE results in the creation of a 5' overhang that is 1 nucleotide long. In some embodiments, the IISRE sequence is such that cleavage with the corresponding IISRE results in the creation of a 5' overhang that is 2 nucleotides long. In some embodiments, the IISRE sequence is such that cleavage with the corresponding IISRE results in the creation of a 5' overhang that is 3 nucleotides long. In some embodiments, the IISRE sequence is such that cleavage with the corresponding IISRE results in the creation of a 5' overhang that is 4 nucleotides long. In some embodiments, the IISRE sequence is such that cleavage with the corresponding IISRE results in the creation of a 5' overhang that is 5 nucleotides long. In some embodiments, the IISRE sequence is such that cleavage with the corresponding IISRE results in the creation of a 5' overhang that is about 1 nucleotide to about 5 nucleotides long.
[0093] Non-limiting examples of IISRE sequences, along with their corresponding IISREs, and descriptions of the overhangs / blunt ends created by cleavage of the IISRE sequences and corresponding IISREs, are shown in Table 2. Thus, an adamer can include one or more of the IISRE sequences set forth in Table 2.
[0094] [Table 2]
[0095] In some embodiments, the hairpin structure or hairpin (used interchangeably) located at the terminus of the adamer has at least about one, or at least about two, or at least about three, or at least about four, or at least about five, or at least about six, or at least about seven, or at least about eight, or at least about nine, or at least about ten, or at least about eleven, or at least about twelve, or at least about thirteen, or at least about fourteen, or at least about fifteen, or at least about sixteen, or at least about seventeen, or at least about eighteen, or at least about nineteen, or at least about twenty, or at least about twenty-one, or at least about twenty-two, or at least about twenty-three, or at least about twenty-four, or at least about 25, or at least about 26, or at least about 27, or at least about 28, or at least about 29, or at least about 30, or at least about 31, or at least about 32, or at least about 33, or at least about 34, or at least about 35, or at least about 36, or at least about 37, or at least about 38, or at least about 39, or at least about 40, or at least about 41, or at least about 42, or at least about 43, or at least about 44, or at least about 45, or at least about 46, or at least about 47, or at least about 48, or at least about 49, or at least about 50 nucleotides.
[0096] As mentioned above, the adamers are capped at either end by a hairpin structure. The hairpin structure serves several purposes. First, the hairpin provides protection against exonuclease digestion of the adamers. This allows for the removal of unreacted intermediates from a given reaction in the disclosed methods, thereby providing purity to both the developing adamers and the product after adamer extension. Second, the hairpin structure provides a means for attaching the adamers to solid supports. These attachments are generated directly by binding of an aptamer that binds to a specific solid support binding ligand, by ligation with an MCS after digestion with a conventional RE such as BamHI, by ligation with a lambda phage cos site, or by hybridization to a single-stranded solid support binding anchor NA. Finally, the hairpins described herein allow for attachment of the adamers to solid supports (e.g., beads) without the need for non-natural modifications such as biotin. Thus, the adamers of the present disclosure can be synthesized using entirely natural means, obviating the need for small-scale and / or large-scale phosphoramidite synthesis. Thus, the disclosed adamers and methods allow for faster and cheaper synthesis of nucleic acid molecules and can generate less toxic waste products.
[0097] In some embodiments, the hairpin located at the end of the adamer can include a structural sequence that allows for affinity purification of the adamer and / or binding of the adamer to a solid support (eg, a bead).
[0098] In some embodiments, the hairpin located at the end of the adamer can contain an enzymatic sequence (eg, a DNAzyme sequence) that allows for controlled self-cleavage.
[0099] In some embodiments, the hairpin located at the end of the adamer can contain one or more restriction enzyme sites. Without wishing to be bound by theory, the one or more restriction enzyme sites in the hairpin can be cleaved with a corresponding restriction enzyme to generate at least one single-stranded overhang, which can then be used to hybridize and / or ligate the cleaved adamer to a solid support (e.g., a bead) that contains a nucleic acid complementary to the at least one single-stranded overhang.
[0100] In some embodiments, the hairpin at the end of the adamer can include an aptamer sequence. Without wishing to be bound by theory, the aptamer sequence can be used for affinity purification and / or binding to a solid support (e.g., beads). Non-limiting examples of aptamer sequences are shown in Table 3.
[0101] [Table 3]
[0102] In some embodiments, the adamers can include lambda phage cos sites.
[0103] In some embodiments, an adamer can comprise an "N-mer sequence" that comprises a fragment of a nucleic acid synthesized using one of the methods described herein. The terms "N-mer sequence," "payload," "payload sequence," "N-mer payload," and "double-stranded insert" are used interchangeably herein.
[0104] In some embodiments, the N-mer sequence can be about 3 nucleotides in length. In some embodiments, the N-mer sequence is about 3 nucleotides in length. An N-mer sequence that is 3 nucleotides in length is referred to herein as a 3-mer.
[0105] In some embodiments, the N-mer sequence can be about 4 nucleotides in length. In some embodiments, the N-mer sequence is about 4 nucleotides in length. An N-mer sequence that is 4 nucleotides in length is referred to herein as a 4-mer.
[0106] In some embodiments, the N-mer sequence can be about 5 nucleotides in length. In some embodiments, the N-mer sequence is about 5 nucleotides in length. An N-mer sequence that is 5 nucleotides in length is referred to herein as a 5-mer.
[0107] In some embodiments, the N-mer sequence can be about 6 nucleotides in length. In some embodiments, the N-mer sequence is about 6 nucleotides in length. An N-mer sequence that is 6 nucleotides in length is referred to herein as a 6-mer.
[0108] In some embodiments, the N-mer sequence can be any number of nucleotides in length. In some embodiments, the N-mer sequence can be about at least 25 nucleotides, or at least about 50 nucleotides, or at least about 75 nucleotides, or at least about 100 nucleotides, or at least about 125 nucleotides, or at least about 150 nucleotides, or at least about 175 nucleotides, or at least about 200 nucleotides, or at least about 225 nucleotides, or at least about 250 nucleotides, or at least about 275 nucleotides, or at least about 300 nucleotides in length.
[0109] In some embodiments, the N-mer sequence can be any number of nucleotides in length. In some embodiments, the N-mer sequence can be about 25 nucleotides, or about 50 nucleotides, or about 75 nucleotides, or about 100 nucleotides, or about 125 nucleotides, or about 150 nucleotides, or about 175 nucleotides, or about 200 nucleotides, or about 225 nucleotides, or about 250 nucleotides, or about 275 nucleotides, or about 300 nucleotides in length.
[0110] In some embodiments, an adamer can include an MCS sequence, a first IISRE sequence, an N-mer sequence, and at least a second IISRE sequence. In some embodiments, an adamer can include an MCS sequence, followed by a first IISRE sequence, followed by an N-mer sequence, followed by at least a second IISRE sequence. Exemplary schematics of the aforementioned adamers are shown in FIG. 8 as adamer design numbers 1-4. In the non-limiting examples of adamer design numbers 1-3 shown in FIG. 8, the first IISRE sequence is an IISRE sequence that creates a 4 nucleotide long 5' overhang when cleaved, the N-mer sequence is a 3-mer sequence, and at least a second IISRE sequence is an IISRE sequence that creates a blunt end when cleaved. In a non-limiting example of adamer design number 4 shown in Figure 8, the first IISRE sequence is an IISRE sequence that creates a blunt end when cleaved, the N-mer sequence is a 3-mer sequence, and at least the second IISRE sequence is an IISRE sequence that creates a 3 nucleotide long 5' overhang when cleaved.
[0111] In some embodiments, the adamer can include an MCS sequence, a first IISRE sequence, a second IISRE sequence, an N-mer sequence, a third IISRE sequence, and at least a fourth IISRE sequence. In some embodiments, the adamer can include an MCS sequence, followed by a first IISRE sequence, followed by a second IISRE sequence, followed by an N-mer sequence, followed by a third IISRE sequence, followed by at least a fourth IISRE sequence. An exemplary schematic of the adamer described above is shown in FIG. 8 as adamer design number 5. In a non-limiting example of adamar design number 5 shown in FIG. 8, the first IISRE sequence is an IISRE sequence that creates a 4-nucleotide 5' overhang when cleaved, the second IISRE sequence is an IISRE sequence that creates a 4-nucleotide 5' overhang when cleaved, the N-mer sequence is a 3-mer sequence, the third IISRE sequence is an IISRE sequence that creates a 4-nucleotide 5' overhang when cleaved, and at least the fourth IISRE sequence is an IISRE sequence that creates a blunt end when cleaved.
[0112] In some embodiments, an adamer can include a first MCS sequence, a first IISRE sequence, an N-mer sequence, at least a second IISRE sequence, and at least a second MCS sequence. In some embodiments, an adamer can include a first MCS sequence, followed by a first IISRE sequence, followed by an N-mer sequence, followed by at least a second IISRE sequence, followed by at least a second MCS sequence. An exemplary schematic of the aforementioned adamer is shown in FIG. 8 as adamer design number 6. In a non-limiting example of adamer design number 6 shown in FIG. 8, the first IISRE sequence is an IISRE sequence that creates a blunt end when cleaved, the N-mer sequence is a 3-mer sequence, and the at least a second IISRE sequence is an IISRE sequence that creates a 4 nucleotide long 5-overhang when cleaved.
[0113] In some embodiments, the adamer can include a first MCS sequence, a first IISRE sequence, a second IISRE sequence, an N-mer sequence, a third IISRE sequence, at least a fourth IISRE sequence, and at least a second MCS sequence. In some embodiments, the adamer can include a first MCS sequence, followed by a first IISRE sequence, followed by a second IISRE sequence, followed by an N-mer sequence, followed by a third IISRE sequence, followed by at least a fourth IISRE sequence, followed by at least a second MCS sequence. An exemplary schematic of the aforementioned adamer is shown in FIG. 8 as adamer design number 7. In a non-limiting example of adamer design number 7 shown in FIG. 8, the first IISRE sequence is an IISRE sequence that creates a blunt end when cleaved, the second IISRE sequence is an IISRE sequence that creates a 4-nucleotide long 5' overhang when cleaved, the N-mer sequence is a 3-mer sequence, the third IISRE sequence is an IISRE sequence that creates a 4-nucleotide long 5' overhang when cleaved, and at least the fourth IISRE sequence is an IISRE sequence that creates a blunt end when cleaved.
[0114] In some embodiments, an adamer can include a hairpin that includes an aptamer sequence, a first IISRE sequence, an N-mer sequence, at least a second IISRE sequence, and an MCS sequence. In some embodiments, an adamer can include a hairpin that includes an aptamer sequence, followed by a first IISRE sequence, followed by an N-mer sequence, followed by at least a second IISRE sequence, followed by an MCS sequence. An exemplary schematic of the aforementioned adamer is shown in FIG. 8 as adamer design number 7. In a non-limiting example of adamer design number 7 shown in FIG. 8, the aptamer sequence is a thrombin aptamer sequence, the first IISRE sequence is an IISRE sequence that creates a 4 nucleotide long 5' overhang when cleaved, the N-mer sequence is a 3-mer sequence, and at least a second IISRE sequence is an IISRE that creates a blunt end when cleaved.
[0115] When two IISRE sequences are included adjacent to each other within an adamer, these IISRE sequences can be referred to as "nested IISRE sequences" or "nested IISRE sites." In some embodiments, a nested IISRE sequence can include two IISRE sequences that are directly adjacent to each other. In some embodiments, a nested IISRE sequence can include two IISRE sequences that are adjacent to each other but separated by about 1 to about 10 nucleotides.
[0116] Without wishing to be bound by theory, it is possible to nest IISRE sites, since some IISRE sites have cleavage sites that are far enough away from their recognition sites to fit into other IISRE sites between the first site and the payload. Thus, in some embodiments, the adamer can include a unique blunt cleavage site and a 4-base overhang site on either side of the payload (N-mer sequence). Without wishing to be bound by theory, this significantly reduces the number of adamer reagents required to perform routine nucleic acid production.
[0117] Without wishing to be bound by theory, the inclusion of nested IISRE sequences in the adamer provides several options for cleaving at the same position with two separate sites in the methods of the present disclosure. The option to cleave at the same position with two separate sites can reduce the number of separate adamers required in a library (see below) for general nucleic acid synthesis.
[0118] In some aspects, an adumer can include any element known in the art to facilitate cloning, including, but not limited to, cognate sequences of an amplification primer. Without wishing to be bound by theory, including the cognate sequences of an amplification primer in an adumer may allow recovery of a particular adumer design for clonal expansion.
[0119] In some embodiments, an adamer can include any element known in the art to facilitate large-scale production of an adamer by fermentation in a plasmid or bacteriophage, including, but not limited to, a sequence corresponding to a DNAzyme scar and / or a sequence that facilitates smooth folding of the adamer following excision using a particular DNAzyme (see, e.g., Praetorius et al., Nature, 2017, 552, 84-87, which is incorporated herein by reference in its entirety).
[0120] Nucleic acid synthesis method using the disclosed adamer
[0121] The adamers described herein can be used in the methods described herein to synthesize nucleic acid molecules containing any target nucleic acid sequence, also referred to herein as a "target nucleic acid" or "gene."
[0122] In some embodiments, the target nucleic acid sequence can be at least about 100, or at least about 200, or at least about 300, or at least about 500, or at least about 600, or at least about 700, or at least about 800, or at least about 900, or at least about 1000, or at least about 1500, or at least about 2000, or at least about 2500, or at least about 3000, or at least about 3500, or at least about 4000, or at least about 4500, or at least about 5000 nucleotides in length. In some embodiments, the target double-stranded nucleic acid can comprise at least one homopolymer sequence.
[0123] In some embodiments, the target nucleic acid sequence can include at least one homopolymeric sequence. As used herein, the term homopolymeric sequence is used to refer to any type of repetitive nucleic acid sequence, including, but not limited to, repeats of a single nucleotide or repeats of a small motif. In some embodiments, the homopolymeric sequence can be at least about 10 nucleotides, or at least about 20 nucleotides, or at least about 30 nucleotides, or at least about 40 nucleotides, or at least about 50 nucleotides, or at least about 60 nucleotides, or at least about 70 nucleotides, or at least about 80 nucleotides, or at least about 90 nucleotides, or at least about 100 nucleotides in length.
[0124] In some embodiments, a target nucleic acid sequence can have a GC content of at least about 10%, or at least about 20%, or at least about 50%, or at least about.
[0125] As part of the synthetic methods of the present disclosure, one or more adamers can be immobilized to a solid support. The solid support can be any solid support known in the art, including, but not limited to, at least one bead. In some aspects, the at least one bead can comprise polyacrylamide, polystyrene, agarose, or any combination thereof. In some aspects, the at least one bead can be magnetic. In some aspects, the solid support comprises a well or chamber. In some aspects, the solid support can comprise a plurality of wells or chambers. In some aspects, the plurality of wells comprises a multi-well plate. In some aspects, the solid support can comprise glass. In some aspects, the solid support can comprise a glass slide. In some aspects, the solid support can comprise quartz. In some aspects, the solid support can comprise a quartz slide. In some aspects, the solid support can comprise polystyrene. In some aspects, the solid support can comprise a polystyrene slide. In some aspects, the solid support can comprise a coating, where the coating prevents non-specific binding of undesired proteins, undesired nucleic acids, or other undesired biomolecules. In some embodiments, the coating can include polyethylene glycol (PEG). In some embodiments, the coating can include triethylene glycol (TEG).
[0126] In some embodiments where the adamer comprises a hairpin comprising an aptamer sequence, the adamer can be immobilized to a solid support via binding to the aptamer sequence. That is, the solid support can comprise at least one moiety that binds to the aptamer sequence on the adamer. Thus, in a non-limiting example where the adamer comprises a hairpin comprising one of the aptamer sequences listed in Table 3, the solid support can comprise the corresponding ligand listed in Table 3.
[0127] In some embodiments where the adamers comprise an MCS sequence, the adamers can be immobilized to a solid support by a method comprising: a) contacting the adamers with at least one corresponding restriction endonuclease to cleave the MCS sequence, thereby generating a 5' or 3' overhang, and b) hybridizing the 5' or 3' overhang to a complementary single-stranded nucleic acid molecule on a solid support, thereby immobilizing the adamers to the solid support. The foregoing method can further comprise contacting the adamers hybridized to the complementary single-stranded nucleic acid molecule on the solid support with a ligase, thereby ligating the adamers to the complementary single-stranded nucleic acid molecule on the solid support.
[0128] In some embodiments in which the adammer comprises an MCS sequence, the adammer can be immobilized to a solid support by a method comprising: a) contacting the adammer with at least one corresponding restriction endonuclease to cleave the MCS sequence, thereby generating a blunt end; and b) ligating the blunt end of the adammer to a nucleic acid molecule located on a solid support, thereby immobilizing the adammer to the solid support.
[0129] In some embodiments, an adamer bound to a solid support may be referred to herein as a "binding stud."
[0130] A schematic diagram of the disclosed hadamer-based nucleic acid assembly / synthesis method is shown in FIG.
[0131] In the first step of the method, an adammer is provided that is immobilized on a solid support (shown as a bead or surface in FIG. 9). This adammer, referred to herein as a "binding stud," is connected at one end to a solid support using any of the methods described above and is capped at the other end with a hairpin. The binding stud also includes an MCS sequence. In the next step of the method, the binding stud is contacted with a restriction endonuclease that cleaves the MCS sequence, thereby generating a 3' overhang, a 5' overhang, or a blunt end. In the next step of the method, the first adammer, including the MCS sequence, a first IISRE sequence (shown as "L1" in FIG. 9), a first N-mer sequence (shown as "Payload number 1" in FIG. 9), and a second IISRE sequence (shown as "R1" in FIG. 9), is contacted with a restriction endonuclease that cleaves the MCS sequence, thereby generating a 3' overhang, a 5' overhang, or a blunt end.
[0132] In the next step of the method, the cleaved first adamer is ligated to the cleaved binding stud by contacting the cleaved binding stud, the cleaved adamer, and a ligase enzyme, thereby generating a first ligation product that is immobilized to a solid support and includes a MCS sequence, a first IISRE sequence, a payload number 1 sequence, and a second IISRE sequence (see left side of FIG. 9). The first ligation product is then treated with an exonuclease to remove any unligated binding studs and / or first adamers.
[0133] The above steps are then repeated using another adumer immobilized on a solid support and a second adumer comprising an MCS sequence, a third IISRE sequence (shown as "R2" in Figure 9), a second N-mer sequence (shown as "Payload Number 2" in Figure 9), and a fourth IISRE sequence (shown as "L2" in Figure 9) to generate a second ligation product immobilized on a solid support and comprising an MCS sequence, a third IISRE sequence, a payload number 2 sequence, and a fourth IISRE sequence (see the right side of Figure 9).
[0134] In the next step of the method, the first ligation product is contacted with an IISRE (shown as "R1 enzyme in FIG. 9") that cleaves the second IISRE sequence (R1), thereby generating a 3' overhang, a 5' overhang, or a blunt end, thereby generating a) a first cleaved product that is immobilized on a solid support and includes the MCS sequence, the first IISRE sequence (R1), the payload number 1 sequence, and a 3' overhang, a 5' overhang, or a blunt end, and b) a second cleaved product that includes the second IISRE sequence (R1). The second cleaved product is then discarded by washing.
[0135] In the next step of the method, the second ligation product is contacted with an IISRE (shown as "L2 enzyme" in Figure 9) that cleaves the fourth IISRE sequence (L2), thereby generating a 3' overhang, a 5' overhang, or a blunt end, thereby creating a) a third cleaved product that is released into solution and comprises a hairpin, the third IISRE sequence (R2), the payload number 2 sequence, and a 3' overhang, a 5' overhang, or a blunt end at one end, and b) a fourth cleaved product that is immobilized to a solid support and comprises the MCS sequence and the fourth IISRE sequence (L2).
[0136] In the next step, the first cleaved product and the third cleaved product are ligated together by contacting the first cleaved product, the third cleaved product, and a ligase enzyme (e.g., a solution containing the third cleaved product is transferred to a solution containing the first cleaved product immobilized on a solid surface and a ligase enzyme is added to the solution), thereby generating a third ligation product that is immobilized on a solid surface and includes the MCS sequence, the first IISRE sequence (L1), the payload number 1 sequence, the payload number 2 sequence, and the third IISRE sequence (R2). The ligation reaction is then treated with an exonuclease to remove any unligated first cleaved product and / or the third cleaved product.
[0137] The above steps can be repeated until the target nucleic acid sequence is synthesized.
[0138] A schematic of the synthesis of an exemplary 27 nucleotide long target nucleic acid sequence is shown in Figures 10A-10H. The sequence to be synthesized is shown at the top of Figure 10A. The sequence is subdivided into eleven 6-mer fragments that overlap with either three or four nucleotides that are incorporated into the adamers that are ligated together to synthesize the target nucleic acid sequence. Figure 10B shows an assembly tree of an exemplary target nucleic acid sequence that maps the order in which adamers containing 6-mer fragments are ligated to efficiently synthesize the target nucleic acid sequence. There are several different methods that conflict with the assembly and placement of odd vs. even overhangs, but this assembly order should be determined by the compatibility of the IISRE enzyme site with the sequence to be generated. In Figure 10B, the numbered 6-mers (1)-(11) correspond to the numbered 6-mers in Figures 10C-10H. The numbers at each node of the tree correspond to the payload length at each step of the assembly. The "4" and "3" indicate the length of the overhangs used. The length of the resulting payload sequence is length=a+bn, where "a" and "b" are the lengths of the input payload and "n" is the length of the overhang.
[0139] The first step in the synthesis of a target nucleic acid sequence is shown in Figure 9C, showing the loading of an adammer containing a 3-mer sequence of GAC and an adammer containing a 3-mer sequence of ATC to form an adammer containing a GACATG 6-mer, which is 6-mer number 1 in Figure 10B. To generate the GACATG hexamer, a first binding stud containing the MCS sequence and a first adammer containing the MCS sequence, a first IISRE sequence (shown as "L1" in Figure 10C), a 3-mer sequence containing the sequence GAC, and a second IISRE sequence (shown as "R1" in Figure 10C) are contacted with one or more restriction endonucleases to cleave the MCS sequence, thereby generating a complementary overhang. These complementary overhangs are then hybridized, and the adamer and binding stud are ligated together by contacting the hybridized complex with a ligase enzyme to obtain ligation product number 1, which is immobilized on a solid surface and includes the MCS sequence, a first IISRE sequence (L1), a 3-mer sequence GAC, and a second IISRE sequence (R1). The same process is repeated with a second binding stud including the MCS sequence and a second adamer including the MCS sequence, a third IISRE sequence (shown as "L2" in FIG. 10C), a 3-mer sequence including the sequence ATG, and a fourth IISRE sequence (shown as "R2" in FIG. 10C) to obtain ligation product number 2, which is immobilized on a solid surface and includes the MCS sequence, a third IISRE sequence (L2), a 3-mer sequence ATG, and a fourth IISRE sequence (R2). Ligation product number 1 is then contacted with an IISRE (shown as "R1 enzyme" in Figure 10C) that cleaves the second IISRE sequence (R1), thereby generating cleaved product number 1, which is immobilized on a solid surface and contains the MCS sequence, the first IISRE sequence (L1), and the 3-mer sequence GAC, followed by a blunt end.Similarly, ligation product number 2 is contacted with an IISRE (shown as "L2 enzyme" in FIG. 10C) that cleaves the third IISRE sequence (L2), thereby generating cleaved product number 2 that is released into solution and contains a blunt end, a 3-mer sequence ATG, and a fourth IISRE sequence (R2). Cleaved product number 1 and cleaved product number 2 are then ligated together using a ligase enzyme to obtain ligation product number 3 that is immobilized on a solid support and contains a MCS sequence, a first IISRE sequence (L1), a 6-mer sequence GACATG, and a fourth IISRE sequence (R2). These products can be optionally treated with an exonuclease to remove any unligated cleaved product number 1 and / or cleaved product number 2. The steps described in this paragraph can be repeated with additional adamers containing different 3-mer sequences to generate adamers containing 6-mer SEQ ID NOs: 2-11 as shown in FIG. 10B.
[0140] The method continues in Figure 10D, showing the ligation of an adammer containing a 6-mer SEQ ID NO:1 to an adammer containing a 6-mer SEQ ID NO:2 (see Figure 10B). Adamer number 1 is immobilized on a solid surface and contains an MCS sequence, a first IISRE site (L1) of Figure 10C, a 6-mer sequence GACATG (6-mer SEQ ID NO:1 of Figure 10B), and a fourth IISRE sequence (R2) of Figure 10C. Adamer number 2 is immobilized on a solid surface and contains an MCS sequence, a fifth IISRE site ("L3" in Figure 10D), a 6-mer sequence ATGAGG (6-mer SEQ ID NO:2 of Figure 10B), and a sixth IISRE site (shown as "R3" in Figure 10D). Adamer number 1 contacts the IISRE cleaving the fourth IISRE site (R2) to generate a single-stranded overhang on the N-mer sequence, thereby generating cleaved product number 3, which is immobilized on a solid surface and includes the MCS sequence, the first IISRE sequence (L1), and the N-mer sequence with a single-stranded overhang. Adamer number 2 contacts the IISRE cleaving the fifth IISRE site (L3) to generate a single-stranded overhang on the N-mer sequence, thereby generating cleaved product number 4, which is released into solution and includes the N-mer sequence with a single-stranded overhang and the sixth IISRE sequence (R3). The cleaved product number 3 and the cleaved product number 4 are then ligated together using a ligase enzyme to obtain ligation product number 4, which is immobilized on a solid surface and includes a MCS sequence, a first IISRE sequence (L1), an N-mer sequence GACATGAGG (the first nine nucleotides in the target nucleic acid sequence to be synthesized), and a sixth IISRE sequence (R3). Ligation product number 4 can be optionally treated with an exonuclease to remove any unligated cleaved product number 1 and / or cleaved product number 2.
[0141] The method continues in FIG. 10E, where ligation product number 4 and an adumer containing the 6-mer sequence number 3 (see FIG. 10B) are treated with the corresponding IISRE to generate cleaved products which are then ligated together to generate an adumer containing the N-mer sequence GACATGAGGGT (SEQ ID NO: 116), which is immobilized on a solid surface and represents the first 11 nucleotides of the target nucleic acid sequence to be synthesized.
[0142] Successive IISRE digestions and ligations are repeated in Figures 10F-10H according to the assembly map shown in Figure 10B until an adammer containing an N-mer sequence corresponding to the 27 nucleotide long target nucleic acid sequence is synthesized. In a final step, the 27 nucleotide long target nucleic acid sequence can be excised from the final synthetic adammer by treating the final synthetic adammer with an IISRE that cleaves IISRE sequences adjacent to the 27 nucleotide long target nucleic acid sequence.
[0143] The method can be described as follows: a) providing a first adamer of the present disclosure immobilized on a solid support, the first adamer comprising a first IISRE sequence, followed by a first N-mer sequence, followed by a second IISRE sequence, followed by a hairpin structure; b) providing a second adamer of the present disclosure immobilized on a solid support, the second adamer comprising a third IISRE sequence, followed by a second N-mer sequence, followed by a fourth IISRE sequence, followed by a hairpin structure; and c) contacting the first adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer, thereby generating a first cleaved product that is immobilized on a solid support and comprises the first IISRE sequence, the first N-mer sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end. d) contacting the second adamer with an IISRE that cleaves a third IISRE sequence located in the second adamer, thereby generating second cleaved products that are released into solution and include a second N-mer sequence, a fourth IISRE sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end, wherein the second cleaved products are capped at one end by a hairpin structure; e) ligating the first cleaved product and the second cleaved product using a ligase enzyme to generate a first ligation product; f) treating the product of step (e) with an exonuclease, thereby removing unligated first cleaved product and / or second cleaved product; and g) repeating steps (a)-(f) until a nucleic acid molecule containing the target nucleic acid sequence is synthesized.
[0144] The method can be described as follows: a) providing a first adamer of the present disclosure immobilized to a solid support, the first adamer comprising a first IISRE sequence, followed by a first N-mer sequence, followed by a second IISRE sequence, followed by a hairpin structure; b) providing a second adamer of the present disclosure immobilized to a solid support, the second adamer comprising a third IISRE sequence, followed by a second N-mer sequence, followed by a fourth IISRE sequence, followed by a hairpin structure; c) contacting the first adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer, thereby generating a first cleaved product that is immobilized to a solid support and comprises the first IISRE sequence, the first N-mer sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end; and d) contacting a second adamer with an IISRE that cleaves a third IISRE sequence located in the second adamer, thereby generating second cleaved products that are released into solution and include a second N-mer sequence, a fourth IISRE sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end, wherein the second cleaved products are capped at one end by a hairpin structure; e) ligating the first cleaved product and the second cleaved product using a ligase enzyme to generate a first ligation product; f) treating the product of step (e) with an exonuclease, thereby removing unligated first cleaved product and / or second cleaved product; and g) repeating steps (c) to (f) until a nucleic acid molecule comprising the target nucleic acid sequence is synthesized.
[0145] The method can be described as follows: a) providing a first adamer of the present disclosure immobilized on a solid support, the first adamer comprising a first IISRE sequence, followed by a first N-mer sequence, followed by a second IISRE sequence, followed by a hairpin structure; b) providing a second adamer of the present disclosure immobilized on a solid support, the second adamer comprising a third IISRE sequence, followed by a second N-mer sequence, followed by a fourth IISRE sequence, followed by a hairpin structure; c) contacting the first adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer, thereby generating a first cleaved product that is immobilized on a solid support and comprises the first IISRE sequence, the first N-mer sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end; and d) contacting the second adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer. contacting the second adamer with an IISRE that cleaves a third IISRE sequence located in the second adamer, thereby generating a second cleaved product that is released into solution and includes a second N-mer sequence, a fourth IISRE sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end, wherein the second cleaved product is capped at one end by a hairpin structure; e) ligating the first cleaved product and the second cleaved product using a ligase enzyme to generate a first ligation product; f) treating the product of step (e) with an exonuclease, thereby removing unligated first cleaved product and / or second cleaved product; and g) repeating steps (a)-(f) with one or more additional adamers until a nucleic acid molecule containing the target nucleic acid sequence is synthesized.
[0146] The method can be described as follows: a) providing a first adamer of the present disclosure immobilized on a solid support, the first adamer comprising a first IISRE sequence, followed by a first N-mer sequence, followed by a second IISRE sequence, followed by a hairpin structure; b) providing a second adamer of the present disclosure immobilized on a solid support, the second adamer comprising a third IISRE sequence, followed by a second N-mer sequence, followed by a fourth IISRE sequence, followed by a hairpin structure; c) contacting the first adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer, thereby generating a first cleaved product that is immobilized on a solid support and comprises the first IISRE sequence, the first N-mer sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end; and d) contacting the second adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer. a) contacting a second cleaved product with an IISRE that cleaves a third IISRE sequence located within the second adamer, thereby generating second cleaved products that are released into solution and include a second N-mer sequence, a fourth IISRE sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end, wherein the second cleaved product is capped at one end by a hairpin structure; b) ligating the first cleaved product and the second cleaved product using a ligase enzyme to generate a first ligation product; c) treating the product of step (e) with an exonuclease, thereby removing unligated first cleaved product and / or second cleaved product; and g) repeating steps (c)-(f) with one or more additional adamers until a nucleic acid molecule comprising the target nucleic acid sequence is synthesized.
[0147] The method can be described as follows: a) providing a first adamer of the present disclosure immobilized on a solid support, the first adamer comprising a first IISRE sequence, followed by a first N-mer sequence, followed by a second IISRE sequence, followed by a hairpin structure; b) providing a second adamer of the present disclosure immobilized on a solid support, the second adamer comprising a third IISRE sequence, followed by a second N-mer sequence, followed by a fourth IISRE sequence, followed by a hairpin structure; c) contacting the first adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer, thereby generating a first cleaved product that is immobilized on a solid support and comprises the first IISRE sequence, the first N-mer sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end; and d) contacting the second adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer. contacting the second cleaved product with an IISRE that cleaves a third IISRE sequence located within the second adamer, thereby generating second cleaved products that are released into solution and include a second N-mer sequence, a fourth IISRE sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end, wherein the second cleaved product is capped at one end by a hairpin structure; e) ligating the first cleaved product and the second cleaved product using a ligase enzyme to generate a first ligation product; f) treating the product of step (e) with an exonuclease, thereby removing unligated first cleaved product and / or second cleaved product; and g) repeating steps (a)-(f) using the product of step (f) and one or more additional adamers until a nucleic acid molecule comprising the target nucleic acid sequence is synthesized.
[0148] The method can be described as follows: a) providing a first adamer of the present disclosure immobilized on a solid support, the first adamer comprising a first IISRE sequence, followed by a first N-mer sequence, followed by a second IISRE sequence, followed by a hairpin structure; b) providing a second adamer of the present disclosure immobilized on a solid support, the second adamer comprising a third IISRE sequence, followed by a second N-mer sequence, followed by a fourth IISRE sequence, followed by a hairpin structure; c) contacting the first adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer, thereby generating a first cleaved product that is immobilized on a solid support and comprises the first IISRE sequence, the first N-mer sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end; and d) contacting the second adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer. a) contacting a nucleic acid molecule comprising a target nucleic acid sequence with an IISRE that cleaves a third IISRE sequence located within the target nucleic acid sequence, thereby generating second cleaved products that are released into solution and include a second N-mer sequence, a fourth IISRE sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end, wherein the second cleaved products are capped at one end by a hairpin structure; b) ligating the first cleaved product and the second cleaved product using a ligase enzyme to generate a first ligation product; c) treating the product of step (e) with an exonuclease, thereby removing unligated first cleaved product and / or second cleaved product; and g) repeating steps (c)-(f) using the product of step (f) and one or more additional adumers until a nucleic acid molecule comprising the target nucleic acid sequence is synthesized.
[0149] The method can be described as follows: a) providing a first adamer of the present disclosure immobilized on a solid support, the first adamer comprising a first IISRE sequence, followed by a first N-mer sequence, followed by a second IISRE sequence, followed by a hairpin structure; b) providing a second adamer of the present disclosure immobilized on a solid support, the second adamer comprising a third IISRE sequence, followed by a second N-mer sequence, followed by a fourth IISRE sequence, followed by a hairpin structure; c) contacting the first adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer, thereby generating a first cleaved product that is immobilized on a solid support and comprises the first IISRE sequence, the first N-mer sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end; and d) contacting the second adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer. contacting the N-mer with an IISRE that cleaves a third IISRE sequence located within the N-mer, thereby generating second cleaved products that are released into solution and include a second N-mer sequence, a fourth IISRE sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end, wherein the second cleaved products are capped at one end by a hairpin structure; e) ligating the first cleaved product and the second cleaved product using a ligase enzyme to generate a first ligation product; f) treating the product of step (e) with an exonuclease, thereby removing unligated first cleaved product and / or second cleaved product; and g) repeating any combination of steps (a)-(f) using the product of step (f) and / or one or more additional adumers until a nucleic acid molecule comprising the target nucleic acid sequence is synthesized.
[0150] The method can be described as follows: a) providing a first adamer of the present disclosure immobilized on a solid support, the first adamer comprising a first IISRE sequence, followed by a first N-mer sequence, followed by a second IISRE sequence, followed by a hairpin structure; b) providing a second adamer of the present disclosure immobilized on a solid support, the second adamer comprising a third IISRE sequence, followed by a second N-mer sequence, followed by a fourth IISRE sequence, followed by a hairpin structure; c) contacting the first adamer with an IISRE that cleaves the second IISRE sequence located within the first adamer, thereby generating a first cleaved product that is immobilized on a solid support and comprises the first IISRE sequence, the first N-mer sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end; and d) contacting the second adamer with an IISRE that cleaves the second IISRE sequence located within the second adamer. a) contacting a third IISRE sequence with an IISRE that cleaves the third IISRE sequence, whereby the third IISRE sequence is released into solution and produces second cleaved products comprising a second N-mer sequence, a fourth IISRE sequence, and at least one of a 3' overhang, a 5' overhang, and a blunt end, wherein the second cleaved products are capped at one end by a hairpin structure; b) ligating the first cleaved product and the second cleaved product using a ligase enzyme to produce a first ligation product; c) treating the product of step (e) with an exonuclease, thereby removing unligated first cleaved product and / or second cleaved product; and g) repeating any combination of steps (c)-(f) using the product of step (f) and / or one or more additional adumers until a nucleic acid molecule comprising the target nucleic acid sequence is synthesized.
[0151] Another adamer-based gene synthesis method, referred to as the "pooled synthesis method," is shown in a schematic diagram in FIG. 6. As shown in FIG. 6, in the pooled synthesis method, multiple adamers, including multiple different adamer species, are provided in a common volume. Each of the adamer species includes a hairpin, followed by a first IISRE sequence, followed by a payload sequence, followed by a second IISRE sequence, followed by a hairpin. Two of the adamer species include a "terminal IISRE sequence" that is cleaved at the end of the method to release the fully assembled / synthesized target nucleic acid sequence. A schematic diagram of these adamer species is shown in FIG. 6. The IISRE sequences are shown as dotted boxes, and the terminal IISRE sequences are specifically labeled. The payload sequences are indicated as "A," "B," "C," "D," and "E." That is, in the non-limiting example shown in FIG. 6, the target gene is split into five fragments for assembly process purposes. After bringing the multiple adamers into a common volume, the adamers are contacted with one or more IISREs that cleave each of the IISRE sequences, except for the terminal IISRE sequences, thereby resulting in single-stranded overhang sites, which are shown in FIG. 6 by the striped boxes labeled "1," "2," "3," "4," "5," and "6." In a next step, the complementary single-stranded overhang sites are hybridized together, and the resulting hybridized complex is contacted with a ligase to ligate the digested adamers together. As shown in FIG. 6, this ligation step can be performed in the presence of an IISRE. After ligation, the ligation product can be treated with T7 exonuclease to remove any unligated adamers. The result of the ligation and T7 exonuclease digestion steps is shown at the bottom of FIG. 6, i.e., the target nucleic acid sequence ("ABCDE") is fully assembled and flanked by hairpins on both sides. The target nucleic acid sequence can then be cleaved by this product using one or more IISREs that cleave the terminal IISRE sequences.
[0152] In the pooled synthesis method, the selected IISRE and IISRE sequences are designed to generate a set of complementary single-stranded overhangs that result in the assembly of the target nucleic acid sequence after hybridization of the single-stranded overhang sequences. In some embodiments of the pooled synthesis method, the number of available IISRE sequences in the multiple adamers to be used to assemble the target nucleic acid is two, with a first IISRE sequence being used to achieve the initial assembly and a second IISRE sequence to release the assembled gene for subsequent cloning in a plasmid or bacteriophage. An additional consideration in the design is the total length of the assembly as well as the length of each fragment. Sequences flanking the site need to be considered, as secondary structures such as G-quadruplexes may form that interfere with ligation. The number of fragments used to assemble the gene is also an important consideration. A 400 bp gene assembly, for example, requires only two adamer payloads of 200 bp each. Similarly, a 1 kb gene would require at least five adamer payloads of 200 bp.
[0153] Thus, the present disclosure provides a method for synthesizing a nucleic acid molecule comprising a target nucleic acid sequence, the target nucleic acid sequence being divided into two or more sequence fragments, the method comprising: a) providing a plurality of adamers, the plurality of adamers comprising a plurality of different adamer species, each of the adamer species comprising a hairpin structure followed by a first IISRE sequence followed by a payload sequence followed by a second IISRE sequence followed by a hairpin structure, the payload sequence of the adamer species corresponding to one of the sequence fragments of the target nucleic acid sequence, the plurality of adamers comprising at least one adamer for every sequence fragment; b) contacting the plurality of adamers with at least one IISRE that cleaves at least one IIRSRE sequence in each of the adamers, thereby generating at least one single-stranded overhang; c) ligating the cleaved adamers together by contacting the cleaved adamers with at least one ligase enzyme, thereby synthesizing a nucleic acid molecule comprising the target nucleic acid sequence. In some embodiments, the above method may further comprise a MutS enzyme treatment step. In some embodiments, the above method may further comprise treating the product of step (c) with an exonuclease, thereby purifying the properly ligated product.
[0154] The adumer design for synthesis of a specific target nucleic acid involves analyzing the sequence to be assembled / synthesized, verifying that no IISRE sites are present, and then selecting optimal junction sites for ligation of the individual gene fragments. The junction sites are derived from a collection of compatible groups discovered during the development of the gSynth assembly process. Compatible groups range from as few as five four-base 5' overhang sites (see PCT Application No. PCT / US2020 / 051838, published as WO 2021055962A1) to as many as 35 four-base 5' overhang sites (Potapov V et al., Comprehensive Profiling of Four Base Overhang Ligation Fidelity by T4 DNA Ligase and Application to DNA Assembly. ACS Synth. Biol. 2018, 7, 11, 2665-2674 and Pryor JMet al., Enabling one-pot Golden Gate assemblies of unprecedented complexity using data-optimized assembly design. PLoS ONE 15(9):e0238592. 2020) that have been used in Golden Gate assemblies.
[0155] In some aspects of the disclosed method, the ligase enzyme can be human DNA ligase III (hLig3). As will be appreciated by those skilled in the art, hLig3 exhibits high blunt-end ligation efficiency (greater than 60%). In some aspects of the disclosed method, the ligase enzyme can be T4 DNA ligase. As will be appreciated by those skilled in the art, T4 DNA ligase exhibits high ligation efficiency (greater than 80%) of nucleic acid fragments containing 3' or 5' overhangs of 2, 3, or 4 nucleotides in length. The ligase enzyme can be any ligase enzyme known in the art.
[0156] In some embodiments of the disclosed methods, the exonuclease can be T7 exonuclease.
[0157] It has been shown that modified Cas9 enzymes can be used to generate nicks instead of complete double-strand breaks (Mali P et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology volume 31, p833-838, 2013, and Ann Ran, F. et al., Double Nicking by RNA-Guided CRISPR Cas9 for Enhanced Genome Editing Specificity. Cell, volume 154, issue 6, p1380-1389, September 12, 2013). Such Cas9 nickase activity can be directed to a specific site in the target DNA by a single guide RNA (sgRNA). In addition, this nickase activity can be directed to separate plus and minus DNA strands in an opposing manner, resulting in nicks that expose the overhanging single-stranded DNA after digestion. Such studies of opposite nicking sites show that 10-15 base overhangs give the best ligation results, and interestingly, 10 base pairs apart is the distance of one turn of the double helix (Wang.RY et al., DNA Fragments Assembly Based on Nicking Enzyme System. PLoS One, March 2013, Volume 8, Issue 3, e57943). Thus, in some embodiments of the adamer-based synthesis method, a mutant Cas9 combined with an sgRNA that targets the end of the payload within the adamer can be used to generate an overhang that can be ligated to assemble the gene of interest, rather than the IREE sequence and IRES.
[0158] In some aspects of the disclosed methods, the synthesized target nucleic acid sequence has a purity of at least about 80%, or at least about 85%, or at least about 90%, or at least about 95%, or at least about 99%.
[0159] In some aspects, the purity of the synthesized target nucleic acid sequence refers to the percentage of the total ligation products formed as part of a single ligation reaction or multiple ligation reactions that correspond to the correct / desired ligation product. Without wishing to be bound by theory, the disclosed method involving ligation of nucleic acid molecules can generate multiple ligation products, some of which correspond to the correct / desired ligation product and some of which are undesired (side reactions, incorrect ligation, etc.). The purity of the ligation product or the synthesized target molecule can be expressed as a percentage that corresponds to the percentage of the total ligation products formed that correspond to the correct / desired ligation product.
[0160] gSynth Methods and Compositions of the Disclosure
[0161] The pooled oligonucleotide synthesis methods and compositions disclosed herein can be used in combination with the gSynth method described in detail in PCT Application No. PCT / US2020 / 051838, published as International Publication No. WO2021055962A1. The gSynth method is also referred to as "double-stranded geometric synthesis (gSynth)" and related compositions for the synthesis of any long double-stranded nucleic acid sequence. In the double-stranded gSynth assembly reaction, the target sequence (i.e., the sequence to be synthesized) is computationally decomposed into a set of adjacent double-stranded nucleic acid fragments, and these adjacent double-stranded nucleic acid fragments are then ligated together in a pair of parallel ligation reactions in a systematic assembly method. These fragments have 3' and / or 5' overhanging single-stranded N-mer sites and have the following three properties: 1) the Nmer sites are not self-hybridizing or self-reactive in the ligation reaction. 2) the N-mer site on one end of the fragment does not cross-hybridize or cross-react with the N-mer site on the other end. Finally, 3) there is one N-mer site on each fragment of the adjacent fragment pair that hybridizes and ligates with the adjacent fragment in a ligation reaction resulting in a new longer double-stranded fragment. PCT Application No. PCT / US2020 / 051838, published as International Publication No. WO2021055962A1, provides preferred N-mer sites that facilitate more efficient and accurate ligation reactions, thereby allowing the double-stranded gSynth method of the present disclosure to be used to synthesize nucleic acid sequences of unprecedented length that cannot be achieved using existing nucleic acid assembly and synthesis techniques. The double-stranded fragments of the present disclosure can be generated using the methods described herein.
[0162] 11A-11G show a non-limiting example of a double-stranded gSynth assembly reaction. FIG. 11A is the target sequence (designated "5050Seq03") to be synthesized using the double-stranded gSynth method of the present disclosure. The bold and underlined portions of the sequence correspond to the selected 4-mer overhangs and thus define the fragments used to synthesize the entire sequence. FIG. 11B shows the individual double-stranded nucleic acid fragments of the sequence shown in FIG. 11A that are used in the double-stranded gSynth method of the present disclosure to construct 5050Seq03. FIG. 11D is a schematic diagram of the first round of ligation in the double-stranded gSynth method to synthesize the sequence shown in FIG. 11A. In the first ligation round, fragments 1 and 2, fragments 3 and 4, fragments 5 and 6, fragments 7 and 8, fragments 9 and 10, fragments 11 and 12, and fragments 13 and 14 are hybridized via their complementary 5' overhangs and then ligated together to generate fragments 1+2, fragment 3+4, fragment 5+6, fragment 7+8, fragment 9+10, fragment 11+12, and fragment 13+14. Figure 1E is a schematic diagram of the second round of ligation in the double-stranded gSynth method for synthesizing the sequence shown in Figure 11A. In the second ligation round, fragments 1+2 and 3+4, fragments 5+6 and 7+8, and fragments 11+12 and 13+14 are hybridized through their complementary 5' overhangs and then ligated together to generate fragments 1+2+3+4, fragments 5+6+7+8, and fragments 11+12+13+14. Figure 1F is a schematic diagram of the third round of ligation in the double-stranded gSynth method for synthesizing the sequence shown in Figure 11A. In the third ligation round, fragments 1+2+3+4 and 5+6+7+8, and fragments 9+10 and 11+12+13+14 are hybridized through their complementary 5' overhangs and then ligated together to generate fragments 1+2+3+4+5+6+7+8 and fragments 9+10+11+12+13+14. FIG. 1G is a schematic diagram of the fourth and final round of ligation in the double-stranded gSynth method for synthesizing the sequence shown in FIG. 11A.In the fourth ligation round, fragments 1+2+3+4+5+6+7+8 and 9+10+11+12+13+14 are hybridized via their complementary 5' overhangs and ligated together, thereby generating the sequence shown in Figure 11A.
[0163] Thus, the pooled oligonucleotide synthesis methods and compositions disclosed herein can be used to generate the double-stranded fragments used in the gSynth method described above and in PCT Application No. PCT / US2020 / 051838, published as WO2021055962A1.
[0164] Additional Exemplary Embodiments
[0165] 1. A double-stranded adamer comprising: a) a first sequence that allows for the generation of an overhang that is capable of ligating to a complementary overhang; b) payload array; c) a second sequence that allows for the generation of an overhang capable of ligating to a complementary overhang; and d) a double-stranded adamer, comprising an adamer, at least one end of the adamer comprising a hairpin structure.
[0166] 2. The adammer of claim 1, comprising hairpin structures at both termini.
[0167] 3. The adammer of claim 1, wherein the first sequence and the second sequence that enable the generation of an overhang capable of ligating to a complementary overhang are restriction endonuclease nickase recognition sites.
[0168] 4. The adammer of claim 1, wherein the first sequence and the second sequence that enable the generation of an overhang capable of ligating to a complementary overhang are sites that support the nicking activity of the mutant Cas9 in the presence of an sgRNA.
[0169] 5. The adumer of claim 1, wherein the generated overhang is 0, 1, 2, 3, 4 or 5 bases.
[0170] 6. The adumer of claim 1, wherein the overhang generated is 10 to 15 bases.
[0171] 7. The adumer of claim 1, wherein the generated overhang corresponds to a four-base overhang of a specific compatible group.
[0172] 8. A pool of oligonucleotides, each of which is an adameric strand but is not directly complementary to any other oligonucleotide in the pool.
[0173] 9. Two, several or many pools of oligonucleotides, each oligonucleotide being a strand of adamer, and a mixture of any two, three, four or five pools will uniquely lead to the formation of an adamer or group of adamers containing a gene or fragment of a gene.
[0174] 10. A method for producing an adamer, comprising: a mixture of two or more pools of oligonucleotides, under conditions favorable for hybridization, leading to the formation of an adamer structure; a) Hybridization b) Ligation to resolve unique nicks c) A method for producing adumers, comprising the step of treatment with an exonuclease to remove non-adumer DNA.
[0175] 11. The method of claim 10, wherein the mixture is treated with His-tagged Taq MutS and Ni-NTA agarose to remove adamers containing mismatched base pairs.
[0176] 12. A gene assembly method, comprising: a group of adamers having internal and external sites for digestion of a common volume; a) treating with an endonuclease enzyme or set of enzymes or enzymes in combination with a short nucleic acid to remove one or both hairpins at the internal sites of each adamer; b) treated with ligase in a suitable buffer; c) A gene assembly method in which the gene is treated with an exonuclease to remove unligated material. EXAMPLES
[0177] Working Example
[0178] Example 1 - Use of pooled oligonucleotides to generate adamers and long sequence assemblies
[0179] In this example, a pool of oligonucleotides (also referred to herein as multiple pools) is used to generate a population of adamers, which can then be assembled into a full-length target nucleic acid sequence.
[0180] In this example, a graph theoretical method was used to programmatically split the 431 bp sequence into five fragments (F1-F5), optimizing for approximately equal GC content across the fragments and selecting fragments approximately 100 bp in size (see FIG. 12). The predicted fragments had compatible single-stranded overhangs (CATC, AACG, TTGA, CAGA) with an estimated ligation fidelity of 100% (see FIG. 13).
[0181] Each predicted fragment sequence was used to generate an adammer containing a type IIS restriction endonuclease site (IISRE) on either end of the payload. Terminal fragments F1 and F5 also contain PCR primer sites for amplification and restriction endonuclease (RE) sites for subsequent cloning steps (see Figure 13).
[0182] Each adamer was formed from three oligonucleotide sequences provided from the pool(s) of nucleic acid molecules described herein. The adamers were then assembled using the pooled synthesis method described in FIG. 6 and in more detail above. As shown in FIG. 6, each of the adamers was uncapped using an IISRE, leaving a 4-nucleotide single-stranded overhang (in this example, the IISRE is BsaI). The caps at the beginning of F1 and at the end of F5 remain intact, thus assembling the fragments into a single large adamer that is the product of the assembly. This larger adamer sequence was designed to assemble into a longer sequence in a later step via the overhangs left by treatment with a different IISRE (in this example, BbsI). The larger adamer product was amplified, cut with REs XbaI and EcoRI, and ligated into a cloning vector. The plasmid clones were subjected to capillary sequencing. Finally, the assembled genes were confirmed to have the desired target sequences by aligning the capillary sequencing results (see Figure 14).
[0183] The above protocol was also used to generate additional target sequences to demonstrate the generalizability of the method. The design of each of these additional target sequence assembly approaches is shown in FIG. 15, which includes a breakdown of the number of obtained fragments of the target sequence ("Number of Fragments" column) and how many individual nucleic acid molecules were used to generate the adamers corresponding to each fragment ("Oligos per Fragment" column). These experiments show that up to six fragments can be combined to generate larger sequences using adamers. In addition, the number of oligonucleotide sequences used to generate the component adamers can be up to four, and as with sequence F, a mixed pool of up to 22 different oligonucleotide sequences can be used simultaneously (see FIG. 15).
[0184] These results demonstrate that using the compositions and methods described herein, target nucleic acid sequences can be assembled and synthesized using an adumer-based assembly method, with individual adumers being generated using pools of nucleic acid molecules generated using pooled oligonucleotide synthesis.
Claims
1. 1. A composition comprising two or more plurality of nucleic acid molecules, each of the plurality of nucleic acid molecules comprises two or more nucleic acid molecules; The nucleic acid molecules of different species contain different nucleic acid sequences; A composition, wherein at least one set of corresponding pluralities is present within two or more pluralities of nucleic acid molecules such that when the at least one set of corresponding pluralities is combined in a single reaction volume, nucleic acid molecules from different pluralities within the sets hybridize together to form at least one type of hybridized complex.
2. The composition comprises about a) 6 types, b) 10 types, c) 15 types, d) 20 species, or e) 50 types The composition of claim 1 comprising a plurality of:
3. The composition of any one of claims 1-2, wherein each plurality of nucleic acid molecules comprises at least about 25 different species of nucleic acid molecules.
4. The composition of claim 1 , wherein each set of corresponding pluralities includes the same number of pluralities.
5. The composition of claim 1 , wherein said at least one hybridized complex comprises one nucleic acid species from each of said plurality within a corresponding plurality of said sets.
6. The composition of claim 1 , wherein the two or more plurality of nucleic acid molecules are present in separate volumes.
7. The at least one set of corresponding plurality comprises at least about a) 2 types, b) 3 types, c) 4 types, or d) 5 types The composition of claim 1 comprising a plurality of:
8. The number of corresponding sets of [0010] 2. The composition of claim 1, wherein X is equal to the total number of nucleic acid molecules and Y is equal to the number of species of nucleic acid that hybridize together to form a single hybridized complex.
9. a) said at least one set of corresponding pluralities comprises two pluralities and said at least one hybridized complex comprises two nucleic acid molecules; b) said at least one set of corresponding pluralities comprises a plurality of three, and said at least one hybridized complex comprises three nucleic acid molecules; or c) the at least one set of corresponding pluralities comprises a plurality of four, and the at least one hybridized complex comprises four nucleic acid molecules.
10. The composition of claim 1, wherein corresponding sets are combined in a single reaction to form at least about five different hybridized complex species.
11. The composition of claim 1 , wherein the different species of nucleic acid molecules within a single plurality of nucleic acid molecules are not complementary to one another.
12. 1. A method for producing at least one hadamur, comprising the steps of: a) providing a composition according to claim 1; b) combining at least one set of corresponding pluralities of nucleic acid molecules in a single reaction volume such that at least one hybridized complex is formed; c) contacting said at least one hybridized complex with at least one ligase enzyme to form at least one adamer capped at both ends by a hairpin.
13. 13. The method of claim 12, further comprising treating the product of step (c) with an exonuclease, thereby purifying properly ligated adamers.
14. 14. The method of claim 12 or claim 13, further comprising contacting said at least one hybridized complex with a MutS enzyme.
15. The said adamant is a) a first type II S restriction endonuclease (IISRE) sequence, b) a payload sequence; c) comprising at least a second IISRE sequence; The method of claim 12 , wherein at least one end of the adamer comprises a hairpin structure.
16. The method of claim 12 , wherein the adamer comprises a hairpin structure at both ends of the adamer.
17. The said adamar a) a first IISRE sequence, b) a second IISRE sequence; c) a payload sequence, and d) at least a third IISRE sequence.
18. The said adamar a) a first IISRE sequence, b) a second IISRE sequence; c) a payload sequence; d) a third IISRE sequence, and The method of claim 12, further comprising: e) at least a fourth IISRE sequence.
19. 13. The method of claim 12, wherein the adamer further comprises a multiple cloning site (MCS) sequence, the MCS sequence comprising one or more restriction endonuclease sequences.
20. 13. The method of claim 12, wherein at least one of the IISRE sequences is selected from the group consisting of MlyI, NgoAVII, SspD5I, AlwI, BccI, BcefI, PleI, BceAI, BceSIV, BscAI, BspD6I, FauI, EarI, BspQI, BfuAI, PaqCI, Esp3I, BbsI, BbvI, BtgZI, FokI, BsmFI, BsaI, BcoDI and HgaI sequences.
21. The method of claim 12 , wherein at least one hairpin structure comprises an aptamer sequence.
22. 13. The method of claim 12, wherein the aptamer sequence is selected from a pL1 aptamer sequence, a thrombin 29-mer aptamer sequence, a S2.2 aptamer sequence, an ART1172 aptamer sequence, a R12.45 aptamer sequence, a Rb008 aptamer sequence, and a 38NT SELEX aptamer sequence.
23. 1. A method for synthesizing a nucleic acid molecule comprising a target nucleic acid sequence, comprising: a) providing a composition according to claim 1, said composition comprising corresponding pluralities of sets of nucleic acid molecules such that when corresponding pluralities of the sets are combined in a single reaction volume, nucleic acid molecules from different pluralities within the sets hybridize together to form two or more hybridized complexes, said two or more hybridized complexes comprising fragments of said target nucleic acid sequence; b) combining in a single reaction volume corresponding pluralities of said sets of nucleic acid molecules such that said two or more hybridized complexes are formed; c) contacting the two or more hybridized complexes with at least one ligase enzyme to form two or more adamers capped at both ends by hairpins; d) assembling said two or more species of adumers to synthesize said nucleic acid molecule comprising said target nucleic acid sequence.
24. Assembling the two or more types of adamars comprises: a) one or more restriction enzymes, and b) using one or more ligases, 24. The method of claim 23, comprising either simultaneous or sequential processing to thereby assemble said nucleic acid molecules comprising said target nucleic acid sequences.
25. Assembling the two or more types of adamars comprises: a) a modified Cas9 that exhibits nickase activity in combination with at least one guide RNA; and b) using one or more ligases, 24. The method of claim 23, comprising either simultaneous or sequential processing to thereby assemble said nucleic acid molecules comprising said target nucleic acid sequences.