Compositions and methods for efficient nucleic acid assembly and tissue engineering
The Double Selection Ligation approach addresses the limitations of current DNA synthesis technologies by reducing errors and increasing efficiency, enabling rapid and cost-effective production of long DNA sequences with high fidelity.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PRESIDENT & FELLOWS OF HARVARD COLLEGE
- Filing Date
- 2025-11-14
- Publication Date
- 2026-05-21
AI Technical Summary
Current DNA synthesis technologies are limited by high error rates, low throughput, and high costs, making it difficult to efficiently produce long DNA sequences, such as those required for genome-based medicines and applications like cell therapies, which exceed the capabilities of existing methods.
A Double Selection Ligation (DSL) approach is employed, utilizing reversible immobilization and ligation-based synthesis with double-stranded nucleic acids featuring single-strand overhangs, combined with additional washing steps to achieve near-perfect synthesis of genomic-length DNA by reducing errors and increasing efficiency.
This method achieves a >1000-fold reduction in errors, enabling rapid and cost-effective production of megabase-scale nucleic acids with error rates as low as <1x10^-7, allowing for the synthesis of long DNA sequences with high fidelity.
Smart Images

Figure US2025055572_21052026_PF_FP_ABST
Abstract
Description
COMPOSITIONS AND METHODS FOR EFFICIENT NUCLEIC ACID ASSEMBLY AND TISSUE ENGINEERINGRELATED APPLICATIONS
[0001] This application claims priority under 35 U. S. C. § 119(e) to U. S. Provisional Application, U. S. S. N. 63 / 720,589, filed November 14, 2024, and U. S. Provisional Application, U. S. S. N. 63 / 752,612, filed January 31, 2025, each of which is incorporated herein by reference.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0002] The contents of the electronic sequence listing (H082470442WO00-SEQ-GJM.xml; Size: 28,748 bytes; and Date of Creation: November 13, 2025) is herein incorporated by reference in their entirety.FEDERALLY SPONSORED RESEARCH
[0003] This invention was made with government support under DE-FG02-02ER63445 awarded by U. S. Department of Energy (DOE). The government has certain rights in this invention.BACKGROUND
[0004] Current DNA synthesis technologies (DSTs) are parameterized by the metrics of i) throughput, ii) length, iii) error rate, and iv) cost / base pair (bp) at a given length. Achieving extrema for each of these metrics can impact the future of molecular scale technologies and genome-based medicines.
[0005] Currently long DNA manufacturing typically begins with single nucleotide monomers and one-by-one addition via phosphoramidite chemistry up to 300-mers or via enzymatic synthesis up to 1000-mers, which effects the efficiency of linear synthesis methods. Although oligomers can be combined to form longer DNA molecules, they must first be error-corrected as there is roughly one error per DNA molecule (>1 error: 1000 nucleotides). Current DSTs employ error correction strategies to eliminate incorrect products, either as in-line error correction that cleans error-containing DNA during synthesis or post-synthesis error correction via sequencing of the synthesized DNA in plasmids and in vitro clonal forms. This increases the cost and complexity of DNA synthesis.
[0006] To synthesize larger DNA sequences, current methods and systems use various combinations of exonucleases, polymerases, and ligases to assemble two or more gene fragments (e.g., up to 52 DNA fragments at once (see Golden Gate Assembly and Gibson assembly)). To reach scales in excess of 100 kilobases (kb), assembly can occur in the final 1 / 112#14588836v1H0824.70442WO00target cells using combinations of CRISPR, recombineering, site-specific recombination, and homologous recombination with DNA transfer between cells accomplished by yeast spheroplast fusion or conjugation. See, e.g., the methods described in: Fredens, J. et al. Total synthesis of Escherichia coli with a recoded genome. Nature 569, 514-518 (2019); Carr, P. A. Protein-mediated error correction for de novo DNA synthesis. Nucleic Acids Research 32, el62-el62 (2004); Hughes, R. A. & Ellington, A. D. Synthetic DNA Synthesis and Assembly: Putting the Synthetic in Synthetic Biology. Cold Spring Harb Perspect Biol 9, a023812 (2017); Hoose, A., Vellacott, R., Storch, M., Freemont, P. S. & Ryadnov, M. G. DNA synthesis technologies to close the gene writing gap. Nat Rev Chem 1, 144-161 (2023); Gibson, J. J. The Senses Considered as Perceptual Systems. (Greenwood Press, Westport, Conn, 1983); Song, L.-F., Deng, Z.-H., Gong, Z.-Y., Ei, L.-L. & Li, B.-Z. Large-Scale de novo Oligonucleotide Synthesis for Whole-Genome Synthesis and Data Storage: Challenges and Opportunities. Front. Bioeng. Biotechnol. 9, 689797 (2021); and Son, W. & Chung, K. W. Targeted recombination of homologous chromosomes using CRISPR-CAS9. FEBS Open Bio 13, 1658-1666 (2023)
[0007] Chip-based phosphoramidite oligomers coupled with protein-based error correction for gene synthesis has helped to enable the current commercial production of DNA with sequence lengths up to ~ 15 kb with costs from ~$0.07 / bp to $0.45 / bp, and delivery times of one to several weeks, both dependent on length. An exponential increase in the number of natural sequences that are of interest to synthesize, powered by the continued geometric scaling of nextgeneration sequencing, as well as a host of applications ranging from concatenated enzyme systems (i.e., polyketides synthases PKS’s) in the range of ~ 100 kb to cell therapies that exceed 3 Gigabases (Gb), now greatly outpace the capabilities of current DNA synthesis technology. For example, using existing technologies, a 3 Gb mammalian genome would cost a minimum of ~$200M to synthesize with 300,000 x 10 kb fragments and would likely require years to complete.SUMMARY
[0008] Compositions, systems, apparatuses, and methods provided in this application utilize a Double Selection Ligation (DSL) approach for nucleic acid production. Implementing a series of reversible immobilization steps in DSL can achieve > 1000-fold reduction in errors during the synthesis, bringing down the error rate (e.g., from <lxl0-4(enzymatic synthesis) to <lxl0-7), approaching near-perfect synthesis even for genomic length DNA. Compositions, systems, apparatuses, and methods provided in this application utilize ligation-based synthesis to achieve a megabase scale nucleic acid production rapidly and efficiently. Expediency can be achieved, at 2 / 112#14588836v1H0824.70442WO00least in some embodiments, by eliminating one-by-one additions, and, instead, the size of the constructed nucleic acid increases by a factor of 2 at each step, utilizing only 28 steps to reach the genome scale for either a plant or animal. Rather than starting with dNTP monomers, compositions, systems, apparatuses, and methods described herein can utilize double- stranded nucleic acids (which may be referred to herein as adapters) comprising single-strand overhangs (e.g., 1-mer adapters) that are capable of being immobilized on one end and ligated on the other end.
[0009] For example, though the nucleic acids can be designed to add only 1 bp initially, subsequent steps exhibit a dramatic increase in nucleic acid length. This exponential growth (2n) occurs because each step incorporates a same-sized nucleic acid fragment from aggregate contributions of the nucleic acids utilized in the ligations. Initially, nucleic acids can be produced via clonal copies with the highest fidelity and sequence accuracy. In this example, only 16 unique nucleic acids comprising 1-mer overhangs (approximately 40 nucleotides in length) can be utilized to achieve whole-genome synthesis.
[0010] Compositions, systems, apparatuses, and methods provided in this application can achieve high-fidelity nucleic acid synthesis by implementing reversible immobilization and additional washing steps (e.g., two additional washing steps). This immobilization can be mediated by, for example, an anchor binding to a binding molecule (e.g., binding interactions involving nucleic acid-peptide or nucleic acid-nucleic acid interactions). Each cycle involves: a) one side of the newly synthesized nucleic acid being temporarily immobilized at one end via an anchor binding to a binding molecule located on a surface, b) the incoming nucleic acid being ligated at the other end (z.e., the non-immobilized end), creating a new overhang for the next ligation cycle, c) the newly ligated nucleic acid being immobilized via a reversible immobilization, and d) a restriction endonuclease cleaves the initially immobilized end of the nucleic acid that effectively switches the immobilized side of the nucleic acid for continued synthesis (see, e.g., FIGs. 2A-2B and FIG. 15B). Switching of the immobilization ends selects for successfully ligated DNA fragments that have been extended. Unextended nucleic acids are released via the reversible immobilization technique and then washed away. This is the first selection in DSE, as washing away unligated nucleic acid strands will mitigate nucleic acid synthesis errors (e.g., missing base errors). After washing, nucleic acid fragments are prepared for transfer to a new surface (e.g., by subjecting the nucleic acids to conditions that promote binding to the other binding molecule); only successfully flipped strands will be able to transfer (i.e., nucleic acids bound after ligation to a binding molecule after that is different relative to the3 / 112#14588836v1H0824.70442WO00binding molecule involved in the immobilization prior to ligation), constituting the second selection (see, e.g., FIGs. 3A-3B).
[0011] Aspects of the application relate to a methods of producing a nucleic acid a method of producing a nucleic acid comprising (i) ligating: (a) a nucleic acid comprising a first anchor bound to a first binding molecule on a surface, a cleavage site for a first nuclease, and a fragment of a portion of a target sequence comprising an overhang; and (b) a nucleic acid comprising a second anchor bound to a second binding molecule on a surface, a cleavage site for a second nuclease, and a fragment of a portion of a target sequence comprising an overhang that is reverse complementary to the overhang in the nucleic acid comprising the first anchor, thereby generating a ligated nucleic acid comprising the first anchor, the cleavage site for the first nuclease, the portion of the target sequence formed by ligating the overhangs, the cleavage site for the second nuclease, and the second anchor.
[0012] The methods can further comprise, for example, (ii) cleaving the ligated nucleic acid with the first nuclease, thereby generating: a cleavage fragment comprising the first anchor bound to the first binding molecule and the cleavage site for the first nuclease; and a cleaved nucleic acid comprising the second anchor bound to the second surface, the cleavage site for the second nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang generated by the first nuclease, and the fragment is elongated relative to the fragments of the target sequence in the nucleic acids ligated in a preceding ligation step. The methods can further comprise subjecting the first binding molecule to conditions that inhibit binding to the first anchor after the cleavage step and then repeating step (i) by performing a subsequent ligation step, wherein the subsequent ligation step comprises ligating: a nucleic acid comprising the first anchor bound to the first binding molecule on a surface, the cleavage site for the first nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang that is reverse complementary to the overhang in the cleaved nucleic acid; and the cleaved nucleic acid, thereby generating a ligated nucleic acid comprising the portion of the target sequence formed by ligating the overhangs, wherein the portion of the target sequence is elongated relative to the portion of the target sequence in the ligated nucleic acid generated in the preceding ligation step. The methods can further comprise repeating step (ii) by performing a subsequent cleavage step, wherein the subsequent cleavage step comprises cleaving the ligated nucleic acid with the first nuclease or the second nuclease, thereby generating: a cleavage fragment; and a cleaved nucleic acid comprising a fragment of a portion of the target sequence. In some embodiments, the ligation and cleavage steps (i)-(ii) are repeated two (2) times or more. In some embodiments, each of the ligation steps comprise ligating (a) a 4 / 112#14588836v1H0824.70442WO00cleaved nucleic acid generated in a preceding cleavage step by cleaving a ligated nucleic acid with the second nuclease, and (b) cleaved nucleic acid generated in a preceding cleavage step by cleaving a ligated nucleic acid with the first nuclease, wherein a final ligation step generates a ligated nucleic acid comprising the target sequence. In some embodiments, the target sequence is generated with an error rate of 1 x 10’3per base pair or lower. In some embodiments, the error rate is 3 x 10’9per base pair or lower.
[0013] Aspects of the application further relate to a cell comprising the nucleic acid produced by a method described herein. In some embodiments, the cell is an animal cell (e.g., a mammalian cell, such as a human cell). A cell population comprising a cell produced by a method described herein is also provided herein.
[0014] Aspects of the application further relate to a surface comprising a binding molecule bound to a nucleic acid, wherein the nucleic acid comprises: two different anchors located at the terminal ends of the nucleic acid, wherein one of the anchors is bound to the binding molecule; and two different nuclease cleavage sites. The nucleic acid can comprise any nucleic acid described herein including, but not limited to, any nucleic acid comprising a target sequence or a portion thereof.
[0015] Aspects of the application further relate to an apparatus or a system comprising: a surface comprising a first binding molecule bound to a first anchor on a nucleic acid comprising a cleavage site for a first nuclease; and a surface comprising a second binding molecule bound to a second anchor on a nucleic acid comprising a cleavage site for a second nuclease, wherein the nucleic acid comprising the first anchor and the nucleic acid comprising the second anchor comprise reverse complementary overhangs.BRIEF DESCRIPTION OF DRAWINGS
[0016] FIGs. 1A-1D show an overview of whole genome synthesis by ligation, showing examples of steps involved in whole genome synthesis. FIG. 1A and 1C show examples of DNA Ink Libraries. FIG. IB and ID show examples of methods utilizing dCas9 immobilization and supercoiling.
[0017] FIGs. 2A-2B show an example of the detailed synthesis steps by hierarchical ligation (FIG. 2A) and examples of steps for removing and / or preventing sequence errors, such as nucleotide substitutions, nucleotide deletions, and / or nucleotide insertions in a target sequence, which could arise from hierarchical ligation.
[0018] FIGs. 3A-B shows possible factors affecting synthesis efficiency (e.g. potential causes of faulty ligation) (FIG. 3A) and examples of steps for removing and / or preventing sequence 5 / 112#14588836v1H0824.70442WO00errors, such as nucleotide substitutions, nucleotide deletions, and / or nucleotide insertions in a target sequence, which could arise from a faulty ligation step (FIG. 3B).
[0019] FIG. 4 shows an example of a nucleic acid that can be cleaved to generate acceptor and donor nucleic acids, wherein the bolded “X” indicate positions that can be methylated and / or demethylated during target sequence production and / or that can comprise chemically modified nucleotides, and the unbolded positions “n” indicate any nucleotides.
[0020] FIGs. 5A-5C show schematic diagrams of an example of droplet loading and droplet merging. FIG. 5A shows an example of an array comprising 5384 x 5384 droplets (13 pm spot and space). FIG. 5B show representative images of plate-to-plate transfer. FIG. 5C show side views of contact transfer droplet between Substrate A and Substrate B and an array of three (3) such droplets (FIG. 5B).
[0021] FIG. 6 is a table showing SynGOAT technology compared to current methods.SynGOAT technology has two steps that do not exist in the current synthesis methods due to their intrinsic nature of linear ligation. These steps help to keep very low error rates throughout ligation processes, enabling whole genome synthesis. S: Solid, M: Mobile.
[0022] FIG. 7 is a schematic depicting a site- specific recombinase reaction mechanism. (Top) Serine integrases generates double- stranded breaks on both DNA molecules. (Bottom) Tyrosine recombinases generates sequential single stranded nicks.
[0023] FIGs. 8A-8C show methods using bivalent dCas and recombinase. Bivalent dCas technology improves the efficiency and accuracy of recombinase reactions. FIG. 8A and FIG.8B are different illustrations of the same representative process. FIG. 8C is an expanded view of the bivalent dCas (dCas9 and dCasl2a) in the overall process.
[0024] FIG. 9 shows example of a method for reconstituted pronuclei assembly utilizing one or more recombinases and bivalent Cas proteins.
[0025] FIGs. 10A-10E shows examples of steps to cleave and release immobilized DNA dualhairpin loops. FIG. 10A-10C shows a first example of steps to cleave and release immobilized DNA dual-hairpin loops. The top panel (FIG. 10A) shows an ink. The middle panel (FIG. 10B) shows that after ligation, only products with both the shaded loops will pass double selection. In the lower panel (FIG. 10C), endonuclease cleaves the DNA dual-hairpin and toehold strand displaces the attachment. FIG. 10D-10E shows a second example of steps to cleave and release immobilized DNA dual-hairpin loops. Steps to cleave and release immobilized DNA dualhairpin loops. As shown in the top panel (FIG. 10D), after ligation, only products with both the shaded loops will pass double selection. As shown in the lower panel (FIG. 10E), endonuclease cleaves DNA dual-hairpin and the toehold strand displaces the attachment. Mobile solo-hairpins 6 / 112#14588836v1H0824.70442WO00are washed away, and then the immobile remainder can be used as is or released by an antitoehold and become mobile. The toehold can regenerate the anchor.
[0026] FIGs. 11A-11C show ligation process optimization results. FIGs. 11A-11B show results from a nucleic acid production method decreased inefficiency to 6x1 O’2. FIG. 11C shows examples of misligation rates.
[0027] FIG. 12 shows washing efficiency as a graphical representation showing the average of two experimental repeats in logarithmic scale. On average 1.98xlO10DNA molecules were shown to be effectively washed away.
[0028] FIGs. 13A-13E depict an example of a stepwise assembly of a synthesized DNA sequence using Alwl and BccI restriction enzymes. “X” indicates 1-mer Cut Sites. These positions indicate where the enzymes (Alwl and BccI) cleave the DNA, generating 1-base overhangs. Uppercase and lowercase letters represent complementary bases, highlighting the formation of the synthesized product as the process progresses. Highlighted sections showcase the specific recognition sequences targeted by Alwl and BccI enzymes.
[0029] FIGs. 14A-14B show effects of a toehold system. FIG. 14A is a bar graph showing the release efficiency of three types of toehold systems (toehold, triplex toehold, and hairpin-toehold systems). Release efficiency was calculated with the following formula: (Released / (Unreleased + Released)). FIG. 14B show diagrams of toehold (top panel), triplex toehold (middle panel), and hairpin-toehold (lower panel) systems. A coded legend is provided in the lower panel.
[0030] FIGs. 15A-15B show exemplary methods for producing nucleic acids for ligation (FIG.15A) and exemplary methods for synthesizing a target sequence (FIG. 15B).
[0031] FIG. 16 shows results from an accelerated recombinase assay.DETAILED DESCRIPTION
[0032] Aspects of the application relate to compositions, systems, apparatuses, and methods for producing a desired nucleic acid comprising a target sequence (e.g., a sequence of interest) (see, e.g., FIGs. 1A-1D). Producing a desired nucleic acid typically entails a plurality of ligation and cleavage steps (e.g., a plurality of cycles, each cycle comprising a ligation step and a cleavage step) that result in exponential elongation of a nucleic acid, ultimately generating a desired nucleic acid with a target sequence (see, e.g., FIGs. 2A-2B). One or more additional steps and / or conditions can be utilized to select for nucleic acids that were successfully ligated. For example, selecting for ligated nucleic acids may include a double selection method described herein (see, e.g., FIGs. 2A-2B, FIGs. 3A-3B, and FIGs. 10A-10B). Double selection can comprise reversibly binding nucleic acids to surfaces and removing unligated nucleic acids (see, e.g., FIGs. 2A-2B,7 / 112#14588836v1H0824.70442WO00FIGs. 3A-3B, and FIGs. 10A-10B). Double selection can also include enriching for nucleic acids that were elongated by ligation, and that have overhangs generated by a nuclease such that the nucleic acids can be utilized in a subsequent ligation (see, e.g., FIGs. 2A-2B, FIGs. 3A-3B, and FIGs. 13A-13E). Double selection can be particularly useful for, as an example, reducing error rates in a target nucleic acid. Further aspects of the invention relate to compositions, apparatuses, systems, and methods for high-throughput nucleic acid production (see, e.g., FIGs.1A-1D, FIGs. 2A-2B, FIGs. 3A-3B, FIG. 5, FIGs. 13A-13E, and FIGs. 15A-15B). In some aspects, the application relates to compositions, apparatuses, systems, and methods for introducing a nucleic acid comprising a target sequence into a cell. Accordingly, compositions, apparatuses, systems, and methods described herein can be utilized to generate nucleic acids ranging in length from tens or hundreds of nucleotides to hundreds of millions of nucleotides or longer, including nucleic acids having lower error rates that can be achieved with conventional nucleic acid synthesis methods alone.Definitions
[0033] The terms “acceptor nucleic acid” and “donor nucleic acid” refer to nucleic acids that can be ligated to produce a ligated nucleic acid. Acceptor nucleic acids will typically comprise, in relative order, an anchor, at least one nuclease cleavage site (e.g., two nuclease cleavage sites), and an overhang. Donor nucleic acids will typically comprise, in relative order, an overhang that is reverse complementary to the overhang on the acceptor nucleic acid, at least one nuclease cleavage site (e.g., two nuclease cleavage sites), and an anchor, wherein the anchors and the nuclease cleavage sites typically differ between the acceptor nucleic acids and the donor nucleic acids.
[0034] The term “anchor” refers to a molecule that can reversibly bind to a binding molecule. An anchor may be referred to as the “first” anchor or “second” anchor. However, it should be appreciated that such descriptions are relative and do not necessarily refer to an order of method steps (e.g., the order that nucleic acids are loaded and immobilized on a surface) unless otherwise indicated herein.
[0035] The term “binding molecule” refers to a molecule that can bind reversibly to an anchor. Binding molecules may be referred to as the “first” binding molecule or “second” binding molecule. However, it should be appreciated that such descriptions are relative and do not necessarily refer to an order of method steps (e.g., the order that nucleic acids are loaded and immobilized on a surface) unless otherwise indicated herein.8 / 112#14588836v1H0824.70442WO00
[0036] The terms “cleavage fragment” and “cleaved nucleic acid” refer to nucleic acids that are generated by cleaving a ligated nucleic acid with a nuclease. The nuclease will typically cleave the nucleic acid at a cleavage site located in one of the nucleic acids (e.g., an acceptor nucleic acid or a donor nucleic acid) that was used in producing the ligated nucleic acid. A cleavage fragment and cleaved nucleic acid will typically both comprise an anchor, wherein the anchors are different and bind to different binding molecules. A cleaved nucleic acid will typically comprise a portion of a target sequence that is elongated relative to a portion of the target sequence that was located in either of the nucleic acids (e.g., the acceptor nucleic acid or the donor nucleic acid) used to produce the corresponding ligated nucleic acid. Accordingly, a cleaved nucleic acid will typically have a longer length than a cleavage fragment.
[0037] The term “cleavage site” refers to a nucleotide sequence or an amino acid sequence that can be involved in a cleavage reaction catalyzed by a nuclease or a protease, respectively.Cleavage sites may be referred to as the “first” cleavage site, “second” cleavage site, “third” cleavage site, or “fourth” cleavage site. However, it should be appreciated that such descriptions are relative and do not necessarily refer to an order of method steps (e.g., the order that cleavage steps are performed in) unless otherwise indicated herein.
[0038] As used herein, the term “click chemistry” refers to a chemical philosophy introduced by K. Barry Sharpless of The Scripps Research Institute, describing chemistry tailored to generate covalent bonds quickly and reliably by joining small units comprising reactive groups together. Click chemistry does not refer to a specific reaction, but to a concept including reactions that mimic reactions found in nature. See, e.g., Kolb, Finn and Sharpless Angewandte Chemie International Edition (2001) 40: 2004-2021; Evans, Australian Journal of Chemistry (2007) 60: 384-395, and Joerg Lahann, Click Chemistry for Biotechnology and Materials Science, 2009, John Wiley & Sons Ltd, ISBN 978-0-470-69970-6, the entire contents of each of which are incorporated herein by reference. In some embodiments, click chemistry reactions are modular, wide in scope, give high chemical yields, generate inoffensive byproducts, are stereospecific, exhibit a large thermodynamic driving force > 84 kJ / mol to favor a reaction with a single reaction product, and / or can be carried out under physiological conditions. A distinct exothermic reaction makes a reactant “spring loaded”. In some embodiments, a click chemistry reaction exhibits high atom economy, can be carried out under simple reaction conditions, use readily available starting materials and reagents, uses no toxic solvents or use a solvent that is benign or easily removed (preferably water), and / or provides simple product isolation by non-chromatographic methods (crystallization or distillation).9 / 112#14588836v1H0824.70442WO00
[0039] The term “crowding agent” refers to a molecule added to a solution that alters the interactions of other components of the solution. A crowding agent can be, but is not limited to, a molecule that is inert with respect to other components in the same solution (e.g., the crowding agent does not specifically or stably bind to other constituent components in the solution), a molecule that forces other constituent components into uneven concentrations (e.g., crowding them) through displacement, which can increase and decrease the concentrations of displaced components throughout the solution in pockets and permit interactions and reactions which would not likely occur in a uniform solution (e.g., interactions unlikely to occur in the absence of the crowding agent), a molecule, such as a volume excluder, that reduces the volume of solvent that is available for other macromolecules, and / or a molecule that promotes phase separation.
[0040] The term “ink” refers to a volume of liquid comprising one or more components of a solution utilized in producing a nucleic acid comprising a target sequence. The one or more components can include, for example, a nucleic acid utilized in ligation (e.g., an acceptor nucleic acid or a donor nucleic acid), a protein (e.g., a ligase, a nuclease, a methyltransferase, a demethylase, a recombinase, or a Cas protein), an agent that can inhibit binding of an anchor to a binding molecule (e.g., a competitive binding agent, such as a nucleic acid utilized in toehold-mediated immobilization, imidazole, or maltose), a buffer (e.g., a wash buffer or enzyme buffer), or a crowding agent.
[0041] The term “ligase” refers to an enzyme an that catalyzes the formation of a phosphodiester bond between the 3 '-hydroxyl end of one nucleic acid strand and the 5 '-phosphate end of another, thereby “sealing” breaks or joining fragments of DNA or RNA. The disclosed methods and compositions may use any ligase, e.g., T4 DNA ligase, T7 DNA ligase, E. coli DNA ligase, HiFi Taq DNA ligase, etc.
[0042] The term “ligated nucleic acid” refers to a nucleic acid produced by ligating reverse complementary overhangs on two nucleic acids (e.g., an acceptor nucleic acid and a donor nucleic acid) and that comprises an anchor at opposite terminal ends of the nucleic acid, wherein each of the two anchors are different from one another.
[0043] The term “overhang” refers to one or more nucleotides or nucleobases that are located at a terminal end of a nucleic acid, and that are not base paired with a nucleotide or nucleobase of the nucleic acid.
[0044] The terms “peptides” and “proteins” refer to a polymer of amino acid residues linked together by peptide bonds, such as naturally occurring peptide bonds or chemically modified peptide bonds. Amino acids in peptides and proteins can include naturally occurring amino acids 10 / 112#14588836v1H0824.70442WO00and non-naturally occurring amino acids, such as amino acid analogs and / or chemically modified amino acids. For example, amino acids can be modified by the addition of a carbohydrate group, a hydroxyl group, a phosphate group, a farnesyl group, an isofamesyl group, a fatty acid group, a linker for conjugation or functionalization, or other modification known in the art.
[0045] The term “percentage sequence identity” or “% sequence identity” may be used to refer to the percentage of nucleotides, nucleosides, or nucleobases that are identical between two nucleotide sequences, or to the percentage of amino acids that are identical between two amino acid sequences. The two sequences can be, for example, a reference sequence and a query sequence, such as wherein the query sequence is a target sequence. Percentage sequence identity can be determined by aligning the sequences. Gaps, overhangs, and / or other parameters used to optimize an alignment can also be introduced during alignments, if necessary. Alignment for purposes of determining percentage sequence identity can be achieved in various ways that are within the means of one of ordinary skill in the art. Examples of computer software programs for alignment include those utilized in BLAST, BLAST-2, ALIGN, Megalign (DNASTAR), COBALT, OPAL, Multlin, Clustal Omega, Clustal W2.0, or Clustal X2.0 software.
[0046] The terms “polynucleotide,” “nucleic acid,” and “oligonucleotide” refer to a series of two or more nucleotides in DNA and / or RNA. The terms polynucleotide and nucleic acid can be used interchangeably herein to refer to a single- stranded molecule, a double-stranded molecule, or a molecule comprising one or more single- stranded regions and one or more double- stranded regions. The term oligonucleotide will typically be used to refer to a single- stranded molecule unless stated otherwise. Nucleic acids, polynucleotides, and oligonucleotides can comprise naturally occurring base moieties, naturally occurring sugar moieties, naturally occurring internucleotide linkages (e.g., phosphate backbone moieties), non-naturally occurring base moieties (e.g., chemically modified base moieties), non-naturally occurring sugar moieties (e.g., chemically modified sugar moieties), non-naturally occurring internucleotide linkages (e.g., chemically modified phosphate backbone moieties), or any combination thereof.
[0047] The terms “recognition site” or “recognition sequence” refers to a nucleotide sequence or amino acid sequence that binds to a peptide or protein (e.g., enzyme, such as a nuclease, recombinase, methyltransferase, etc.).
[0048] The term “reference sequence” refers to a segment of a nucleic acid, or a contiguous assembly of nucleic acid segments that are each derived from non-contiguous segments located in a different nucleic acid and / or derived from a plurality of different nucleic acids, that can be aligned against a query sequence (e.g., a target sequence).11 / 112#14588836v1H0824.70442WO00
[0049] The term “repeating” or “repeated” may be used to refer performing a ligation step after a preceding ligation step and to refer to performing a cleavage step after a preceding cleavage step. If a ligation step is not preceded by a ligation step or if a cleavage step is not preceded by another cleavage step, the steps may be referred to as “initial ligation” and “initial cleavage” steps, respectively. The terms “subsequent ligation” and “subsequent cleavage” steps may be used to refer to the immediately preceding ligation step and the immediately preceding cleavage step, respectively. The term “cycle” may be used to refer to performance of a ligation step and a cleavage step, wherein the cleavage step utilizes the ligated nucleic acid produced in the ligation step. For example, if a method comprises, in relative order, a first ligation and first cleavage step, a second ligation step, a second cleavage step, a third ligation step, and a third cleavage step, then the method may be referred to as comprising three cycles of ligation and cleavage wherein the first ligation and cleavage steps are the initial ligation and cleavage steps, the second ligation and cleavage steps are subsequent ligation and cleavage steps relative to the first ligation and cleavage steps, and the third ligation and cleavage steps are subsequent ligation and cleavage steps relative to the preceding second ligation and cleavage steps.
[0050] The term “reverse complementary” refers to a nucleic acid, or a portion thereof, that comprises a sequence that can hybridize with another sequence at a different location in the same nucleic acid (e.g., wherein the hybridization occurs intramolecularly) or located in a different nucleic acid (e.g., wherein the hybridization occurs intermolecularly), thereby forming a stable duplex or triplex. The ability to hybridize will depend on both the degree of reverse complementarity and the length of the reverse complementary sequence. Generally, the longer the hybridizing nucleic acid, the greater number of base mismatches can be contained in the nucleic acid and still form a stable duplex (or triplex, as the case may be). One skilled in the art can ascertain a tolerable degree of mismatch by use of standard procedures, such as determining the melting point of the hybridized complex.
[0051] The term “subject” to a human or non-human animal of any age group (e.g., infant, child, adolescent, young adult, middle-aged adult, or senior adult). In certain embodiments, the non-human animal is a mammal (e.g., primate (e.g., cynomolgus monkey or rhesus monkey)), commercially relevant mammal (e.g., cattle, pig, horse, sheep, goat, cat, or dog), or bird (e.g., commercially relevant bird, such as chicken, duck, goose, or turkey)).
[0052] The term “surface” refers to a solid or semi-solid material comprising at least one binding molecule. For example, a surface can comprise two binding molecules, wherein each binding molecules binds a different anchor. The terms “acceptor surface” and “donor surface” may be used to refer to surfaces bound by an acceptor nucleic acid and a donor nucleic acid,12 / 112#14588836v1H0824.70442WO00respectively. It should be appreciated that referring to a surface in this way does not necessarily mean that the surface can bind only to the acceptor nucleic acid or to the donor nucleic acid unless stated otherwise. For example, in some embodiments, an acceptor surface comprises two binding molecules, wherein one of the binding molecules can bind the anchor of an acceptor nucleic acid located on the acceptor surface and the other binding molecule can bind the anchor of a donor nucleic acid that is not located on the acceptor surface. As a further example, in some embodiments, a donor surface comprises two binding molecules, wherein one of the binding molecules can bind the anchor of a donor nucleic acid located on the donor surface and the other binding molecule can bind the anchor of an acceptor nucleic acid that is not located on the donor surface. Surfaces described as “separate surfaces” may refer to surfaces that are located on physically separate objects (e.g., physically separate planar platforms, such as two-dimensional platforms including, but not limited to, plates, slides, or wafers) or on the same object (e.g., wherein each of the surfaces is a physically separate spot, such as a droplet, located on the same planar platform, such as a two-dimensional platforms including, but not limited to, plates, slides, or wafers).
[0053] The term “target sequence” refers to a nucleotide sequence having a high percentage sequence identity to, or that is identical to, a reference sequence.
[0054] The term “toehold” refers to a nucleic acid that base pairs with a reverse complementary nucleic acid to promote strand displacement or hybridization with the reverse complementary nucleic acid.Nucleic Acid Production
[0055] Compositions and method described in this application relate to production of nucleic acids that involves elongating a target sequence by a series of ligation reactions. Each ligation step will typically involve at least two nucleic acids, such as an acceptor and a donor nucleic acid.
[0056] Initially, during a ligation step, each nucleic acid (e.g., a donor nucleic acid and an acceptor nucleic acid) is immobilized on a surface (see, e.g., FIGs. 2A-2B) (e.g., wherein a donor nucleic acid immobilized on a donor surface and an acceptor nucleic acid is immobilized on an acceptor surface). The immobilization will typically involve an anchor that is located at the end of the nucleic acid, and that binds to a binding molecule on the surface (e.g., wherein the surface is coated with the binding molecule). The other end (the free end) of the nucleic acids will typically comprise an overhang. Generally, the goal of each ligation step is to ligate the overhangs to generate a ligated nucleic acid that comprises two anchors, wherein each anchor is 13 / 112#14588836v1H0824.70442WO00bound to a different binding molecule located on different surfaces. Following ligation, ligated nucleic acids are immobilized at each end to a different surface while unligated nucleic acids remain bound to a surface at only one end (see, e.g., FIGs. 2A-2B and FIGs. 3A-3B). The ligated nucleic acids comprise a sequence (e.g., a target sequence or a portion thereof) that is flanked by the cleavage sites and that is elongated relative to the portion of the sequence that was located in the corresponding nucleic acids utilized in the ligation (e.g., the acceptor and donor nucleic acids) (see, e.g., FIG. 13C).
[0057] Acceptor nucleic acids and donor nucleic acids will typically comprise, in relative order, the components outlined in the Acceptor Nucleic Acid Example and Donor Nucleic Acid Example shown below, respectively.Acceptor Nucleic Acid Example*[Anchor] [Nuclease Cleavage Site(s)]- [Overhang]*See also, e.g., (i)(a) in Method Examples 1-2 below.Donor Nucleic Acid Example*[Overhang\-[Nuclease Cleavage Site(s) -\Anchor\*See also, e.g., (i)(b) in Method Examples 1-2 below.
[0058] Acceptor nucleic acids and donor nucleic acids will typically comprise at least 5 nucleotides in length, not including the overhangs and sequences located in the portion of the target sequence. In some embodiments, acceptor and donor nucleic acids are 5-200 nucleotides in length (e.g., 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 10-20 nucleotides, 20-50 nucleotides, 50-100 nucleotides, 100-150 nucleotides, or 150-200 nucleotides). In some embodiments, acceptor and donor nucleic acids comprise 5-100 nucleotides in length (e.g., 5-10 nucleotides, 10-15 nucleotides, 15-20 nucleotides, 20-25 nucleotides, 25-30 nucleotides, 30-35 nucleotides, 35-40 nucleotides, 40-45 nucleotides, 45-50 nucleotides, 50-55 nucleotides, 55-60 nucleotides, 60-65 nucleotides, 65-70 nucleotides, 70-75 nucleotides, 75-80 nucleotides, 80-85 nucleotides, 85-90 nucleotides, 90-95 nucleotides, or 95-100 nucleotides). In some embodiments, acceptor and donor nucleic acids comprise 20-80 nucleotides in length, such as 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, 25 nucleotides, 26 nucleotides, 27 nucleotides, 28 nucleotides, 29 nucleotides, 30 nucleotides, 31 nucleotides, 32 nucleotides, 33 nucleotides, 34 nucleotides, 35 nucleotides, 36 nucleotides, 37 nucleotides, 38 nucleotides, 39 nucleotides, 40 nucleotides, 41 nucleotides, 42 nucleotides, 43 nucleotides, 44 nucleotides, 45 nucleotides, 46 nucleotides, 47 nucleotides, 4814 / 112#14588836v1H0824.70442WO00nucleotides, 49 nucleotides, 50 nucleotides, 51 nucleotides, 52 nucleotides, 53 nucleotides, 54 nucleotides, 55 nucleotides, 56 nucleotides, 57 nucleotides, 58 nucleotides, 59 nucleotides, 60 nucleotides, 61 nucleotides, 62 nucleotides, 63 nucleotides, 64 nucleotides, 65 nucleotides, 66 nucleotides, 67 nucleotides, 68 nucleotides, 69 nucleotides, 70 nucleotides, 71 nucleotides, 72 nucleotides, 73 nucleotides, 74 nucleotides, 75 nucleotides, 76 nucleotides, 77 nucleotides, 78 nucleotides, 79 nucleotides, or 80 nucleotides. An acceptor nucleic acid and a donor nucleic acid can be of equal lengths. An acceptor or donor nucleic acid comprising an anchor that, for example, is a polynucleotide (e.g., a single- stranded oligonucleotide) may comprise a greater number of nucleotides in length than, for instance, an acceptor or donor nucleic acid comprising a peptide as an anchor. However, in some embodiments, an acceptor nucleic acid and a donor nucleic acid are of non-equal lengths (e.g., wherein one of the nucleic acids comprises a peptide as an anchor and the other nucleic acid comprises a polynucleotide as an anchor). The regions comprising the nuclease cleavage sites will typically be double stranded. However, one or more gaps of single- stranded sequence can be present in the regions comprising the nuclease cleavage site(s) provided that such regions do not inhibit binding of the corresponding nuclease(s). In addition to the overhangs, single- stranded regions, if present, may be located in or near the anchor (e.g., wherein the anchor is a single- stranded oligonucleotide or polynucleotide comprising one or more regions of single- stranded sequence and one or more regions of doublestranded sequence). Examples of acceptor nucleic acids and donor nucleic acids, and overhangs thereof, are shown in FIG. 4 and FIGs. 13A-13E.
[0059] Typically, in a ligation, the anchor in an acceptor nucleic acid will be different from the anchor in a donor nucleic acid. For example, the acceptor nucleic acid can comprise a first anchor and the donor nucleic acid can comprise a second anchor. The one or more cleavage sites in the acceptor nucleic acid will also typically be different from the one or more nuclease cleavage sites in the donor nucleic acid. For example, the acceptor nucleic acid can comprise a cleavage site for a first nuclease and the donor nucleic acid can comprise a cleavage site for a second nuclease. In some embodiments, the one or more nuclease cleavages sites comprise two nuclease cleavage sites. For example, the acceptor nucleic acid can comprise a cleavage site for a first nuclease and a third nuclease, and the donor nucleic acid can comprise a cleavage site for a second nuclease and a fourth nuclease. Thus, the acceptor nucleic acids and donor nucleic acids can bind to different binding molecules and can be cleaved by different nucleases. Further, the overhangs will by typically be equal in length but will be reverse complementary in sequence.15 / 112#14588836v1H0824.70442WO00
[0060] A ligated nucleic acid will typically comprise, in relative order, the components outlined in the Ligated Nucleic Acid Example shown below.Ligated Nucleic Acid Example*[Anchor] -[Nuclease Cleavage Site(s)]- [Target Sequence or a Portion Thereof* *]-[Nuclease Cleavage Si te( j |-[ Anchor*See also, e.g., step (i) in Method Examples 1-2 below.**Before a final ligation reaction is performed, this component of ligated nucleic acids will typically comprise a portion of the target sequence, wherein the portion of the target sequence comprises nucleotides provided in the corresponding acceptor nucleic acid and nucleotides provided in the corresponding donor nucleic acid.
[0061] In some embodiments, a method of producing a nucleic acid comprises a plurality of ligation steps (e.g., an initial ligation step and one or more subsequent ligation steps and / or a final ligation, such as wherein a plurality of ligation and cleavage cycles are performed), wherein each ligation step forms a bond between overhangs on two nucleic acids (e.g., an acceptor nucleic acid and a donor nucleic acid) comprising a portion of a target sequence. The portion of the target sequence can be one or more nucleotides in length depending on the number of ligation and cleavage cycles performed. The one or more nucleotides can be located in the strand that comprises the overhang, and / or in the strand that is hybridized to the strand comprising the overhang.
[0062] The portion of the target sequence present in a ligated nucleic acid which is generated in an initial ligation reaction can be relatively short, and comprises the nucleotides that were located in the overhangs prior to ligation. For example, in some embodiments, an initial ligation step ligates overhangs on two nucleic acids (e.g., an acceptor nucleic acid and a donor nucleic acid) that each comprise a portion of a target sequence, wherein the portion of the target sequence in each of the nucleic acids used in the initial ligation comprises AA, GA, CA, TA, AG, GG, CG, TG, AC, GC, CC, TC, AT, GT, CT, or TT, and wherein the second nucleotide in the sequence is in the overhang (see, e.g., FIG. 4 and FIGs. 13A-13E).
[0063] A final ligated nucleic acid that is generated in a final ligation step can comprise hundreds, thousands, hundreds of thousands, millions, or hundreds of millions of nucleotides in length. In certain embodiments, the final ligated nucleic acid that is generated in a final ligation step can comprise between 1-1000, 500-2000, 1000-3000, 1500-4000, 2000-5000, 2500-6000, 3000-7000, 3500-8000, 4000-9000, 4500-10,000, 5000-11,000, 5500-12,000, 6000-13,000, 6500-14,000, 7000-15,000, 7500-16,000, 8000-17,000, 8500-18,000, 9000-19,000, 9500- 16 / 112#14588836v1H0824.70442WQ0020,000, 10,000-21,000, 10,500-22,000, 11,000-23,000, 11,500-24,000, 12,000-25,000, 12, SOO-26, 000, 13,000-27,000, 13,500-28,000, 14,000-29,000, 14,500-30,000, 15,000-31,000, 15, SOO-32, 000, 16,000-33,000, 16,500-34,000, 17,000-35,000, 17,500-36,000, 18,000-37,000, 18, SOO-38, 000, 19,000-39,000, 19,500-40,000, 20,000-41,000, 20,500-42,000, 21,000-43,000, 21, SOO-44, 000, 22,000-45,000, 22,500-46,000, 23,000-47,000, 23,500-48,000, 24,000-49,000, 24,500-50,000, 25,000-51,000, 25,500-52,000, 26,000-53,000, 26,500-54,000, 27,000-55,000, 27, SOO-56, 000, 28,000-57,000, 28,500-58,000, 29,000-59,000, 29,500-60,000, 30,000-61,000, 30, SOO-62, 000, 31,000-63,000, 31,500-64,000, 32,000-65,000, 32,500-66,000, 33,000-67,000, 33, SOO-68, 000, 34,000-69,000, 34,500-70,000, 35,000-71,000, 35,500-72,000, 36,000-73,000, 36, SOO-74, 000, 37,000-75,000, 37,500-76,000, 38,000-77,000, 38,500-78,000, 39,000-79,000, 39,500-80,000, 40,000-81,000, 40,500-82,000, 41,000-83,000, 41,500-84,000, 42,000-85,000, 42, SOO-86, 000, 43,000-87,000, 43,500-88,000, 44,000-89,000, 44,500-90,000, 45,000-91,000, 45, SOO-92, 000, 46,000-93,000, 46,500-94,000, 47,000-95,000, 47,500-96,000, 48,000-97,000, 48, SOO-98, 000, 49,000-99,000, 49,500-100,000, 50,000-500,000, 55,000-1,000,000, 60,000-2,000,000, 65,000-3,000,000, 70,000-4,000,000, 75,000-5,000,000, or up to hundreds of millions of nucleotides or more. In some embodiments, the number of nucleotides in a target sequence or portion thereof that is located in a ligated nucleic acid can be characterized by X=2n, wherein “X” is the number of nucleotides and “n” is the number of ligations performed (e.g., the number of ligations performed in a plurality of ligation and cleavage cycles).
[0064] In some embodiments, ligating the nucleic acids comprises merging separate liquid droplets. In some embodiments, merging the droplets comprises contacting two plates (e.g., plates comprising one or more, such as a plurality, of donor and acceptor surfaces). In some embodiments, merging the droplets comprises contacting two plates which are oriented faceface, such as wherein one surface is oriented above or below the other surface (see, e.g., FIGs.2A-2B and FIG. 5) or wherein one of the surfaces is oriented horizontally to the other surface. In some embodiments, merging the droplets comprises utilizing an instrument that uses an electromagnetic field and / or sound energy to join the droplets (e.g., wherein the liquid droplets are located on the same two-dimensional slide, and the instrument uses the electromagnetic field and / or sound energy to fuse the liquid droplets). In some embodiments, the two droplets comprising the nucleic acids are provided in separate droplets, and a third droplet is provided to form a bridge between the other two droplets (e.g., a third droplet comprising a solution, such as a buffer which can include, but is not limited to, a solution comprising ligase).
[0065] Nucleic acids (e.g., acceptor and donor nucleic acids) can be incubated in the presence of ligase for at least 5 seconds (e.g., at least 5 seconds, at least 6 seconds, at least 7 seconds, at least 17 / 112#14588836v1H0824.70442WO008 seconds, at least 9 seconds, at least 10 seconds, at least 11 seconds, at least 12 seconds, at least 13 seconds, at least 14 seconds, at least 15 seconds, at least 16 seconds, at least 17 seconds, at least 18 seconds, at least 19 seconds, at least 20 seconds, at least 30 seconds, at least 40 seconds, at least 50 seconds, at least 60 seconds, or longer than 60 seconds, such as 1-2 minutes, 2-3 minutes, 3-4 minutes, or longer than 4 minutes). In some embodiments, nucleic acids are incubated in the presence of ligase for at least 15 seconds (e.g., 15-20 seconds, 15-30 seconds, 15-40 seconds, 15-50 seconds, or 15-60 seconds). However, in some embodiments, ligation involves incubating nucleic acids (e.g., acceptor and donor nucleic acids) in the presence of a ligase for less than 5 seconds, such as 4 seconds, 3 seconds, 2 seconds, 1 second, or less than 1 second (e.g., ligating nucleic acids via instantaneous ligation or for sub-second time frames, such as 0.1-0.25 seconds, 0.25-0.5 seconds, 0.5-0.75 seconds, or 0.75-0.99 seconds).
[0066] In some embodiments, nucleic acids (e.g., acceptor and donor nucleic acids) are ligated in the presence of a crowding agent. Examples of crowding agents include, but are not limited to, Polysorbate 20 (IUPAC: Polyoxyethylene (20) sorbitan monolaurate; commercially known as “Tween 20”TM), polysorbate 80 (IUPAC: Polyoxyethylene (20) sorbitan monooleate; commercially known as “Tween 80” TM), (l,l,3,3-Tetramethylbutyl)phenyl-polyethylene glycol, Polyethylene glycol tert- octylphenyl ether (commercially known as “Triton X-114”TM), sodium dodecylsulfate (SDS), deoxycholate sodium, (3-((3-cholamidopropyl) dimethylammonio)- 1 -propanesulfonate) (commonly known as CHAPS detergent), benzalkonium chloride, polyethylene glycol (PEG) (e.g., PEG 200, PEG 400, PEG 1000, PEG 1550, PEG 2000, PEG 3350, PEG 6000, PEG 8000, PEG 20000, or PEG 50000), Ficoll 70, Ficoll 400, serum albumin, and / or dextran (for example, but not limited to 50 k or 500 k). In some embodiments, the concentration of a crowding agent is at least 1% w / v (e.g., at least 2% w / v, at least 3% w / v, at least 4% w / v, at least 5% w / v, at least 6% w / v, at least 7% w / v, at least 8% w / v, at least 9% w / v, at least 10% w / v, at least 15% w / v, at least 20% w / v, at least 30% w / v, at least 40% w / v, at least 50% w / v, or more than 50% w / v). In some embodiments, a crowding agent is polyethylene glycol (PEG). In some embodiments, the concentration of PEG is at least 1 mM (e.g., at least 2 mM, at least 3 mM, at least 4 mM, at least 5 mM, at least 6 mM, at least 7 mM, at least 8 mM, at least 9 mM, at least 10 mM, at least 15 mM, at least 20 mM, at least 30 mM, at least 40 mM, at least 50 mM, at least 60 mM, at least 70 mM, at least 80 mM, at least 90 mM or more than 100 mM). In some embodiments, the concentration of PEG is at least 10 mM (e.g., 10-15 mM, 15-20 mM, 20-25 mM, 25-30 mM, 30-35 mM, 35-40 mM, 40-45 mM, 45-50 mM, 50-60 mM, 60-70 mM, 70-80 mM, 80-90 mM, 90-100 mM, or more than 100 mM). In some embodiments, the concentration of PEG is between 10 mM to 50 mM (e.g., 10-15 mM,18 / 112#14588836v1H0824.70442WO0015-20 mM, 20-25 mM, 25-30 mM, 30-35 mM, 35-40 mM, 40-45 mM, or 45-50 mM). In some embodiments, the concentration of PEG is between 25-50 mM (e.g., 25 mM, 26 mM, 27 mM, 28 mM, 29 mM, 30 mM, 31 mM, 32 mM, 33 mM, 34 mM, 35 mM, 36 mM, 37 mM, 38 mM, 39 mM 40 mM, 41 mM, 42 mM, 43 mM, 44 mM, 45 mM, 46 mM, 47 mM, 48 mM, 49 mM, or 50 mM). In some embodiments, ligation is performed in the presence of at least 10 mM PEG (e.g., 10-50 mM PEG, such as 25 mM PEG) for at least 5 seconds (e.g., 5-60 seconds, such as 15 seconds). In some embodiments, automated pipetting is utilized to contact nucleic acids with a crowding agent (e.g., PEG).
[0067] In some embodiments, nucleic acids (e.g., acceptor and donor nucleic acids) are ligated in the presence of glycerol. In some embodiments, the concentration of glycerol is at least 1% w / v (e.g., at least 2% w / v, at least 3% w / v, at least 4% w / v, at least 5% w / v, at least 6% w / v, at least 7% w / v, at least 8% w / v, at least 9% w / v, at least 10% w / v, at least 15% w / v, at least 20% w / v, at least 30% w / v, at least 40% w / v, at least 50% w / v, at least 60% w / v, at least 70% w / v, at least 80% w / v, or more than 80% w / v). In some embodiments, the concentration of glycerol is at least 40% w / v. In some embodiments, the concentration of glycerol is between 10% w / v to 80% w / v (e.g., 10-15% w / v, 15-20% w / v, 20-25% w / v, 25-30% w / v, 30-35% w / v, 35-40% w / v, 40-45% w / v, or 45-50% w / v, 50-60% w / v, 60-70% w / v, or 70-80% w / v). In some embodiments, the concentration of glycerol is 1% w / v to 50% w / v. In some embodiments, glycerol is utilized at a given concentration continuously to reduce or eliminate evaporation. In some embodiments, ligation is performed in the presence of glycerol and a crowding agent (e.g., PEG). In some embodiments, automated pipetting is utilized to contact nucleic acids with glycerol.
[0068] In some embodiments, ligation (e.g., ligation of acceptor and donor nucleic acids) is performed at a temperature between 1 °C to 50 °C (e.g., 1-10 °C, 1-14 °C, 10-20 °C, 20-30 °C, 30-40 °C, 40-50 °C, or 18-50 °C). In some embodiments, ligation (e.g., ligation of acceptor and donor nucleic acids) is performed at a temperature between 14 °C to 18 °C (e.g., 14 °C, 15 °C, 16 °C, 17 °C, or 18 °C). In some embodiments, surfaces (e.g., acceptor and donor surfaces) have a temperature between 1 °C to 50 °C (e.g., l-10°C, 1-14°C, 10-20 °C, 20-30 °C, 30-40 °C, 40-50 °C, or 18-50 °C) during ligation. In some embodiments, surfaces (e.g., acceptor and donor surfaces) have a temperature between 14 °C to 18 °C (e.g., 14 °C, 15 °C, 16 °C, 17 °C, or 18 °C) during ligation.
[0069] A DNA ligase or an RNA ligase can be utilized to ligate nucleic acids. In some embodiments, ligating the nucleic acids comprises contacting the nucleic acids with a DNA ligase. In some embodiments, a DNA ligase is a T4 DNA ligase. In some embodiments, ligating the nucleic acids comprises contacting the nucleic acids with an RNA ligase. In some19 / 112#14588836v1H0824.70442WO00embodiments, an RNA ligase is a T4 RNA ligase. In some embodiments, automated pipetting is utilized to contact nucleic acids with a ligase.
[0070] In some embodiments, a ligation step described herein exhibits an error rate in the range of approximately 1% to 10%. This ligation accuracy can be achieved in a method of producing a nucleic acid that utilizes an assembly pipeline comprising a plurality of ligation steps and selection steps (e.g., wherein the ligation and selection steps are performed in cycles). See, e.g., FIGs. 1A-1B, FIGs. 2A-2B, and FIGs. 3A-3B. Subsequent processing steps (e.g., subsequent selection steps) substantially reduce the accumulation of undesired products.
[0071] Following each ligation (e.g., ligation of an acceptor nucleic acid and a donor nucleic acid), one or more selection steps are performed to select for nucleic acids that were successfully ligated.
[0072] In some embodiments, surface(s) (e.g., acceptor surfaces and / or donor surfaces) are subjected to one or more washes before, during, and / or after ligation steps (e.g., to remove acceptor and / or donor nucleic acids that were not ligated). In some embodiments, automated pipetting is utilized to perform the one or more washes. In some embodiments, the one or more washes comprises one or more mechanical washes and / or one or more chemical (e.g., enzymatic) washes. In some embodiments, the one or more mechanical washes achieves 104error rate by removing unwanted or non-specifically bound material. In some embodiments, exonuclease-mediated washing achieves removal up to 10 * error rates, effectively eliminating exposed or non-ligated DNA strands with high specificity. In some embodiments, silanization to improve overall assembly fidelity, signal-to-noise ratio, and / or reproducibility across multiple synthesis cycles. In some embodiments, silanization of the substrate surface is employed prior to or after contacting a surface with a binding molecule (e.g., prior to or after functionalizing a surface with a binding molecule) to minimize nonspecific adsorption of enzymes, nucleic acids, and other biomolecules. In some embodiments, silanization is used to form a self-assembled silane monolayer that provides a uniform chemical landscape and enhances the specificity of coupling between click chemistry handles.
[0073] The one or more selection steps will typically comprise cleaving the ligated nucleic acids (see, e.g., FIGs. 3A-3B and FIGs. 13C-13D). Non-ligated nucleic acids can also be cleaved. A cleavage step can comprise contacting the nucleic acids with a nuclease that uses a cleavage site in the portion of the ligated nucleic acid corresponding to the acceptor nucleic acid or the portion of the ligated nucleic acid corresponding to the donor nucleic acid. In some embodiments, a nuclease binds a second recognition sequence and catalyzes a reaction that cleaves a second nuclease cleavage site (e.g., wherein the recognition sequence and the cleavage site is located in 20 / 112#14588836v1H0824.70442WO00a sequence provided by the corresponding donor nucleic acid utilized in the preceding ligation state) (see, e.g., Method Example 1 below). In some embodiments, a nuclease binds a first recognition sequence and catalyzes a reaction that cleaves a first nuclease cleavage site (e.g., wherein the recognition sequence and the cleavage site is located in a sequence provided by the corresponding acceptor nucleic acid utilized in the preceding ligation step) (see, e.g., Method Example 2 below).
[0074] Cleaving the ligated nucleic acids, thereby generates cleavage fragments and cleaved nucleic acids (see, e.g., FIGs. 2A-2B and FIGs. 3A-3B). At one end, cleaved nucleic acids comprise a sequence (e.g., a portion of a target sequence) that is elongated relative to the sequence that provided the overhang utilized in the preceding ligation), wherein the sequence in the cleaved nucleic acid provides an overhang for a subsequent ligation (see, e.g., FIGs. 13B-13C). At the other end, cleaved nucleic acids comprise one of the anchors from the ligated nucleic acid and are immobilized by binding to a binding molecule on one of the surfaces (e.g., the acceptor surface or the donor surface from the preceding ligation) (see, e.g., FIGs. 2A-2B and FIGs. 3A-3B). The cleavage fragments comprise the other anchor from the corresponding ligated nucleic acid and are bound to the other surface (see, e.g., FIGs. 2A-2B and FIGs. 3A-3B). For example, a cleaved nucleic acid will typically comprise, in relative order, the components outlined in the Cleaved Nucleic Acid Example 1 or Cleaved Nucleic Acid Example 2 shown below. In some embodiments, a cleaved nucleic acid comprising the components set forth in Cleaved Nucleic Acid Example 1 is utilized as an acceptor nucleic acid in a subsequent ligation step. In some embodiments, a cleaved nucleic acid comprising the components set forth in Cleaved Nucleic Acid Example 2 is utilized as a donor nucleic acid in a subsequent ligation step. In some embodiments, a method comprises generating a first cleaved nucleic acid comprising the components set forth in Cleaved Nucleic Acid Example 1 and a second cleaved nucleic acid comprising the components set forth in Cleaved Nucleic Acid Example 2 (e.g., wherein the two cleavage steps are performed concurrently), wherein a subsequent ligation step utilizes the first cleaved nucleic acid as an acceptor nucleic acid and the second cleaved nucleic acid as a donor nucleic acid.Cleaved Nucleic Acid Example 1*[Anchor] -[Nuclease Cleavage Site(s)]- [Fragment of a Portion of a Target Sequence +Overhang]*Produced by cleaving a ligated nucleic acid using a nuclease that binds the cleavage site which was located in the donor nucleic acid utilized in the preceding ligation step. See also, e.g., step (ii) in Example Method 1 below.21 / 112#14588836v1H0824.70442WO00Cleaved Nucleic Acid Example 2*[Fragment of a Portion of a Target Sequence + Overhang\-\Nuclease Cleavage Site(s)]- [Anchor]*Produced by cleaving a ligated nucleic acid using a nuclease that binds the cleavage site which was located in the acceptor nucleic acid utilized in the preceding ligation step. See also, e.g., step (ii) in Example Method 2 below.
[0075] Nucleic acids (e.g., acceptor and donor nucleic acids) can be incubated in the presence of a nuclease for at least 30 seconds (e.g., at least 35 seconds, at least 40 seconds, at least 50 seconds, at least 60 seconds, at least 1 minute, at least 2 minutes, at least 3 minutes, at least 4 minutes, at least 5 minutes, at least 6 minutes, at least 7 minutes, at least 8 minutes, at least 9 minutes, at least 10 minutes, at least 11 minutes, at least 12 minutes, at least 13 minutes, at least 14 minutes, at least 15 minutes, or more than 15 minutes). In some embodiments, nucleic acids are incubated in the presence of a nuclease for at 1-15 minutes (e.g., 1-2 minutes, 2-3 minutes, 3-4 minutes, 4-5 minutes, 5-6 minutes, 6-7 minutes, 7-8 minutes, 8-9 minutes, 9-10 minutes, 10-11 minutes, 11-12 minutes, 12-13 minutes, 13-14 minutes, or 14-15 minutes). In some embodiments, automated pipetting is utilized to contact nucleic acids with a nuclease.
[0076] In some embodiments, nucleic acids (e.g., acceptor and donor nucleic acids) are cleaved in the presence of glycerol. In some embodiments, the concentration of glycerol is at least 1% w / v (e.g., at least 2% w / v, at least 3% w / v, at least 4% w / v, at least 5% w / v, at least 6% w / v, at least 7% w / v, at least 8% w / v, at least 9% w / v, at least 10% w / v, at least 15% w / v, at least 20% w / v, at least 30% w / v, at least 40% w / v, at least 50% w / v, at least 60% w / v, at least 70% w / v, at least 80% w / v, or more than 80% w / v). In some embodiments, the concentration of glycerol is at least 40% w / v. In some embodiments, the concentration of glycerol is between 10% w / v to 80% w / v (e.g., 10-15% w / v, 15-20% w / v, 20-25% w / v, 25-30% w / v, 30-35% w / v, 35-40% w / v, 40-45% w / v, or 45-50% w / v, 50-60% w / v, 60-70% w / v, or 70-80% w / v). In some embodiments, the concentration of glycerol is 1% w / v to 50% w / v. In some embodiments, glycerol is utilized at a given concentration continuously to reduce or eliminate evaporation. In some embodiments, automated pipetting is utilized to contact nucleic acids with glycerol.
[0077] In some embodiments, cleavage (e.g., cleavage of acceptor and donor nucleic acids) is performed at a temperature of 15 °C to 70 °C (e.g., 15 °C, 16 °C, 17 °C, 18 °C, 19 °C, 20 °C, 21°C, 22 °C, 23 °C, 24 °C, 25 °C, 26 °C, 27 °C, 28 °C, 29 °C, 30 °C, 31 °C, 32 °C, 33 °C, 34 °C, 35 °C, 36 °C, 37 °C, 38 °C, 39 °C, 40 °C, 40-45 °C, 45-50 °C, 50-60 °C, or 60-70 °C). In some embodiments,22 / 112#14588836v1H0824.70442WO00cleavage (e.g., cleavage of acceptor and donor nucleic acids) is performed at a temperature of 37 °C. In some embodiments, a surface has a temperature of 15 °C to 70 °C (e.g., a temperature of 15 °C, 16 °C, 17°C, 18°C, 19°C, 20 °C, 21°C, 22 °C, 23 °C, 24 °C, 25°C, 26 °C, 27°C, 28 °C, 29 °C, 30 °C, 31 °C, 32 °C, 33 °C, 34 °C, 35 °C, 36 °C, 37 °C, 38 °C, 39 °C, 40 °C, 40-45 °C, 45-50 °C, 50-60 °C, or 60-70 °C) during a cleavage step. In some embodiments, a surface has a temperature of 37 °C during a cleavage step.
[0078] In some embodiments, surface(s) (e.g., acceptor surfaces and / or donor surfaces) are subjected to one or more washes before, during, and / or after cleavage steps. In some embodiments, automated pipetting is utilized to perform the one or more washes. In some embodiments, the one or more washes comprises one or more mechanical washes and / or one or more chemical (e.g., enzymatic) washes. In some embodiments, the one or more mechanical washes achieves 104error rate by removing unwanted or non-specifically bound material. In some embodiments, exonuclease-mediated washing achieves removal up to 10 * error rates, effectively eliminating exposed or non-ligated DNA strands with high specificity. In some embodiments, silanization to improve overall assembly fidelity, signal-to-noise ratio, and / or reproducibility across multiple synthesis cycles. In some embodiments, silanization of the substrate surface is employed prior to or after contacting a surface with a binding molecule (e.g., prior to or after functionalizing a surface with a binding molecule) to minimize nonspecific adsorption of enzymes, nucleic acids, and other biomolecules.
[0079] The one or more selection steps will also typically comprise removing cleavage fragments. For example, nucleic acids can be subjected to one or more conditions that inhibit binding between the anchor of the cleavage fragment and the binding molecule, such as wherein nucleic acids are subjected to conditions that inhibit binding between a first anchor and a first binding molecule and subsequently subjecting nucleic acids to conditions that allow the first anchor to re-bind to the first binding molecule and / or wherein nucleic acids are subjected to conditions that inhibit binding of a second anchor and a second binding molecule and subsequently subjecting nucleic acids to conditions that allow the second anchor to re-bind to the second binding molecule (see, e.g., FIGs. 2A-2B and FIG. 15B). For example, in some embodiments, the anchor comprises a nucleic acid and the binding molecule comprises a toehold, which can be used in a toehold mediated capture and release system. In some embodiments, a toehold mediated capture and release system has over 90% accuracy (e.g., approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% accuracy, which can results in an error reduction to 104in double selection).23 / 112#14588836v1H0824.70442WO00
[0080] In some embodiments, automated pipetting is utilized to subject nucleic acids to one or more conditions that inhibiting between anchors and binding molecules. In some embodiments, removing cleavage fragments comprises subjecting surface(s) (e.g., acceptor surfaces and / or donor surfaces) to one or more washes before, during, and / or after subjecting the nucleic acids to one or more conditions that inhibit binding between the anchor of the cleavage fragment and the binding molecule. In some embodiments, automated pipetting is utilized to subject nucleic acids to the one or more washes.
[0081] Producing a nucleic acid can comprise repeating one or more method steps described herein. For example, a subsequent ligation step can be performed (e.g., wherein the subsequent ligation utilizes only the nucleic acids that were elongated in a preceding ligation step and then subjected to one or more selection steps) which is followed by a subsequent cleavage step. For example, in some embodiments, a method comprises performing steps (i)-(v) set forth in Method Example 1 below and / or performing steps (i)-(v) Method Example 2 below. In some embodiments, a method comprises performing steps (i)-(v) set forth in Method Example 1, wherein step (iv) in Method Example 1 utilizes a cleaved nucleic acid generated by performing steps (i)-(iii) set forth in Method Example 2 (e.g., wherein steps (i)-(iii) of Method Example 1 and steps (i)-(iii) of Method Example 2 are performed concurrently). In some embodiments, a method comprises performing steps (i)-(v) set forth in Method Example 2, wherein step (iv) in Method Example 2 utilizes a cleaved nucleic acid generated by performing steps (i)-(iii) set forth in Method Example 1 (e.g., wherein steps (i)-(iii) of Method Example 1 and steps (i)-(iii) of Method Example 2 are performed concurrently). In some embodiments, a method comprises performing steps (i)-(v) set forth in Method Example 1 and / or performing steps (i)-(v) set forth in Method Example 2, wherein a step (vi) is utilized, which comprises subjecting the nucleic acids to one or more conditions that inhibit binding of one of the anchors to its cognate binding molecule. For example, wherein the first nuclease cleaves the nucleic acid in step (v), the nucleic acids can be subjected to similar conditions or the same conditions utilized in step (iii) of Method Example 1 that inhibit binding of the first anchor and the first binding molecule to select for cleaved nucleic acids comprising the second anchor. As a further example, wherein the second nuclease cleaves the nucleic acid in step (v), the nucleic acids can be subjected to similar conditions or the same conditions utilized in step (iii) of Method Example 2 that inhibit binding of the second anchor and the second binding molecule to select for cleaved nucleic acid comprising the first anchor.Method Example 1:24 / 112#14588836v1H0824.70442WO00(i) ligating(a) a nucleic acid comprising a first anchor bound to a first binding molecule on a surface, a cleavage site for a first nuclease, and a fragment of a portion of a target sequence comprising an overhang, and(b) a nucleic acid comprising a second anchor bound to a second binding molecule on a surface, a cleavage site for a second nuclease, and a fragment of a portion of a target sequence comprising an overhang that is reverse complementary to the overhang in the nucleic acid comprising the first anchor, thereby generating a ligated nucleic acid comprising the first anchor, the cleavage site for the first nuclease, the portion of the target sequence formed by ligating the overhangs, the cleavage site for the second nuclease, and the second anchor;(ii) cleaving the ligated nucleic acid with the second nuclease, thereby generating a cleaved nucleic acid comprising the first anchor bound to the first binding molecule, the cleavage site for the first nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang generated by the second nuclease and the fragment is elongated relative to the fragments in the nucleic acids ligated in a preceding ligation step, anda cleavage fragment comprising the second anchor bound to the second binding molecule and the cleavage site for the second nuclease;(iii) subjecting the second binding molecule to conditions that inhibit binding to the second anchor;(iv) repeating step (i) by performing a subsequent ligation step, wherein the subsequent ligation step comprises ligating(a) the cleaved nucleic acid, and(b) a nucleic acid comprising the second anchor bound to the second binding molecule on a surface, the cleavage site for the second nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang that is reverse complementary to the overhang in the cleaved nucleic acid, thereby generating a ligated nucleic acid comprising the portion of the target sequence formed by ligating the overhangs, wherein the portion of the target sequence is elongated relative to the portion of the target sequence in the ligated nucleic acid generated in the preceding ligation step; and25 / 112#14588836v1H0824.70442WO00(v) repeating step (ii) by performing a subsequent cleavage step, wherein the subsequent cleavage step comprises cleaving the ligated nucleic acid with the first nuclease or the second nuclease, thereby generatinga cleavage fragment, anda cleaved nucleic acid comprising a fragment of a portion of the target sequence.Method Example 2:(i) ligating(a) a nucleic acid comprising a first anchor bound to a first binding molecule on a surface, a cleavage site for a first nuclease, and a fragment of a portion of a target sequence comprising an overhang; and(b) a nucleic acid comprising a second anchor bound to a second binding molecule on a surface, a cleavage site for a second nuclease, and a fragment of a portion of a target sequence comprising an overhang that is reverse complementary to the overhang in the nucleic acid comprising the first anchor, thereby generating a ligated nucleic acid comprising the first anchor, the cleavage site for the first nuclease, the portion of the target sequence formed by ligating the overhangs, the cleavage site for the second nuclease, and the second anchor;(ii) cleaving the ligated nucleic acid with the first nuclease, thereby generating a cleavage fragment comprising the first anchor bound to the first binding molecule and the cleavage site for the first nuclease, anda cleaved nucleic acid comprising the second anchor bound to the second surface, the cleavage site for the second nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang generated by the first nuclease and the fragment is elongated relative to the fragments of the target sequence in the nucleic acids ligated in a preceding ligation step;(iii) subjecting the first binding molecule to conditions that inhibit binding to the first anchor;(iv) repeating step (i) by performing a subsequent ligation step, wherein the subsequent ligation step comprises ligating(a) a nucleic acid comprising the first anchor bound to the first binding molecule on a surface, the cleavage site for the first nuclease, and a fragment of a26 / 112#14588836v1H0824.70442WO00portion of the target sequence, wherein the fragment comprises an overhang that is reverse complementary to the overhang in the cleaved nucleic acid; and (b) the cleaved nucleic acid,thereby generating a ligated nucleic acid comprising the portion of the target sequence formed by ligating the overhangs, wherein the portion of the target sequence is elongated relative to the portion of the target sequence in the ligated nucleic acid generated in the preceding ligation step; and(v) repeating step (ii) by performing a subsequent cleavage step, wherein the subsequent cleavage step comprises cleaving the ligated nucleic acid with the first nuclease or the second nuclease, thereby generatinga cleavage fragment, anda cleaved nucleic acid comprising a fragment of a portion of the target sequence.
[0082] In some embodiments, a method comprises repeating the ligation and cleavage steps (e.g., by performing a cycle of ligation and cleavage steps) two times or more (e.g., 2 times or more, 3 times or more, 4 times or more, 5 times or more, 6 times or more, 7 times or more, 8 times or more, 9 times or more, 10 times or more, 11 times or more, 12 times or more, 13 times or more, 14 times or more, 15 times or more, 16 times or more, 17 times or more, 18 times or more, 19 times or more, 20 times or more, 21 times or more, 22 times or more, 23 times or more, 24 times or more, 25 times or more, 26 times or more, 27 times or more, 28 times or more, 29 times or more, 30 times or more, 31 times or more, 32 times or more, 33 times or more, 34 times or more, 35 times or more, 36 times or more, 37 times or more, 38 times or more, 39 times or more, or 40 times or more). In some embodiments, a method comprises repeating the ligation and cleavage steps (e.g., by performing a cycle of ligation and cleavage steps) two times to fifty times (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 times). In some embodiments, one or more steps for selection are performed before a subsequent ligation and / or a subsequent cleavage (e.g., before performing a subsequent cycle obligation and cleavage). In some embodiments, the one or more steps for selection comprise subjecting a surface to conditions that remove cleavage fragments and / or subjecting a surface to one or more washes. In some embodiments, one or more steps for selection are performed before performing a final ligation.
[0083] In some embodiments, a method comprises multiplexed nucleic acid production, wherein the multiplexing comprises repeating a ligation step followed by a cleavage step and one or 27 / 112#14588836v1H0824.70442WO00more steps for double selection as described herein (e.g., by performing a cycle of ligation and cleavage steps) for a plurality of times (e.g., 2-50 times), wherein each repetition comprises a plurality of ligations performed in parallel (e.g., concurrently) and then a plurality of double selection steps performed in parallel (e.g., concurrently) (see, e.g., Example 1 below). For example, in some embodiments, a plurality of surfaces (e.g., a plurality of acceptor surfaces and a plurality of donor surfaces), such as two more surfaces, are utilized (e.g., utilized concurrently) during each of steps (i)-(v) according to Method Example 1 and during each of steps (i)-(v) according to Method Example 2. In some embodiments, the plurality of surfaces utilized during steps (i)-(iii) comprises hundreds of surfaces (e.g., 100-999 surfaces), thousands of surfaces (e.g., 1,000-99,999 surfaces, hundreds of thousands of surfaces (e.g., 100,000-999,999 surfaces), millions of surfaces (e.g., 1,000,000-99,000,000 surfaces), or hundreds of millions of surfaces (e.g., 100,000,000-999,000,000 surfaces), or more. In some embodiments, wherein cleaved nucleic acids generated according to steps (i)-(iii) in Method Example 1 are ligated to cleaved nucleic acids generated according to steps (i)-(iii) in Method Example 2, the plurality of surfaces utilized during steps (iv)-(v) can be equal to ½ x the plurality of surfaces utilized in the preceding steps (e.g., the preceding ligation step according to (i) and the preceding cleavage step according to (iii)). In some embodiments, multiplexing comprises performing a plurality of cycles (e.g., 2-50 cycles), wherein each cycle comprises performing a ligation step followed by a cleavage step and one or more steps for selection as described herein, wherein each ligation step comprises performing a plurality of ligations (e.g., hundreds to hundreds of millions of ligations) in parallel (e.g., concurrently), wherein each cleavage step comprises performing a plurality of cleavages (e.g., hundreds to hundreds of millions of cleavages) in parallel (e.g., concurrently), and wherein the one or more steps for selection comprise subjecting one or more surfaces (e.g., hundreds to hundreds of millions of surfaces) to conditions that remove cleavage fragments and / or to one or more washes in parallel (e.g., concurrently). In some embodiments, a method comprising multiplexed nucleic acid production can be used to produce a plurality of nucleic acids, each comprising the same target sequence (e.g., a plurality of nucleic acids, wherein two or more nucleic acids in the plurality comprise the same target sequence (e.g., wherein two or more nucleic acids comprise a target sequence that is human chromosome I) and / or wherein two more nucleic acids in the plurality comprise different target sequences (e.g., wherein two or more nucleic acids comprise a target sequence that is a first chromosome, such as human chromosome 7 and two and more nucleic acids comprise a target sequence that is a second chromosome, such as human chromosome 9).28 / 112#14588836v1H0824.70442WO00
[0084] In some embodiments, a final ligation step generates a ligated nucleic acid comprising a target sequence. In some embodiments, each of the ligation steps comprise ligating a cleaved nucleic acid generated in a preceding cleavage step (e.g., generated by cleaving a ligated nucleic acid with the second nuclease, and cleaved nucleic acid generated in a preceding cleavage step by cleaving a ligated nucleic acid with the first nuclease), wherein a final ligation step generates a ligated nucleic acid comprising the target sequence. In some embodiments, the target sequence is produced in a final ligation step, wherein cleavage step is not performed following the final ligation step.
[0085] Table 1 shows example of a nucleic acid production method comprising “n” steps of ligation (32 ligation steps in this example) followed by cleavage (e.g., steps (i)-(v) according to Method Example 1 described herein and / or steps (i)-(v) according to Method Example 2 described herein) and the final product size as a result. The number of ligations performed during each of the “n” steps will typically be determined the by the length of the target sequence, as further described herein.
[0086] Table 1. Example of Nucleic Acid Production Comprising 32 Ligation StepsSteps Exponential Grow Synthesized BasesStep 1 Starting material = 2 basesStep 2 2 + 2 - 1 = 3 basesStep 3 3 + 3 - 1 = 5 basesStep 4 5 + 5 - 1 = 9 basesStep 5 9 + 9 - 1 = 17 basesStep 6 17 + 17 - 1 = 33 basesStep 33 2A(33-1) + 1 = 4.29 10A9 basesStep n 2A(n-l) + 1 = x bases
[0087] Compositions, apparatuses, systems, kits, and methods described in the application can be utilized to produce target nucleic acid sequences with complete sequence flexibility, encompassing both functional genetic elements and arbitrary nucleotide sequences devoid of biological significance. This versatility enables the creation of a wide range of target nucleic acid sequences, including but not limited to gene fragments, regulatory elements, and synthetic DNA sequences for diverse applications. The length of a target sequence can range from tens of nucleotides to hundreds of millions of nucleotides. For example, in some embodiments, a target sequence or a portion of a target sequence located in a ligated nucleic acid comprises “X” bases, wherein X=2n, wherein “n” is equal to the number of cycles performed, wherein each cycle29 / 112#14588836v1H0824.70442WO00involves a cleavage step followed by a ligation step. As a further example, in some embodiments, wherein overhangs comprising one nucleotide in length are utilized in the ligation steps, a fragment of a target sequence or a fragment of a portion of a target sequence located in a cleaved nucleic acid comprises “Y” bases, wherein Y=2(n l)+1, “n” is equal to the number of cycles performed, wherein each cycle involves a ligation step followed by a cleavage step. In some embodiments, wherein overhangs comprising three nucleotides in length are utilized in the ligation steps, a fragment of a target sequence or a fragment of a portion of a target sequence located in a cleaved nucleic acid comprises “Y” bases, wherein Y=3 x (2(n l)+l), “n” is equal to the number of cycles performed, wherein each cycle involves a ligation step followed by a cleavage step.
[0088] In some embodiments, producing a nucleic acid comprises one or more steps (e.g., a ligation, a cleavage, or a recombination; see, e.g., FIGs. 1A-1D. 2A-2B, and 15B), or using a surface that is configured, to control the spatial proximity of nucleic acids. Controlling the spatial proximity of nucleic acids can involve reducing the diffusion distance or localized concentration of reactants to promote more frequent productive encounters between components of a composition (e.g., nucleic acids and proteins, such as enzymes, including ligases, nucleases, and recombinases). For example, in some embodiments, a crowding agent is used in an enzymatic step of a method described herein (e.g., a ligation step, a cleavage step, or a recombination step). In some embodiments, the crowding agent comprises Polysorbate 20 (IUPAC: Polyoxyethylene (20) sorbitan monolaurate; commercially known as “Tween 20”TM), polysorbate 80 (IUPAC: Polyoxyethylene (20) sorbitan monooleate; commercially known as “Tween 80” TM), (l,l,3,3-Tetramethylbutyl)phenyl-polyethylene glycol, Polyethylene glycol tert- octylphenyl ether (commercially known as “Triton X-114”TM), sodium dodecylsulfate (SDS), deoxycholate sodium, (3-((3-cholamidopropyl) dimethylammonio)-l -propanesulfonate) (commonly known as CHAPS detergent), benzalkonium chloride, polyethylene glycol (PEG) (e.g., PEG 200, PEG 400, PEG 1000, PEG 1550, PEG 2000, PEG 3350, PEG 6000, PEG 8000, PEG 20000, or PEG 50000), Ficoll 70, Ficoll 400, serum albumin, and / or dextran (for example, but not limited to 50 k or 500 k). In some embodiments, the concentration of the crowding agent used in an enzymatic step is at least at least 1% w / v (e.g., 1-5% w / v, 10-15% w / v, 15-20% w / v, 20-25% w / v, 25-30% w / v, 30-35% w / v, 35-40% w / v, 40-45% w / v, or 45-50% w / v, 50-60% w / v, 60-70% w / v, or 70-80% w / v). In some embodiments, the crowding agent is PEG (e.g., wherein the concentration of PEG is between 10 mM PEG to 50 mM PEG). As a further example, in some embodiments, nucleic acids are confined or utilized in a surface-immobilized system (e.g., during a ligation, a cleavage, or a recombination reaction). In some embodiments, a surface 30 / 112#14588836v1H0824.70442WO00comprises one or more molecular spacers, which can be used to increase efficiency of enzymatic events (e.g., a ligation, a cleavage, or a recombination reaction) on solid surfaces. In some embodiments, nucleic acids are contacted with a protein (e.g., a nucleic acid-binding protein, such as a Cas protein or an catalytically inactive variant thereof) and form interactions that tether or localize nucleic acids within confined reaction environments.
[0089] Target sequences can also be generated with low error rates (see, e.g., FIG. 6). Error rates will typically be quantified relative to a reference sequence. For example, if the target sequence is human chromosome 1, the reference sequence can be a wild- type human chromosome 1 nucleotide sequence. Accordingly, in some embodiments, reference sequences can be obtained from a publicly available database, such as NCBI GenBank. Conventional chemical and enzymatic nucleic acid synthesis methods exhibit error rates of approximately 10’3errors per base and 10’4errors per base, respectively, relative to a reference sequence.Compositions, apparatuses, systems, kits, and methods described herein can be utilized to produce target sequences comprising 1 x 10’4errors per base or lower relative to a reference sequence. In some embodiments, a target sequence comprises IxlO-4to IxlO-9(e.g., 9xl0-4, 8xl0’4, 7xl0’4, 6xl0’4, 5xl0’4, 4xl0’4, 3xl0’4, 2xl0’4, IxlO’4, 9xl0’5, 8xl0’5, 7xl0’5, 6xl0’5, 5xl0’5, 4xl0’5, 3xl0’5, 2xl0’5, IxlO’5, 9xl0’6, 8xl0’6, 7xl0’6, 6xl0’6, 5xl0’6, 4xl0’6, 3xl0’6, 2xl0’6, IxlO’6, 9xl0’7, 8xl0’7, 7xl0’7, 6xl0’7, 5xl0’7, 4xl0’7, 3xl0’7, 2xl0’7, IxlO’7, 9xl0’8, 8xl0’8, 7xl0’8, 6xl0’8, 5xl0’8, 4xl0’8, 3xl0’8, 2xl0’8, IxlO’8, 9xl0’9, 8xl0’9, 7xl0’9, 6xl0’9, 5xl0-9, 4xl0-9, 3xl0-9, 2xl0-9, or IxlO-9) errors per base or lower relative to a reference sequence. In some embodiments, a target sequence comprises 3 x 10’9errors per base or lower relative to a reference sequence.
[0090] Further embodiments of the compositions, apparatuses, systems, kits, and methods of this application are described below.Anchors and Binding Molecules
[0091] Examples of anchors include small molecules, click-chemistry handles, peptides, proteins, and nucleic acids (e.g., double- stranded nucleic acids or single- stranded nucleic acid, such as oligonucleotides, or nucleic acids comprising one or more double- stranded regions of sequence and one or more single- stranded regions of sequence), which can include naturally occurring nucleic acids and chemically modified nucleic acids. The acceptor and donor nucleic acids can comprise different anchors within the same group of molecules (e.g., wherein the first and second anchors are peptides that bind to different binding molecules) or different groups of molecules (e.g., wherein the first anchor is an oligonucleotide that binds a reverse31 / 112#14588836v1H0824.70442WO00complementary oligonucleotide and the second anchor is a peptide). Anchors can be connected to nucleic acids through covalent bonds and / or non-covalent bonds (e.g., via interactions involving van der Waals interactions, pi-pi interactions, hydrophobic interactions, hydrogen bonding, ionic bonds, or any combination thereof). Anchors will typically bind to cognate binding molecules reversibly such that, for example, nucleic acids can be subjected to one or more conditions that inhibit the binding between the anchor and the binding molecule, such as by contacting the anchor or binding molecule with a competitive binding agent (e.g., an agent that has higher affinity for the anchor or binding molecule and / or an agent that is provided in a concentration sufficient to reverse binding between the anchor and the binding molecule).
[0092] In some embodiments, an anchor and / or a binding molecule comprises a peptide or protein. Examples of peptides or proteins that can be utilized as anchors and / or a binding molecule include a carbohydrate-binding peptide or protein (e.g., peptides or proteins that bind to maltose, lactose, or galactose), peptides or proteins used for affinity purification (e.g., a polyHis tag, a glutathione S-transferase (GST) tag, a poly Arg tag, a FLAG tag, a streptavidin-binding tag, a calmodulin-binding tag, a chitin-binding tag, and a cellulose-binding tag), and antibodies or antigen-binding fragments thereof. In some embodiments, an anchor comprises a maltose-binding protein tag and a binding molecule comprises maltose (e.g., wherein a nucleic acid is subjected to maltose to inhibit binding of the polyHis tag to NTA). In some embodiments, an anchor comprises a polyHis tag (e.g., a 6x-His tag, a 9x-His tag, or a 12x-His tag) and a binding molecule comprises Nickel-Nitriloacetic acid (NTA) (e.g., wherein a nucleic acid is subjected to imidazole to inhibit binding of the polyHis tag to NTA). In some embodiments, the first anchor comprises a maltose-binding protein (and the first anchor comprises maltose) and / or the second anchor comprises a polyHis tag (and the second binding molecule comprises Nickel-Nitriloacetic acid (NTA)).
[0093] In some embodiments, an anchor comprises a nucleic acid and a binding molecule comprises a reverse complementary nucleic acid. In some embodiments, both the first and second anchors comprise nucleic acids that hybridize with respective nucleic acids that are utilized as binding molecules. In some embodiments, a nucleic acid utilized as an anchor and a reverse complementary nucleic acid utilized as a binding molecule comprise at least ten nucleotides (e.g., at least ten contiguous nucleotides) that promote hybridization between the anchor nucleic acid and the binding molecule nucleic acid, such as wherein the at least ten nucleotides comprise at least fifteen nucleotides, at least twenty nucleotides, at least twenty-five nucleotides, at least thirty nucleotides, at least thirty-five nucleotides, at least forty nucleotides, at least fifty nucleotides, or more than fifty nucleotides. In some embodiments, a nucleic acid 32 / 112#14588836v1H0824.70442WO00utilized as an anchor and a reverse complementary nucleic acid utilized as a binding molecule comprises ten to forty nucleotides (e.g., ten to forty contiguous nucleotides) that promote hybridization between the anchor nucleic acid and the binding molecule nucleic acid, such as ten to fifteen nucleotides, fifteen to twenty nucleotides, twenty to twenty-five nucleotides, twenty-five to thirty nucleotides, thirty to thirty-five nucleotides, or thirty-five to forty nucleotides.
[0094] In some embodiments, a nucleic acid utilized as an anchor comprises a reverse complementary nucleic acid to a portion of an nucleic acid utilized as a binding molecule (e.g., wherein the anchor nucleic acid hybridizes to ten to forty nucleotides located in the portion of the binding molecule nucleic acid), and a further nucleic acid is contacted with the binding molecule to reverse the binding between the anchor nucleic acid and the binding molecule nucleic acid.
[0095] In some embodiments, wherein the anchor comprises a nucleic acid and the binding molecule comprises a reverse complementary nucleic acid, toehold displacement is utilized to inhibit the hybridization of the anchor nucleic acid and the binding molecule nucleic acid, such as wherein a further nucleic acid binds to a single- stranded region or “toehold” located in the binding molecule nucleic acid and displaces the anchor nucleic acid (see, e.g., the toehold systems diagramed in FIGs. 14B). Generally, a toehold comprises a single- stranded nucleic acid segment that is typically positioned at the end or within a larger nucleic acid that is designed to be specifically recognized and hybridized by a complementary nucleic acid sequence. A toehold may be composed of DNA, RNA, chemically modified nucleotides, or any functional analog thereof. The toehold may be located at the 5' end, 3' end, or an internal position of a hairpin, dumbbell, circular DNA, or other nucleic acid construct. The toehold provides an initial binding site that facilitates the hybridization, enabling subsequent strand exchange, strand displacement, or structural rearrangement via anti-toehold strand. In some embodiments, the single- stranded region or toehold comprises at least ten nucleotides (e.g., wherein the at least ten nucleotides comprises at least fifteen nucleotides, at least twenty nucleotides, at least twenty-five nucleotides, at least thirty nucleotides, at least thirty-five nucleotides, at least forty nucleotides, at least fifty nucleotides, or more than fifty nucleotides, or wherein the at least ten nucleotides comprises ten to fifteen nucleotides, fifteen to twenty nucleotides, twenty to twenty-five nucleotides, twenty-five to thirty nucleotides, thirty to thirty-five nucleotides, or thirty-five to forty nucleotides). In some embodiments, nucleic acids comprising eight nucleotides or more in length (e.g., at least nine nucleotides, at least ten nucleotides, at least 11 nucleotides, at least twelve nucleotides, at least thirteen nucleotides, at least fourteen nucleotides, at least fifteen nucleotides, at least twenty nucleotides, at least twenty-five nucleotides, at least thirty 33 / 112#14588836v1H0824.70442WO00nucleotides, at least thirty-five nucleotides, at least forty nucleotides, at least fifty nucleotides, or more than fifty nucleotides, or wherein the eight nucleotides or more comprises ten to fifteen nucleotides, fifteen to twenty nucleotides, twenty to twenty-five nucleotides, twenty-five to thirty nucleotides, thirty to thirty-five nucleotides, or thirty-five to forty nucleotides) are utilized in a triplex toehold system. In some embodiments, nucleic acids utilized in a triplex toehold system comprises more purine residues relative pyrimidine residues. Regular triplex forming rules can also be used in triplex toehold systems (see, e.g., the descriptions of triplex nucleic acids in: Seidman & Glazer (2003). The potential for gene repair via triple helix formation. J Clin Invest, 112(4): 487-494; Rapozzi et al. (2002). Antigene Effect in K562 Cells of a PEG-Conjugated Triplex-Forming Oligonucleotide Targeted to the bcr / abl Oncogene. Biochemistry, 41(2): 502-510; and Zain & Sun (2003). Do natural DNA triple-helical structures occur and function in vivo? Cell Mol Life Sci, 60: 862-870, which are incorporated by reference herein for descriptions of triplex nucleic acids). In some embodiments, wherein toehold displacement is utilized, an anchor nucleic acid comprises a sequence having reverse complementarity to the binding molecule nucleic acid and that hybridizes with lower affinity (e.g., wherein the sequence is shorter in length and / or has less than 100% reverse complementarity) to the binding molecule nucleic acid relative to the nucleic acid that is contacted with the binding molecule nucleic acid to displace the anchor nucleic acid). In some embodiments, wherein toehold displacement is utilized, the anchor nucleic acid is a double- stranded nucleic acid. In some embodiments, wherein toehold displacement is utilized, the nucleic acid utilized to disrupt the binding between the anchor nucleic acid and the binding molecule nucleic acid is a single- stranded nucleic acid (see, e.g., the toehold system diagrammed in the top panel of FIG. 14B). In some embodiments, wherein toehold displacement is utilized, the nucleic acid utilized to disrupt the binding between the anchor nucleic acid and the binding molecule nucleic acid is a double- stranded nucleic acid (see, e.g., the triplex toehold system diagrammed in the middle panel of FIG. 14B). In some embodiments, wherein toehold displacement is utilized, the nucleic acid utilized to disrupt the binding between the anchor nucleic acid and the binding molecule nucleic acid comprises a double- stranded region (e.g., a nucleic acid comprising a hairpin) (see, e.g., the toehold system diagramed in the bottom panel of FIG. 14B).
[0096] In some embodiments, one or more chemical modifications are located in a nucleic acid utilized as an anchor, a nucleic acid utilized as a binding molecule, and / or a nucleic acid utilized to inhibit binding between an anchor nucleic acid and a binding molecule nucleic acid. The one or more chemical modifications can include modifications located at the base moiety, sugar moiety, or an internucleotide linkage. The one or more chemical modifications can be located at 34 / 112#14588836v1H0824.70442WO00a purine residue (e.g., adenine, guanine, xanthine, caffeine, uric acid, isoguanine, theobromine, theophylline, or hypoxanthine) and / or a pyrimidine residue (e.g., thymine, cytosine, uracil, fluorouracil, barcituric acid, or orotic acid). Examples of chemically modified base moieties include, but are not limited to, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-hyrdoxymethylcytosine, 5-(carboxyhydroxylmethyl) uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1 -methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N4-methylcytosine, N6-methyladenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosine, 5 ’’-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methylester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl) uracil, a thioguanine, and 2,6-diaminopurine. Examples of chemically modified sugar moieties include, but are not limited to, sugar moieties comprising a modification at a 2' carbon, such as 2'-O-alkyl (including 2'-O-methyl and 2'-O-ethyl), 2'-alkoxy, 2'-amino, 2'-S-alkyl, 2'-halo (including 2'-fluoro), 2'-2-0-methoxyethoxy, 2'-allyloxy (-OCH2CH=CH2), 2'-propargyl, 2'-propyl, 2'-ethynyl, 2'-ethenyl, 2'-propenyl, and 2'-cyano. Other 2' modifications include substitutions of a 2'-OH group with H, OR, R, F, Cl, Br, I, SH, SR, NH, NHR, NR, COOR, or OR, wherein “R” is a substituted or unsubstituted aliphatic group or alkyl group. The term “aliphatic,” as used herein, includes both saturated and unsaturated, straight chain (i.e., unbranched), branched, acyclic, cyclic, or polycyclic aliphatic hydrocarbons, which are optionally substituted with one or more functional groups. As will be appreciated by one of ordinary skill in the art, “aliphatic” is intended herein to include, but is not limited to, alkyl, alkenyl, alkynyl, cycloalkyl, cycloalkenyl, and cycloalkynyl moieties. Examples of chemically modified internucleotide linkages include, but are not limited to, substitution of an oxygen atom with a sulfur atom as well as linkages comprising phosphorothioate, borano-phosphate, and alkyl phosphonate acid. This includes single-and double- stranded molecules, i.e., DNA-DNA, DNA-RNA and RNA-RNA hybrids, as well as “protein nucleic acids” or “peptide nucleic acids” (PNAs) formed by conjugating bases to an amino acid backbone. Chemically modified nucleic acids can also be modified to comprise carbohydrate or lipids. Further examples of chemical modifications include Romesburg modifications (e.g., also known as third base pairing X-Y, artificial base pairings) (see, e.g., Malyshev et al. (2014). Nature, 509: 385-388, which is incorporated by 35 / 112#14588836v1H0824.70442WO00reference for its descriptions of Romesburg modifications and nucleic acid engineering involving Romesburg modifications). Chemical modifications can be utilized to, for example, improve stability of the molecule and / or hybridization parameters.
[0097] In some embodiments, an anchor and / or a binding molecule comprises a click-chemistry handle. In some embodiments, a click-chemistry handle is capable of reacting with an amino group, a sulfhydryl group, a hydroxyl group, an imidazole group, an indole group, a guanidinium, or a carboxylic acid group. In some embodiments, a click-chemistry handle comprises an ester functional group, such as N-hydroxy succinimidyl (NHS) esters or a functional group in succinimidyl valeric acid (SVA). In some embodiments, a click-chemistry handle comprises acrylate or methacrylate. In some embodiments, a click-chemistry handle comprises n-acetylgalactosamine. In some embodiments, a click-chemistry handle comprises maleimide. In some embodiments, a click-chemistry handle comprises an azide group. In some embodiments, a click-chemistry handle comprises an alkyne group. In some embodiments, a click-chemistry handle comprises a hexynyl group.
[0098] In some embodiments, an anchor and / or a binding molecule comprises a small molecule. A small molecule can be naturally- occurring or artificially created (e.g., via chemical synthesis). Typically, a small molecule is an organic compound (e.g., it contains carbon). The small molecule may contain multiple carbon-carbon bonds, stereocenters, and other functional groups (e.g., amines, hydroxyl, carbonyls, and heterocyclic rings, etc.). In certain embodiments, the molecular weight of a small molecule is not more than about 1,000 g / mol, not more than about 900 g / mol, not more than about 800 g / mol, not more than about 700 g / mol, not more than about 600 g / mol, not more than about 500 g / mol, not more than about 400 g / mol, not more than about 300 g / mol, not more than about 200 g / mol, or not more than about 100 g / mol. In certain embodiments, the molecular weight of a small molecule is at least about 100 g / mol, at least about 200 g / mol, at least about 300 g / mol, at least about 400 g / mol, at least about 500 g / mol, at least about 600 g / mol, at least about 700 g / mol, at least about 800 g / mol, or at least about 900 g / mol, or at least about 1,000 g / mol. Combinations of the above ranges (e.g., at least about 200 g / mol and not more than about 500 g / mol) are also possible. In certain embodiments, the small molecule is a therapeutically active agent such as a drug (e.g., a molecule approved by the U. S. Food and Drug Administration as provided in the Code of Federal Regulations (C. F. R.)). The small molecule may also be complexed with one or more metal atoms and / or metal ions. In this instance, the small molecule is also referred to as a “small organometallic molecule.” Preferred small molecules are biologically active in that they produce a biological effect in animals, preferably mammals, more preferably humans. Small molecules include, but are not limited to,36 / 112#14588836v1H0824.70442WO00radionuclides and imaging agents. In certain embodiments, the small molecule is a drug.Preferably, though not necessarily, the drug is one that has already been deemed safe and effective for use in humans or animals by the appropriate governmental agency or regulatory body. For example, drugs approved for human use are listed by the FDA under 21 C. F. R. §§ 330.5, 331 through 361, and 440 through 460, incorporated herein by reference; drugs for veterinary use are listed by the FDA under 21 C. F. R. §§ 500 through 589, incorporated herein by reference. All listed drugs are considered acceptable for use in accordance with the present application.Nucleases and Cleavage Sites Thereof
[0099] Nuclease cleavage sites can be used for liberation of nucleic acids from the clonally sequenced plasmids before the nucleic acids are processed further to generate acceptor and donor nucleic acids. Nuclease cleavage sites can also be used to remove 5’ or 3’ends to, for example, further prepare for generating acceptor and donor nucleic acids. Nuclease cleavage sites can also be used to facilitate attachment of the nucleic acids to one or more additional molecules (e.g., anchors and / or DNA adapters). Nuclease cleavage sites can also be used to introduce overhangs (e.g., 1-mer overhangs) to prepare for subsequent ligation steps.
[0100] Nucleases can include exonucleases and endonucleases including, but not limited to, nucleases having double-strand cutting activity or single-strand cutting activity (e.g., nickase activity). Further examples of nucleases include restriction enzymes (e.g., type II restriction enzymes or type V restrictions enzymes, such as Cas proteins including, but not limited to, Cas9 (e.g., SpCas9, SaCas9, StCas9, NmCas9, CjCas9, and SpyCas9) and variants thereof, Casl2 (e.g., Casl2a (Cpfl), such as AsCasl2a, FnCasl2a, LbCasl2a, and PaCasl2a, or Casl2b) and variants thereof, Cas 13 and variants thereof, Cas 14 and variants thereof, CasX, CasY, zinc-finger nucleases, and transcription activator-like effector nucleases). The first and second nucleic acids can comprise different nuclease cleavage sites that are recognized by nucleases within the same group (e.g., wherein the first and second nucleases are different restriction enzymes) or different groups (e.g., wherein the first nuclease is a restriction enzyme and the second nuclease is a Cas protein bound to a guide RNA (gRNA), wherein the nucleic acid comprises a sequence that hybridizes to the gRNA and a protospacer adjacent motif (PAM) that can bind to the Cas protein). In some embodiments, cleaving nucleic acids comprises contacting nucleic acids with a nuclease by utilizing automatic pipetting.
[0101] In some embodiments, a nucleic acid comprises a cleavage site for a restriction enzyme. In some embodiments, a restriction enzyme is a type IIS restriction enzyme. In some 37 / 112#14588836v1H0824.70442WO00embodiments, a restriction enzyme generates single- stranded breaks (e.g., wherein the restriction enzyme is a nickase). In some embodiments, a restriction enzyme generates double- stranded breaks. In some embodiments, one or more nicking enzymes can be employed as restriction enzymes to introduce single- stranded breaks in the DNA to achieve double- stranded breaks. In some embodiments, a restriction enzyme is Alwl, Bccl, Piel, BciVI, BmrI, HphI, HpryAv, Mboll, Mull, Sbfl, or PflMI. In some embodiments, a nuclease is an isoschizomer of a restriction enzyme (e.g., an isoschizomer of Alwl, Bccl, Piel, BciVI, BmrI, HphI, HpryAv, Mboll, Mull, Sbfl, or PflMI, such as AclWI or BspPI). In some embodiments, a modified restriction enzyme is utilized, such as a restriction enzyme engineered to cleave or nick DNA sequences outside of their canonical recognition sites. For example, in some embodiments, a Bcgl restriction enzyme, which typically cleaves both upstream and downstream of DNA within its recognition site, is modified to cleave only one side of a nucleic acid. In some embodiments, a restriction enzyme is Nt. AlwI, Nt. BbvCI, Nt. BsmAI, or Nt. BspQI. In some embodiments, the first nuclease is Alwl and / or the second nuclease is Bccl.
[0102] In some embodiments, a nucleic acid comprises two nuclease cleavage sites. For example, in some embodiments, a nucleic acid comprises, in relative order, an anchor, a cleavage site for one nuclease and a cleavage site for another nuclease, and an overhang that is utilized in the ligation. In some embodiments, one or both of the nucleases are selected from Alwl, Bccl, Piel, BciVI, BmrI, HphI, HpryAv, Mboll, Mull, Sbfl, and PflMI. In some embodiments, a nucleic acid comprises and an Alwl cleavage site and a Sbfl cleavage site. In some embodiments, a nucleic acid comprises a Bccl cleavage site and a PflMI cleavage site. Examples of restriction enzymes and sequences comprising cleavage sites thereof are shown in Table 2.Table 2. Examples of Restriction enzymes and Sequences Comprising Cleavage Sites Thereof Restriction Enzyme List5' cutters 5' to 3'Alwl GGATCNNNNVNBccl CCATCNNNNVNAPiel GAGTCNNNNVN3' cuttersBciVI GTATCCNNNNNANVBmrI ACTGGGNNNNANVHphI GGTGANNNNNNNANVHpyAV CCTTCNNNNNANVMboll GAAGANNNNNNNANVMnll CCTCNNNNNNANVAdapter Enzymes38 / 112#14588836v1H0824.70442WO00Sbfl CCTGCAvGGPflMI CCANNNNVNTGG*Aand v indicate positions cut by the corresponding nucleaseOverhangs
[0103] Reverse complementary overhangs utilized in ligation will typically comprise the same number of nucleotides in length. However, in some embodiments, overhangs that differ in length by one or more nucleotides can be ligated. For example, in some embodiments, a DNA ligase capable of repairing single-nucleotide gaps in double stranded DNA is utilized and / or a additional enzymes (e.g., DNA polymerase) are utilized when overhangs of different lengths are ligated. Accordingly, complementary overhangs used in ligation will typically, but not necessarily, comprise the same number of nucleotides in length, and ligation can still proceed under conditions where the lengths are not identical. In some embodiments, overhangs comprising odd number of nucleotides in length are ligated, such as overhangs of one nucleotide, three nucleotides, five nucleotides, seven nucleotides, or nine nucleotides in length. In some embodiments, overhangs comprising one nucleotide in length are ligated.
[0104] However, in some embodiments, overhangs comprising an even number of nucleotides in length are ligated, such as two nucleotides, four nucleotides, or six nucleotides in length.
[0105] However, in some embodiments, a blunt end of a nucleic acid is utilized in a ligation.Target Nucleic Acids
[0106] The ability to synthesize target nucleic acids, including large target nucleic acids (e.g., 50 kb or more), with near-perfect fidelity substantially expands the range of biological and industrial applications accessible through the compositions, apparatuses, systems, kits, and methods described herein. These large target nucleic acids bridge the current gap between genelength and genome-scale assemblies, and can be used for both de novo genome synthesis and precise, modular engineering of complex biological systems. For example, in some embodiments, compositions, apparatuses, systems, kits, and methods described herein can be used for the hierarchical assembly of megabase- or gigabase- scale synthetic chromosomes, allowing de novo design, recoding, or refactoring of viral, microbial, plant, or mammalian genomes. These synthetic genomes can serve as chassis for programmable cells, xenobiotic metabolism, or viral resistance. In some embodiments, compositions, apparatuses, systems, and methods described herein can be used for the synthesis of custom or personalized organelle 39 / 112#14588836v1H0824.70442WO00DNA, including but not limited to mitochondrial DNA or chloroplast DNA. These constructs can be generated with high sequence fidelity and structural precision, supporting applications in organelle engineering, functional genomics, and therapeutic development. In some embodiments, compositions, apparatuses, systems, kits, and methods described herein can be scaled to megabase or gigabase for a full virus, microbial, archaea, plant, fungi, mammalian, or any other life forms’ genome dimensions and facilitate the generation of entire genomes within a continuous high-fidelity assembly framework. In some embodiments, compositions, apparatuses, systems, and methods described herein can_constructs within the 50-500 kb range, which can accommodate entire biosynthetic gene clusters, such as those encoding nonribosomal peptide synthetases, polyketide synthases, or metabolic pathways for antibiotics, pigments, and fine chemicals. This facilitates rapid prototyping and optimization of complex enzymatic pathways in both native and heterologous hosts. In some embodiments, compositions, apparatuses, systems, kits, and methods described herein can be used to reconstitute minimal cells, synthetic organelles, or virus-like particles, where long contiguous genetic circuits are required to maintain structural or regulatory integrity. In some embodiments, compositions, apparatuses, systems, kits, and methods described herein can produce long synthetic DNA molecules can serve as backbones for gene therapy vectors, oncolytic viruses, or cellular immunotherapies, allowing customized or personalized genetic payloads that integrate regulatory networks, safety switches, and multi-gene expression cassettes in a single construct. Additionally, in some embodiments, compositions, apparatuses, systems, kits, and methods described herein the synthesis of low error rate DNA molecules exceeding 10 kb, which can be used for high-density molecular data storage with dramatically reduced error correction overhead, supporting both archival and dynamic rewriting applications.
[0107] A target sequence can comprise DNA and / or RNA. Examples of DNAs include single- stranded DNA (ssDNA), double- stranded DNA (dsDNA), plasmid DNA (pDNA), genomic DNA (gDNA), complementary DNA (cDNA), antisense DNA, non-coding DNA (ncDNA), DNA comprising coding sequences, chloroplast DNA (ctDNA or cpDNA), microsatellite DNA, mitochondrial DNA (mtDNA or mDNA), kinetoplast DNA (kDNA), provirus, lysogen, repetitive DNA, satellite DNA, and viral DNA. In some embodiments, a target sequence comprises a gene sequence or a portion of a gene. In some embodiments, a target sequence comprises a plurality of gene sequences (e.g., 2-10 genes, 10-100 genes, 100-1,000 genes, or more than 1,000 genes). Examples of RNAs encoded by genes include a smallhairpin RNA (shRNA), a short-interfering RNA (siRNA), a prokaryotic-interfering RNA (prosiRNA), a micro-RNA (miRNA), a long non-coding RNA (IncRNA), a Piwi-interacting RNA 40 / 112#14588836v1H0824.70442WO00(piRNA), an exon-skipping RNA, an enzymatic RNA, a guide RNA (gRNA), a small nuclear RNA (snRNA), a small nucleolar RNA (snoRNA), a ribosomal RNA (rRNA), a transfer RNA (tRNA), an RNA aptamer, and a messenger RNA (mRNA). In some embodiments, an RNA is an inhibitory RNA, such as a shRNA, an siRNA, a miRNA, a IncRNA, an exon-skipping RNA, an RNA aptamer, or an enzymatic RNA. In some embodiments, a nucleic acid comprising a target sequence is administered to a subject characterized as having or suspected of having a disease, a disorder, or a condition described herein.
[0108] In some embodiments, a target sequence comprises a sequence encoding an mRNA, wherein the mRNA encodes a peptide or a protein. Examples of peptides and proteins include a cell-, organelle-, or tissue-targeting peptide or protein, an antibody, an antigen-binding fragment, an antigen, an enzyme or an enzymatic domain, a proteinaceous enzyme substrate, a glycoprotein, a lipoprotein, a secreted protein, an extracellular matrix protein or a fragment thereof, a viral coat protein, a transmembrane receptor or a fragment thereof, a toxin or a fragment thereof, a hormone, receptors, a peptibody, a growth factor, a clotting factor, a cytokine, a chemokine, an activating or inhibitory peptide, a thrombolytic, a bone morphogenetic protein, an Fc-fusion protein, an anticoagulant, a signaling protein, a cell surface protein, a nucleic acid-binding protein, and a reporter. In some embodiments, a reporter is mNeonGreen, GFP, EGFP, Superfold GFP, Azami Green, mWasabi, TagGFP, TurboGFP, acGFP, zsGreen, T-sapphire, EBFP, EBFP2, mTagBFP, ECFP, Cerulean, mTurquoise, CyPet, AmCyanl, TagCFP, Mtfpl, EYFP, mCitrine, TagYFP, phiYFP, zsYellowl, mBanana, Kusabira Orange, mOrange, dTomato, DsRed, mTangerine, mRuby, mApple, mStrawberry, AsRed2, Mrfpl, mCherry, HcRedl, or smURFP. In some embodiments, an antibody is a monoclonal antibody, a polyclonal antibody, a nanobody, or a single-chain antibody (e.g., an scFv). In some embodiments, a protein comprising an antigen-binding fragment is a chimeric antigen receptor, such as one comprising an antibody. In some embodiments, an enzyme is a protease, a signaling protein, a polymerase, a transcriptional regulator, a nuclease, an RNA-guided nuclease, a metabolic enzyme, a kinase, a phosphatase, a lipid-transferase, a glycosylase, a DNA ligase, a ubiquitin ligase, a methyltransferase, an acetyltransferase, a SUMO transferase, a foldase, a reductase, a lyase, a dehydrogenase, a phosphorylase, a decarobxylase, a dephosphorylase, a kinase, a synthase, or a hydrolase. In some embodiments, a nucleic acid-binding protein is a transcription regulator (e.g., a transcription factor), a splicing regulator, a translation regulator (e.g., a ribosome-binding protein, a signaling protein that post-translationally modifies RNA subunits, etc.), a nuclease (e.g., an RNA-guided nuclease, a zinc finger nuclease, or a transcription activator-like effector nuclease (TAEEN). In some embodiments, a signaling 41 / 112#14588836v1H0824.70442WO00protein is an enzyme (e.g., a kinase), a ligand (e.g., a ligand of a receptor), a secreted protein, a morphogen, a cell differentiation regulator. In some embodiments, a cell surface protein is a membrane protein, such as a receptor, a channel, a lineage- specific antigen, an extracellular matrix protein, or protein comprising a sequence that binds to an antigen-binding fragment, an antibody, or a chimeric antigen receptor. In some embodiments, a hormone is insulin, glucagon, growth hormone, thyroid-stimulating hormone, adrenocorticotropic hormone, follicle-stimulating hormone, luteinizing hormone, prolactin, oxytocin, vasopressin, parathyroid hormone, cortisol, erythropoietin, luteinizing hormone p, and aldosterone. In some embodiments, an antigen is a protein or a fragment thereof on a pathogen (e.g., an immunogenic peptide) or a tumor antigen. In some embodiments, a pathogen is a virus, a bacterium, a fungus, or a parasite. In some embodiments, a pathogen is an agricultural pathogen. In some embodiments, the pathogen is pathogenic (e.g., infects and / or causes one or more symptoms of a disease, disorder, or condition) to mammals (e.g., humans). In some embodiments, an immunogenic protein or an immunogenic fragment thereof (e.g., an immunogenic peptide) comprises an antigen which is derived from a pathogen.
[0109] In some embodiments, a nucleic acid comprising a target sequence (e.g., a nucleic acid that can be administered to a subject) comprises a sequence having one or more positions that lack a mutation characterized as being associated with a disease, disorder, or condition when the one or more mutated positions are present in a subject, such as a human subject. For example, in some embodiments, the target sequence comprises a wild-type version of the sequence. In some embodiments, the one or more comprise nucleotide positions located in a coding sequence and / or nucleotide positions located in a non-coding sequence that are associated with a disease, disorder, or condition.
[0110] In some embodiments, a target sequence comprises a sequence encoding a therapeutic RNA and / or a therapeutic peptide or protein. The term “therapeutic” RNA, peptide, or protein refers to an RNA, peptide, or protein molecule that leads to a physiological change in a cell that can improve a biological process in a cell and / or the organism comprising the cell. In some embodiments, a therapeutic RNA, peptide, or protein is associated with or expected to at least partially, if not fully, prevent and / or remedy at least one symptom associated with a disease, disorder, or condition. In some embodiments, the disease, disorder, or condition comprises a genetic disease, cancer, inflammatory disease or an inflammatory condition, autoimmune disease, spleen disease, lung disease, hematological disease, neurological disease, painful condition, psychiatric disorder, metabolic disorder, immune disorder, infection of a pathogen, a kidney disease, cardiovascular disease, pancreatic disease, intestinal disease, retinal 42 / 112#14588836v1H0824.70442WO00disease, neuromuscular disease, musculoskeletal disease, lysosomal storage disease, or other disease, or any combination thereof. In some embodiments, the disease, disorder, or condition can arise from an infection of a pathogen (e.g., wherein the RNA results in inhibition of pathogen infection and / or propagation, or wherein the protein is a variant of a protein utilized by the pathogen during infection and / or propagation). In some embodiments, the pathogen is pathogenic (e.g., infects and / or causes one or more symptoms of a disease, disorder, or condition) to mammals (e.g., humans). Examples of pathogens and pathogen-associated conditions include, but are not limited to, Adenoviridae, Picornaviridae, Herpesviridae, Hepadnaviridae, Coronaviridae, Flaviviridae, Retroviridae, Orthomyxoviridae, Paramyxoviridae, Papovaviridae, Polyomavirus, Poxviridae, Rhabdoviridae, Togaviridae, Mycobacterium tuberculosis, Streptococcus, Pseudomonas, Shigella, Campylobacter, Salmonella, Candida, Aspergillus, Cryptococcus, Histoplasma, Pneumocytis, Stachybotrus, Bacillus anthracis, Clostridium botulinum, Mycobacterium leprae, Yersinia pestis, Rickettsia prowazekii, Bartonella spp., malaria, amoebiasis, babesiosis, giardiasis, toxoplasmosis, cryptosporidiosis, trichomoniasis, Chagas disease, leishmaniasis, African trypanosomiasis (sleeping sickness), Acanthamoeba keratitis, or primary amoebic meningoencephalitis (naegleriasis).
[0111] In some embodiments, a target sequence comprises a regulatory sequence. In some embodiments, a target sequence comprises a plurality of regulatory sequences (e.g., 2-10 regulatory sequences, 10-100 regulatory sequences, 100-1,000 regulatory sequences, or more than 1,000 regulatory sequences). Examples of regulatory sequences include sequences capable of regulating transcription (e.g., a promoter, an enhancer, a transcription factor binding sequence, a transcriptional start sequence, or a transcription termination sequence), regulating translation (e.g., a 5’ UTR, a translation initiation regulatory sequence, a Kozack sequence, a Shine-Dalgarno sequence, a start codon, a ribosome binding site, a 3’ UTR, a translation termination sequence, or a stop codon), regulating splicing (e.g., binding sites for small nuclear ribonucleoproteins, splicing acceptor sites, or splicing donor sites), or regulating post-translational modifications of a peptide or protein (e.g., an enzyme recognition sequence, a protease cleavage, a self-cleaving peptide, an intein, or a degron).
[0112] In some embodiments, a reference sequence for a target sequence is derived from a nucleic acid that is naturally present in a eukaryote (e.g., a yeast, a plant, or a mammal) or an organelle thereof (e.g., a mitochondria or chloroplast). In some embodiments, a reference sequence for a target sequence is derived from a nucleic acid that is naturally present in an animal (e.g., a mammal, such as a human). In some embodiments, a reference sequence for a 43 / 112#14588836v1H0824.70442WO00target sequence is derived from a nucleic acid that is naturally present in a plant or an organelle in a plant. In some embodiments, a reference sequence for a target sequence is derived from a nucleic acid that is naturally present in bacteria. In some embodiments, a reference sequence for a target sequence is derived from a nucleic acid that is naturally present in archaea. In some embodiments, a reference sequence for a target sequence is derived from a nucleic acid that is naturally present in a virus.
[0113] In some embodiments, a target sequence comprises the sequence of a chromosome or a portion thereof. In some embodiments, the reference sequence for the target sequence is a chromosome that is naturally present in a eukaryote (e.g., a yeast, a plant, or a mammal) or an organelle thereof (e.g., a mitochondria or chloroplast). In some embodiments, the chromosome is naturally present in a genome of an animal (e.g., a mammal, such as a human). In some embodiments, the reference sequence for the target sequence is a chromosome that is naturally present in a genome of a plant or an organelle in a plant. In some embodiments, the reference sequence for the target sequence is a nucleic acid that is naturally present in a bacterial genome. In some embodiments, the reference sequence for the target sequence is naturally present that is naturally present in an archaea genome. In some embodiments, the reference sequence for the target sequence is a nucleic acid that is naturally present in a viral genome.
[0114] In some embodiments, a method comprises isolating a nucleic acid comprising a target sequence. In some embodiments, isolating the nucleic acid comprises subjecting one or more surfaces comprising a binding molecule to conditions that inhibit binding to one or more anchors. In some embodiments, isolating the nucleic acid comprises cleaving a ligated nucleic acid generated in a final ligation step or a cleaved nucleic acid generated in a final cleavage step with one or more nucleases (e.g., one or more nucleases that cleave sites located closest to the anchors). In some embodiments, a nucleic acid comprising a target sequence is purified following isolation.
[0115] In some embodiments, a nucleic acid comprising a target sequence is introduced into a cell. In some embodiments, a cell comprising a nucleic acid comprising a target sequence is an animal cell. In some embodiments, a cell comprising a nucleic acid comprising a target sequence is a mammalian cell. In some embodiments, a cell comprising a nucleic acid comprising a target sequence is a human cell. In some embodiments, a cell comprising a nucleic acid comprising a target sequence is a plant cell. In some embodiments, a cell comprising a nucleic acid comprising a target sequence is a bacterium. In some embodiments, a cell comprising a nucleic acid comprising a target sequence is an archaeon.44 / 112#14588836v1H0824.70442WO00
[0116] However, in some embodiments a nucleic acid comprising a target sequence can be located in a virus, such as a virus utilized for delivering an exogenous nucleic acid to a cell (e.g., adeno-associated virus or lentivirus), such as a mammalian cell (e.g., a human cell, such as a human patient cell).Modifications of Target Sequences and Nucleic Acids Thereof
[0117] In some embodiments, the nucleic acid is subjected to the one or more conditions and / or one or more additional steps prior to performing a final ligation step. In some embodiments, a nucleic acid comprising a target sequence is subjected to one or more conditions and / or one or more additional steps after isolation and / or before being introduced into a cell. In some embodiments, the one or more conditions and / or one or more additional steps comprise contacting the nucleic acid with one or more agents, such as chemical agents or proteins (e.g., proteins involved in chromatin formation and / or enzymes, such as enzymes that chemically modify nucleic acids). In some embodiments, the one or more conditions and / or one or more additional steps comprise nucleic acid sequencing. In some embodiments, the one or more conditions and / or one or more additional steps comprise increasing the copy number of the nucleic acid. In some embodiments, the one or more conditions and / or one or more additional steps comprise recombinant engineering (e.g., wherein the nucleic acid is engineered to produce a vector, such as a plasmid).Chemical Modification of Target Sequences
[0118] In some embodiments, a nucleic acid comprising a target sequence is subjected to one or more conditions (e.g., contacted with one or more enzymes and / or one or more agents, such as agents comprising a click-chemistry handle) following a ligation step and / or following a cleavage step that generate one or more chemical modifications located in the nucleic acid. The one or more chemical modifications can be generated at a purine residue (e.g., adenine, guanine, xanthine, caffeine, uric acid, isoguanine, theobromine, theophylline, or hypoxanthine) and / or a pyrimidine residue (e.g., thymine, cytosine, uracil, fluorouracil, barcituric acid, or orotic acid). Examples of chemically modified base moieties include, but are not limited to, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-hyrdoxymethylcytosine, 5-(carboxyhydroxylmethyl) uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1 -methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N4-methylcytosine, N6- 45 / 112#14588836v1H0824.70442WO00methyladenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosine, 5’-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methylester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl) uracil, a thio-guanine, and 2,6-diaminopurine. Examples of chemically modified sugar moieties include, but are not limited to, sugar moieties comprising a modification at a 2' carbon, such as 2'-O-alkyl (including 2'-O-methyl and 2'-O-ethyl), i.e., 2'-alkoxy, 2'-amino, 2'-S-alkyl, 2'-halo (including 2'-fluoro), 2'-2-0-methoxyethoxy, 2'-allyloxy (-OCH2CH=CH2), 2'-propargyl, 2'-propyl, 2'-ethynyl, 2'-ethenyl, 2'-propenyl, and 2'-cyano. Other 2' modifications include substitutions of a 2'-OH group with H, OR, R, F, Cl, Br, I, SH, SR, NH, NHR, NR, COOR, or OR, wherein “R” is a substituted or unsubstituted aliphatic group or alkyl group. Examples of chemically modified internucleotide linkages include, but are not limited to, substitution of an oxygen atom with a sulfur atom as well as linkages comprising phosphorothioate, borano-phosphate, and alkyl phosphonate acid. This includes single-and double- stranded molecules, i.e., DNA-DNA, DNA-RNA and RNA-RNA hybrids, as well as protein nucleic acids or peptide nucleic acids (PNAs) formed by conjugating bases to an amino acid backbone. Chemically modified nucleic acids can also be modified to comprise carbohydrate or lipids. Further examples of chemical modifications include Romesburg modifications (e.g., also known as third base pairing X-Y, artificial base pairings) (see, e.g., Malyshev et al. (2014). Nature, 509: 385-388, which is incorporated by reference for its descriptions of Romesburg modifications and nucleic acid engineering involving Romesburg modifications).
[0119] In some embodiments, a nucleic acid (e.g., a nucleic acid comprising a target sequence) is contacted with one or more methyltransferases, such as a DNA methyltransferase or an RNA methyltransferase. In some embodiments, a methyltransferase is contacted with a nucleic acid to methylate one or more nucleotide positions located in a target sequence. One or more methyltransferases can be utilized to methylate any nucleotide position or combination of nucleotide positions in a nucleic acid, such as a purine position(s) (e.g., guanine (G) or adenine (A)) and / or a pyrimidine position(s) (e.g., cytosine (C), thymine (T), or uracil (U)). In some embodiments, a methyltransferase is contacted with a nucleic acid comprising a target sequence to methylate a guanine (G), an adenine (A), a cytosine (C), and / or a thymine (T) located in a sequence comprising AA, AC, AG, AT, CA, CC, CG, CT, GA, GC, GG, GT, TA, TC, TG, or TT. In some embodiments, a methyltransferase is contacted with a nucleic acid comprising a target sequence to methylate a guanine (G) located in a sequence comprising AG, CG, GA, GC,46 / 112#14588836v1H0824.70442WO00GG, GT, or TG. In some embodiments, a methyltransferase is contacted with a nucleic acid comprising a target sequence to methylate an adenine (A) located in a sequence comprising AA, AC, AG, AT, CA, GA, or TA. In some embodiments, a methyltransferase is contacted with a nucleic acid comprising a target sequence to methylate a cytosine (C) located in a sequence comprising AC, CA, CC, CG, CT, GC, or TC. In some embodiments, a methyltransferase is contacted with a nucleic acid comprising a target sequence to methylate a thymine (T) located in a sequence comprising AT, CT, GT, TA, TC, TG, or TT. In some embodiments, a methylated nucleobase is a methylated adenine (e.g., N6-methyladenine (6mA), N1 -methyladenine (1mA), 2'-O-methyladenine (2'-0-mA), or N7 -methyladenine (7mA)), a methylated cytosine (e.g., 5-methylcytosine (5mC), N4-methylcytosine (4mC), or N3-methylcytosine (3mC)), or a methylated guanine (e.g., N7-methylguanine (7mG), N2-methylguanine (2mG), N1-methylguanine (1mG), or 2'-O-methylguanine (2'-0-mG)). In some embodiments, allowing for the creation of methylation or modified nucleic acid patterns. This capability enables the synthesis of nucleic acids with precise methylations or modifications, without the need for postsynthetic methylation reactions. This allows expanding known methylation patterns.
[0120] For example, DNA methylation is an epigenetic mark that regulates various cellular functions, including regulating gene expression patterns. The human genome harbors approximately 30 million methylation sites. In some embodiments, nucleotide positions in a nucleic acid (e.g., nucleotide positions located in a target sequence) are methylated to replicate the epigenetic landscape during nucleic acid production, such as during production of a chromosome or a portion thereof. In some embodiments, nucleotide positions comprising methylated nucleobases (e.g., N6-methyladenine (m6A) and 5-methylcytosine (5mC)) are generated in a target sequence based on epigenetic modifications present a human methylome.
[0121] As a further example, in some embodiments, nucleotide positions in a nucleic acid (e.g., nucleotide positions located in a target sequence) are methylated to prevent unintended cleavage after multiple ligations. In some embodiments, nucleotide positions comprising methylated nucleobases (e.g., N6-methyladenine (m6A)) are generated in a target sequence to reduce or prevent cleavage by nucleases that produce overhangs utilized in ligation (e.g., to block recognition of the sequence GGaTC by Alwl and / or recognition of the sequence CCaTC by BccI, wherein the position indicated by lower case “a” is modified to comprise N6-methyladenine (m6A)). In some embodiments, a methyltransferase is contacted with a nucleic acid comprising a target sequence to methylate nuclease recognition sequences.
[0122] In some embodiments, a nucleic acid (e.g., a nucleic acid comprising a target sequence) is subjected to one or more conditions that demethylates one or more positions 47 / 112#14588836v1H0824.70442WO00located in the nucleic acid. For example, in some embodiments, a nucleic acid can be subjected to one or more conditions that demethylates one or more positions located in a target sequence, such as wherein the one or more conditions comprise contacting the nucleic acid with one or more de-methylases. In some embodiments, a demethylase removes a methyl group from a methylated adenine, a methylated cytosine, or a methylated guanine, such as a methylated adenine, methylated cytosine, or methylated guanine described herein.Compaction, Condensation, Circularization, and Supercoiling
[0123] In some embodiments, a method comprises compacting, condensing, circularizing, and / or supercoiling a nucleic acid comprising a target sequence. In some embodiments, contacting, condensing, circularizing, and / or supercoiling a nucleic acid comprises contacting the nucleic acid with one or more proteins (e.g., a gyrase or gyrase-like enzyme, a Cas protein or a variant thereof, a ligase, a nuclease, and / or a recombinase) and / or one or more nucleic acid condensation agents (e.g., a protamine, a histone variant, a polyamine (e.g., spermidine, spermine), an engineered nucleoprotein, or a combination thereof ). In some embodiments, nucleic acids are compacted, condensed, circularized, and / or supercoiled to generate chromosome like- structures (e.g., wherein the nucleic acids comprise synthetic chromosomes or genome-scale DNA constructs), which are suitable for downstream biological or biophysical applications (e.g., for use in a method described herein, such as a method of introducing the nucleic acids into a cell or a subject). In some embodiments, nucleic acids are compacted, condensed, circularized, and / or supercoiled to generate artificially compacted nuclei, which can serve as functional substitutes for native nuclei in in vitro or ex vivo systems. In some embodiments, compacted, condensed, circularized, and / or supercoiled nucleic acids are directly used in somatic cell nuclear transfer (SCNT)-like operations (or other high-throughput operations, such as transformation), wherein the compacted synthetic chromosome or genome is introduced into an enucleated oocyte, zygote, or recipient cell to reconstitute a viable or semisynthetic cell-like entity. Accordingly, contacting, condensing, circularizing, and / or supercoiling nucleic acids as described herein can be used to generate artificial nuclei that are physically and functionally compatible with cellular reconstitution and nuclear transplantation workflows.
[0124] Circularization can be accomplished by enzymatic ligation (e.g., T4 DNA ligase), by restriction-enzyme-based strategies, such as using Sbfl or PflMI to generate compatible sticky ends, or via homologous recombination.
[0125] In some embodiments, compacting, condensing, circularizing, and / or supercoiling the nucleic acid comprises contacting the nucleic acid with one or more48 / 112#14588836v1H0824.70442WO00recombinases. Examples of recombination events generated by recombinases include an inversion, an excision, a translocation, an integration, a duplication, a conversion, and a circularization. Recombinases can be provided as purified enzymes, as in-vitro translated proteins, or expressed in situ within a host cell or cell-free system.
[0126] Site-specific recombinases (SSRs) are named after the amino acid residue that forms a covalent protein-DNA linkage in the reaction intermediate, with two major families, the tyrosine recombinases and the serine recombinases, allowing for integration, excision, and inversion of defined DNA segments.
[0127] Serine recombinases (SRs) are unidirectional, making them highly efficient. Serine recombinases function via first specifically binding two ~ 30 - 50 bp DNA attachment sites (attB and attP), each as a homodimer. Once SR-bound att sites are in proximity, tetramerization of SR protomers occurs (FIG. 7). At this point, the complex may initiate concerted DNA cleavage, yielding i) phosphoserine bonds between a DNA strand and each protomer’s catalytic serine and ii) di-nucleotide overhangs at the centers of both att sites.Cleaved DNA permits subunit rotation, where one pair of protomers rotates an equivalent of 180 degrees around the other. If the central di-nucleotides across att sites match, DNA re-ligation can occur, and recombinant products form. Due to asymmetry in the product sites (attL and attR), an LSR DNA-binding domain is positioned such that the reverse reaction (yielding attB and attP again) can only proceed with a recombinase directionality factor. In some embodiments, a recombinase is a serine recombinase.
[0128] Tyrosine recombinases (YRs), are well characterized and can also be engineered. Compared to SRs, YRs yield the same end result for recombining two specific ~20 - 30 bp recognition sites, though they have a distinct phylogeny and reaction mechanism. After YR tetramerization of one or two unique subunits joins two target sites flanking a 6-8 bp crossover site, the catalytic tyrosines sequentially nick pairs of single DNA strands at the crossover site boundaries (FIG. 7). Unlike SRs, natural YR recombination sites are functionally symmetric, allowing excision and integration activity. In some embodiments, a recombinase is a tyrosine recombinase.
[0129] SRs and YRs can be used to synthesize megabase (Mb) DNA fragments in vitro due to their shared characteristics of multiplexability, specificity, and efficiency in recombining large. SSRs can perform multiplex recombination by exploiting orthogonality between i) orthologs and their cognate recognition sites and / or ii) recognition sites of the same ortholog with differing core sequences. Demonstrating multiplexing, SSTRA (site-specific recombination-based tandem assembly) and SIRA (serine integrase recombinational assembly)49 / 112#14588836v1H0824.70442WO00systems use linearized DNA fragments flanked with attB and attP sites, mutated to recombine orthogonally for assembly of multiple biosynthetics. Both SRs and YRs have been used for inversions, deletions, and insertions in cellulo up to the Mb scale, underscoring that SSR function is likely invariant to size of the DNA substrates alone. When applying SSRs to in vitro assembly of large synthetic DNA, various methods can be exploited to determine SSR specificity profiles. These profiles can then be used to synthesize large DNA containing preferred recombination sites at desired loci while excluding pseudosite sequences to minimize off-target recombination.
[0130] In some embodiments, a recombinase is a CRE recombinase.
[0131] In some embodiments, a recombinase is Bbxl, PhiC31, PaOl, Pa03, Ec03, Nm60, LI Int, Cre R32M, YR1, YR2, YR11, Xerl, a homolog of Xerl (e.g., eXerl, eXer2, or eXer3), Dn29, Si74, Bm99, Bt24, Cbl6, Fm04, uCb64, Ec03, Nm60, LI Int, or an engineered variant thereof with modified attP / attB recognition specificities.
[0132] In some embodiments, compacting, condensing, circularizing, and / or supercoiling a nucleic acid comprising a target sequence comprises contacting the nucleic acid with one or more DNA-binding molecules and one or more recombinases. For example, the one or more DNA-binding molecules can comprise a DNA-binding protein, or a fusion protein thereof, and / or one or more reverse complementary nucleic acids (e.g., single- stranded polynucleotides, such as single- stranded DNA or single- stranded RNA). In some embodiments, the one or more DNA-binding molecules bind to a sequence comprising a binding site for a recombinase. In some embodiments, the one or more DNA-binding molecules are contacted with the nucleic acid comprising the target sequence prior to contacting the nucleic acid with the one or more recombinases, such as to prevent or reduce binding of the one or more recombinases to certain recombinase binding sites located in the nucleic acid comprising the target sequences.
[0133] For example, in some embodiments, compacting, condensing, circularizing, and / or supercoiling the nucleic acid comprises contacting the nucleic acid with a first DNA binding molecule, contacting the nucleic acid with a recombinase that generates a recombination event at one or more recombinase binding sites, contacting the nucleic acid with a second DNA-binding molecule (e.g., a single- stranded polynucleotide), and contacting the nucleic acid with a recombinase that generates a recombination event at one or more other recombinase binding sites. In some embodiments, the nucleic acid comprising the target sequence is subjected to one or more conditions that inhibit binding to the one or more DNA-binding molecules before, during, and / or after contacting the nucleic acid with the recombinase, such as to remove the 50 / 112#14588836v1H0824.70442WO00DNA-binding molecules from binding sites for the recombinase. In some embodiments, the one or more conditions that inhibit binding to the one or more DNA-binding molecules comprise one or more washes (e.g., washes that dilute out the reverse complementary nucleic acids) and / or contacting the nucleic acid comprising the target sequence with one or more proteins (e.g., histone proteins) that competitively bind with the one or more DNA-binding molecules.
[0134] In some embodiments, a method comprises compacting, condensing, supercoiling, and / or circularizing a nucleic acid comprising a target sequence, wherein the nucleic acid is contacted with a recombinase and one or more DNA-binding comprising a guide RNA (gRNA) and a Cas protein.
[0135] A gRNA typically comprises a sequence that hybridizes with a sequence of interest (e.g., a sequence in an acceptor nucleic acid, a donor nucleic acid, a cleavage product, a cleavage fragment, a ligated nucleic acid, or a target sequence) and a sequence that binds to a Cas protein, which may be referred to as a spacer sequence and a tracrRNA, respectively. In some embodiments, a gRNA comprises at least 10 (e.g., 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35) nucleotides (e.g., wherein the at least 10 nucleotides are contiguous nucleotides) that hybridizes with the sequence of interest. In some embodiments, the at least 10 nucleotides are 100% identical to the reverse complement of the sequence of interest.
[0136] In some embodiments, the gRNA comprises one or more naturally occurring nucleotides and / or one or more chemically modified nucleotides. The one or more chemical modifications can be generated at a purine residue (e.g., adenine, guanine, xanthine, caffeine, uric acid, isoguanine, theobromine, theophylline, or hypoxanthine) and / or a pyrimidine residue (e.g., thymine, cytosine, uracil, fluorouracil, barcituric acid, or orotic acid). Examples of chemically modified base moieties include, but are not limited to, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-hyrdoxymethylcytosine, 5-(carboxyhydroxylmethyl) uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1 -methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N4-methylcytosine, N6-methyladenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosine, 5 '-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methylester, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)51 / 112#14588836v1H0824.70442WO00uracil, a thio-guanine, and 2,6-diaminopurine. Examples of chemically modified sugar moieties include, but are not limited to, sugar moieties comprising a modification at a 2' carbon, such as 2 -O-alkyl (including 2 -O-methyl and 2 -O-ethyl), i.e., 2 -alkoxy, 2 -amino, 2 -S-alkyl, 2 -halo (including 2'-fluoro), 2 -2-O-methoxyethoxy, 2'-allyloxy (-OCH2CH=CH2), 2 -propargyl, 2'-propyl, 2'-ethynyl, 2 '-ethenyl, 2 '-propenyl, and 2 '-cyano. Other 2' modifications include substitutions of a 2'-OH group with H, OR, R, F, Cl, Br, I, SH, SR, NH, NHR, NR, COOR, or OR, wherein “R” is a substituted or unsubstituted aliphatic group or alkyl group. Examples of chemically modified internucleotide linkages include, but are not limited to, substitution of an oxygen atom with a sulfur atom as well as linkages comprising phosphorothioate, borano-phosphate, and alkyl phosphonate acid. This includes single-and double- stranded molecules, i.e., DNA-DNA, DNA-RNA and RNA-RNA hybrids, as well as protein nucleic acids or peptide nucleic acids (PNAs) formed by conjugating bases to an amino acid backbone. Chemically modified nucleic acids can also be modified to comprise carbohydrate or lipids. Further examples of chemical modifications include Romesburg modifications (e.g., also known as third base pairing X-Y, artificial base pairings) (see, e.g., Malyshev et al. (2014). Nature, 509: 385-388, which is incorporated by reference for its descriptions of Romesburg modifications and nucleic acid engineering involving Romesburg modifications).
[0137] Examples of Cas proteins include Cas9 (e.g., SpCas9, SaCas9, StCas9, NmCas9, CjCas9, and SpyCas9) and variants thereof, Casl2 (e.g., Casl2a (Cpfl), such as AsCasl2a, FnCasl2a, LbCasl2a, and PaCas 12a, or Cas 12b) and variants thereof, Cas 13 and variants thereof, Cas 14 and variants thereof, CasX, and CasY. In some embodiments, a Cas protein is a catalytically inactive Cas protein or a partially catalytically inactive protein. In some embodiments, a Cas protein binds to a protospacer adjacent motif (PAM) in the sequence of interest. In some embodiments, a Cas protein binds one or more of the following PAM sequences: NGG; NNAGAAW; NNGRRT; NNNNGATT; NVNDCCY; BRTTTTT; NR(A or G)TTTT; NNAAAR(G or A); N(N or A)G; NAAN; NAAAAY; NHDTCCA; NNNVRYM; NNNNRYAC; NAA; GNNNNCNNA; NNGTGA; NNNNGTA; NNGGG; NNNCAT;NNRHHHY; NRRNAT; NNNNCNAA; NNNNCMCA; NNNNCRAA; NNNNGMAA;NNNCC; NGGNG; NNNNCNDD; NYAAA; NRGNN; N(C or D)GGN(T or A or G or C)NN; NRTAW; N(C or K or A)AARC; NAAAG; NV(A or G or C)R(A or G)ACCN; NNGAC;NATGNT; N(T or V)NTAAW(A or T); NNGW(A or T)AY(T or C); NCAA(H(Y or A)B(Y or G); NH(T or C or A)AAAA; NNNATTT; NATAWN(A or T or S); NATARCH; B(T or G or C)GGD(A or T or G)TNN; N(G or T or M)GGAH(T or A or C)N(A or C or K)N; NRG; N(B or A)GG; NGGD(A or K)W(T or A); N(T or C or R)AGAN(A or K or C)NN; NGGD(A or T or 52 / 112#14588836v1H0824.70442WO00G)H(T or M); NGGDT; NGGD(A or T or G)GNN; NNGTAM(A or C)Y; NNGH(W or C)AAA; NTGAR(G or A)N(A or Y or G)N(Y or R); NNGAAAN; NNGAD; NHARMC; NNAAAG; NHGYNAN(A or B); NNAGAAA; NHAAAAA; NH(T or M)AAAAA; NHGYRAA;NNAAACN; NN(H or G)D(A or K)GGDN(A or B); NNNNCTA; NNNNCVGAA;NNNNGYAA; NNNNATN(W or S)ANN; NNWHR(G or A)TA(not G)AA; YHHNGTH;NNNNCDAANN; NNNNCTAA; N(C or D)NNTCCN; NNNNCCAA; NAGRGN(T or V)N(T or C); NNAH(T or M)ACN; CN(C or W or G)AV(A or S)GAC; NAR(G or A)H(W or C)H(A or T or C)GN(C or T or R); NAGNGC; NATCCTN; NGTGANN; HGCNGCR; NAR(A or G)W(T or A)AC; N(C or D)M(A or C)RN(A or B)AY(C or T); NNNCAC; BGGGTCD; NNRRCC; NRRNTT; KARDAT; BRRTTTW; NARNCCN; NAR(A or G)TC; NAAN(A or T or S)RCN; HHAAATD; NNNNGNA; TTV; TTTV; YYV; KKYV; TTTM; TTYV; TTTN; TTTTA; TTN; BTTV; YTV; YTN; NYTV; DTTD; ATTN; RTTNT; HATT; ATTW; RTTN; TVT; TG; TN; TR; TA; TTCN; TTAT; TTTR; TTR; YTTR; YTTN; CTT; TTC; CCD; RTR; VTTR; TBN; VTTN; NGTT; CGTT; AGG; CGG; GTT; or RGTG, wherein: “N” is any nucleotide or base; “W” is adenine (A) or thymine (T); “R” is A or guanine (G); “V” is A, cytosine (C), or G; “Y” is C or T; “M” is C or A; “K” is G or T; “D” is A, G, or T; “S” is G or C; “B” is C, G, or T; and “H” is A, C, or T.
[0138] In some embodiments, Cas protein (e.g., dCas9) or analogous programmable DNA-binding proteins are used to position or tether nucleic acids without introducing strand cleavage. As shown in FIGs. 8A-8B, in some embodiments, a single RNA strand with two sgRNAs sequences targeting two distal nucleic acid sequences is utilized and permits the formation of dCas9-sgRNA-DNA complexes that physically bring those nucleic acid segments into close proximity.
[0139] In some embodiments, a Cas protein is a bivalent Cas protein (e.g., a fusion protein, such as a protein comprising two copies of the same Cas proteins or a protein comprising two different Cas proteins, for example wherein the Cas proteins are connected directly or connected by a linker, such as a flexible linker, including, but not limited to, a serineglycine linker). In some embodiments, wherein a bivalent Cas protein is utilized, each of the Cas proteins bind to different gRNAs that bind different sites on the nucleic acid comprising the target sequence. In some embodiments, a bivalent Cas protein is utilized to bring sequences in the nucleic acid comprising the target sequence into close physical proximity, thereby increasing the efficiency of recombination. An example of a method for reconstituted pronuclei assembly utilizing one or more recombinases and bivalent Cas proteins is shown in FIG. 9. The recombinase sites on the nucleic acid comprising the target sequence can be generated such that 53 / 112#14588836v1H0824.70442WO00the sites are not occluded by the Cas protein. In some embodiments, a TAL protein or zinc finger protein, or a fusion protein thereof, can be utilized as an alternative to a Cas protein.
[0140] In some embodiments, compacting, condensing, supercoiling, and / or circularizing a nucleic acid comprises contacting the nucleic acid with one or more (e.g., one, two, three, four, or more than four) nucleic acid condensation agents. In some embodiments, the one or more nucleic acid condensation agents can be used to fine-tune chromatin density, accessibility, and stability of nucleic acids. The resulting nucleic acids can be further stabilized by cross-linking, ionic tuning, or encapsulation within lipidic or polymeric vesicles to enhance mechanical integrity during transfer. In some embodiments, the one or more nucleic acid condensation agents comprise a protamine, a histone variant, a poly amine (e.g., spermidine, spermine), an engineered nucleoprotein, or a combination thereof. In some embodiments, the one or more nucleic acid condensation agents comprise a protamine. Protamines are small, highly basic proteins naturally involved in sperm chromatin condensation and inducing tight supercoiling and nucleoprotein complex formation. When applied to synthetic or reconstituted chromosomes, protamine-mediated condensation yields densely packed DNA-protein assemblies that mimic the structural and mechanical properties of a natural nucleus or pronucleus.Surfaces
[0141] Surfaces can be utilized in the compositions (e.g., systems) and methods (e.g., nucleic acid production methods) described in the application. A surface will typically comprise one or more binding molecules that can be contacted with and / or can be bound to one or more nucleic acids comprising an anchor.
[0142] In some embodiments, a surface is generated by contacting a nucleic acid comprising two different anchors located at the terminal ends of the nucleic acid, and that flank two different nuclease cleavage sites, with a surface comprising a binding molecule, wherein one of the anchors can bind to the binding molecule located on the surface. In some embodiments, a surface is generated by contacting a nucleic acid comprising one binding molecule at a terminal end, two nuclease cleavage sites, and an overhang at the other terminal end, with a surface comprising a binding molecule, wherein the anchor can bind to the binding molecule located on the surface. The binding molecules and nucleic acids can include those involved in ligation and cleavage steps as described in the application (see, e.g., the description of anchors and binding molecules comprising small molecules, click-chemistry handles, peptides, proteins, or polynucleotide provided herein).54 / 112#14588836v1H0824.70442WO00
[0143] Surfaces can include glass surfaces, metal surfaces, plastic surfaces, ceramic surfaces, and other polymeric surfaces. In some embodiments, a surface is glass. Functionalized surfaces and surface functionalization services, including those that utilize and / or are compatible with click-chemistry handles, are commercially available (see, e.g., services and products available through SuSoS Surface Technology).
[0144] Liquid droplets can be loaded onto surfaces to provide components utilized in methods described herein, including nucleic acids, enzymes (e.g., enzymes that bind to nucleic acids, such as nucleases, ligases, recombinases, methylases, and demethylases), binding molecules, agents that inhibit binding between anchors and the binding molecules, compositions utilized for washing surfaces, crowding agents (e.g., a crowding agent described herein, such as PEG), and glycerol. For example, in some embodiments, a liquid droplet comprises an ink. In some embodiments, a method comprises utilizing a plurality of unique inks (e.g., wherein each unique ink in the plurality of unique inks are provided in a liquid droplet) including, but not limited to, one or more inks comprising nucleic acids that are not chemically modified (e.g., one or more inks comprising an acceptor nucleic acid and one or more inks comprising a donor nucleic acid), one or more inks for chemically modifying a nucleic acid (e.g., one or more inks for methylation of a nucleic acid, such as one or more inks comprising a methyltransferase), one or more inks for cleaving a nucleic acid (e.g., one or more inks comprising a nuclease, such as a restriction enzyme), one or more inks for providing ligase (e.g., one or more inks comprising ligase, a crowding agent, such as polyethylene glycol, and / or glycerol), and / or one or more inks for inhibiting binding of an anchor to a binding molecule (e.g., one or more inks comprising an agent that binds to a binding molecule, such as a nucleic acid including, but not limited to, a nucleic acid utilized for toehold displacement). In some embodiments, a method comprises utilizing 16 inks for non-methylated DNA synthesis (e.g., inks comprising acceptor and inks comprising donor nucleic acids). In some embodiments, a method comprises utilizing 12 inks for introducing methylation patterns (e.g., inks comprising methyltransferases and / or inks comprising methylated nucleic acids). In some embodiments, a method comprises utilizing 4 inks for cleaving nucleic acids (e.g., inks comprising a restriction enzyme). In some embodiments, a method comprises utilizing 1 ink for providing a ligase (e.g., an ink comprising ligase, a crowding agent, such as polyethylene glycol, and / or glycerol). In some embodiments, a method comprises utilizing 2 inks for inhibiting binding of an anchor to a binding molecule (e.g., an ink comprising an agent that binds to a binding molecule, such as a nucleic acid including, but not limited to, a nucleic acid utilized for toehold displacement). In some embodiments, a method comprises utilizing 35 unique inks, which includes 16 inks for non- 55 / 112#14588836v1H0824.70442WO00methylated DNA synthesis (e.g., inks comprising acceptor and inks comprising donor nucleic acids), 12 inks for introducing methylation patterns (e.g., inks comprising methyltransferases and / or inks comprising methylated nucleic acids), 4 inks for cleaving nucleic acids (e.g., inks comprising a restriction enzyme), 1 ink for providing a ligase (e.g., an ink comprising ligase, a crowding agent, such as polyethylene glycol, and / or glycerol), and 2 inks for inhibiting binding of an anchor to a binding molecule (e.g., an ink comprising an agent that binds to a binding molecule, such as a nucleic acid including, but not limited to, a nucleic acid utilized for toehold displacement).
[0145] In some embodiments, liquid droplets are loaded on surfaces using automated pipetting. In some embodiments, liquid droplets are 10 fL or more in volume. In some embodiments, liquid droplets are 10 fL to 10 pL in volume (e.g., 10 fL-100 fL, 100 fL-1 pL, 1 pL-10 pL, 10 pL-100 pL, 100 pL-1 nL, 1 nL-10 nL, 10 nL-100 nL, 100 nL-1 pL, or 1 pL-10 pL in volume). In some embodiments, liquid droplets are 100 fL in volume. In some embodiments, liquid droplets are 1 pL in volume. In some embodiments, liquid droplets are 10 pL in volume. In some embodiments, liquid droplets are 100 pL in volume. In some embodiments, liquid droplets are 1 nL in volume. In some embodiments, liquid droplets are 10 nL in volume. In some embodiments, liquid droplets are 1 pL in volume. In some embodiments, liquid droplets are spaced 1 pm to 2 mm apart (e.g., 1 pm- 10 pm, 10 pm- 100 pm, 100 pm-1 mm, or 1 mm-2 mm apart). In some embodiments, liquid droplets are spaced 50 pm apart. In some embodiments, wherein liquid droplets are 100 fL in volume, the liquid drops are spaced 5 pm to 10 pm (e.g., about 6 pm) apart. In some embodiments, wherein liquid droplets are 1 pL in volume, the liquid drops are spaced 10 pm to 20 pm (e.g., about 13 pm) apart. In some embodiments, wherein liquid droplets are 10 pL in volume, the liquid drops are spaced 20 pm to 40 pm (e.g., about 27 pm) apart. In some embodiments, wherein liquid droplets are 100 pL in volume, the liquid drops are spaced 40 pm to 100 pm (e.g., about 58 pm) apart. In some embodiments, wherein liquid droplets are 1 nL in volume, the liquid drops are spaced 100 pm to 200 pm (e.g., about 120 pm) apart. In some embodiments, wherein liquid droplets are 10 nL in volume, the liquid drops are spaced 200 pm to 300 pm (e.g., about 266 pm) apart. In some embodiments, wherein liquid droplets are 1 pL in volume, the liquid drops are spaced 1 mm to 1.5 mm (e.g., about 1.24 mm) apart. In some embodiments, each liquid droplet on a surface (e.g., each liquid droplet comprising nucleic acids involved in ligation and / or cleavage steps) comprise an equal volume and / or are spaced an equal distance apart.
[0146] Liquid droplets can be loaded on a surface such that the droplets are organized into columns and rows, such as wherein the droplets are arranged in an array (e.g., wherein one 56 / 112#14588836v1H0824.70442WO00or more slides, chips, or wafers are utilized, each comprising an array or sub-array of the liquid droplets). The number of rows and columns of liquid droplets can be equal or non-equal. The number of droplets in each row and column can be equal or non-equal.
[0147] Various liquid droplet seeding geometries can be employed, including square, triangle, or hexagonal arrangements. In some embodiments, a symmetrical arrangement of liquid droplets is utilized (e.g., an array comprising 256 rows of liquid droplets and 256 columns of liquid droplets, or an array comprising 1,398 rows of liquid droplets and 1,398 columns of liquid droplets). In some embodiments, liquid droplets are loaded onto a surface in a hexagonal geometry having a central liquid droplet in order increase droplet seeding density. In some embodiments, wherein a hexagonal geometry is utilized, the number of liquid droplets within a given area can be increased by approximately 4 / (31 / 2) = 2.3 times relative to an array having an equal number of rows and columns and / or an equal number of liquid droplets in each row and column.
[0148] Generally, the number of liquid droplets utilized is related to how many cycles of ligations will be performed and the length of the nucleic acid product. The number of liquid droplets utilized can be limited by the size of the stage movement (e.g., the platform comprising the surface(s)), the size of the rows, the number of slides, the liquid droplet size, the linear dimension divided by the droplet distance (or other computational geometry techniques). For example, to produce a nucleic acid comprising 1024 base pairs, ten hierarchical steps and 32 x 32 liquid droplets can be utilized. In this example, a rectangular liquid droplet loading format can be utilized, wherein the liquid droplets are loaded on slides (or chips) having ~0.5 mm by 0.5 mm surface area (e.g., 416 pm edge length square surface). As further example, to produce a nucleic acid comprising 1 giga (G) base pairs (109base pairs), thirty hierarchical steps and 31,623 liquid droplets by liquid 31,623 droplets can be utilized. In this example, a rectangular liquid droplet loading format can be utilized, wherein the liquid droplets are loaded on slides (or chips) having -0.41 meter by 0.41 meter surface area (e.g., 411096 pm edge length square surface) (e.g., wherein the surface area can be achieved by using several -0.03142 meter square surface wafers or with six 10 cm radius wafers, this synthesis could be done on 6 wafers). As a further example, to produce a nucleic acid comprising 3.2 G base pairs (3.2 x 109base pairs), 32 hierarchical steps and approximately 17 wafers can be utilized.
[0149] In some embodiments, a surface is subjected to one or more conditions that, and / or the surface is configured to, control the spatial proximity of nucleic acids during one or more steps of nucleic acid production. In some embodiments, controlling the spatial proximity of nucleic acids involves reducing the diffusion distance or localized concentration of reactants 57 / 112#14588836v1H0824.70442WO00to promote more frequent productive encounters between components of a composition (e.g., nucleic acids and proteins, such as enzymes, including ligases, nucleases, recombinases, methylases, and demethylases), such as a liquid droplet.
[0150] For example, in some embodiments, a crowding agent (e.g., a crowding agent described herein) is contacted with a surface to reduce the effective reactive volume of a liquid droplet, such as wherein the liquid droplet comprises components for an enzymatic step of a method described herein (e.g., a ligation, a cleavage, or a recombination). In some embodiments, the crowding agent comprises Polysorbate 20 (IUPAC: Polyoxyethylene (20) sorbitan monolaurate; commercially known as “Tween 20”TM), polysorbate 80 (IUPAC:Polyoxyethylene (20) sorbitan monooleate; commercially known as “Tween 80” TM), ( 1,1, 3,3-Tetramethylbutyl)phenyl-polyethylene glycol, Polyethylene glycol tert- octylphenyl ether (commercially known as “Triton X-114”TM), sodium dodecylsulfate (SDS), deoxycholate sodium, (3-((3-cholamidopropyl) dimethylammonio)- 1 -propanesulfonate) (commonly known as CHAPS detergent), benzalkonium chloride, polyethylene glycol (PEG) (e.g., PEG 200, PEG 400, PEG 1000, PEG 1550, PEG 2000, PEG 3350, PEG 6000, PEG 8000, PEG 20000, or PEG 50000), Ficoll 70, Ficoll 400, serum albumin, and / or dextran (for example, but not limited to 50 k or 500 k). In some embodiments, the concentration of the crowding agent used in an enzymatic step is at least at least 1% w / v (e.g., 1-5% w / v, 10-15% w / v, 15-20% w / v, 20-25% w / v, 25-30% w / v, 30-35% w / v, 35-40% w / v, 40-45% w / v, or 45-50% w / v, 50-60% w / v, 60-70% w / v, or 70-80% w / v). In some embodiments, the crowding agent is PEG (e.g., wherein the concentration of PEG is between 10 mM PEG to 50 mM PEG). In some embodiments, the crowding agent is PEG (e.g., wherein the concentration of PEG is at least 10 mM, such as 10 mM-50 mM).
[0151] As a further example, in some embodiments, a surface is used to confine nucleic acids or to utilize a surface-immobilized system (e.g., during a ligation, a cleavage, or a recombination reaction).
[0152] In some embodiments, a surface comprises an immobilization system comprising a polynucleotide and / or a protein (e.g., wherein the polynucleotide and / or the protein comprises a binding molecule described herein) that binds to nucleic acids (e.g., that binds to acceptor nucleic acids, donor nucleic acids, cleaved nucleic acids, cleavage fragments, or ligated nucleic acids), or a molecule bound or attached thereto (e.g., a binding molecule or an anchor), and that uses the binding interaction to tether and / or localize nucleic acids within confined reaction environments on the surface.58 / 112#14588836v1H0824.70442WO00
[0153] In some embodiments, a surface comprises an immobilization system comprising a protein, wherein the protein comprises a Cas protein or an catalytically inactive variant thereof. In some embodiments, spatial proximity between nucleic acids are controlled using Cas (e.g., dCas9) or analogous DNA-binding proteins. In some embodiments, the Cas protein is bound to a guide RNA that hybridizes to a nucleic acid used in a step of a method described herein (e.g., a nucleic acid-binding protein, such as a Cas protein or an catalytically inactive variant thereof). The protein can be strategically programmed through guide RNAs or custom-designed DNA-recognition domains to bind predetermined genomic or synthetic nucleic acid sequences (e.g., target nucleic acids). In some embodiments, these protein-DNA interactions are strategically designed to tether or localize DNA molecules within confined reaction environments (e.g., environments that promote ligation, recombination, or hybridization), thereby enhancing the frequency and efficiency of enzymatic processes. In some embodiments, Cas protein (e.g., dCas9) or analogous programmable DNA-binding proteins are used to position or tether nucleic acids without introducing strand cleavage. As shown in FIGs. 8A-8B, in some embodiments, a single RNA strand with two sgRNAs sequences targeting two distal nucleic acid sequences is utilized and permits the formation of dCas9-sgRNA-DNA complexes that physically bring those nucleic acid segments into close proximity. In some embodiments, an immobilization system uses protein-mediated tethering to organize DNA substrates, promote spatial confinement, and facilitate reactions such as ligation, recombination, or higher-order genome assembly.
[0154] In some embodiments, a surface comprises an immobilization system comprising a toehold, wherein the toehold is employed for controlled capture, release, and directional assembly during nucleic acid production. In some embodiments, a surface is used in a toehold-mediated capture strategy, wherein polynucleotides are selectively hybridized and anchored to a surface (e.g., a functionalized surface) in a manner that permits precise, sequential, and design-directed synthesis of the intended construct. In some embodiments, this programmable toehold capture enables the controlled assembly of DNA modules according to a predefined sequence architecture, allowing the generation of large, structured nucleic acid molecules with chromosomal or organellar organization. In some embodiments, a toehold is employed to target hairpin structures during ligation steps. In some embodiments, a toehold is employed to target triplex-forming DNA sequences during recombination steps.
[0155] In some embodiments, surface immobilization systems can be applied to the synthesis of organelle genomes, chromosomes, and entire genomes.
[0156] In some embodiments, a surface comprises one or more molecular spacers. In some embodiments, the efficiency of enzymatic events (e.g., a ligation, a cleavage, or a 59 / 112#14588836v1H0824.70442WO00recombination reaction) on solid surfaces can be significantly improved through the introduction of molecular spacers between the substrate and the immobilized DNA constructs. Molecular spacers can physically separate an enzymatically active region (e.g., a region in a droplet comprising the components for a ligation, a cleavage, or a recombination reaction) from the surface, thereby reducing steric hindrance, improving enzyme accessibility, and enhancing hybridization and ligation kinetics within the reaction pocket. Spacer elements are positioned between the surface (e.g., chip, wafer, glass, slide, or other solid support) and the substrate. These spacers serve to extend and elevate the substrate away from the surface by a distance corresponding to the spacer length. By increasing the separation between the active nucleic acid region and the solid support, the spacer reduces steric hindrance, enhances molecular accessibility, and improves the efficiency of downstream reactions, such as hybridization, ligation, recombination, or strand displacement. A molecular spacer may be positioned at any location permitted by the underlying chemistry and design of the system. In some embodiments, spacers are positioned in separate regions (e.g., separate regions of a surface or separate regions of a plate comprising a plurality of surfaces) from where ligation or cleavage is performed. A molecular spacer may be attached directly to the nucleic acid, in between two nucleic acids, to a conjugated modification on the nucleic acid, or to the surface itself. In some embodiments, wherein the spacer is attached directly to the surface, a nucleic acid comprising a toehold is coupled to the distal end of the spacer. In some embodiments, wherein the spacer is incorporated within a nucleic acid, the nucleic acid itself can be attached directly to the surface (e.g., through a terminal DBCO-azide reaction or an amine-reactive silane), with the spacer providing additional distance, flexibility, and steric clearance from the underlying solid support.Accordingly, in some embodiments, spacers can function as physical separation elements, increasing the distance between the functional nucleic acid domain (e.g., a toehold) and the surface or neighboring components, thereby enhancing hybridization efficiency, enzyme accessibility, molecular mobility, and reaction performance. In some embodiments, the spacer comprises a flexible linker or polymeric moiety capable of maintaining the orientation and mobility of the attached nucleic acid. Exemplary spacers include, but are not limited to, polymers, polyethylene glycol (PEG) linkers, alkyl chains, or nucleotide-based flexible sequences, such as in poly tails, and combinations thereof. In some embodiments, a spacer is covalently linked to a binding molecule (e.g., a toehold sequence) through a click-chemistry handle (e.g., an NHS-ester to amine or maleimide-thiol reactions) or integrated directly during surface functionalization step (e.g., through azide-PEG- silane coatings prior to DBCO-azide coupling). For example, commercially available spacers (e.g., PEG-based linkers or other 60 / 112#14588836v1H0824.70442WO00nucleic-acid-compatible modifiers, such spacers available through Integrated DNA Technologies linkers, including C3 Spacer, Hexanediol, 1’, 2’ -Dideoxyribose (dSpacer), Photo-Cleavable Spacer, Spacer 9, and Spacer 18) can be installed within a nucleic acid sequence, at the terminus of an oligonucleotide, or between two nucleic acid segments. In some embodiments, via click chemistry an azide-functionalized surface is attached to DBCO moiety that is attached to a nucleic acid and a spacer (e.g., Spacer 18) is incorporated to the nucleic acid, followed by the nucleic acid region containing the toehold sequence (e.g., wherein the spacer acts as a physical separator between the surface and the functional toehold domain).
[0157] In some embodiments, a surface is subjected to silanization. In some embodiments, silanization is used to improve overall assembly fidelity, signal-to-noise ratio, and / or reproducibility across multiple synthesis cycles. In some embodiments, silanization is employed prior to or after contacting a surface with a binding molecule (e.g., prior to or after functionalizing a surface with a binding molecule, such as an azide) to minimize nonspecific adsorption of enzymes, nucleic acids, and other biomolecules. In some embodiments, silanization is used to form a self-assembled silane monolayer that provides a uniform chemical landscape and enhances the specificity of coupling between click chemistry handles (e.g., DBCO-azide coupling). Representative silanes suitable for this purpose include, but are not limited to 3-azidopropyl)triethoxysilane, and other alkoxy- or amino-silanes that introduce reactive handles while passivating non-reactive regions.
[0158] Surfaces can be configured to have a temperature, or a temperature range, that promotes one or more steps of a nucleic acid production method. The temperature or temperature range can be selected to be specific for the ligase utilized in the ligation step(s), the nuclease utilized in the cleavage step(s), and / or to promote or disrupt binding between anchors and binding molecules. The temperature or temperature range can be maintained across the surface for seconds (e.g., during a ligation step), minutes (e.g., during a cleavage step), or hours, depending on the step in nucleic acid production. The temperature or temperature on a surface can also be altered (e.g., cycled) throughout nucleic acid production depending on the step in nucleic acid production. In some embodiments, a surface has a temperature of between 4 °C and 95 °C (e.g., 4 °C to 15 °C, 15 °C to 20 °C, 20 °C to 25 °C, 25 °C to 30 °C, 30 °C to 35 °C, 35 °C to 40 °C, 40 °C to 45 °C, 45 °C to 65 °C, 65 °C to 75 °C, or 75 °C to 95 °C). In some embodiments, a surface has a temperature of between 15 °C and 45 °C (e.g., 16 °C, °C 17 °C, 18°C, 33 °C, 34 °C, 35 °C, 36 °C, 37 °C, 38 °C, 39 °C, 40 °C, 41 °C, 42 °C, 43 °C, 44 °C, or 45 °C). In some embodiments, a surface can undergo a temperature change from a first temperature 61 / 112#14588836v1H0824.70442WO00(e.g., a temperature utilized during a ligation step) to a second temperature (e.g., a temperature utilized during a cleavage step). In some embodiments, a surface can undergo a temperature change from a first temperature to a second temperature, wherein the first temperature and the second temperature are both between 4 °C and 95 °C (e.g., wherein the first temperature and the second temperature are both between 4 °C to 15 °C, 15 °C to 20 °C, 20 °C to 25 °C, 25 °C to 30 °C, 30 °C to 35 °C, 35 °C to 40 °C, 40 °C to 45 °C, 45 °C to 65 °C, 65 °C to 75 °C, or 75 °C to 95 °C). In some embodiments, a surface can undergo a temperature change from a first temperature to a second temperature that are both between 15 °C and 70 °C (e.g., wherein the first and second temperatures are 1 °C to 55 °C, 1 °C to 50 °C, 1 °C to 40 °C, 1 °C to 30 °C, 1 °C to 25 °C, 1 °C to 20 °C, 1 °C to 15 °C, 1 °C to 10 °C apart, or 1 to 5 °C apart). In some embodiments, a surface has a temperature between 18 °C to 25 °C during a ligation step (e.g., 18 °C, 19 °C, 20 °C, 21 °C, 22 °C, 23 °C, 24 °C, or 25 °C). In some embodiments, a surface has a temperature of 30 °C to 70 °C (e.g., a temperature of 37 °C or 65 °C) during a cleavage step. In some embodiments, a surface undergoes a temperature decrease or a temperature increase to reduce the activity of one or more enzymes involved in nucleic acid production (e.g., wherein the temperature is decreased to about 4 °C, or wherein the temperature is increased to about 80 °C, 90 °C, 100 °C, or more than 100 °C) (e.g., wherein the decreased temperature or the increased temperature is maintained for seconds or minutes).Preparing Nucleic Acids for Ligation
[0159] In some embodiments, a method comprises preparing nucleic acids ligated in an initial ligation step. An example of a method of preparing nucleic acids for ligation as described herein is provided in Example 2.
[0160] In some embodiments, preparing nucleic acids ligated in an initial ligation step comprises contacting one or more polynucleotides (e.g., one or more vectors, such as one or more plasmids) with one or more nucleases, thereby generating a plurality of nucleic acids, wherein each nucleic acid of the plurality comprises a unique portion of the target sequence which is selected from the group consisting of: AA, GA, CA, TA, AG, GG, CG, TG, AC, GC, CC, TC, AT, GT, CT, and TT, wherein the portion of the target sequence is flanked by a sequence comprising the cleavage site for the third nuclease and the cleavage site for the first nuclease, and a sequence comprising the cleavage site for the second nuclease and the cleavage site for the fourth nuclease.
[0161] In some embodiments, the one or more polynucleotides comprise a vector (e.g., a plasmid). Suitable vectors can range from low to high copy number vectors. In some62 / 112#14588836v1H0824.70442WO00embodiments, vectors are grown in antimutator strains with high copy numbers or with strategies to elevate plasmid copy number, such as inducible replication origins. Other strategies to decrease mutation rates during DNA replication can also be employed to enhance yield without increasing the risk of mutational errors. See Narayan et al. (2014). Applied and Environmental Microbiology, 80(23): 7154-7160; Fijalkowska & Schaaper (1993). Genetics, 134(4): 1039-1044; Pandey et al. (2025). Nucleic Acids Research, 53(19): gkafl035, which are each incorporated by reference herein for descriptions related to engineering nucleic acids and using bacteria to produce engineered nucleic acids. In some embodiments, strategies to enhance plasmid yield, including the use of inducible origins of replication or high-copy-number plasmids, can be employed. Typically, a unique host strain is generated for each construct. However, in some embodiments, pooled production strategies can also be implemented where appropriate.
[0162] In some embodiments, nucleic acids of the plurality are further cleaved with nucleases that generate the overhangs that are utilized in the initial ligation step. In some embodiments, the nucleases exonucleases (e.g., T5 Exonuclease, T7 Exonuclease, and Exonuclease III, which act on nicked or blunt end DNA) and / or endonucleases (e.g., nicking endonucleases, such as Nt. BspQI and / or Nt. BstNBI). In some embodiments, the nucleases comprise a blunt end-cutter. In some embodiments, the nucleases are used to generate singlestranded regions that, under controlled physicochemical conditions (e.g., temperature, ionic strength), promote the formation of secondary structures through intramolecular base pairing, which can be stabilized and sealed using a DNA ligase. In some embodiments, the nucleases are used to selectively degrade residual vector backbone and non-target DNA species. In some embodiments, nucleic acids of the plurality are purified following nuclease cleavage using standard nucleic acid purification techniques, including organic extraction (e.g., phenolchloroform), non-organic methods (e.g., salting out, silica-based columns), or physical separation methods (e.g., agarose gel electrophoresis, magnetic bead-based capture). In some embodiments, selective purification is achieved by tagging the nucleic acid for ligation or the vector backbone with affinity labels (e.g., nucleic acid, protein, or small molecule tags).
[0163] In some embodiments, nucleic acids of the plurality are produced by assembly in a modular fashion rather than as a single contiguous sequence. In some embodiments, nucleic acids of the plurality are generated as discrete single- stranded or double- stranded DNA fragments, each representing a portion of a nucleic acid comprising a target sequence. For example, in some embodiments, nucleic acids of the plurality are synthesized or cloned separately and subsequently joined through enzymatic ligation. In some embodiments, terminal 63 / 112#14588836v1H0824.70442WO00hairpin structures may be ligated to both ends of a double- stranded DNA fragment to generate a closed, topologically constrained molecule. In other embodiments, nucleic acids of the plurality are generated as discrete single- stranded or double- stranded DNA fragments and then extended by ligation to one or more additional nucleic acid segments to complete.
[0164] In some embodiments, nucleic acids of the plurality are further fused to anchors. In some embodiments, nucleic acids fused to anchors are loaded onto surfaces comprising binding molecules.
[0165] As an example, to ensure an aggregate error rate of less than IxlO-7, initial nucleic acid fragments can be cloned into vectors and then sequenced. Subsequently, specific amplicons corresponding to each 1-mer overhang are generated from the respective vectors. After achieving sufficient DNA quantities, end modifications are introduced to facilitate immobilization onto surfaces coated with binding molecules. Surfaces can be strategically set up with geometric layouts of spots, so there is no need to re-array across the subsequent 32 assembly steps.
[0166] However, in some embodiments, preparing nucleic acids ligated in an initial ligation step comprises using an alternative method. For example, alternative methods include using a one-pot reaction systems comprising some or all of the enzymes described herein or incorporation of intermediate purification steps between reaction stages, such as isolation of the desired DNA sequence from its backbone or removal of enzymatic components from the reaction mixture.Pharmaceutical Compositions, Methods of Administration, and Uses
[0167] In some aspects, the present application relates to pharmaceutical compositions, methods of administration to a subject, and compositions for use in administration to a subject (e.g., wherein the administration comprises performing one or more administrations subcutaneously, intraretinally, intraocularly, subretinally, intravitreally, parenterally, subcutaneously, intravenously, intracerebroventricularly, intramuscularly, intracranially, intrathecally, intraperitoneally, or by direct injection to one or more cells, tissues, or organs). In some embodiments, a composition that can be administered to a subject comprises a nucleic acid comprising a target sequence. In some embodiments, a composition that can be administered to a subject comprises one or more cells (e.g., a population of cells) comprising a nucleic acid comprising a target sequence (e.g., wherein the cells are allogeneic to the subject or wherein the cells are autologous to the subject). In some embodiments, a composition that can be administered to a subject comprises a nucleic acid comprising a target sequence, wherein the 64 / 112#14588836v1H0824.70442WO00target sequence comprises one or more sequences (e.g., one or more gene sequences and / or one or more non-coding sequences) that are therapeutic for a disease, disorder, or condition described herein. In some embodiments, a composition that can be administered to a subject comprises one or more cells (e.g., a population of cells, such as wherein the cells are allogeneic to the subject or wherein the cells are autologous to the subject) that comprise a nucleic acid comprising a target sequence, wherein the one or more cells and / or the target sequence is therapeutic for a disease, disorder, or condition described herein.
[0168] In some embodiments, a subject is a mammalian subject. In some embodiments, a mammalian subject is a human subject. In other embodiments, a mammalian subject is nonhuman, such as a mouse, a rat, or a non-human primate. In some embodiments, a subject (e.g., a subject in need thereof) is a subject that has, is suspected of having, or at risk of developing a disease, disorder, or condition. In some embodiments, the disease, disorder, or condition comprises a genetic disease, cancer, inflammatory disease or an inflammatory condition, autoimmune disease, spleen disease, lung disease, hematological disease, neurological disease, painful condition, psychiatric disorder, metabolic disorder, immune disorder, infection of a pathogen, a kidney disease, cardiovascular disease, pancreatic disease, intestinal disease, retinal disease, neuromuscular disease, musculoskeletal disease, lysosomal storage disease, or other disease, or any combination thereof. In some embodiments, a composition (e.g., a nucleic acid) is administered to a subject to treat a disease, disorder, or condition.
[0169] In some embodiments, a composition further comprises a pharmaceutical excipient. Pharmaceutically acceptable excipients (excipients) are substances other than a therapeutic agent that are intentionally included in a delivery system. In some embodiments, excipients do not exert or are not intended to exert a therapeutic effect. In some embodiments, excipients may act to a) aid in processing of the therapeutic agent during manufacture, b) protect, support or enhance stability, bioavailability or patient acceptability of the API, c) assist in product identification, and / or d) enhance any other attribute of the overall safety, effectiveness, or delivery of the therapeutic agent during storage or use. In some embodiments, a pharmaceutically acceptable excipient may be an inert substance. In some embodiments, a pharmaceutically acceptable excipient may not be an inert substance. Excipients include, but are not limited to, absorption enhancers, anti- adherents, anti-foaming agents, anti-oxidants, binders, buffering agents, carriers, coating agents, colors, delivery enhancers, delivery polymers, dextran, dextrose, diluents, disintegrants, emulsifiers, extenders, fillers, flavors, glidants, humectants, lubricants, oils, polymers, preservatives, saline, salts, solvents, sugars, suspending agents,65 / 112#14588836v1H0824.70442WO00sustained release matrices, sweeteners, thickening agents, tonicity agents, vehicles, waterrepelling agents, and wetting agents.
[0170] In some embodiments, a composition further comprises additional components commonly found in pharmaceutical compositions. Such additional components can include, but are not limited to: anti-pruritic s, astringents, local anesthetics, or anti-inflammatory agents (e.g., antihistamine, diphenhydramine).
[0171] In some embodiments, a composition (e.g., a pharmaceutical composition) further comprises additional components commonly found in compositions for delivering nucleic acids to cells, such as nanoparticles (lipid nanoparticles, silver nanoparticles, gold nanoparticles, etc.) and / or vesicles.Kits
[0172] In some aspects, the present application relates to kits that can be utilized in a nucleic acid production method (e.g., kits comprising surfaces, nucleic acids, enzymes, crowding agents, or any combination thereof, as described herein) or that comprises a pharmaceutical composition (e.g., a compositions for use in administration to a subject). In some embodiments, the kits described herein may include one or more containers housing components for performing the methods described herein, and optionally instructions for use. In some embodiments, the components may be prepared sterilely, packaged in a syringe, and shipped refrigerated. Alternatively, in some embodiments, they may be housed in a vial or other container for storage. In some embodiments, a second container may have other components prepared sterilely. Alternatively, in some embodiments, the kits may include the active agents premixed and shipped in a vial, tube, or other container. In some embodiments, the kits may also include other components, depending on the specific application, for example, containers, cell media, salts, buffers, reagents, syringes, needles, a fabric, such as gauze, for applying or removing a disinfecting agent, disposable gloves, a support for the agents prior to administration, etc.
[0173] In some embodiments, any of the kits described herein may further comprise components provided in liquid form (e.g., in solution) or in solid form (e.g., a dry powder). In some embodiments, some of the components may be reconstitutable or otherwise processible (e.g., to an active form), for example, by the addition of a suitable solvent or other species (e.g., water), which may or may not be provided with the kit.
[0174] In some embodiments, a kit further comprises a set of instructions for carrying out the methods described herein. As used herein, “instructions” can define a component of 66 / 112#14588836v1H0824.70442WO00instruction and / or promotion, and typically involve written instructions on or associated with packaging of this disclosure. In some embodiments, instructions also can include any oral or electronic instructions provided in any manner such that a user will clearly recognize that the instructions are to be associated with the kit, for example, audiovisual (e.g., videotape, DVD, etc.), internet, and / or web-based communications, etc. In some embodiments, the written instructions may be in a form prescribed by a governmental agency regulating the manufacture, use, or sale of pharmaceuticals or biological products, which can also reflect approval by the agency of manufacture, use or sale for animal administration.
[0175] Additionally, in some embodiments, the kits may include other components depending on the specific application, as described herein. In some embodiments, the kits may have a variety of forms, such as a blister pouch, a shrink-wrapped pouch, a vacuum sealable pouch, a sealable thermoformed tray, or a similar pouch or tray form, with the accessories loosely packed within the pouch, one or more tubes, containers, a box, or a bag. In some embodiments, the kits may be sterilized after the accessories are added, thereby allowing the individual accessories in the container to be otherwise unwrapped. In some embodiments, the kits, or any of its components, can be sterilized using any appropriate sterilization techniques, such as radiation sterilization, heat sterilization, or other sterilization methods known in the art.EXAMPLESExample 1: Hierarchical DNA Synthesis using Double Selection Ligation (DSL)
[0176] The present example describes a genome ligation method named Double Section Ligation (DSL) which uses ligation-based synthesis to achieve a megabase scale rapidly and efficiently (FIGs. 1A-1D and FIGs. 2A-2B) with the need for a DNA polymerase. Expediency is achieved by eliminating one-by-one additions, and, instead, the size of constructed DNA increases by a factor of 2 at each step, needing only 28 steps to reach the genome scale for either a plant or animal. Rather than starting with dNTP monomers, the process starts with doublestrand DNA with 1 base single-strand overhangs that are capable of being immobilized on one end and ligated on the other end, referred to from here onwards as 1-mer adapters (see FIGs. 2A-2B and FIG. 4). Though each adapter adds only 1 bp initially, subsequent steps exhibit a dramatic increase in DNA length. This exponential growth (2n) occurs because each step incorporates a same- sized DNA fragment from aggregate contributions of the previous top and bottom surfaces. Initially, adapters are produced via clonal copies with the highest fidelity and sequence accuracy. This approach utilizes only 16 unique 1-mer adapters (approximately 40 nucleotides in length) to achieve whole-genome synthesis (see FIGs. 2A-2B, FIG. 4, and FIGs.67 / 112#14588836v1H0824.70442WO0013A-13E). This genome ligation method may be referred to herein as a “hierarchical DNA synthesis” or a “hierarchical DNA assembly” method.
[0177] The hierarchical assembly described herein offers a significant advantage in reaction speed due to incorporating polyethylene glycol (PEG), a known molecular crowding agent proven to accelerate ligation reactions. In the system described herein, PEG promotes rapid ligation and substantial yield within 15 seconds. Further optimization, including enzyme concentration and reaction time, has demonstrated ligation efficiencies in excess of 90% within minutes. The dramatic reduction in synthesis time for large-scale genomes (28 steps) and in reaction time per step underscores the capabilities of this technology.
[0178] The DSL method described here achieves high-fidelity DNA synthesis by implementing two additional washing steps through reversible immobilization (FIGs. 2A-2B).This immobilization can be mediated by either nucleic acid-peptide or nucleic acid-nucleic acid interactions, for example, toehold-mediated strand displacement. Each cycle involves: a) one side of the newly synthesized DNA is temporarily immobilized at its adapter end via an immobilization agent (i.e., Toehold DNA), b) incoming DNA is ligated at the other end, creating a new one (1) base overhang for the next ligation cycle, c) the newly ligated DNA is immobilized via a reversible immobilization, and d) a restriction endonuclease cleaves the initially immobilized end of the DNA that effectively switches the immobilized side for continued synthesis (FIGs. 2A-2B). Switching of the immobilization ends filters for successfully ligated DNA fragments that have been extended. Un-extended DNA are released via the reversible immobilization technique and then washed away (FIGs. 3A-3B). This is the first selection in DSL, as washing away un-ligated DNA strands will mitigate missing base errors. After washing, DNA fragments on a spot are prepared for transfer to a new spot on the opposing wafer; only successfully flipped strands will be able to transfer, constituting the second selection. Consequently, the DSL approach enables whole-genome synthesis while adhering to acceptable mutation rate thresholds.
[0179] The DSL process prioritizes achieving the highest purity in the final product by actively mitigating potential errors. These errors could include unsuccessful ligation between two DNA molecules; incomplete digestion of a DNA molecule; difficulty releasing a fragment from the synthesis surface; and ineffective washing of an immobilized agent (FIGs. 3A-3B).
[0180] To address these potential issues, the strategy described herein incorporates error removal mechanisms within each ligation step wash. If an error persists on one surface after washing, subsequent wash step removes it (FIGs. 2A-2B and FIGs. 3A-3B). This iterative approach helps ensure the final synthesized DNA is highly pure.68 / 112#14588836v1H0824.70442WO00
[0181] To ensure an aggregate error rate <lxl0-7, synthesis begins with “perfect parts,” (see, e.g., Example 2 (below) and FIG. 15A) initial gene fragments will be cloned into vectors and then sequenced. Subsequently, specific amplicons corresponding to each 1-mer are generated from the respective vectors. After achieving sufficient DNA quantities, end modifications are introduced to facilitate immobilization onto specially coated substrates. This coating minimizes non-specific binding and promotes double negative selection. Substrates are strategically set up with geometric layouts of spots, so there is no need to re-array across the subsequent 32 assembly steps (see, e.g., FIGs. 5A-5C).Toehold Reversible Immobilization
[0182] Toehold reversible immobilization leverages the well-established principle of Watson-Crick (WC) base pairing used in toehold displacement. During DSL synthesis, this method is utilized to achieve a controlled release of synthesized DNA. The process begins with deposition via automated pipetting on every wafer in a grid array of spots containing toehold DNA strands covalently bound to the surface. Covalent binding may be accomplished via typical lithography or mask-based approaches. The toeholds constitute reversible immobilization. To add the DNA assembly fragments to the surface, 16 inks are dispensed via automated pipetting or another similar precision dispense method onto the existing grid array of toeholds. The control from an immobilized to mobilized DNA state is achieved by adding an excess of new oligos that replace the WC base pairs (WCbp) between the immobilization bases and the assembly DNA, thereby releasing the DNA assembly fragments (see, e.g., FIGs. 5A-5B and FIGs. 10A-10B).
[0183] Toehold reversal immobilization can be used in the DSL process described above, as it allows a) bridging of successfully ligated DNA assembly fragments and b) transferring DNA assembly fragments inter-wafer.
[0184] In some embodiments, a total of 35 unique inks are utilized, which includes 16 inks for non-methylated DNA synthesis, 12 inks for introducing methylation patterns, 4 inks for restriction enzymes, 1 ink for the ligase mix, and 2 inks for toehold DNA sequences.Ligation Review
[0185] Based on previous experience with genome synthesis and editing, the DSL method is expected to surpass current standards in terms of accuracy, efficiency, and cost.Regarding ink quality, the best mean deletion frequency per raw nucleotide (0.02%) is E=2xl0-4, and the best consensus read error rate per bp for MGI & Illumina is Q=lxl0-7. Starting with ten billion molecules per drop, over 20 ligation cycles, the target yield is 50% per cycle69 / 112#14588836v1H0824.70442WO00(1x10*0.520= -10,000 molecules), and the error rate goal is 3xl0-7over all cycles (an adequate error for many viable diploid cells). Using the DSL approach, C cloned molecules are sequenced per candidate DNA ink and then cloned to check for mutants in the initial synthesis of the F=50 bp, flanking the 1-mer regions. If one clone (C= 1) is sequenced, the chance that it will be correctly rejected is E, and the chance of an undetectable mutation is E*(F*Q)c=lxlO’9. Each ligation stage has a double error-correction step, including incomplete ligations, which assures faithful construction of the desired genome (see, e.g., FIGs. 3A-3B).
[0186] The DSL approach was compared to the three most frequently used reactions in DNA synthesis: serial, multiplex, and recombinase (see FIG. 6). Serial chemical and enzymatic syntheses typically require 200 to 1000 linear assembly steps, represented as Solid and Mobile. At each step, there can be a failure to cap or deprotect a DNA strand, yet no method is built into these processes designed to mitigate this failure. Additionally, this scenario may wash away the solid phase, leaving only the mobile phase. Specifically, 3 to 52 DNA assembly fragments are simultaneously recombined in yeast, via Gibson Assembly or Golden Gate Assembly. These techniques do not have selection steps and depend entirely on the fidelity and efficiency of each reaction, leaving tremendous room for error and inefficiency. Finally, recombination-based current methods, or any process with side reactions beyond a failure to react, run the risk of off-target recombination events. Using DSL, requiring both solid and mobile phases, is a significant improvement over the current methodology.Results
[0187] FIGs. 11A-11C show ligation process optimization results. FIGs. 11A-11B show results from a nucleic acid production method decreased inefficiency to 6xl0-2(Goal: 5xl0-2).FIG. 11C shows examples of mis-ligation rates.
[0188] FIG. 12 shows washing efficiency as a graphical representation showing average of two experimental repeats in logarithmic scale. On average 1.98xlO10DNA molecules were shown to be effectively washed away.
[0189] FIGs. 13A-13E depict an example of a stepwise assembly of a synthesized DNA sequence using Alwl and BccI restriction enzymes. “X” indicates 1-mer Cut Sites. These positions indicate where the enzymes (Alwl and BccI) cleave the DNA, generating 1-base overhangs. Uppercase and lowercase letters represent complementary bases, highlighting the formation of the synthesized product as the process progresses. Highlighted sections showcase the specific recognition sequences targeted by Alwl and BccI enzymes.70 / 112#14588836v1H0824.70442WO00
[0190] FIGs. 14A-14B show effects of a toehold system. FIG. 14A is a bar graph showing the release efficiency of three types of toehold systems (toehold, triplex toehold, and hairpin-toehold systems). Release efficiency was calculated with the following formula:(Released / (Unreleased + Released)). FIG. 14B show diagrams of toehold, triplex toehold, and hairpin-toehold systems. A coded legend is provided in the lower panel.Example 2: Production of Perfect Parts
[0191] This Example relates to a method of producing “perfect parts,” which refers to nucleic acids exhibiting extremely low error rates (e.g., on the order of lower than 1 in 103errors, such as on the order of 1 in 104, 1 in 105, 1 in 106, 1 in 107, 1 in 108, or lower) and that can be produced using a method described herein and / or used (e.g., as an acceptor or donor nucleic acid) in performing a method described herein (FIG. 15A).
[0192] A suitable vector was first selected. Following the incorporation of the sequence of interest, intended to be produced with high fidelity, into a plasmid vector using conventional molecular biology techniques (e.g., cloning, synthesis), the resulting construct was introduced into host cells, such as Escherichia coli (E. coli), via transformation. The plasmid DNA was subsequently isolated and subjected to in vitro treatment with nicking endonucleases preferably Nt. BspQI and Nt. BstNBI under the described conditions. This step facilitated the generation of single- stranded regions that, under controlled physicochemical conditions (e.g., temperature, ionic strength), promote the formation of secondary structures through intramolecular base pairing. These structures were then stabilized and sealed using a DNA ligase, resulting in the formation of closed, dumbbell-like DNA constructs. To enrich for the desired product, the reaction mixture was treated with a combination of endonucleases and exonucleases that selectively degrade residual vector backbone and non-target DNA species. The resulting constructs were then purified using standard nucleic acid purification techniques.
[0193] The final product was an enriched, clonal DNA construct exhibiting high sequence fidelity, suitable for use as a DNA synthesis ink (see, e.g., Example 1 and FIGs. 1A-1D, FIGs. 2A-2B, FIGs. 3A-3B, FIGs. 13A-13E, or FIG. 15B), or for other downstream applications requiring error-free nucleic acid templates.Example 3: Exemplary Nucleic Acid Synthesis
[0194] This Example relates to an example of a method for producing a nucleic acid, such as a genome. The method initiated with a surface such as a glass slide, a silica-based wafer, or any alternative material capable of serving as a support for nucleic acid attachment (see, e.g., FIG. 1A, FIG. 1C, FIGs. 2A-2B, FIG. 5A-5C, and FIG. 15B). The surface was first 71 / 112#14588836v1H0824.70442WO00functionalized with azide groups, enabling subsequent covalent attachment of nucleic acid toehold sequences via DBCO-azide click chemistry. The immobilized toehold sequences served as anchoring points for the hybridization of incoming DNA inks (see, e.g., FIGs. 10A-10B).
[0195] The DNA inks, dumbbell-like DNA structures were deposited in a programmable pattern, allowing arbitrary combinations and spatial arrangements. In some embodiments, the deposition happened in the form of droplets (see, e.g., FIGs. 5A-5C). Each droplet was compositionally defined and corresponds to a specific DNA 1-mer used in subsequent synthesis steps (see, e.g., FIG. 4, FIGs. 13A-13E, and FIG. 15A). Depending on the design, from a few to millions or billions of droplets can be printed on the surface (see FIG. 5A). The total number of droplets ultimately determined the maximum attainable length of the synthesized nucleic acid product.
[0196] At the initiation of synthesis, droplets containing DNA inks were printed onto the azide-coated substrate bearing toehold single- stranded DNAs (ssDNAs) installed through DBCO-azide coupling. These toeholds hybridized with the complementary hairpin region of the incoming DNA ink. In some embodiments, hybridization occurs through the single- stranded structure. In some embodiments, hybridization occurs through the double- stranded structure. The toehold sequence was designed to selectively recognize and attach to a specific site on the hairpin.
[0197] Subsequently, a nicking endonuclease, such as Alwl or BccI, was introduced to process the DNA ink, converting it into an exposed dumbbell configuration as illustrated in Step 1 of FIG. 15B. Once all droplets with defined configurations were deposited and processed, a ligation step followed. During this step, a ligase enzyme and an anti-toehold sequence, a complementary DNA that mediates strand displacement and releases the hairpin from the toehold hybridization, were introduced to neighboring droplets to form localized reaction pockets (see FIGs. 5B-5C). The ligase catalyzed covalent joining of the diffusing donor hairpin with the surface-bound acceptor hairpin, generating an extended construct.
[0198] After ligation, a washing sequence was performed to remove residual reagents and minimize background signal. The washing protocol typically comprised (i) treatment with T5 exonuclease to degrade unligated or exposed DNA strands, followed by (ii) a rinse with water, buffer, or another suitable aqueous medium.
[0199] In subsequent stages, anti-toehold sequences were again introduced via droplet-to-droplet reaction pockets to mobilize the DNA constructs from one surface to another, enabling immobilization from the opposite end of the molecule. This alternating immobilization strategy constituted the double- selection mechanism. Following another wash cycle, a 1-mer 72 / 112#14588836v1H0824.70442WO00cutter endonuclease (e.g., Alwl) was applied to excise undesired regions of the newly ligated dumbbell structures, while preserving the additional nucleotide incorporated during the prior ligation event (see FIGs. 13A-13E and FIG. 15B). The resulting construct, bearing a 1-mer sticky end, became a new acceptor ready for hybridization and ligation in the next synthesis cycle (corresponding to Step 2 in FIG. 15B).
[0200] This iterative process was repeated until the desired sequence length was achieved up to Step 15 (see FIG. 15B) while in each step the synthesized product increased in hierarchically. At this stage, the nascent DNA was circularized.
[0201] Beyond Step 15 (see FIG. 15B), the system transitioned to recombinase-mediated assembly (see FIGs. 1A-1D, FIG. 7, and FIGs. 8A-8C). Reaction pockets were again formed between droplets containing recombinases and anti-toehold sequences targeted to mobilize the circular DNA. After recombination and washing, the newly formed construct underwent a double-selection mobilization and immobilization from the alternate end, which prepared it for the next extension cycle. This stepwise process can be iterated through approximately 33 cycles (or more), enabling the synthesis of DNA molecules of genomic scale.Example 4: Assembling Nucleic Acids Using Recombinases
[0202] This Example relates to methods comprising contacting a nucleic acid (e.g., a nucleic acid comprising a target sequence) with a recombinase (see, e.g., FIGs. 1A-1D, FIG. 7, FIGs. 8A-8C, and FIG. 15B). In some embodiments, the assembly of large or circularized DNA constructs proceeded through recombinase-mediated synthesis, leveraging the high fidelity and sequence- specific recombination capabilities of large serine or tyrosine recombinases. Following the initial assembly and circularization stages (e.g., after Step 15 as illustrated in FIG. 15B), the nascent DNA molecules served as substrates for site-specific recombination events that expand, reorganize, or concatenate genomic segments in a controlled manner.Accelerated Recombinase Assay and Double-Selection Mechanism for Recombinase-mediated Synthesis
[0203] The lower panel of FIG. 15B schematically represents the step-by-step operation of the recombinase-mediated double- selection system. This configuration enables both spatial control of recombination events and acceleration of reaction kinetics via the use of optimized buffer composition and molecular crowding agents, as described herein.
[0204] Each synthesis cycle involved droplet-based confinement of recombination reactions between two DNA constructs, typically a donor and an acceptor circle, each bearing recombination sites of defined sequence and spacing. Reaction pockets were formed by merging 73 / 112#14588836v1H0824.70442WO00droplets containing recombinase, buffer, and anti-toehold sequences that facilitate strand displacement and spatial rearrangement of the substrate DNA. Two double- stranded DNA (dsDNA) constructs, each containing a recombination recognition site, served as substrates. These constructs can be circular DNA molecules, one bearing an attP site and the other an attB site corresponding to a specific large serine recombinase. As shown in FIG. 5 and FIG. 15B, droplets containing these circular DNAs were positioned in an upper (donor) and lower (acceptor) configuration. Upon merging the two droplets, the anti-toehold sequence complementary to the toehold of one immobilized DNA circle was delivered from the opposite droplet. This exchange mobilized one circular DNA molecule while the second remained surface-anchored, thereby bringing both DNA substrates into close spatial proximity.
[0205] Once proximity was established, the recombinase enzyme was introduced, catalyzing site-specific recombination between the two circles to yield a recombined DNA construct. Recombination proceeded under optimized temperature and ionic conditions until recombination events generated the new DNA molecules of the desired configuration. If the resulting construct contained additional recombinase-specific recognition sites (e.g., for orthogonal recombinases), the same reaction cycle was repeated, enabling hierarchical multirecombination assembly. Because each DNA molecule can be selectively immobilized or released through the triplex toehold capture-and-release system, the platform supported a double- selection mechanism, ensuring directional and sequential recombination.
[0206] Reaction acceleration was achieved through the use of a specialized buffer system designed to optimize enzyme activity and macromolecular interactions. Thee reaction buffer can comprise: Tris (e.g., wherein the Tris is at a concentration of 1-1000 mM Tris), EDTA (e.g., wherein the EDTA is at a concentration of 1-100 mM EDTA), Glycerol (e.g., wherein the glycerol is at a concentration of 1-20% glycerol), one or more salts (e.g., wherein the one or more salts are at a concentration of 0-1000 mM salt and / or wherein the one or more salts comprise NaCl), and / or one or more crowding agents (e.g., wherein the concentration of the one or more crowding agents is 1-60% and / or wherein the one or more crowding agents comprise PEG-200, PEG-400, PEG- 1000, PEG-2000, PEG-6000, PEG-8000, PEG- 10000, PEG-20000, or another PEG variant). Reactions were typically conducted at temperatures between 4 °C and 50 °C, depending on the recombinase species.
[0207] After completion of each recombination event, washing and double- selection steps were carried out as described previously, ensuring removal of residual enzyme and unreacted substrates. The double- selection system allowed alternating immobilization from opposite ends of the DNA, minimizing carryover and error propagation across cycles. In some 74 / 112#14588836v1H0824.70442WO00embodiments, the recombined DNA molecules are further processed by additional recombination events to minimize the carry over recombination sites.Results
[0208] Experimental data associated with these recombinase-based synthesis steps demonstrate high efficiency and fidelity under accelerated reaction times (FIG. 16). Complete recombination events per cycle have been observed for recombinases under the conditions described herein. The conditions shown here improved recombinase efficiencies over 200-fold, including 50% efficiency in an hour and almost 100% efficiency in 24 hours (FIG. 16).Additionally, under the conditions described herein, recombination events proceeded at markedly accelerated rates relative to reactions lacking one or more of these buffer components. The molecular crowding environment created by PEG and related polymers enhanced local DNA concentration and recombinase-substrate encounter frequency, thereby reducing reaction time while maintaining high site-specific fidelity. The efficiency of recombination events was confirmed through qPCR.
[0209] In some embodiments, recombinase-mediated assembly can be iteratively performed for up to 33 sequential steps or more, enabling the construction of megabase- to genome-scale DNA molecules with defined architecture and minimal sequence deviation from the intended design.
[0210] Collectively, these findings demonstrate that engineered buffer composition synergistically enable high-speed, high-fidelity recombinase-mediated DNA assembly suitable for accelerated, iterative, genome-scale synthesis applications.NON-LIMITING EMBODIMENTS
[0211] The following numerated embodiments represent non-limiting aspects of the invention.1. A method of producing a nucleic acid comprising (i) ligating:(a) a nucleic acid comprising a first anchor bound to a first binding molecule on a surface, a cleavage site for a first nuclease, and a fragment of a portion of a target sequence comprising an overhang; and(b) a nucleic acid comprising a second anchor bound to a second binding molecule on a surface, a cleavage site for a second nuclease, and a fragment of a portion of a target75 / 112#14588836v1H0824.70442WO00sequence comprising an overhang that is reverse complementary to the overhang in the nucleic acid comprising the first anchor,thereby generating a ligated nucleic acid comprising the first anchor, the cleavage site for the first nuclease, the portion of the target sequence formed by ligating the overhangs, the cleavage site for the second nuclease, and the second anchor.2. The method of embodiment 1 further comprising (ii) cleaving the ligated nucleic acid with the second nuclease, thereby generating:a cleaved nucleic acid comprising the first anchor bound to the first binding molecule, the cleavage site for the first nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang generated by the second nuclease and the fragment is elongated relative to the fragments in the nucleic acids ligated in a preceding ligation step; and a cleavage fragment comprising the second anchor bound to the second binding molecule and the cleavage site for the second nuclease.3. The method of embodiment 2 further comprising subjecting the second binding molecule to conditions that inhibit binding to the second anchor.4. The method of embodiment 2 or 3, wherein reverse immobilization is utilized to inhibit the binding of the second anchor to the second binding molecule.5. The method of any one of embodiments 2-4 further comprising repeating step (i) by performing a subsequent ligation step, wherein the subsequent ligation step comprises ligating:(a) the cleaved nucleic acid; and(b) a nucleic acid comprising the second anchor bound to the second binding molecule on a surface, the cleavage site for the second nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang that is reverse complementary to the overhang in the cleaved nucleic acid,thereby generating a ligated nucleic acid comprising the portion of the target sequence formed by ligating the overhangs, wherein the portion of the target sequence is elongated relative to the portion of the target sequence in the ligated nucleic acid generated in the preceding ligation step.76 / 112#14588836v1H0824.70442WO006. The method of embodiment 1 further comprising (ii) cleaving the ligated nucleic acid with the first nuclease, thereby generating:a cleavage fragment comprising the first anchor bound to the first binding molecule and the cleavage site for the first nuclease; anda cleaved nucleic acid comprising the second anchor bound to the second surface, the cleavage site for the second nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang generated by the first nuclease and the fragment is elongated relative to the fragments of the target sequence in the nucleic acids ligated in a preceding ligation step.7. The method of embodiment 6 further comprising subjecting the first binding molecule to conditions that inhibit binding to the first anchor.8. The method of embodiment 6 or 7, wherein reverse immobilization is utilized to inhibit the binding of the first anchor to the first binding molecule.9. The method of embodiment 6 or 7 further comprising repeating step (i) by performing a subsequent ligation step, wherein the subsequent ligation step comprises ligating:(a) a nucleic acid comprising the first anchor bound to the first binding molecule on a surface, the cleavage site for the first nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang that is reverse complementary to the overhang in the cleaved nucleic acid; and(b) the cleaved nucleic acid,thereby generating a ligated nucleic acid comprising the portion of the target sequence formed by ligating the overhangs, wherein the portion of the target sequence is elongated relative to the portion of the target sequence in the ligated nucleic acid generated in the preceding ligation step.10. The method of embodiment 5 or 9 further comprising repeating step (ii) by performing a subsequent cleavage step, wherein the subsequent cleavage step comprises cleaving the ligated nucleic acid with the first nuclease or the second nuclease, thereby generating:a cleavage fragment; anda cleaved nucleic acid comprising a fragment of a portion of the target sequence.77 / 112#14588836v1H0824.70442WO0011. The method of embodiment 10, wherein the ligation and cleavage steps (i)-(ii) are repeated two (2) times or more.12. The method of embodiment 11, wherein each of the ligation steps comprise ligating (a) a cleaved nucleic acid generated in a preceding cleavage step by cleaving a ligated nucleic acid with the second nuclease, and (b) cleaved nucleic acid generated in a preceding cleavage step by cleaving a ligated nucleic acid with the first nuclease, wherein a final ligation step generates a ligated nucleic acid comprising the target sequence.13. The method of embodiment 12, wherein the portion of the target sequence in each of the ligated nucleic acids comprises “Y” bases, wherein Y=2(n l)+1, wherein “n” is equal to the number of ligation steps performed.14. The method of any one of embodiments 10-13, wherein the target sequence is generated with an error rate of 1 x 10’3per base pair or lower.15. The method of embodiment 14, wherein the error rate is 3 x 10’9per base pair or lower.16. The method of any one of embodiments 10-15 further comprising methylating a ligated nucleic acid.17. The method of any one of embodiments 10-16 further comprising methylating a cleaved nucleic acid.18. The method of embodiment 16 or 17, wherein the methylation comprises contacting the nucleic acid with a methyltransferase.19. The method of any one of embodiments 16-18, wherein a guanine (G), an adenine (A), a cytosine (C), and / or a thymine (T) located in a sequence comprising AA, AC, AG, AT, CA, CC, CG, CT, GA, GC, GG, GT, TA, TC, TG, or TT is methylated.20. The method of any one of embodiments 16-19, wherein the methylation generates a N6-methyladenosine at a nucleotide position.78 / 112#14588836v1H0824.70442WO0021. The method of any one of embodiments 16-20, wherein the methylation generates a 5-methylcytosine at a nucleotide position.22. The method of embodiment 20 or 21, wherein the nucleotide position is located in a portion of the target sequence.23. The method of any one of embodiments 20-22, wherein the nucleotide position is located in an overhang.24. The method of any one of embodiments 16-23, wherein the ligated nucleic acid is methylated prior to utilizing the ligated nucleic acid in a subsequent cleavage step, wherein the methylation inhibits cleavage of the target sequence in the subsequent cleavage step.25. The method of any one of embodiments 17-24, wherein the cleaved nucleic acid is methylated prior to utilizing the cleaved nucleic acid in a subsequent ligation step, wherein the methylation inhibits cleavage of the target sequence in a subsequent cleavage step that utilizes a ligated nucleic acid generated in the subsequent ligation step.26. The method of any one of embodiments 16-25 further comprising demethylating a ligated nucleic acid.27. The method of any one of embodiments 17-26 further comprising demethylating a cleaved nucleic acid.28. The method of embodiments 26 or 27, wherein the demethylation comprises contacting the nucleic acid with a demethylase.29. The method of any one of embodiments 26-28, wherein a sequence comprising a guanine (G), an adenine (A), a cytosine (C), and / or a thymine (T) located in a sequence comprising AA, AC, AG, AT, CA, CC, CG, CT, GA, GC, GG, GT, TA, TC, TG, or TT is demethylated.30. The method of any one of embodiments 26-29, wherein a N6-methyladenosine is demethylated.79 / 112#14588836v1H0824.70442WO0031. The method of any one of embodiments 26-30, wherein a 5-methylcytosine is demethylated.32. The method of any one of embodiments 26-31, wherein the ligated nucleic acid is demethylated and utilized in a subsequent cleavage step.33. The method of any one of embodiments 10-32, wherein the target sequence comprises a gene sequence or a portion of a gene.34. The method of any one of embodiments 10-33, wherein target the sequence comprises a plurality of gene sequences.35. The method of any one of embodiments 10-34, wherein the target sequence comprises the sequence of a chromosome or a portion thereof.36. The method of embodiment 35, wherein the chromosome is naturally present in a genome of an animal.37. The method of embodiment 36, wherein the animal is a mammal.38. The method of embodiment 37, wherein the mammal is a human.39. The method of any one of embodiments 10-34, wherein the target sequence comprises a bacterial genome or a portion of a bacterial genome.40. The method of any one of embodiments 10-35, wherein the target sequence comprises a viral genome or a portion of a viral genome.41. The method of embodiment 36, wherein the chromosome is naturally present in a genome of a plant.42. The method of any one of embodiments 10-41 further comprising isolating a nucleic acid comprising the target sequence.80 / 112#14588836v1H0824.70442WO0043. The method of embodiment 42, wherein isolating the nucleic acid comprises subjecting the surface comprising the first binding molecule to conditions that inhibit binding to the first anchor.44. The method of embodiment 42 or 43, wherein isolating the nucleic acid comprises subjecting the surface comprising the second binding molecule to conditions that inhibit binding to the second anchor.45. The method of any one of embodiments 42-44, wherein isolating the nucleic acid comprises cleaving the ligated nucleic acid generated in the final ligation step or a cleaved nucleic acid generated in a final cleavage step with a third nuclease and / or a fourth nuclease, wherein a cleavage site for the third nuclease is flanked by the first anchor and the cleavage site for the first nuclease,wherein a cleavage site for the fourth nuclease is flanked by the second anchor and the cleavage site for the second nuclease.46. The method of embodiment 45 further comprising preparing the nucleic acids ligated in an initial ligation step, wherein preparing the nucleic acids comprises:(i) contacting one or more polynucleotides with one or more nucleases, thereby generating a plurality of nucleic acids, wherein each nucleic acid of the plurality comprises a unique portion of the target sequence which is selected from the group consisting of: AA, GA, CA, TA, AG, GG, CG, TG, AC, GC, CC, TC, AT, GT, CT, and TT, wherein the portion of the target sequence is flanked by(a) a sequence comprising the cleavage site for the third nuclease and the cleavage site for the first nuclease, and(b) a sequence comprising the cleavage site for the second nuclease and the cleavage site for the fourth nuclease;(ii) cleaving one or more nucleic acids of the plurality with the first nuclease and cleaving nucleic acids of the plurality with the second nuclease; and(iii) fusing the first anchor to the one or more nucleic acids cleaved with the second nuclease and fusing the second anchor to the one or more nucleic acids cleaved with the first nuclease.81 / 112#14588836v1H0824.70442WO0047. The method of embodiment 46 further comprising (iv) contacting at least one of the nucleic acids comprising the first anchor with the surface comprising the first binding molecule and contacting at least one of the nucleic acids comprising the second anchor with the surface comprising the second binding molecule.48. The method of any one of embodiments 1-47 further comprising washing the surface comprising the first binding molecule before, during, and / or after ligating the overhangs.49. The method of any one of embodiments 1-48 further comprising washing the surface comprising the first binding molecule before, during, and / or after cleaving the ligated nucleic acid.50. The method of any one of embodiments 1-49 further comprising washing the surface comprising the second binding molecule before, during, and / or after ligating the overhangs.51. The method of any one of embodiments 1-50 further comprising washing the surface comprising the first binding molecule before, during, and / or after cleaving the ligated nucleic acid.52. The method of any one of embodiments 1-51, wherein the first anchor comprises a small molecule, a click-chemistry handle, a peptide, a protein, or an oligonucleotide;.53. The method of any one of embodiments 1-52, wherein the first anchor comprises a maltose-binding protein tag and the first binding molecule comprises maltose.54. The method of any one of embodiments 1-53, wherein the first anchor comprises an oligonucleotide and the first binding molecule comprises a reverse complementary oligonucleotide.55. The method of embodiment 54, wherein toehold displacement is utilized to inhibit the hybridization of the oligonucleotides.56. The method of any one of embodiments 1-55, wherein the second anchor comprises a small molecule, a click-chemistry handle, a peptide, a protein, or an oligonucleotide.82 / 112#14588836v1H0824.70442WO0057. The method of any one of embodiments 1-56, wherein the second anchor comprises a poly His tag and the second binding molecule comprises Nickel-Nitriloacetic acid (NTA).58. The method of any one of embodiments 1-56, wherein the second anchor comprises an oligonucleotide and the second binding molecule comprises a reverse complementary oligonucleotide.59. The method of embodiment 58, wherein toehold displacement is utilized to inhibit the hybridization of the oligonucleotides.60. The method of any one of embodiments 1-59, wherein at least one of the nucleases is a restriction enzyme.61. The method of embodiment 60, wherein the restriction enzyme is a type IIS restriction enzyme.62. The method of any one of embodiments 1-61, wherein at least one of the nucleases is selected from the group consisting of Alwl; Bccl; Piel; BciVI; BmrI; HphI; HpryAv; Mboll; Mnll; Sbfl; and PflMI.63. The method of any one of embodiments 1-62, wherein the first nuclease is Alwl.64. The method of any one of embodiments 1-63, wherein the second nuclease is Bccl.65. The method of any one of embodiments 45-47, wherein the third nuclease is Sbfl.66. The method of any one of embodiments 45-47 or 65, wherein the fourth nuclease is PflMI.67. The method of any one of embodiments 1-66, wherein each of the overhangs comprise an odd number of nucleotides.83 / 112#14588836v1H0824.70442WO0068. The method of any one of embodiments 1-67, wherein each of the overhangs is one (1) nucleotide in length.69. The method of any one of embodiments 1-68, wherein each of the overhangs is two (2), three (3), four (4), or five (5) nucleotides in length.70. The method of any one of embodiments 1-69, wherein ligating the nucleic acids comprises contacting the nucleic acids with a T4 DNA ligase.71. The method of any one of embodiments 1-70, wherein the nucleic acids are ligated in the presence of a crowding agent.72. The method of embodiment 71, wherein the crowding agent is polyethylene glycol (PEG).73. The method of embodiment 72, wherein the concentration of PEG is between 1% w / v to 50% w / v.74. The method of embodiment 72 or 73, wherein the concentration of PEG is at least 25% w / v.75. The method of any one of embodiments 1-74, wherein the nucleic acids are in the presence of glycerol during the ligation step.76. The method of any one of embodiments 1-75, wherein the nucleic acids are in the presence of glycerol during the cleavage step.77. The method of embodiment 75 or 76, wherein the concentration of glycerol is between 10% w / v to 80% w / v.78. The method of any one of embodiments 75-77, wherein the concentration of glycerol is at least 40% w / v.84 / 112#14588836v1H0824.70442WO0079. The method of any one of embodiments 1-78, wherein one or both of the surfaces is glass.80. The method of any one of clams 1-79, wherein one or both of the surfaces has a temperature of between 4 °C and 95 °C.81. The method of any one of embodiments 1-80, wherein one or both of the surfaces undergo a temperature change from a first temperature to a second temperature, wherein the first temperature and the second temperature are both between 4 °C and 95 °C.82. The method of any one of embodiments 1-81 further comprising providing the nucleic acids in liquid droplets on the surfaces, wherein the liquid droplets are merged during each ligation step.83. The method of embodiment 82, wherein merging the liquid droplets is automated.84. The method of embodiment 82 or 83, wherein the liquid droplets are loaded on the surfaces using automated pipetting.85. The method of any one of embodiments 82-84, wherein each liquid droplet is 10 pL or more in volume.86. The method of any one of embodiments 81-84, wherein each liquid droplet is spaced 50 pm apart.87. The method of any one of embodiments 82-86, wherein the liquid droplets are in an array on a slide comprising an equal number of rows and columns of liquid droplets.88. The method of embodiment 87, wherein the array comprises 256 rows of liquid droplets and 256 columns of liquid droplets.89. The method of embodiment 87 or 88, wherein the array comprises 1,398 rows of liquid droplets and 1,398 columns of liquid droplets.85 / 112#14588836v1H0824.70442WO0090. The method of any one of embodiments 1-89, wherein a ligase is contacted with the nucleic acid comprising the first anchor and / or the nucleic acid comprising the second anchor using automated pipetting.91. The method of any one of embodiments 2-90, wherein the nuclease contacted with the ligated nucleic using automated pipetting.92. The method of any one of embodiments 10-91 further comprising compacting, condensing, circularizing, and / or supercoiling the nucleic acid comprising the target sequence.93. The method of any one of embodiments 10-92 further comprising compacting, condensing, circularizing, and / or supercoiling the nucleic acid comprising the target sequence by:(i) contacting the nucleic acid comprising the target sequence with a Cas protein and a first single- stranded polynucleotide, wherein the target sequence comprises a first binding site for a recombinase and a second binding site for the recombinase, wherein the first single- stranded polynucleotide hybridizes to the first binding site for the recombinase;(ii) contacting the nucleic acid with the recombinase, thereby generating a recombination event at the second binding site for the recombinase;(iii) contacting the nucleic acid with a Cas protein and a second single- stranded polynucleotide that hybridizes to the second binding site for the recombinase; and (iv) contacting the nucleic acid with the recombinase, thereby generating a recombination event at the first binding site for the recombinase.94. The method of embodiment 93 further comprising subjecting the nucleic acid to conditions that inhibit binding of the first polynucleotide and the first binding site for the recombinase before contacting the nucleic acid with the second polynucleotide in step (iii).95. The method of embodiment 93 or 94 further comprising subjecting the nucleic acid to conditions that inhibit binding of the second polynucleotide and the second binding site for the recombinase after contacting the nucleic acid with the recombinase in step (iv).86 / 112#14588836v1H0824.70442WO0096. The method of any one of embodiments 93-95 further comprising washing the nucleic acid before, during, and / or after contacting the nucleic acid with the first single- stranded polynucleotide in step (i), the second single- stranded polynucleotide in step (iii), or both.97. The method of any one of embodiments 93-96 further comprising washing the nucleic acid before, during, and / or after contacting the nucleic acid with the recombinase in step (ii), step (iv), or both.98. The method of any one of embodiments 93-97 further comprising repeating steps (i)-(iv) one or more times.99. The method of any one of embodiments 93-98, wherein the recombinase is a serine recombinase.100. The method of any one of embodiments 93-98, wherein the recombinase is a tyrosine recombinase.101. The method of any one of embodiments 93-98, wherein the recombinase is selected from the group consisting of: Bbxl; PhiC31; PaOl; Pa03; Ec03; Nm60; LI Int; Cre R32M; YR1; YR2; YR11; eXerl; eXer2; eXer3; Dn29; Si74; Bm99; Bt24; Cbl6; Fm04; uCb64; Ec03;Nm60; LI Int; and an engineered variant thereof with modified attP / attB recognition specificities.102. The method of any one of embodiments 93-98, wherein the recombinase is Bxbl.103. The method of any one of embodiments 93-98, wherein the recombinase is PhiC31.104. The method of any one of embodiments 93-98, wherein the recombinase is PaOl.105. The method of any one of embodiments 93-98, wherein the recombinase is Cre32M.106. The method of any one of embodiments 93-98, wherein the recombinase is YR1.107. The method of any one of embodiments 93-98, wherein the recombinase is Xerl.87 / 112#14588836v1H0824.70442WO00108. The method of any one of embodiments 93-107, wherein the first single- stranded polynucleotide is a single- stranded DNA.109. The method of any one of embodiments 93-108, wherein the second single- stranded polynucleotide is a single- stranded DNA.110. The method of any one of embodiments 93-109, wherein the Cas protein comprises a catalytically inactive Cas protein.111. The method of any one of embodiments 93-110, wherein the Cas protein comprises Cas9.112. The method of any one of embodiments 93-111, wherein the Cas protein comprises Casl2.113. The method of any one of embodiments 93-112, wherein the Cas protein is a bi-valent Cas protein.114. The method of any one of embodiments 10-113 further comprising introducing the nucleic acid comprising the target sequence into a cell.115. The method of embodiment 114, wherein the cell is an animal cell.116. The method of embodiment 115, wherein the animal cell is a mammalian cell.117. The method of embodiment 116, wherein the mammalian cell is a human cell.118. A cell comprising the nucleic acid produced by the method of any one of embodiments 1-117.119. The cell of embodiment 118, wherein the cell is an animal cell.120. The cell of embodiment 119, wherein the animal cell is a mammalian cell.88 / 112#14588836v1H0824.70442WO00121. The cell of embodiment 120, wherein the mammalian cell is a human cell.122. A cell population comprising the cell of any one of embodiments 108-121.123. A surface comprising a binding molecule bound to a nucleic acid, wherein the nucleic acid comprises:two different anchors located at the terminal ends of the nucleic acid, wherein one of the anchors is bound to the binding molecule; andtwo different nuclease cleavage sites.124. The surface of embodiment 123, wherein the anchors comprise a small molecule, a clickchemistry handle, a peptide, a protein, or an oligonucleotide.125. The surface of embodiment 123 or 124, wherein one of the anchors comprise a maltose-binding protein and the binding molecule comprises maltose.126. The surface of any one of embodiments 123-125, wherein one of the anchors comprise a polyHis tag and the binding molecule comprises Nickel-Nitriloacetic acid (NTA).127. The nucleic acid of any one of embodiments 123-126, wherein the nucleic acid comprises four different nuclease cleavage sites.128. The surface of any one of embodiments 123-127, wherein one or more of the nuclease cleavage sites are endonuclease cleavage sites.129. The surface of embodiment 128, wherein the endonuclease cleavage sites comprise type IIS restriction enzyme cleavage sites.130. The surface of any one of embodiments 123-129, wherein one or more of the nuclease cleavage sites are selected from the group consisting of: a Alwl cleavage site; a Bccl cleavage site; a Piel cleavage site; a BciVI cleavage site; a BmrI cleavage site; a HphI cleavage site; a HpryAv cleavage site; a Mboll cleavage site; a Mull cleavage site; a Sbfl cleavage site; and a PflMI cleavage site.89 / 112#14588836v1H0824.70442WO00131. The surface of any one of embodiments 123-130, wherein the surface is glass.132. A system comprising:(a) a surface comprising a first binding molecule bound to a first anchor on a nucleic acid comprising a cleavage site for a first nuclease; and(b) a surface comprising a second binding molecule bound to a second anchor on a nucleic acid comprising a cleavage site for a second nuclease,wherein the nucleic acid comprising the first anchor and the nucleic acid comprising the second anchor comprise reverse complementary overhangs.133. The system of embodiment 132, wherein the first and second anchors comprise a small molecule, a click-chemistry handle, a peptide, a protein, or a polynucleotide.134. The system of embodiment 132 or 133, wherein the first anchor comprises a maltose-binding protein tag and the first binding molecule comprises maltose.135. The system of any one of embodiments 132-134, wherein the second anchor comprises a poly His tag and the second binding molecule comprises Nickel-Nitriloacetic acid (NTA).136. The system of any one of embodiments 132-135, wherein the nucleic acid comprising the first anchor comprises a cleavage site for a third nuclease which is located between the first anchor and the cleavage site for the first nuclease.137. The system of any one of embodiments 132-136, wherein the nucleic acid comprising the second anchor comprises a cleavage site for a fourth nuclease which is located between the first anchor and the cleavage site for the second nuclease.138. The system of any one of embodiments 132-137, wherein one or more of the nuclease cleavage sites are endonuclease cleavage sites.139. The system of embodiment 138, wherein the endonuclease cleavage sites comprise type IIS restriction enzyme cleavage sites.90 / 112#14588836v1H0824.70442WO00140. The system of any one of embodiments 132-139, wherein one or more of the nuclease cleavage sites are selected from the group consisting of: a Alwl cleavage site; a Bccl cleavage site; a Piel cleavage site; a BciVI cleavage site; a BmrI cleavage site; a HphI cleavage site; a HpryAv cleavage site; a Mboll cleavage site; a Mull cleavage site; a Sbfl cleavage site; and a PflMI cleavage site.141. The system of any one of embodiments 132-140, wherein the overhangs comprise an odd number of nucleotides.142. The system of any one of embodiments 132-141, wherein the overhangs comprise one (1) nucleotide in length.143. The system of any one of embodiments 132-139, wherein the overhangs comprise two (2), three (3), four (4), or five (5) nucleotides in length.144. The system of any one of embodiments 132-143, wherein the surface comprising the first binding molecule and / or the surface comprising the second binding molecule is glass.145. The system of any one of embodiments 132-144, wherein the nucleic acids are provided in liquid droplets on the surfaces.146. The system of embodiment 145, wherein each liquid droplet is 10 pL or more in volume.147. The system of embodiment 145 or 146, wherein each liquid droplet is spaced 50 pm apart.148. The system of any one of embodiments 145-147, wherein the system comprises one or more slides, wherein each slide comprises an array of the liquid droplets, wherein each array comprises an equal number of rows and columns of liquid droplets.149. The system of embodiment 148, wherein the array comprises 256 rows of liquid droplets and 256 columns of liquid droplets.91 / 112#14588836v1H0824.70442WO00150. The system of embodiment 148 or 149, wherein the array comprises 1,398 rows of liquid droplets and 1,398 columns of liquid droplets.151. The system of any one of embodiments 132-150, wherein one or more of the nucleic acids comprising the first anchor further comprise a methylated nucleobase located in the overhang.152. The system of any one of embodiments 132-151, wherein one or more of the nucleic acids comprising the second anchor further comprise a methylated nucleobase located in the overhang153. The system of embodiment 151 or 152, wherein the methylated nucleobase is a guanine (G), an adenine (A), a cytosine (C), or a thymine (T) located in a sequence comprising AA, AC, AG, AT, CA, CC, CG, CT, GA, GC, GG, GT, TA, TC, TG, or TT.154. The system of any one of embodiments 151-153, wherein the methylated nucleobase is N6-methyladenosine.155. The system of any one of embodiments 151-154, wherein the methylated nucleobase is 5-methylcytosine.EQUIVALENTS AND SCOPE
[0212] While several inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents 92 / 112#14588836v1H0824.70442WO00thereto, inventive embodiments may be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.
[0213] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0214] All references, patents and patent applications disclosed herein are incorporated by reference with respect to the subject matter for which each is cited, which in some cases may encompass the entirety of the document.
[0215] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”
[0216] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0217] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting 93 / 112#14588836v1H0824.70442WO00essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.
[0218] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0219] It should also be understood that, unless clearly indicated to the contrary, in any methods claimed herein that include more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited.
[0220] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03. It should be appreciated that embodiments described in this document using an open-ended transitional phrase (e.g., “comprising”) are also contemplated, in alternative embodiments, as “consisting of’ and “consisting essentially of’ the feature described by the open-ended transitional phrase. For example, if the disclosure describes “a composition comprising A and B”, the disclosure also contemplates the alternative embodiments “a composition consisting of A and B” and “a composition consisting essentially of A and B”.94 / 112#14588836v1H0824.70442WO00
Claims
CLAIMSWhat is claimed is:
1. A method of producing a nucleic acid comprising (i) ligating:(a) a nucleic acid comprising a first anchor bound to a first binding molecule on a surface, a cleavage site for a first nuclease, and a fragment of a portion of a target sequence comprising an overhang; and(b) a nucleic acid comprising a second anchor bound to a second binding molecule on a surface, a cleavage site for a second nuclease, and a fragment of a portion of a target sequence comprising an overhang that is reverse complementary to the overhang in the nucleic acid comprising the first anchor,thereby generating a ligated nucleic acid comprising the first anchor, the cleavage site for the first nuclease, the portion of the target sequence formed by ligating the overhangs, the cleavage site for the second nuclease, and the second anchor.
2. The method of claim 1 further comprising (ii) cleaving the ligated nucleic acid with the second nuclease, thereby generating:a cleaved nucleic acid comprising the first anchor bound to the first binding molecule, the cleavage site for the first nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang generated by the second nuclease and the fragment is elongated relative to the fragments in the nucleic acids ligated in a preceding ligation step; and a cleavage fragment comprising the second anchor bound to the second binding molecule and the cleavage site for the second nuclease.
3. The method of claim 2 further comprising subjecting the second binding molecule to conditions that inhibit binding to the second anchor.
4. The method of claim 2 or 3, wherein reverse immobilization is utilized to inhibit the binding of the second anchor to the second binding molecule.
5. The method of any one of claims 2-4 further comprising repeating step (i) by performing a subsequent ligation step, wherein the subsequent ligation step comprises ligating:(a) the cleaved nucleic acid; and95 / 112#14588836v1H0824.70442WO00(b) a nucleic acid comprising the second anchor bound to the second binding molecule on a surface, the cleavage site for the second nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang that is reverse complementary to the overhang in the cleaved nucleic acid,thereby generating a ligated nucleic acid comprising the portion of the target sequence formed by ligating the overhangs, wherein the portion of the target sequence is elongated relative to the portion of the target sequence in the ligated nucleic acid generated in the preceding ligation step.
6. The method of claim 1 further comprising (ii) cleaving the ligated nucleic acid with the first nuclease, thereby generating:a cleavage fragment comprising the first anchor bound to the first binding molecule and the cleavage site for the first nuclease; anda cleaved nucleic acid comprising the second anchor bound to the second surface, the cleavage site for the second nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang generated by the first nuclease and the fragment is elongated relative to the fragments of the target sequence in the nucleic acids ligated in a preceding ligation step.
7. The method of claim 6 further comprising subjecting the first binding molecule to conditions that inhibit binding to the first anchor.
8. The method of claim 6 or 7, wherein reverse immobilization is utilized to inhibit the binding of the first anchor to the first binding molecule.
9. The method of claim 6 or 7 further comprising repeating step (i) by performing a subsequent ligation step, wherein the subsequent ligation step comprises ligating:(a) a nucleic acid comprising the first anchor bound to the first binding molecule on a surface, the cleavage site for the first nuclease, and a fragment of a portion of the target sequence, wherein the fragment comprises an overhang that is reverse complementary to the overhang in the cleaved nucleic acid; and(b) the cleaved nucleic acid,thereby generating a ligated nucleic acid comprising the portion of the target sequence formed by ligating the overhangs, wherein the portion of the target sequence is 96 / 112#14588836v1H0824.70442WO00elongated relative to the portion of the target sequence in the ligated nucleic acid generated in the preceding ligation step.
10. The method of claim 5 or 9 further comprising repeating step (ii) by performing a subsequent cleavage step, wherein the subsequent cleavage step comprises cleaving the ligated nucleic acid with the first nuclease or the second nuclease, thereby generating:a cleavage fragment; anda cleaved nucleic acid comprising a fragment of a portion of the target sequence.
11. The method of claim 10, wherein the ligation and cleavage steps (i)-(ii) are repeated two (2) times or more.
12. The method of claim 11, wherein each of the ligation steps comprise ligating (a) a cleaved nucleic acid generated in a preceding cleavage step by cleaving a ligated nucleic acid with the second nuclease, and (b) cleaved nucleic acid generated in a preceding cleavage step by cleaving a ligated nucleic acid with the first nuclease, wherein a final ligation step generates a ligated nucleic acid comprising the target sequence.
13. The method of claim 12, wherein the portion of the target sequence in each of the ligated nucleic acids comprises “Y” bases, wherein Y=2(n l)+1, wherein “n” is equal to the number of ligation steps performed.
14. The method of any one of claims 10-13, wherein the target sequence is generated with an error rate of 1 x 10’3per base pair or lower.
15. The method of claim 14, wherein the error rate is 3 x 10’9per base pair or lower.
16. The method of any one of claims 10-15 further comprising methylating a ligated nucleic acid.
17. The method of any one of claims 10-16 further comprising methylating a cleaved nucleic acid.97 / 112#14588836v1H0824.70442WO0018. The method of claim 16 or 17, wherein the methylation comprises contacting the nucleic acid with a methyltransferase.
19. The method of any one of claims 16-18, wherein a guanine (G), an adenine (A), a cytosine (C), and / or a thymine (T) located in a sequence comprising AA, AC, AG, AT, CA, CC, CG, CT, GA, GC, GG, GT, TA, TC, TG, or TT is methylated.
20. The method of any one of claims 16-19, wherein the methylation generates a N6-methyladenosine at a nucleotide position.
21. The method of any one of claims 16-20, wherein the methylation generates a 5-methylcytosine at a nucleotide position.
22. The method of claim 20 or 21, wherein the nucleotide position is located in a portion of the target sequence.
23. The method of any one of claims 20-22, wherein the nucleotide position is located in an overhang.
24. The method of any one of claims 16-23, wherein the ligated nucleic acid is methylated prior to utilizing the ligated nucleic acid in a subsequent cleavage step, wherein the methylation inhibits cleavage of the target sequence in the subsequent cleavage step.
25. The method of any one of claims 17-24, wherein the cleaved nucleic acid is methylated prior to utilizing the cleaved nucleic acid in a subsequent ligation step, wherein the methylation inhibits cleavage of the target sequence in a subsequent cleavage step that utilizes a ligated nucleic acid generated in the subsequent ligation step.
26. The method of any one of claims 16-25 further comprising demethylating a ligated nucleic acid.
27. The method of any one of claims 17-26 further comprising demethylating a cleaved nucleic acid.98 / 112#14588836v1H0824.70442WO0028. The method of claims 26 or 27, wherein the demethylation comprises contacting the nucleic acid with a demethylase.
29. The method of any one of claims 26-28, wherein a sequence comprising a guanine (G), an adenine (A), a cytosine (C), and / or a thymine (T) located in a sequence comprising AA, AC, AG, AT, CA, CC, CG, CT, GA, GC, GG, GT, TA, TC, TG, or TT is demethylated.
30. The method of any one of claims 26-29, wherein a N6-methyladenosine is demethylated.
31. The method of any one of claims 26-30, wherein a 5-methylcytosine is demethylated.
32. The method of any one of claims 26-31, wherein the ligated nucleic acid is demethylated and utilized in a subsequent cleavage step.
33. The method of any one of claims 10-32, wherein the target sequence comprises a gene sequence or a portion of a gene.
34. The method of any one of claims 10-33, wherein target the sequence comprises a plurality of gene sequences.
35. The method of any one of claims 10-34, wherein the target sequence comprises the sequence of a chromosome or a portion thereof.
36. The method of claim 35, wherein the chromosome is naturally present in a genome of an animal.
37. The method of claim 36, wherein the animal is a mammal.
38. The method of claim 37, wherein the mammal is a human.
39. The method of any one of claims 10-34, wherein the target sequence comprises a bacterial genome or a portion of a bacterial genome.99 / 112#14588836v1H0824.70442WO0040. The method of any one of claims 10-35, wherein the target sequence comprises a viral genome or a portion of a viral genome.
41. The method of claim 36, wherein the chromosome is naturally present in a genome of a plant.
42. The method of any one of claims 10-41 further comprising isolating a nucleic acid comprising the target sequence.
43. The method of claim 42, wherein isolating the nucleic acid comprises subjecting the surface comprising the first binding molecule to conditions that inhibit binding to the first anchor.
44. The method of claim 42 or 43, wherein isolating the nucleic acid comprises subjecting the surface comprising the second binding molecule to conditions that inhibit binding to the second anchor.
45. The method of any one of claims 42-44, wherein isolating the nucleic acid comprises cleaving the ligated nucleic acid generated in the final ligation step or a cleaved nucleic acid generated in a final cleavage step with a third nuclease and / or a fourth nuclease,wherein a cleavage site for the third nuclease is flanked by the first anchor and the cleavage site for the first nuclease,wherein a cleavage site for the fourth nuclease is flanked by the second anchor and the cleavage site for the second nuclease.
46. The method of claim 45 further comprising preparing the nucleic acids ligated in an initial ligation step, wherein preparing the nucleic acids comprises:(i) contacting one or more polynucleotides with one or more nucleases, thereby generating a plurality of nucleic acids, wherein each nucleic acid of the plurality comprises a unique portion of the target sequence which is selected from the group consisting of: AA, GA, CA, TA, AG, GG, CG, TG, AC, GC, CC, TC, AT, GT, CT, and TT, wherein the portion of the target sequence is flanked by(a) a sequence comprising the cleavage site for the third nuclease and the cleavage site for the first nuclease, and100 / 112#14588836v1H0824.70442WO00(b) a sequence comprising the cleavage site for the second nuclease and the cleavage site for the fourth nuclease;(ii) cleaving one or more nucleic acids of the plurality with the first nuclease and cleaving nucleic acids of the plurality with the second nuclease; and(iii) fusing the first anchor to the one or more nucleic acids cleaved with the second nuclease and fusing the second anchor to the one or more nucleic acids cleaved with the first nuclease.
47. The method of claim 46 further comprising (iv) contacting at least one of the nucleic acids comprising the first anchor with the surface comprising the first binding molecule and contacting at least one of the nucleic acids comprising the second anchor with the surface comprising the second binding molecule.
48. The method of any one of claims 1-47 further comprising washing the surface comprising the first binding molecule before, during, and / or after ligating the overhangs.
49. The method of any one of claims 1-48 further comprising washing the surface comprising the first binding molecule before, during, and / or after cleaving the ligated nucleic acid.
50. The method of any one of claims 1-49 further comprising washing the surface comprising the second binding molecule before, during, and / or after ligating the overhangs.
51. The method of any one of claims 1-50 further comprising washing the surface comprising the first binding molecule before, during, and / or after cleaving the ligated nucleic acid.
52. The method of any one of claims 1-51, wherein the first anchor comprises a small molecule, a click-chemistry handle, a peptide, a protein, or an oligonucleotide.
53. The method of any one of claims 1-52, wherein the first anchor comprises a maltose-binding protein tag and the first binding molecule comprises maltose.101 / 112#14588836v1H0824.70442WO0054. The method of any one of claims 1-53, wherein the first anchor comprises an oligonucleotide and the first binding molecule comprises a reverse complementary oligonucleotide.
55. The method of claim 54, wherein toehold displacement is utilized to inhibit the hybridization of the oligonucleotides.
56. The method of any one of claims 1-55, wherein the second anchor comprises a small molecule, a click-chemistry handle, a peptide, a protein, or an oligonucleotide.
57. The method of any one of claims 1-56, wherein the second anchor comprises a polyHis tag and the second binding molecule comprises Nickel-Nitriloacetic acid (NTA).
58. The method of any one of claims 1-56, wherein the second anchor comprises an oligonucleotide and the second binding molecule comprises a reverse complementary oligonucleotide.
59. The method of claim 58, wherein toehold displacement is utilized to inhibit the hybridization of the oligonucleotides.
60. The method of any one of claims 1-59, wherein at least one of the nucleases is a restriction enzyme.
61. The method of claim 60, wherein the restriction enzyme is a type IIS restriction enzyme.
62. The method of any one of claims 1-61, wherein at least one of the nucleases is selected from the group consisting of Alwl; Bccl; Piel; BciVI; BmrI; HphI; HpryAv; Mboll; Mull; Sbfl; and PflMI.
63. The method of any one of claims 1-62, wherein the first nuclease is Alwl.
64. The method of any one of claims 1-63, wherein the second nuclease is Bccl.
65. The method of any one of claims 45-47, wherein the third nuclease is Sbfl.102 / 112#14588836v1H0824.70442WO0066. The method of any one of claims 45-47 or 65, wherein the fourth nuclease is PflMI.
67. The method of any one of claims 1-66, wherein each of the overhangs comprise an odd number of nucleotides.
68. The method of any one of claims 1-67, wherein each of the overhangs is one (1) nucleotide in length.
69. The method of any one of claims 1-68, wherein each of the overhangs is two (2), three (3), four (4), or five (5) nucleotides in length.
70. The method of any one of claims 1-69, wherein ligating the nucleic acids comprises contacting the nucleic acids with a T4 DNA ligase.
71. The method of any one of claims 1-70, wherein the nucleic acids are ligated in the presence of a crowding agent.
72. The method of claim 71, wherein the crowding agent is polyethylene glycol (PEG).
73. The method of claim 72, wherein the concentration of PEG is between 1% w / v to 50% w / v.
74. The method of claim 72 or 73, wherein the concentration of PEG is at least 25% w / v.
75. The method of any one of claims 1-74, wherein the nucleic acids are in the presence of glycerol during the ligation step.
76. The method of any one of claims 1-75, wherein the nucleic acids are in the presence of glycerol during the cleavage step.
77. The method of claim 75 or 76, wherein the concentration of glycerol is between 10% w / v to 80% w / v.103 / 112#14588836v1H0824.70442WO0078. The method of any one of claims 75-77, wherein the concentration of glycerol is at least 40% w / v.
79. The method of any one of claims 1-78, wherein one or both of the surfaces is glass.
80. The method of any one of clams 1-79, wherein one or both of the surfaces has a temperature of between 4 °C and 95 °C.
81. The method of any one of claims 1-80, wherein one or both of the surfaces undergo a temperature change from a first temperature to a second temperature, wherein the first temperature and the second temperature are both between 4 °C and 95 °C.
82. The method of any one of claims 1-81 further comprising providing the nucleic acids in liquid droplets on the surfaces, wherein the liquid droplets are merged during each ligation step.
83. The method of claim 82, wherein merging the liquid droplets is automated.
84. The method of claim 82 or 83, wherein the liquid droplets are loaded on the surfaces using automated pipetting.
85. The method of any one of claims 82-84, wherein each liquid droplet is 10 pL or more in volume.
86. The method of any one of claims 81-84, wherein each liquid droplet is spaced 50 pm apart.
87. The method of any one of claims 82-86, wherein the liquid droplets are in an array on a slide comprising an equal number of rows and columns of liquid droplets.
88. The method of claim 87, wherein the array comprises 256 rows of liquid droplets and 256 columns of liquid droplets.
89. The method of claim 87 or 88, wherein the array comprises 1,398 rows of liquid droplets and 1,398 columns of liquid droplets.104 / 112#14588836v1H0824.70442WO0090. The method of any one of claims 1-89, wherein a ligase is contacted with the nucleic acid comprising the first anchor and / or the nucleic acid comprising the second anchor using automated pipetting.
91. The method of any one of claims 2-90, wherein the nuclease contacted with the ligated nucleic using automated pipetting.
92. The method of any one of claims 10-91 further comprising compacting, condensing, circularizing, and / or supercoiling the nucleic acid comprising the target sequence.
93. The method of any one of claims 10-92 further comprising compacting, condensing, circularizing, and / or supercoiling the nucleic acid comprising the target sequence by:(i) contacting the nucleic acid comprising the target sequence with a Cas protein and a first single- stranded polynucleotide, wherein the target sequence comprises a first binding site for a recombinase and a second binding site for the recombinase, wherein the first single- stranded polynucleotide hybridizes to the first binding site for the recombinase;(ii) contacting the nucleic acid with the recombinase, thereby generating a recombination event at the second binding site for the recombinase;(iii) contacting the nucleic acid with a Cas protein and a second single- stranded polynucleotide that hybridizes to the second binding site for the recombinase; and (iv) contacting the nucleic acid with the recombinase, thereby generating a recombination event at the first binding site for the recombinase.
94. The method of claim 93 further comprising subjecting the nucleic acid to conditions that inhibit binding of the first polynucleotide and the first binding site for the recombinase before contacting the nucleic acid with the second polynucleotide in step (iii).
95. The method of claim 93 or 94 further comprising subjecting the nucleic acid to conditions that inhibit binding of the second polynucleotide and the second binding site for the recombinase after contacting the nucleic acid with the recombinase in step (iv).105 / 112#14588836v1H0824.70442WO0096. The method of any one of claims 93-95 further comprising washing the nucleic acid before, during, and / or after contacting the nucleic acid with the first single- stranded polynucleotide in step (i), the second single- stranded polynucleotide in step (iii), or both.
97. The method of any one of claims 93-96 further comprising washing the nucleic acid before, during, and / or after contacting the nucleic acid with the recombinase in step (ii), step (iv), or both.
98. The method of any one of claims 93-97 further comprising repeating steps (i)-(iv) one or more times.
99. The method of any one of claims 93-98, wherein the recombinase is a serine recombinase.
100. The method of any one of claims 93-98, wherein the recombinase is a tyrosine recombinase.
101. The method of any one of claims 93-98, wherein the recombinase is selected from the group consisting of: Bbxl; PhiC31; PaOl; Pa03; Ec03; Nm60; LI Int; Cre R32M; YR1; YR2; YR11; eXerl; eXer2; eXer3; Dn29; Si74; Bm99; Bt24; Cbl6; Fm04; uCb64; Ec03; Nm60; LI Int; and an engineered variant thereof with modified attP / attB recognition specificities.
102. The method of any one of claims 93-98, wherein the recombinase is Bxbl.
103. The method of any one of claims 93-98, wherein the recombinase is PhiC31.
104. The method of any one of claims 93-98, wherein the recombinase is PaOl.
105. The method of any one of claims 93-98, wherein the recombinase is Cre32M.
106. The method of any one of claims 93-98, wherein the recombinase is YR1.
107. The method of any one of claims 93-98, wherein the recombinase is Xerl.106 / 112#14588836v1H0824.70442WO00108. The method of any one of claims 93-107, wherein the first single- stranded polynucleotide is a single- stranded DNA.
109. The method of any one of claims 93-108, wherein the second single- stranded polynucleotide is a single- stranded DNA.
110. The method of any one of claims 93-109, wherein the Cas protein comprises a catalytically inactive Cas protein.
111. The method of any one of claims 93-110, wherein the Cas protein comprises Cas9.
112. The method of any one of claims 93-111, wherein the Cas protein comprises Cas 12.
113. The method of any one of claims 93-112, wherein the Cas protein is a bi- valent Cas protein.
114. The method of any one of claims 10-113 further comprising introducing the nucleic acid comprising the target sequence into a cell.
115. The method of claim 114, wherein the cell is an animal cell.
116. The method of claim 115, wherein the animal cell is a mammalian cell.
117. The method of claim 116, wherein the mammalian cell is a human cell.
118. A cell comprising the nucleic acid produced by the method of any one of claims 1-117.
119. The cell of claim 118, wherein the cell is an animal cell.
120. The cell of claim 119, wherein the animal cell is a mammalian cell.
121. The cell of claim 120, wherein the mammalian cell is a human cell.
122. A cell population comprising the cell of any one of claims 108-121.107 / 112#14588836v1H0824.70442WO00123. A surface comprising a binding molecule bound to a nucleic acid, wherein the nucleic acid comprises:two different anchors located at the terminal ends of the nucleic acid, wherein one of the anchors is bound to the binding molecule; andtwo different nuclease cleavage sites.
124. The surface of claim 123, wherein the anchors comprise a small molecule, a clickchemistry handle, a peptide, a protein, or an oligonucleotide.
125. The surface of claim 123 or 124, wherein one of the anchors comprise a maltose-binding protein and the binding molecule comprises maltose.
126. The surface of any one of claims 123-125, wherein one of the anchors comprise a polyHis tag and the binding molecule comprises Nickel-Nitriloacetic acid (NTA).
127. The nucleic acid of any one of claims 123-126, wherein the nucleic acid comprises four different nuclease cleavage sites.
128. The surface of any one of claims 123-127, wherein one or more of the nuclease cleavage sites are endonuclease cleavage sites.
129. The surface of claim 128, wherein the endonuclease cleavage sites comprise type IIS restriction enzyme cleavage sites.
130. The surface of any one of claims 123-129, wherein one or more of the nuclease cleavage sites are selected from the group consisting of: a Alwl cleavage site; a Bccl cleavage site; a Piel cleavage site; a BciVI cleavage site; a BmrI cleavage site; a HphI cleavage site; a HpryAv cleavage site; a Mboll cleavage site; a Mull cleavage site; a Sbfl cleavage site; and a PflMI cleavage site.
131. The surface of any one of claims 123-130, wherein the surface is glass.
132. A system comprising:108 / 112#14588836v1H0824.70442WO00(a) a surface comprising a first binding molecule bound to a first anchor on a nucleic acid comprising a cleavage site for a first nuclease; and(b) a surface comprising a second binding molecule bound to a second anchor on a nucleic acid comprising a cleavage site for a second nuclease,wherein the nucleic acid comprising the first anchor and the nucleic acid comprising the second anchor comprise reverse complementary overhangs.
133. The system of claim 132, wherein the first and second anchors comprise a small molecule, a click-chemistry handle, a peptide, a protein, or a polynucleotide.
134. The system of claim 132 or 133, wherein the first anchor comprises a maltose-binding protein tag and the first binding molecule comprises maltose.
135. The system of any one of claims 132-134, wherein the second anchor comprises a poly His tag and the second binding molecule comprises Nickel-Nitriloacetic acid (NTA).
136. The system of any one of claims 132-135, wherein the nucleic acid comprising the first anchor comprises a cleavage site for a third nuclease which is located between the first anchor and the cleavage site for the first nuclease.
137. The system of any one of claims 132-136, wherein the nucleic acid comprising the second anchor comprises a cleavage site for a fourth nuclease which is located between the first anchor and the cleavage site for the second nuclease.
138. The system of any one of claims 132-137, wherein one or more of the nuclease cleavage sites are endonuclease cleavage sites.
139. The system of claim 138, wherein the endonuclease cleavage sites comprise type IIS restriction enzyme cleavage sites.
140. The system of any one of claims 132-139, wherein one or more of the nuclease cleavage sites are selected from the group consisting of: a Alwl cleavage site; a Bccl cleavage site; a Piel cleavage site; a BciVI cleavage site; a BmrI cleavage site; a HphI cleavage site; a HpryAv109 / 112#14588836v1H0824.70442WO00cleavage site; a Mboll cleavage site; a Mnll cleavage site; a Sbfl cleavage site; and a PflMI cleavage site.
141. The system of any one of claims 132-140, wherein the overhangs comprise an odd number of nucleotides.
142. The system of any one of claims 132-141, wherein the overhangs comprise one (1) nucleotide in length.
143. The system of any one of claims 132-139, wherein the overhangs comprise two (2), three (3), four (4), or five (5) nucleotides in length.
144. The system of any one of claims 132-143, wherein the surface comprising the first binding molecule and / or the surface comprising the second binding molecule is glass.
145. The system of any one of claims 132-144, wherein the nucleic acids are provided in liquid droplets on the surfaces.
146. The system of claim 145, wherein each liquid droplet is 10 pL or more in volume.
147. The system of claim 145 or 146, wherein each liquid droplet is spaced 50 pm apart.
148. The system of any one of claims 145-147, wherein the system comprises one or more slides, wherein each slide comprises an array of the liquid droplets, wherein each array comprises an equal number of rows and columns of liquid droplets.
149. The system of claim 148, wherein the array comprises 256 rows of liquid droplets and 256 columns of liquid droplets.
150. The system of claim 148 or 149, wherein the array comprises 1,398 rows of liquid droplets and 1,398 columns of liquid droplets.
151. The system of any one of claims 132-150, wherein one or more of the nucleic acids comprising the first anchor further comprise a methylated nucleobase located in the overhang.110 / 112#14588836v1H0824.70442WO00152. The system of any one of claims 132-151, wherein one or more of the nucleic acids comprising the second anchor further comprise a methylated nucleobase located in the overhang153. The system of claim 151 or 152, wherein the methylated nucleobase is a guanine (G), an adenine (A), a cytosine (C), or a thymine (T) located in a sequence comprising AA, AC, AG, AT, CA, CC, CG, CT, GA, GC, GG, GT, TA, TC, TG, or TT.
154. The system of any one of claims 151-153, wherein the methylated nucleobase is N6-methyladenosine.
155. The system of any one of claims 151-154, wherein the methylated nucleobase is 5-methylcytosine.111 / 112#14588836v1H0824.70442WO00