Multiplexed scar-less assembly of DNA sequences
The method addresses the challenges of assembling long and complex DNA sequences by using computational segmentation and a fidelity evaluation module to identify high-fidelity overhang sets, achieving efficient, scalable, and accurate DNA assembly with 90% fidelity in a single reaction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- UNIV OF WASHINGTON
- Filing Date
- 2026-01-23
- Publication Date
- 2026-07-30
AI Technical Summary
Current methods for assembling long and complex polynucleotide sequences suffer from high error rates, scalability issues, and high costs, particularly for sequences rich in GC content or containing repetitive motifs, limiting their application in synthetic genomics and gene therapy.
A method involving computational segmentation and an iterative fidelity evaluation module to identify high-fidelity orthogonal overhang sets for simultaneous assembly of multiple long DNA sequences, using a computer system to optimize overhang selection and synthesis, enabling efficient and accurate assembly of target DNA sequences in a single reaction.
Achieves high-throughput, cost-effective, and scalable assembly of long DNA fragments with 90% fidelity, allowing for simultaneous assembly of multiple sequences in a single reaction, overcoming limitations of existing enzymatic and chemical synthesis techniques.
Smart Images

Figure US2026012384_30072026_PF_FP_ABST
Abstract
Description
MULTIPLEXED SCAR- LESS ASSEMBLY OF DNA SEQUENCESCROSS-REFERENCES TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 749395, filed January 24, 2025, the disclosure of which is incorporated herein by reference in its entirety.Field of Technology
[0002] The present disclosure relates to molecular biology, computational biology, and nucleic acid sequencing. More particularly the disclosure relates to methods, compositions, systems, and computer-implemented tools for applications involving assembly of long nucleic acid constructs.STATEMENT OF GOVERNMENT LICENSE RIGHTS
[0003] This invention was made with government support under Grant No. DP5OD036167, awarded by the National Institutes of Health (NIH). The government has certain rights in the invention.BACKGROUND
[0004] The engineering of biological systems in synthetic biology often requires de novo synthesis or assembly of long and complex polynucleotide sequences. However, this has been a particularly challenging problem using conventional polynucleotide synthesis and cloning I assembly techniques. Conventional methods generating chemically-synthesized DNA fragments have very high sequence error rates. Approaches to this problem have included modification of DNA sequences to optimize conventional DNA synthesis shortcomings, as well as development of various technologies to assemble short polynucleotide sequences. However, these approaches exhibit limitations, such as accommodating polynucleotides of less than 300 base pairs, requiring complex tagging and sequencing, relying on large sets of barcodes, involving intricate and time-consuming steps, or failing to synthesize diverse polynucleotide sequences. Furthermore, assembly, synthesis, and cloning of polynucleotides suffer from a high probability of errors.
[0005] More recently, enzymatic polynucleotide synthesis techniques have gained traction, showing promise for de novo assembly of longer and more complex polynucleotides with lower error rates. However, as the length of the DNA sequence increases, the error rate during synthesis and assembly tends to climb. These errors can manifest as insertions, deletions, or substitutions within the sequence, compromising thefunctionality of the resulting assembled gene. Maintaining high fidelity across extended DNA sequences is thus a paramount challenge.
[0006] After achieving high accuracy assembly of a longer and more complex polynucleotide via enzymatic methods, the resulting de novo synthesized polynucleotide must be amplified or cloned with a correspondingly accurate fidelity. Scaling up the synthesis process to assemble large quantities of high-quality, long DNA fragments is another hurdle. Traditional methods often struggle with maintaining consistency and yield as the scale of production expands.
[0007] The cost of synthesizing long genes can be prohibitively high due to the intricate processes and extensive purification steps required. Balancing cost and quality is crucial for making gene synthesis accessible to a broader range of research and industrial applications. Certain DNA sequences, particularly those rich in GC content or containing repetitive motifs, are inherently more difficult to assemble. These sequences can lead to secondary structures or precipitation, further complicating the assembly process. Current in vitro methods of amplifying de novo synthesized long and complex DNA suffer from higher error rates or are time-consuming and difficult to perform with high accuracy, such as for sequences with repeat or GC-rich sequences.
[0008] Overall, there is no cost-effective, fast, scalable, accurate, and flexible source of preparing long or complex synthetic DNA fragments, hindering many applications in enzyme design, synthetic genomics and gene therapy. While short DNA molecules can be synthesized on microarrays in large pools at low cost, these are limited to ~300nt. Synthesizing and assembling large gene fragments is fraught with complexities that significantly elevate the technical demands compared to oligonucleotide synthesis. Newer enzymatic DNA synthesis approaches show promise for long fragment synthesis and assembly but are not yet scalable to thousands of designs. DropSynth, is a promising droplet-based method to assemble thousands of DNA constructs (~1.5kb) directly from microarray-derived oligos. See Sidore, Angus M et al. “DropSynth 2.0: high-fidelity multiplexed gene synthesis in emulsions.” Nucleic acids research vol. 48,16 (2020): e95. However, adoption is limited by the low efficiency of multi-fragment assembly, error rate from input oligos, and specialized equipment or know-how required for generating barcoded beads for droplet microfluidics.
[0009] Therefore, a need exists in the art for a high-throughput method for assembling a plurality of long DNA fragments from pools of short microarray-derived synthetic DNA oligos in a cost-effective, accurate, scalable, and efficient manner.SUMMARY
[0010] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This summary is not intended to identify key features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0011] Aspects of the disclosure relate to methods for identifying high-fidelity overhang sequences for a plurality of target DNA sequences to be assembled in a single reaction simultaneously and for a plurality of such reactions simultaneously. The inventors have advantageously leveraged a fidelity evaluation module to simultaneously identify a plurality of high-fidelity overhangs from a pool of oligonucleotide sequences to enable efficient and accurate assembly of target DNA sequences for high-throughput assembly of a plurality of long target DNA sequences per reaction and across multiple reactions simultaneously.
[0012] Aspects of the disclosure provide a method for identifying a plurality of high-fidelity orthogonal overhang sets for scar-less assembly of a plurality of target DNA sequences in a single reaction. In some embodiments, the method is effective in simultaneously identifying the plurality of high-fidelity orthogonal overhang sets across a plurality of assembly reactions, where each of the plurality of assembly reactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
[0013] In some embodiments, the method comprises selecting a plurality of target DNA sequences to be assembled in a single reaction; and computationally segmenting the plurality of target DNA sequences to be assembled in the single reaction to identify high-fidelity orthogonal overhang sets of DNA sequences along a length of each of the plurality of target DNA sequences to be assembled in the single reaction.
[0014] In certain embodiments, the computational segmentation comprises: (i) defining a plurality of split windows along the length of each of the plurality of target DNA sequences to be assembled in the single reaction; (ii) identifying potential overhang sets of sequences within each of the split windows; and (iii) using an iterative fidelity evaluation module configured to optimize selection of high-fidelity orthogonal overhang sets of DNA sequences.
[0015] In certain embodiments, the step of defining the split window for each of the plurality of target sequences to be assembled is based on the length of the target sequence and the length of the single-stranded DNA fragments representing the target DNA sequence. In some embodiments, each split window comprises at least 12 bp long sequences along the length of each of the plurality of target DNA sequences to be assembled.
[0016] In certain embodiments, the step of identifying potential overhang sets of sequences within each of the split windows is by comparing the sequences within the split window to sequences in an overhang reference dataset.
[0017] In some embodiments, a first iteration of the fidelity evaluation module comprises:(a) creating an initial population of overhang sets of DNA sequences comprising the potential overhang sets of sequences identified in step (ii); (b) selecting a parent population of high-fidelity overhang sets of sequences, wherein the selecting comprises weighing fidelity of the initial population of overhang sets of sequences created in step (a) against fidelity data of a reference overhang dataset; (c) randomly selecting overhang sets of sequences in the parent population to swap / cross within the parent population and crossing the overhang sets of sequences to obtain a first generation of overhang sets of sequences; (d) randomly mutating a subset of the first generation of overhang sets of sequences; (e) weighing fidelity of the resulting overhang sets of sequences from step (d) against the reference overhang dataset to identify high-fidelity and low-fidelity overhang sets of sequences; and (f) providing the high-fidelity overhang sets of sequences identified in step (e) as input for subsequent iterations of the fidelity evaluation module to repeat steps (a) to (e).
[0018] In some embodiments, the computational segmentation is performed using a computer system comprising one or more processors; system memory; and one or more computer readable storage media having stored thereon computer-executable instructions that when executed by the one or more processors, cause the computer system to execute steps (a) -(f).
[0019] In certain embodiments, the method further comprises randomly shuffling the low-fidelity overhang sets of sequences identified in step (e) and providing as input for subsequent iterations of the fitness evaluation module to repeat steps (a) to (e). In an embodiment, the selected high-fidelity overhangs reduce the incorrect ligation events relative to randomly selected overhangs.
[0020] In some embodiments, the method further comprises synthesizing an oligo pool comprising a plurality of single stranded DNA fragments representing each of the plurality of target DNA sequences to be assembled in the single reaction. In certain embodiments, each of the plurality of single-stranded DNA fragments comprise the high-fidelity orthogonal overhang sets of sequences identified for each of the plurality of target DNA sequences to be assembled in the single reaction in step (iii) on each end. In some embodiments, the method further comprises adding: (i) a type Ils restriction site; and (ii) an inner and an outer unique adaptor pair for PCR amplification to each end of the plurality of single stranded DNA fragments.
[0021] In some embodiments, the plurality of single- stranded DNA fragments are synthesized by chemical synthesis, enzymatic DNA synthesis or by assembled DNA products from another assembly method. In some embodiments, the plurality of singlestranded DNA fragments are synthesized by chemical synthesis.
[0022] In some embodiments, the method is performed simultaneously across a plurality of assembly reactions, and wherein each of the plurality of assembly reactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
[0023] In some embodiments, the length of the plurality of target DNA sequences to be assembled is selected from 600 bp, 800 bp, 1000 bp, 1200bp, 1400 bp, 1600 bp, 1800 bp, and 2000bp. In an embodiment, the length of the plurality of DNA sequences is 600 bp and the method achieves 11 assemblies in a single reaction. In some embodiments, the length of the plurality of DNA sequences is 800 bp and the method achieves 7 assemblies in a single reaction. In an embodiment, the length of the plurality of DNA sequences is 1000 bp and the method achieves 5 assemblies in a single reaction. In some embodiments, the length of the plurality of DNA sequences is 1200 bp and the method achieves 4 assemblies in a single reaction. In certain embodiments, the length of the plurality of DNA sequences is 1400 bp- 1600 bp and the method achieves 3 assemblies in a single reaction. In some embodiments, the length of the plurality of DNA sequences is 1800 bp- 2000 bp and the method achieves 2 assemblies in a single reaction.
[0024] In an embodiment, the method is effective in achieving assemblies yielding correct constructs at a fidelity of about 90%. In certain embodiments, the method identifies high-fidelity orthogonal overhang sets that are not self-incompatible or cross-complementary.
[0025] In certain aspects the disclosure provides in vitro method of scar-less assembly of a plurality of target DNA sequences simultaneously in a single reaction. In some embodiments, the method is performed simultaneously across a plurality of assembly reactions, wherein each of the plurality of assembly reactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
[0026] In an embodiment, the method comprises (a) computationally segmenting each of the plurality of target DNA sequences to identify high-fidelity orthogonal overhang sets of sequences along the length of each of the plurality of target DNA sequences to be assembled in a single reaction, wherein the identified high-fidelity overhang sets of sequences are not self-incompatible or cross-complementary; (b) synthesizing an oligo pool comprising a plurality of single stranded DNA fragments representing each of the plurality of target DNA sequences to be assembled in the single reaction, wherein the plurality of single-stranded DNA fragments is obtained by chemical or enzymatic synthesis, and wherein each of the plurality of single stranded DNA fragments comprise the high-fidelity orthogonal overhang sets of sequences identified in step (a) on each end; (c) adding a type Ils restriction site to each end of each of the plurality of DNA fragments synthesized in step (b); (d) adding an inner and an outer unique adaptor pair for PCR amplification to each end of the plurality of DNA fragments of step (c); (e) amplifying the plurality of DNA fragments; (f) assembling multiple longer sequences per reaction from the amplified DNA fragments using type Ils restriction enzyme cutting and ligation; (g) inserting the amplified plurality of DNA fragments into a plurality of compatible destination vectors; and (h) transforming competent host cells with the destination vectors for assembling the plurality of target DNA sequences. In some embodiments, the outer unique adaptor pair is used to amplify oligos for the entire high-throughput reaction, while the inner unique adaptor pair is used to amplify oligos for each individual assembly reaction. In some embodiments, the type Ils restriction site is selected from a site specific for a restriction enzyme selected from Bsal, BsmBI, Esp3I, SapI, PaqCI, Bbslb.
[0027] In some embodiments, the computational segmentation comprises: (i) defining a plurality of split windows along the length of each of the plurality of target DNA sequences to be assembled in a single reaction; (ii) identifying potential overhang sets of sequences within each of the split windows; and (iii) using an iterative fidelity evaluation module to optimize selection of high-fidelity orthogonal overhang sets of sequences.
[0028] In certain embodiments, the step of defining the split window for each of the plurality of target sequences to be assembled is based on the length of the target sequence and the length of the single-stranded DNA fragments representing the target DNA sequence. In some embodiments, each split window comprises at least 12 bp long sequences along the length of each of the plurality of target DNA sequences to be assembled.
[0029] In certain embodiments, the step of identifying potential overhang sets of sequences within each of the split windows is by comparing the sequences within the split window to sequences in an overhang reference dataset.
[0030] In some embodiments, a first iteration of the fidelity evaluation module comprises :(a) creating an initial population of overhang sets of sequences comprising the potential overhang sets of sequences identified in step (ii); (b) selecting a parent population of high-fidelity overhang sets of sequences, wherein the selecting comprises weighing fidelity of the initial population of overhang sets of sequences created in step (a) against fidelity data of a reference overhang dataset; (c) randomly selecting overhang sets of sequences in the parent population to cross within the parent population and crossing the overhang sets of sequences to obtain a first generation of overhang sets of sequences; (d) randomly mutating a subset of the first generation of overhang sets of sequences; (e) weighing fidelity of the resulting overhang sets of sequences from step (d) against the reference overhang dataset to identify high-fidelity and low-fidelity overhang sets of sequences; and (f) providing the high-fidelity overhang sets of sequences identified in step (e) as input for subsequent iterations of the fidelity evaluation module to repeat steps (a) to (e).
[0031] In some embodiments, the computational segmentation is performed using a computer system comprising one or more processors; system memory; and one or more computer readable storage media having stored thereon computer-executable instructions that when executed by the one or more processors, cause the computer system to execute steps (a) -(f).
[0032] In some embodiments, the method further comprises randomly shuffling the low-fidelity overhang sets of sequences identified in step (e) and providing as input for subsequent iterations of the fitness evaluation module to repeat steps (a) to (e).
[0033] In some embodiments, the length of the plurality of target DNA sequences is selected from 600 bp, 800 bp, 1000 bp, 1200bp, 1400 bp, 1600 bp, 1800 bp, and 2000bp.
[0034] In certain embodiments, the length of the plurality of target DNA sequences is 600 bp and the method achieves 11 assemblies in a single reaction. In an embodiment, the length of the plurality of target DNA sequences is 800 bp and the method achieves 7 assemblies in a single reaction.
[0035] In some embodiments, the method is effective in achieving assemblies with about 90% fidelity.
[0036] Aspects of the disclosure provide a system for assembling a plurality of long DNA sequences in a single reaction. In some embodiments, the system comprises (a) a processor configured to identify candidate high-fidelity overhang sequences for each of the plurality of long DNA sequences: and (b) an iterative fidelity evaluation module configured to identify and select high-fidelity overhang sequences for each of the plurality of long DNA sequences to be assembled in the single reaction. In some embodiments, the high-fidelity overhang sequences achieve an assembly fidelity of about 90%. In certain embodiments, the system is configured to generate assembly designs for a plurality of reactions simultaneously, where each of the plurality of reactions comprises a plurality of target DNA sequences.
[0037] In certain aspects, the disclosure provides a non-transitory computer-readable medium storing instructions that, when executed by a processor, causes the processor to perform the methods disclosed herein.DESCRIPTION OF THE DRAWINGS
[0038] The foregoing aspects and many of the attendant advantages of this disclosure will become more readily appreciated as the same become better understood by reference to the following detailed description, when taken in conjunction with the accompanying drawings, wherein:
[0039] FIG. 1 depicts a flowchart of Multiplex scarless assembly protocol in accordance with certain aspects of the disclosure. Prior to ordering single-stranded oligos, target DNA design sequences are defined as a specific number of sequences per reaction. The design process involves defining “split windows”, comprising at least about 12 bp, along the length of each of the target DNA sequence and employing a fidelity evaluation module according to aspects of the disclosure to identify high-fidelity overhangs within those split windows.
[0040] FIG.2 shows the fitness evaluation module workflow according to certain aspects of the disclosure. Optimized DNA fragments are written to FASTA files, andmetadata, including pool assignments and fidelity scores, is saved in CSV files. This ensures both sequence integrity and traceability for downstream applications.
[0041] FIG. 3 shows a predicted fidelity heatmap depicting the relationship between the number of designs per reaction and the average fidelity for sequences of varying lengths. The predicted fidelity is determined based on a comparison to a reference fidelity overhang dataset.
[0042] FIG.4 shows distribution of sequence lengths for a 6821 design library.
[0043] FIG.5 shows nanopore sequencing read count distribution for all designs from the 6821 library assembled using the multi-GGA protocol. Counts here reflect those reads that are able to map to expected designs, without further consideration of unexpected structural variants.
[0044] FIG. 6 shows uniformity and capacity of the methods disclosed herein to assemble 3-8 fragments at 22plex-6plex, respectively.
[0045] FIG. 7 shows -multiplex Golden Gate Assembly using the methods disclosed herein is highly accurate even at longer lengths, with correct reads = <5% length difference from target and <10% mutation rate. Correct rate = Number of correct rcads / Numbcr of reads aligned to any design.
[0046] FIG. 8 shows that assembly rates achieved by the methods disclosed herein are much higher than predicted fidelity calculations. Dotted line = predicted fidelity.
[0047] FIG. 9 shows the efficiency achieved by the methods disclosed herein in assembling target DNA sequences. Non-assembled rate = Number of non-assembled read / Total reads.
[0048] FIG. 10 shows proportion of mis-assemblies is low as measured by nanopore sequencing for 3-8 fragment assemblies. Un assembled vector fraction can be cleaned up via bacterial transformation.
[0049] FIGS. 11A-11B show pools are not highly skewed by some assemblies.FIG. 11A shows the Gini Coefficient of multiplex assembled sequences in one pool. FIG.11B depicts the abundance and correct rate as in FIG. 7 of all members in 2 large pool libraries of approximately 1500 designs each, one from 3 fragment assembly and 8 fragment assembly respectively. Together these data demonstrate that assembled pools are not highly skewed by some assemblies.
[0050] FIG. 12 shows an estimate of cost per design as a function of number of assemblies per reaction and total library size for a 3-fragment assembly.DET AILED DESCRIPTION
[0051] Definitions
[0052] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of any embodiment. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0053] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0054] Unless specifically stated or obvious from context, as used herein, the term “about” in reference to a number or range of numbers is understood to mean the stated number and numbers + / - 10% thereof, or 10% below the lower listed limit and 10% above the higher listed limit for the values listed for a range.
[0055] As used herein, the term “Target sequence" means that the sequence of the polymer is known and chosen before synthesis or assembly of the polymer. In particular, various aspects of the disclosure are described herein primarily with regard to the preparation and assembly of nucleic acids molecules, the sequence of the oligonucleotide or polynucleotide being known and chosen before the assembly of the nucleic acid molecules.
[0056] The term “nucleic acid” as used herein refers broadly to any type of coding or non-coding, long polynucleotide or polynucleotide analog.
[0057] As used herein, the term “complementary” refers to the capacity for precise pairing between two nucleotides. If a nucleotide at a given position of a nucleic acid is capable of hydrogen bonding with a nucleotide of another nucleic acid, then the two nucleic acids are considered to be complementary to one another (or, more specifically in some usage, “reverse complementary”) at that position. Complementarity between two single-stranded nucleic acid molecules may be “partial,” in which only some of the nucleotides bind, or it may be complete when total complementarity exists between the single-stranded molecules. The degree of complementarity between nucleic acid strandshas significant effects on the efficiency and strength of hybridization between nucleic acid strands.
[0058] “Hybridization” and “annealing” refer to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonding between the bases of the nucleotide residues. The term “hybridized” as applied to a polynucleotide is a polynucleotide in a complex that is stabilized via hydrogen bonding between the bases of the nucleotide residues. The hydrogen bonding may occur by Watson Crick base pairing, Hoogsteen binding, or in any other sequence specific manner. The complex may comprise two strands forming a duplex structure, three or more strands forming a multi stranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction may constitute a step in a more extensive process, such as the initiation of a PCR or other amplification reactions, or the enzymatic cleavage of a polynucleotide by a ribozyme. A first sequence that can be stabilized via hydrogen bonding with the bases of the nucleotide residues of a second sequence is said to be “hybridizable” to the second sequence. In such a case, the second sequence can also be said to be hybridizable to the first sequence. In many cases a sequence hybridized with a given sequence is the “complement” of the given sequence.
[0059] In general, a “target nucleic acid” is a desired molecule of predetermined sequence to be assembled, and any fragment thereof.
[0060] The term “primer” refers to an oligonucleotide that is capable of hybridizing (also termed “annealing”) with a nucleic acid and serving as an initiation site for nucleotide (RNA or DNA) polymerization under appropriate conditions (i.e. in the presence of four different nucleoside triphosphates and an agent for polymerization, such as DNA or RNA polymerase or reverse transcriptase) in an appropriate buffer and at a suitable temperature. The appropriate length of a primer depends on the intended use of the primer. In some instances, primers are at least 7 nucleotides long. In some instances, primers range from 7 to 70 nucleotides, 10 to 30 nucleotides, or from 15 to 30 nucleotides in length. In some instances, primers are from 30 to 50 or 40 to 70 nucleotides long. Oligonucleotides of various lengths as further described herein are used as primers or precursor fragments for amplification and / or gene assembly reactions. In this context, “primer length” refers to the portion of an oligonucleotide or nucleic acid that hybridizes to a complementary “target” sequence and primes nucleotide synthesis. Short primer molecules generally require cooler temperaturesto form sufficiently stable hybrid complexes with the template. A primer need not reflect the exact sequence of the template but must be sufficiently complementary to hybridize with a template. The term “primer site” or “primer binding site” refers to the segment of the target nucleic acid to which a primer hybridizes.
[0061] The term “overhang” as used herein refers to a single stranded DNA sequence generated by cleavage of a Type IIS restriction enzyme.
[0062] The term “predicting overhangs” as used herein refers to computational, algorithmic, or rule -based determination of overhang sequences prior to assembly.
[0063] The term “large DNA sequences” as used herein refers to a DNA sequence at least 2 kb, 5 kb, 10 kb, or greater.
[0064] As used herein, the term “split windows” refers to defined sequence regions located at the junctions between adjacent DNA fragments of a target DNA sequence to be assembled after an in silico partitioning of the target DNA sequence into smaller fragments that are easier to process, synthesize, and assemble.
[0065] In certain embodiments, a target DNA sequence is computationally divided in silico into multiple shorter fragments or split windows suitable for DNA ordering and downstream assembly. The computational segmentation is performed for a plurality of target DNA sequences per reaction and for multiple such reactions simultaneously. In some embodiments, each split window comprises a sequence region of approximately at least about 12 to about 50 base pairs spanning the connection site between two neighboring fragments to be assembled.
[0066] The step of defining the split window for each of the plurality of target sequences to be assembled is based on or is a function of the length of the target sequence to be assembled and the length of the single-stranded DNA fragments required for assembling or representing the target DNA sequence. In some embodiments, each split window comprises at least 12 bp long sequences along the length of each of the plurality of target DNA sequences to be assembled.
[0067] In certain embodiments, the step of identifying potential overhang sets of sequences within each of the defined split windows is by comparing the sequences within the split window to sequences in an overhang reference dataset.
[0068] The methods described herein analyze these split window regions to identify potential high-fidelity overhangs or compatible junction sequences, which are then used to determine optimal fragment boundaries and enable accurate, efficient assembly ofthe full-length DNA construct from the ordered fragments. The potential overhang sets are identified by comparison to a reference overhang dataset.
[0069] The term “fidelity” as used herein refers to the degree of exactness or match of a DNA sequence assembled by the methods disclosed herein relative to the target DNA sequence. Fidelity also refers to the predicted ability of the DNA molecules to assemble and to specifically ligate as intended, which is calculated relative to an overhang-to-overhang ligation dataset. Comprehensive overhang to overhang datasets (or ligase fidelity datasets or reference datasets) profile the efficiency and accuracy of DNA endjoining, for example for Golden Gate Assembly, which uses Type IIS restriction enzymes to create 3- or 4-base overhangs.
[0070] The term “high-fidelity overhang sequences” refers to overhang sequences that reduce incorrect ligation events relative to randomly selected overhangs. The term “high-fidelity overhang sequences” also refers to overhang sequences effective in achieving assemblies yielding correct constructs / assemblies at a fidelity at least of about 90%.Target DNA assembly
[0071] CuiTcnt method for assembling DNA longer than 3 kb requires construction of plasmid, fosmid, or BAG libraries. Longer DNA sequences are typically assembled through the “splicing” of shorter oligonucleotides. This process involves synthesizing fragments of less than 100 nucleotides (typically 80nt) with homologous ends, followed by assembly using techniques such as PCA or Gibson assembly. Nevertheless, this approach encounters difficulties with challenging sequences, particularly those containing long repetitive elements or unbalanced base composition.
[0072] The limitations of traditional restriction enzyme and ligase cloning — namely, its multi-step nature, dependency on available restriction sites, and propensity to leave unwanted scar sequences — spurred the development of more efficient, flexible, and cost-effective methods such as Golden Gate Assembly, TA / TOPO-TA®, Gateway®, and Exonuclease-based Seamless Cloning (ESC) techniques (Sorida and Bonasio, 2023a; Biro et al., 2024; Sorida and Bonasio, 2023b). The one -pot assembly of large DNA constructs from smaller component parts is a technology in modern synthetic biology, with common in vitro methods dependent on high-fidelity ligation steps to produce the desired constructs. In restriction enzyme-dependent assembly methods such as BioBricks andGolden Gate cloning, the assembly of large constructs is achieved by the joining of multiple DNA fragments linked by short overhangs.
[0073] Golden Gate Assembly (GGA) and other Type IIS restriction enzymedependent DNA assembly methods enable rapid construction of genes and operons through one-pot, multi-fragment assembly, with the ordering of parts determined by the ligation of Watson-Crick base-paired overhangs. See Engler C, Kandzia R, Marillonnet S (2008) A One Pot, One Step, Precision Cloning Method with High Throughput Capability. PLoS ONE 3(11): e3647. However, GGA has its limitations, for example, to avoid undesired digestion, the Type IIS site should not be present within the fragment to be assembled. If the target or gene of interest or destination vector contains multiple internal restriction sites, alternative methods need to be used. Moreover, this protocol is still quite low throughput, with only one assembly attempted per golden gate reaction. See Lund, Sean et al. “Highly Parallelized Construction of DNA from Low-Cost Oligonucleotide Mixtures Using Data-Optimized Assembly Design and Golden Gate.’’ ACS synthetic biology vol. 13,3 (2024): 745-751.
[0074] Another limitation of GGA is the design of the overhang sequences. Although there arc theoretically 256 distinct flanking sequences, sequences that differ by only one base may result in unintended ligation products. Ligation of mismatched overhangs leads to erroneous assembly, and low-efficiency Watson Crick pairings can lead to truncated assemblies. Using sets of empirically vetted, high-accuracy junction pairs avoids this issue but limits the number of parts that can be joined in a single reaction. Such methods are impractical for assembly of large DNA sequences, such as full-length genes, operons, genomic regions, or constructs exceeding several kilobases.
[0075] To ensure high-fidelity assembly in Golden Gate and derived methods, several rules of thumb have been adopted to minimize the risk of ligating imperfectly basepaired partners during an assembly reaction. See Engler, C. and Marillonnet, S. (2014) Golden Gate cloning. Methods Mol. Biol. 1116, 119- 131, DOI: 10.1007 / 978-1-62703-764-8_9. Following these rules limits the number of four-base overhangs that can be used in a single pot and is particularly constraining when junction sequences are restricted (e.g., when assemblies must break within coding sequences). Several Golden Gate based assembly systems (e.g., MoClo, Golden Braid, Mobius Assembly, Loop Assembly, MIDAS) have further restricted the number of overhangs to standardized, reliable sets to improve efficiency and fidelity.
[0076] Thus, while very large DNA constructs can be produced from successive hierarchical assembly rounds by these methods, the number of fragments that can be assembled in a single pot is limited by the number of allowable overhang pairs (typically six to eight). While these sets have been vetted empirically, there has been a lack of informatics-driven efforts to choose overhang junctions from comprehensive ligase fidelity data.
[0077] The methods of the present disclosure expand the flexibility of Golden Gate and similar methods through the identification of a large number of high-fidelity overhang sets allowing for high-throughput reactions, where each reaction is capable of assembly of many more fragments / genes in a single reaction and further is capable of doing so across a plurality of reactions at the same time.
[0078] Methods and compositions described herein allow assembly of a plurality of large nucleic acid target molecules with a high degree of confidence as to sequence integrity. The target molecules are assembled from precursor nucleic acid fragments.
[0079] Aspects of the disclosure provide high-throughput methods for assembling a plurality of long DNA fragments (target DNA sequences) from short-microarray-dcrivcd synthetic DNA oligos by combining computational segmentation of a plurality of target DNA sequences to be assembled for the purposes of identifying high-fidelity overhang sets of sequences. The methods disclosed herein use a fidelity evaluation module to predict high-fidelity overhang sequences for efficient assembly of the plurality of target large DNA sequences with high fidelity in a single reaction across a plurality of assembly reactions simultaneously.
[0080] In certain aspects the fidelity evaluation module utilizes genetic algorithms. The present disclosure provides a systematic, predictive approach to overhang design that improves efficiency, fidelity, and scalability of existing applications involving assembly of large DNA constructs.Genetic Algorithms
[0081] Genetic algorithms are heuristic techniques that can be used to tackle the DNA Fragment Assembly problem. General steps applying genetic algorithms are as follows: (a) The algorithm randomly generates a pool of solutions; (b) It screens for solutions with a fitness function; (c) Mutation and crossover operations are performed on good solutions to create next generation solutions. Genetic algorithms (GA) have been used to solve complex combinatorial and organizational problems with many variants, byemploying analogy with Nature’s evolution. Genetic algorithms were introduced for the first time in the work of John Holland (Holland 1975). They were further developed by him and other researchers (Goldberg 1989).
[0082] The general steps in a fitness evaluation module of the methods disclosed herein are:
[0083] Generating an initial population of orthogonal overhang sets.
[0084] Evaluating the fitness of each overhang (the accuracy of each model) using a fitness function.
[0085] Select a subset of overhangs based on their fitness to generate a parent population.
[0086] Apply a crossover procedure on the selected overhangs / parent population to create a new generation of a population.
[0087] Apply mutation to directly change low fitness overhangs.
[0088] Continue with the previous procedure until a desired solution (with a desired fitness) is obtained, or the run time is over.
[0089] The fitness evaluation module of the present disclosure comprising genetic algorithms shows a great deal of parallelism. Thus, each of the branches of the search tree for best overhang set can be utilized in parallel with the others. This allows for an easy realization of the genetic algorithms on parallel architectures. Without being bound by theory, it is believed that solutions evolve better for the DNA Fragment Assembly problem from one generation to the next. Having a random initial population, an appropriate fitness function, and suitable mutation and crossover operations allow the fitness evaluation module to converge to good solutions for DNA Fragment Assembly problems.
[0090] Selection of the best overhang set / parent population to continue the process of optimization is based on fitness. A common approach is proportional fitness (roulette wheel selection), e.g., if a model Mx is twice as good as another one, its probability of being selected for the crossover process is twice as high. Roulette wheel selection gives a chance to an overhang set according to their fitness evaluation.
[0091] A feature of the selection procedure is that fitter overhang set (models Mx with higher accuracy) are more likely to be selected.
[0092] The selection procedure can involve also keeping the best overhang set from the previous generation. This operation is called elitism. After the best overhang setsare selected from a population of models, a cross over operation is applied between these individuals.
[0093] Different cross-over operations can be used: one-point cross-over; three-point cross over, or more.
[0094] Mutation can be performed in the following ways: For a binary string, just randomly ‘flip’ a bit. For a more complex “genes” and “chromosomes”, randomly select a gene and change its value.
[0095] Some methods just use mutation (no crossover, e.g. evolutionary strategies). Normally, however, mutation is used to search in a “local search space”, by allowing small changes in the “genotype” (and therefore hopefully in the “phenotype”).
[0096] In certain embodiments, the disclosure provides methods for predicting overhang sequences using a fitness evaluation module to identify overhang sequences suitable for multiplexed scar-less assembly of a plurality of large DNA sequences. The methods may be performed in silico, prior to physical assembly, and may incorporate biological, biochemical, and computational constraints.
[0097] Overhang prediction may be based on one or more of the following criteria: predicted fidelity, sequence orthogonality, thermodynamic properties, secondary structure avoidance, restriction site context, ligation, bias modeling, fragment size and order constraints.
[0098] In some aspects, the disclosure provides a high throughput plate-based protocol for assembling microarray-derived oligos into larger DNA constructs with standard molecular biology reagents using the methods disclosed herein, thereby ensuring ease of adoption, even in resource limited settings, and compatibility with automation.
[0099] Provided herein are comprehensive, high-throughput, and cost-effective methods, systems, and compositions, for multiplexed long-fragment DNA assembly using scar-less assembly of short DNA sequences. The disclosed methods streamline the entire workflow, from computational sequence segmentation to experimental assembly, enabling the efficient synthesis and assembly of thousands of target DNA sequences simultaneously in a single reaction. This approach offers significant advantages in terms of scalability, fidelity, and ease of implementation, making it accessible even in resource-limited settings.
[0100] In one aspect, the present disclosure provides a method for identifying a plurality of high-fidelity orthogonal overhang sets in a plurality of target DNA sequences for scar-less assembly of the plurality of target DNA sequences in a single reaction. Insome embodiments, the method comprises selecting the plurality of target DNA sequences to be assembled in a single reaction; and computationally segmenting the plurality of target DNA sequences to be assembled to identify high-fidelity orthogonal overhang sets of sequences along a length of each of the plurality of target DNA sequences. In certain aspects, the method is performed simultaneously across a plurality of assembly reactions, wherein each of the plurality of assembly reactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
[0101] In some embodiments, the computational segmentation comprises defining a plurality of split windows along the length of each of the plurality of target DNA sequences to be assembled in the single reaction; identifying potential overhang sets of sequences within each of the split windows; and using an iterative fidelity evaluation module to optimize selection of high-fidelity orthogonal overhang sets of sequences.
[0102] In an embodiment, the split windows are defined based on the length of each of the plurality of target DNA sequences to be assembled and the length of each of the plurality of single stranded DNA fragments representing each of the plurality of target DNA sequences to be assembled in the single reaction.
[0103] In certain embodiments, the step of identifying potential overhang sets of sequences within each of the split windows is by comparing the sequences within the split window to sequences in an overhang reference dataset.
[0104] In an embodiment the target DNA is split in silico into roughly even fragment size (-200 nucleotides) and the size of the split window is set as at least 12bp representing the border sequence between two segments of the target DNA. The size of the split window is based on the length of the target sequence and the single-stranded oligo length to be synthesized for assembling the target sequence. Potential overhangs that are available within each of the split windows are then identified. In an embodiment, the identification of potential overhangs is in comparison to a reference overhang data set predicted to yield correctly ligated products in an assembly reaction. In some embodiments, the dataset is the New England Biolab Golden Gate Assembly (NEB GGA) overhang dataset.
[0105] In an embodiment, a first iteration of the fidelity evaluation module comprises; (a) creating an initial population of overhang sets of sequences comprising the potential overhang sets of sequences identified in the split windows; (b) selecting a parent population of high-fidelity overhang sets of sequences, where the selecting comprisesweighing fidelity of the initial population of overhang sets of sequences against fidelity data of a reference overhang dataset; (c) randomly selecting overhang sets of sequences in the parent population to swap / cross within the parent population and crossing the overhang sets of sequences to obtain a first generation of overhang sets of sequences; (d) randomly mutating a subset of the first generation of overhang sets of sequences; (e) weighing fidelity of the resulting overhang sets of sequences against the reference overhang dataset to identify high-fidelity and low-fidelity overhang sets of sequences; and (f) providing the high-fidelity overhang sets of sequences thus identified as input for subsequent iterations of the fidelity evaluation module to repeat steps (a) to (e).
[0106] In certain embodiments, the method further comprises randomly shuffling the low-fidelity overhang sets of sequences identified in step (e) and providing as input for subsequent iterations of the fidelity evaluation module to repeat steps (a) to (e).
[0107] In a particular embodiment, the computational segmentation is performed using a computer system comprising one or more processors; system memory; and one or more computer readable storage media having stored thereon computer-executable instructions that when executed by the one or more processors, cause the computer system to execute steps (a) -(f).
[0108] In some embodiments, the method further comprises synthesizing an oligo pool comprising a plurality of single stranded DNA fragments representing each of the plurality of target DNA sequences to be assembled in the single reaction, wherein the plurality of single-stranded DNA fragments is obtained by chemical synthesis, and where each of the plurality of single- stranded DNA fragments comprise the high-fidelity orthogonal overhang sets of sequences identified by the foregoing method for each of the plurality of target DNA sequences to be assembled in the single reaction on each end.
[0109] In another aspect of the present disclosure, the method is performed simultaneously across a plurality of assembly reactions, wherein each of the plurality of assembly reactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
[0110] In some embodiments, the size of the split window is based on the length of the target sequence and the single- stranded oligo length to be synthesized for assembling the target sequence.
[0111] In an embodiment, each split window comprises at least 12 bp long sequences along the length of each of the plurality of target DNA sequences encoding each of the plurality of target genes.
[0112] In certain embodiments, the length of the plurality of target DNA sequences is selected from 600 bp, 800 bp, 1000 bp, 1200bp, 1400 bp, 1600 bp, 1800 bp, and 2000bp. In some embodiments, the length of the plurality of target DNA sequences is 600 bp and the method achieves 11 assemblies in a single reaction. Tn some embodiments, the length of the plurality of target DNA sequences is 800 bp and the method achieves 7 assemblies in a single reaction. In some embodiments, the length of the plurality of target DNA sequences is 1000 bp and the method achieves 5 assemblies in a single reaction. In some embodiments, the length of the plurality of target DNA sequences is 1200 bp and the method achieves 4 assemblies in a single reaction. In some embodiments, the length of the plurality of target DNA sequences is 1400 bp- 1600 bp and the method achieves 3 assemblies in a single reaction. In some embodiments, the length of the plurality of target DNA sequences is 1800 bp- 2000 bp and the method achieves 2 assemblies in a single reaction.
[0113] In some embodiments, the method is effective in achieving assemblies with about 90% fidelity.
[0114] In some embodiments, the method identifies high-fidelity orthogonal overhang sets that are not self-incompatible or cross-complementary.
[0115] In yet another aspect, the present disclosure provides an in vitro method of scar-less assembly of a plurality of target DNA sequences simultaneously in a single reaction. In some embodiments, the method comprises computationally segmenting each of the plurality of target DNA sequences to identify high-fidelity orthogonal overhang sets of sequences along the length of each of the plurality of target DNA sequences to be assembled in a single reaction, where the identified high-fidelity overhang sets of sequences are not self-incompatible or cross-complementary; synthesizing an oligo pool comprising a plurality of single stranded DNA fragments representing each of the plurality of target DNA sequences to be assembled in the single reaction, where the plurality of single-stranded DNA fragments is obtained by chemical or enzymatic synthesis, and where each of the plurality of single-stranded DNA fragments comprise the high-fidelity orthogonal overhang sets of sequences identified by the disclosed methods on each end; adding a type Ils restriction site to each end of each of the plurality of single-stranded DNAfragments synthesized; amplifying the plurality of DNA fragments; assembling multiple longer DNAs per reaction from the amplified DNA fragments using type Ils restriction enzyme cutting and ligation; inserting the amplified plurality of DNA fragments into a plurality of compatible vectors; and transforming competent host cells for assembling the plurality of target DNA sequences.
[0116] In some embodiments, the method is performed simultaneously across a plurality of assembly reactions, where each of the plurality of assembly reactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
[0117] In some embodiments, the type Ils restriction site is selected from a site specific for a restriction enzyme selected from Bsal, BsmBI, Esp3I, SapI, PaqCI, Bbslb.
[0118] In some embodiments, the computational segmentation comprises: defining a plurality of split windows along the length of each of the target DNA sequences to be assembled in a single reaction; identifying potential overhang sets of sequences within each of the split windows; and using an iterative fidelity evaluation module to optimize selection of high-fidelity orthogonal overhang sets of sequences.
[0119] In an embodiment, a first iteration of the fidelity evaluation module comprises: (a) creating an initial population of overhang sets of sequences comprising the potential overhang sets of sequences identified in each of the split windows; (b) selecting a parent population of high-fidelity overhang sets of sequences, where the selecting comprises weighing fidelity of the initial population of overhang sets of sequences created in step (a) against fidelity data of a reference overhang dataset; (c) randomly selecting overhang sets of sequences in the parent population to cross within the parent population and crossing the overhang sets of sequences to obtain a first generation of overhang sets of sequences; (d) randomly mutating a subset of the first generation of overhang sets of sequences; (e) weighing fidelity of the resulting overhang sets of sequences from step (d) against the reference overhang dataset to identify high-fidelity and low-fidelity overhang sets of sequences; and (f) providing the high-fidelity overhang sets of sequences identified in step (e) as input for subsequent iterations of the fidelity evaluation module to repeat steps (a) to (e).
[0120] The reference overhang data set may be based on rule-based logic, probabilistic scoring, or machine-learning algorithms trained on prior methods of assembling DNA sequences, for example, Golden Gate assembly outcomes.
[0121] In some embodiments, the method further comprising randomly shuffling the low-fidelity overhang sets of sequences identified in step (e) and providing as input for subsequent iterations of the fidelity evaluation module to repeat steps (a) to (e).
[0122] In an embodiment, the computational segmentation is performed using a computer system comprising one or more processors; system memory; and one or more computer readable storage media having stored thereon computer-executable instructions that when executed by the one or more processors, cause the computer system to execute steps(a) to (f).
[0123] In some embodiments, size of the split window is based on the length of the target sequence and the single-stranded oligo length to be synthesized for assembling the target sequence.
[0124] In some embodiments, the length of the plurality of target DNA sequences is selected from 600 bp, 800 bp, 1000 bp, 1200bp, 1400 bp, 1600 bp, 1800 bp, and 2000bp.
[0125] In some embodiments, the length of the plurality of target DNA sequences is 600 bp and the method achieves 11 assemblies in a single reaction. In some embodiments, the length of the plurality of target DNA sequences is 800 bp and the method achieves 7 assemblies in a single reaction. In some embodiments, the length of the plurality of target DNA sequences is 1000 bp and the method achieves 5 assemblies in a single reaction. In some embodiments, the length of the plurality of target DNA sequences is 1200 bp and the method achieves 4 assemblies in a single reaction. In some embodiments, the length of the plurality of target DNA sequences is 1400 bp- 1600 bp and the method achieves 3 assemblies in a single reaction. In some embodiments, the length of the plurality of target DNA sequences is 1800 bp- 2000 bp and the method achieves 2 assemblies in a single reaction. In some embodiments, the method is effective in achieving assemblies with about 90% fidelity.
[0126] In certain embodiments, the predicted overhangs are used to assemble sequence adapters; long range amplicons, genomic regions, and / or barcoded libraries. The resulting constructs may be directly sequenced or further processed.
[0127] In certain embodiments, the disclosure provides a computer-implemented method comprising: (i) receiving a plurality of target DNA sequences; (ii) computationally segmenting the plurality of target DNA sequences into fragments; (iii) generating potential overhang sequences for each junction for the plurality of DNA sequences; (iv) selectinghigh-fidelity orthogonal overhang sets of sequences for assembly; and (v) outputting sequences for synthesis or amplification.
[0128] In some embodiments, the computational segmentation comprises: defining a plurality of split windows along the length of each of the plurality of target DNA sequences to be assembled in a single reaction; identifying potential overhang sets of sequences within each of the split windows; and using the iterative fidelity evaluation module disclosed herein to optimize selection of high-fidelity orthogonal overhang sets of sequences.
[0129] Oligonucleotides serving as target nucleic acids for assembly may be synthesized de novo in parallel. The oligonucleotides may be assembled into precursor fragments which are then assembled into target nucleic acids. In some cases, greater than about 100, 1000, 16,000, 50,000 or 250,000 or even greater than about 1,000,000 different oligonucleotides are synthesized together. In some cases, these oligonucleotides are synthesized in less than 20, 10, 5, 1, 0.1 cm2, or smaller surface area. In some instances, oligonucleotides are synthesized on a support, e.g. surfaces, such as microarrays, beads, miniwells, channels, or substantially planar devices. In some cases, oligonucleotides are synthesized using phosphoramiditc chemistry. Methods of synthesizing oligonucleotides are well-known in the art.
[0130] The DNA and RNA synthesized according to the methods described herein may be used to express proteins in vivo or in vitro. The nucleic acids may be used alone or in combination to express one or more proteins, each having one or more protein activities. Such protein activities may be linked together to create a naturally occurring or non-naturally occurring metabolic / enzymatic pathway. Further, proteins with binding activity may be expressed using the nucleic acids synthesized according to the methods described herein. Such binding activity may be used to form scaffolds of varying sizes.
[0131] Destination vectors used for assembly of the plurality of target nucleic acid are well known in the art and include but are not limited to pucl9, pBR322, and the like.
[0132] The methods and systems described herein may comprise and / or are performed using a software program on a computer system. Accordingly, computerized control for the optimization of design algorithms described herein and the synthesis and assembly of nucleic acids are within the bounds of this disclosure. In an embodiment, the computational segmentation is performed using a computer system comprising one ormore processors; system memory; and one or more computer readable storage media having stored thereon computer-executable instructions that when executed by the one or more processors, cause the computer system to execute steps (a) to (f).
[0133] Aspects of the disclosure provide a system for assembling a plurality of long DNA sequences in a single reaction comprising (a) a processor configured to identify candidate high-fidelity overhang sequences for each of the plurality of long DNA sequences; and (b) an iterative fidelity evaluation module configured to identify and select high-fidelity overhang sequences for each of the plurality of long DNA sequences to be assembled in the single reaction. In some embodiments, the high-fidelity overhang sequences achieve an assembly fidelity of about 90%. In an embodiment, the system is configured to generate assembly designs for a plurality of reactions simultaneously. Workflow For Computational Segmentation
[0134] The computational segmentation using the fidelity evaluation module disclosed herein is depicted in FIG. 2 and described as follows: each split window is treated as a “gene locus,” and possible overhangs within the window are considered “alleles.” The fidelity evaluation module optimizes the overhang selection using the following steps: A “population” of random overhang sets is generated, each with varying fidelity levels. The fidelity is calculated based on a reference dataset, for example, the New England Biolab Golden Gate Assembly (NEB GGA) overhang dataset. See Pryor, .John M et al. “Enabling one-pot Golden Gate assemblies of unprecedented complexity using data-optimized assembly design.” PloS one w\. 15,9 e0238592. 2 Sep. 2020. Parent overhang sets are selected randomly but weighted by their fitness values, where higher-fidelity sets have a greater chance of being chosen for the next generation. Loci from parent overhang sets are randomly selected and swapped to create new combinations. Offspring inherit some overhangs from one parent and the rest from the other. The lowest fidelity overhang is randomly altered / mutated to another overhang in the split window. The fidelity of the resulting overhang sets is recalculated based on the reference dataset, ensuring the highest-quality sets are preserved. This iterative process is repeated for a specified number of generations, gradually improving the fidelity of the overhang sets. Throughout this process, the fidelity evaluation module ensures that reverse complements of overhangs within a set are avoided to prevent redundancy and interference.
[0135] For groups of sequences failing to meet the fidelity threshold within the allowed cycles, the groups of sequences are shuffled and cycled through the fidelity evaluation module.
[0136] Optimized DNA fragments are written to FASTA files, and metadata, including pool assignments and fidelity scores, is saved in CSV files. This ensures both sequence integrity and traceability for downstream applications.Advantages
[0137] By achieving high recall rates and accuracy through optimized assembly protocols, the disclosed method has the potential to greatly accelerate synthetic biology applications such as protein design, enzyme engineering, synthetic genomics, and gene therapy. For example, one round of this pipeline can assemble -3000 genes of ~700bp length from 4,300 nt oligos each, with 3 orthogonal assembly reactions per well, at an estimated ~2 $ / sequence, a >10-fold reduction compared to commercial alternatives. Longer assemblies can be achieved by increasing fragment numbers or using longer input fragments from commercial vendors. Thus, the methods disclosed herein achieve the efficient and accurate assembly of multiple genes in a single reaction utilizing the fidelity evaluation module as disclosed herein to optimize the strategy for pooling different genes together, ensuring high assembly efficiency and fidelity. In contrast, the prior art methods are limited to assembling one gene per reaction.
[0138] The disclosed methods are able to synthesize medium-length (500-1500 bp) DNA fragments in a high-throughput, low-cost manner. These pooled DNA fragments can be used for various high-throughput experiments, such as protein design and genome design.
[0139] The disclosed methods and the fidelity evaluation methods are able to efficiently identify optimal overhangs for assembly, enabling fast and efficient design (~2 mins per 96 well plate). The streamlined assembly methods disclosed herein reduce hands-on time and minimize the use of lab consumables, making it practical for high-throughput settings. The disclosed methods further have been robustly validated by testing assembly of de novo designed proteins and real-world sequence features incorporated, including challenging long repeat regions, to evaluate assembly performance under realistic conditions.
[0140] Advantageously, the methods disclosed herein can be used to assemble a plurality of target fragments within the same assembly reaction, e.g., Golden GateAssembly (GGA) reaction. This approach increases the scale of synthesis and reduces the cost of reaction reagents.Table 1.
[0141] Number of designs assembled per reaction
[0142] *Fragment is -300 bases, assume cost of ~$0.1 per fragmentTable 2. shows per reaction cost
[0143] To achieve efficient and accurate assembly of multiple genes in the same reaction, the disclosed methods utilize a new computational segmentation methodcomprising the fidelity evaluation module disclosed herein to optimize the strategy for pooling different genes together, ensuring high assembly efficiency and fidelity.EXAMPLESExample 1
[0144] The inventors tested the disclosed computational segmentation comprising a fidelity evaluation module on 1,000 randomly generated DNA sequences of varying lengths: 600 bp (3 oligos), 800 bp (4 oligos), and 1000 bp (5 oligos). TATG and CTCG were used as vector overhangs, requiring each overhang set to include these two sequences. The population size and number of generations were both set to 300. For 600 bp sequences, up to 12 designs can be assembled in a single reaction with approximately 80% fidelity. For 800 bp sequences, 8 designs achieve the same fidelity, while for 1000 bp sequences, 6 designs can be assembled with over 80% fidelity. The resulting fidelity is shown in FIG.3.
[0145] In summary, the method allows efficient assembly of multiple DNA sequences in a single reaction with high fidelity. Specifically, 600 bp sequences achieve 12 assemblies, 800 bp sequences achieve 8, and 1000 bp sequences achieve 6, with predicted fidelities exceeding 80% in all cases.Example 2Assembly Protocol
[0146] Following the computational segmentation of each desired sequence as described above, Bsal restriction sites are added to both ends of each fragment, along with two unique adaptor pairs for PCR amplification. The outer adaptor pair is used to amplify oligos for the entire 96-well plate reaction, while the inner pair is used to amplify oligos for individual Golden Gate reactions (FIG. 1).
[0147] Oligos are ordered to be synthesized on a microarray from commercial vendors. Once received, a two-step PCR followed by purification was performed to prepare the oligos for each reaction. These are then combined with the appropriate plasmid backbone and subjected to a GGA reaction to assemble the desired DNA sequences.
[0148] In one pilot experiment, a pool of 6821 DNA sequences ranging in size from 300bp-1600bp (FIG.4) were assembled. These sequences were predicted to have an accuracy of approximately 93.3% across 1476 Golden Gate assembly reactions. After assembly using the protocol described above, we evaluated successful assemblies with Nanopore sequencing. We achieved a sequence recall of about 92.9%, indicating that only482 out of 6821 intended designs were not assembled. The distribution of reads is shown in FIG. 5.
[0149] Overall, the methods disclosed herein provide a comprehensive, high-throughput, and cost-effective solution for long-fragment DNA synthesis using Golden Gate Assembly (GGA). The methods of the present disclosure streamline the entire workflow, from computational sequence segmentation to experimental assembly, enabling the efficient synthesis of thousands of DNA constructs. This approach offers significant advantages in terms of scalability, fidelity, and ease of implementation, making it accessible even in resource-limited settings. By achieving high recall rates and accuracy through optimized assembly protocols, the disclosed methods have the potential to greatly accelerate synthetic biology applications such as protein design, enzyme engineering, synthetic genomics, and gene therapy.NON-LIMITING EMBODIMENTS
[0150] While general features of the disclosure are described and shown and particular features of the disclosure are set forth in the claims, the following non-limiting embodiments relate to features, and combinations of features, that are explicitly envisioned as being part of the disclosure. The following non- limiting Embodiments contain elements that are modular and can be combined with each other in any number, order, or combination to form a new non-limiting Embodiment, which can itself be further combined with other non-limiting Embodiments.
[0151] Embodiment 1. A method for identifying a plurality of high-fidelity orthogonal overhang sets for scar-less assembly of a plurality of target DNA sequences in a single reaction, the method comprising: selecting a plurality of target DNA sequences to be assembled in a single reaction; and computationally segmenting the plurality of target DNA sequences to be assembled in a single reaction to identify high-fidelity orthogonal overhang sets of DNA sequences along the length of each of the plurality of target DNA sequences to be assembled in a single reaction, wherein the computational segmentation comprises: (i) defining a plurality of split windows along the length of each of the plurality of target DNA sequences to be assembled in the single reaction; (ii) identifying potential overhang sets of sequences within each of the split windows; and (iii) using an iterative fitness evaluation module to optimize selection of high-fidelity orthogonal overhang sets of DNA sequences, wherein a first iteration of the fitness evaluation module comprises: (a) creating an initial population of overhang sets of DNA sequences comprising the potentialoverhang sets of sequences identified in step (ii); (b) selecting a parent population of high-fidelity overhang sets of sequences, wherein the selecting comprises weighing fidelity of the initial population of overhang sets of sequences created in step (a) against fidelity data of a reference overhang dataset; (c) randomly selecting overhang sets of sequences in the parent population to swap / cross within the parent population and crossing the overhang sets of sequences to obtain a first generation of overhang sets of sequences; (d)randomly mutating a subset of the first generation of overhang sets of sequences; (e) weighing fidelity of the resulting overhang sets of sequences from step (d) against the reference overhang dataset to identify high-fidelity and low-fidelity overhang sets of sequences; and (f) providing the high-fidelity overhang sets of sequences identified in step (e) as input for subsequent iterations of the fitness evaluation module to repeat steps (a) to (e), wherein the computational segmentation is performed using a computer system comprising one or more processors; system memory; and one or more computer readable storage media having stored thereon computer-executable instructions that when executed by the one or more processors, cause the computer system to execute steps (a) -(f).
[0152] Embodiment 2. The method of Embodiment 1, further comprising randomly shuffling the low-fidelity overhang sets of sequences identified in step (c) and providing as input for subsequent iterations of the fitness evaluation module to repeat steps (a) to (e).
[0153] Embodiment 3. The method of Embodiment 2, wherein the identified high-fidelity overhangs reduce the incorrect ligation events relative to randomly selected overhangs.
[0154] Embodiment 4. The method of Embodiment 1 , further comprising synthesizing an oligo pool comprising a plurality of single stranded DNA fragments representing each of the plurality of target DNA sequences to be assembled in the single reaction, wherein each of the plurality of single-stranded DNA fragments comprise the high-fidelity orthogonal overhang sets of sequences identified for each of the plurality of target DNA sequences to be assembled in the single reaction in step (iii) on each end.
[0155] Embodiment 5. The method of Embodiment 4, wherein the method further comprises adding: (i) a type Ils restriction site; and (ii) an inner and an outer unique adaptor pair for PCR amplification to each end of the plurality of single stranded DNA fragments.
[0156] Embodiment 6. The method of Embodiment 4 or Embodiment 5, wherein the plurality of single-stranded DNA fragments are obtained by chemical synthesis.
[0157] Embodiment 7. The method of any one of Embodiment 1-3, wherein the method is performed simultaneously across a plurality of assembly reactions, and wherein each of the plurality of assembly reactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
[0158] Embodiment 8. The method of any one of Embodiments 1-3, wherein the defining the split window for each of the plurality of target sequences to be assembled is based on the length of the target sequence and the length of the single-stranded DNA fragments representing the target DNA sequence.
[0159] Embodiment 9. The method of Embodiment 8, wherein each split window comprises at least 12 bp long sequences along the length of each of the plurality of target DNA sequences to be assembled.
[0160] Embodiment 10. The method of any one of Embodiments 1-9, wherein the length of the plurality of target DNA sequences to be assembled is selected from 600 bp, 800 bp, 1000 bp, 1200bp, 1400 bp, 1600 bp, 1800 bp, and 2000bp.
[0161] Embodiment 11. The method of Embodiment 10, wherein the length of the plurality of DNA sequences is 600 bp and the method achieves 11 assemblies in a single reaction.
[0162] Embodiment 12. The method of Embodiment 10, wherein the length of the plurality of DNA sequences is 800 bp and the method achieves 7 assemblies in a single reaction.
[0163] Embodiment 1 . The method of Embodiment 10, wherein the length of the plurality of DNA sequences is 1000 bp and the method achieves 5 assemblies in a single reaction.
[0164] Embodiment 14. The method of Embodiment 10, wherein the length of the plurality of DNA sequences is 1200 bp and the method achieves 4 assemblies in a single reaction.
[0165] Embodiment 15. The method of Embodiment 10, wherein the length of the plurality of DNA sequences is 1400 bp- 1600 bp and the method achieves 3 assemblies in a single reaction.
[0166] Embodiment 16. The method of Embodiment 10, wherein the length of the plurality of DNA sequences is 1800 bp- 2000 bp and the method achieves 2 assemblies in a single reaction.
[0167] Embodiment 17. The method of any one of Embodiments 1-16, wherein the method is effective in achieving assemblies with about 90% fidelity.
[0168] Embodiment 18. The method of any one of Embodiment 1-17, wherein the method identifies high-fidelity orthogonal overhang sets that are not selfincompatible or cross-complementary.
[0169] Embodiment 19. An in vitro method of scar- less assembly of a plurality of target DNA sequences simultaneously in a single reaction, the method comprising: (a) computationally segmenting each of the plurality of target DNA sequences to identify high-fidelity orthogonal overhang sets of sequences along the length of each of the plurality of target DNA sequences to be assembled in a single reaction, wherein the identified high-fidelity overhang sets of sequences are not self-incompatible or cross-complementary; (b) synthesizing an oligo pool comprising a plurality of single stranded DNA fragments representing each of the plurality of target DNA sequences to be assembled in the single reaction, wherein the plurality of single-stranded DNA fragments is obtained by chemical or enzymatic synthesis, and wherein each of the plurality of single stranded DNA fragments comprise the high-fidelity orthogonal overhang sets of sequences identified in step (a) on each end; (c) adding a type Ils restriction site to each end of each of the plurality of DNA fragments synthesized in step (b); (d) adding an inner and an outer unique adaptor pair for PCR amplification to each end of the plurality of DNA fragments of step (c); (e) amplifying the plurality of DNA fragments; (f) assembling multiple longer sequences per reaction from the amplified DNA fragments using type Ils restriction enzyme cutting and ligation; (g) inserting the amplified plurality of DNA fragments into a plurality of compatible destination vectors; and (h) transforming competent host cells with the destination vectors for assembling the plurality of DNA sequences, wherein the outer unique adaptor pair is used to amplify oligos for the entire high-throughput reaction, while the inner unique adaptor pair is used to amplify oligos for each individual assembly reaction.
[0170] Embodiment 20. The method of Embodiment 19, wherein the method is performed simultaneously across a plurality of assembly reactions, wherein each of theplurality of assembly reactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
[0171] Embodiment 21. The method of Embodiment 19 or Embodiment 20, wherein the type Ils restriction site is selected from a site specific for a restriction enzyme selected from Bsal, BsmBI, Esp3I, SapI, PaqCI, Bbslb.
[0172] Embodiment 22. The method of any of Embodiments 19-21, wherein the computational segmentation comprises: (i) defining a plurality of split windows along the length of each of the plurality of target DNA sequences to be assembled in a single reaction; (ii) identifying potential overhang sets of sequences within each of the split windows; and (iii) using an iterative fitness evaluation module to optimize selection of high-fidelity orthogonal overhang sets of sequences, wherein a first iteration of the fitness evaluation module comprises: (a) creating an initial population of overhang sets of sequences comprising the potential overhang sets of sequences identified in step (ii); (b) selecting a parent population of high-fidelity overhang sets of sequences, wherein the selecting comprises weighing fidelity of the initial population of overhang sets of sequences created in step (a) against fidelity data of a reference overhang dataset; (c) randomly selecting overhang sets of sequences in the parent population to cross within the parent population and crossing the overhang sets of sequences to obtain a first generation of overhang sets of sequences; (d) randomly mutating a subset of the first generation of overhang sets of sequences; (e) weighing fidelity of the resulting overhang sets of sequences from step (d) against the reference overhang dataset to identify high-fidelity and low-fidelity overhang sets of sequences; and (f) providing the high-fidelity overhang sets of sequences identified in step (e) as input for subsequent iterations of the fitness evaluation module to repeat steps (a) to (e), wherein the computational segmentation is performed using a computer system comprising one or more processors; system memory; and one or more computer readable storage media having stored thereon computer-executable instructions that when executed by the one or more processors, cause the computer system to execute steps (a) -(f).
[0173] Embodiment 23. The method of Embodiment 22, further comprising randomly shuffling the low-fidelity overhang sets of sequences identified in step (e) and providing as input for subsequent iterations of the fitness evaluation module to repeat steps (a) to (e).
[0174] Embodiment 24. The method of Embodiment 22 or Embodiment 23, wherein the defining of the split window for each of the plurality of target sequences to be assembled is based on the length of the target sequence and the length of the single-stranded DNA fragments representing the target DNA sequence.
[0175] Embodiment 25. The method of Embodiment 24, wherein the length of the plurality of target DNA sequences is selected from 600 bp, 800 bp, 1000 bp, 1200bp, 1400 bp, 1600 bp, 1800 bp, and 2000bp.
[0176] Embodiment 26. The method of Embodiment 24, wherein the length of the plurality of target DNA sequences is 600 bp and the method achieves 11 assemblies in a single reaction.
[0177] Embodiment 27. The method of Embodiment 24, wherein the length of the plurality of target DNA sequences is 800 bp and the method achieves 7 assemblies in a single reaction.
[0178] Embodiment 28. The method of Embodiment 24, wherein the length of the plurality of target DNA sequences is 1000 bp and the method achieves 5 assemblies in a single reaction.
[0179] Embodiment 29. The method of Embodiment 24, wherein the length of the plurality of target DNA sequences is 1200 bp and the method achieves 4 assemblies in a single reaction.
[0180] Embodiment 30. The method of Embodiment 24, wherein the length of the plurality of target DNA sequences is 1400 bp- 1600 bp and the method achieves 3 assemblies in a single reaction.
[0181] Embodiment 31. The method of Embodiment 24, wherein the length of the plurality of target DNA sequences is 1800 bp- 2000 bp and the method achieves 2 assemblies in a single reaction.
[0182] Embodiment 32. The method of any one of Embodiments 19-31, wherein the method is effective in achieving assemblies with about 90% fidelity.
[0183] Embodiment 33. A system for assembling a plurality of long DNA sequences in a single reaction comprising: (a) a processor configured to identify candidate high-fidelity overhang sequences for each of the plurality of long DNA sequences; and (b) an iterative fidelity evaluation module configured to identify and select high-fidelity overhang sequences for each of the plurality of long DNA sequences to be assembled in the single reaction, wherein the high-fidelity overhang sequences achieve an assemblyfidelity of about 90%, and wherein the system is configured to generate assembly designs for a plurality of reactions simultaneously.
[0184] Embodiment 34. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1.
[0185] While illustrative embodiments have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the disclosure.
Claims
1. CLAIMSThe embodiments of the invention in which an exclusive property or privilege is claimed are defined as follows:
1. A method for identifying a plurality of high-fidelity orthogonal overhang sets for scar-less assembly of a plurality of target DNA sequences in a single reaction, the method comprising:selecting a plurality of target DNA sequences to be assembled in a single reaction; andcomputationally segmenting the plurality of target DNA sequences to be assembled in a single reaction to identify high-fidelity orthogonal overhang sets of DNA sequences along a length of each of the plurality of target DNA sequences to be assembled in a single reaction, wherein the computational segmentation comprises:(i) defining a plurality of split windows along the length of each of the plurality of target DNA sequences to be assembled in the single reaction;(ii) identifying potential overhang sets of sequences within each of the split windows; and(iii) using an iterative fidelity evaluation module configured to optimize selection of high-fidelity orthogonal overhang sets of DNA sequences, wherein a first iteration of the fidelity evaluation module comprises:(a) creating an initial population of overhang sets of DNA sequences comprising the potential overhang sets of sequences identified in step (ii);(b) selecting a parent population of high-fidelity overhang sets of sequences, wherein the selecting comprises weighing fidelity of the initial population of overhang sets of sequences created in step (a) against fidelity data of a reference overhang dataset;(c) randomly selecting overhang sets of sequences in the parent population to swap / cross within the parent population and crossing the overhang sets of sequences to obtain a first generation of overhang sets of sequences;(d) randomly mutating a subset of the first generation of overhang sets of sequences;(e) weighing fidelity of the resulting overhang sets of sequences from step (d) against the reference overhang dataset to identify high-fidelity and low-fidelity overhang sets of sequences; and(f) providing the high-fidelity overhang sets of sequences identified in step (e) as input for subsequent iterations of the fidelity evaluation module to repeat steps (a) to (e),wherein the computational segmentation is performed using a computer system comprising one or more processors; system memory; and one or more computer readable storage media having stored thereon computer-executable instructions that when executed by the one or more processors, cause the computer system to execute steps (a) -(f).
2. The method of claim 1, further comprising randomly shuffling the low-fidelity overhang sets of sequences identified in step (e) and providing as input for subsequent iterations of the fitness evaluation module to repeat steps (a) to (e).
3. The method of claim 2, wherein the selected high-fidelity overhangs reduce the incorrect ligation events relative to randomly selected overhangs.
4. The method of claim 1, further comprising synthesizing an oligo pool comprising a plurality of single stranded DNA fragments representing each of the plurality of target DNA sequences to be assembled in the single reaction, wherein each of the plurality of single-stranded DNA fragments comprise the high-fidelity orthogonal overhang sets of sequences identified for each of the plurality of target DNA sequences to be assembled in the single reaction in step (iii) on each end.
5. The method of claim 4, wherein the method further comprises adding:(i) a type Ils restriction site; and(ii) an inner and an outer unique adaptor pair for PCR amplification to each end of the plurality of single stranded DNA fragments.
6. The method of claim 4, wherein the plurality of single-stranded DNA fragments is synthesized by chemical synthesis.
7. The method of claim 1, wherein the method is performed simultaneously across a plurality of assembly reactions, and wherein each of the plurality of assemblyreactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
8. The method of claim 4, wherein the defining the split window for each of the plurality of target sequences to be assembled is based on the length of the target sequence and the length of the single-stranded DNA fragments representing the target DNA sequence.
9. The method of claim 8, wherein each split window comprises at least 12 bp long sequences along the length of each of the plurality of target DNA sequences to be assembled.
10. The method of claim 1, wherein the length of the plurality of target DNA sequences to be assembled is selected from 600 bp, 800 bp, 1000 bp, 1200bp, 1400 bp, 1600 bp, 1800 bp, and 2000bp.
11. The method of claim 10, wherein the length of the plurality of DNA sequences is 600 bp and the method achieves 11 assemblies in a single reaction.
12. The method of claim 10, wherein the length of the plurality of DNA sequences is 800 bp and the method achieves 7 assemblies in a single reaction.
13. The method of claim 10, wherein the length of the plurality of DNA sequences is 1000 bp and the method achieves 5 assemblies in a single reaction.
14. The method of claim 10, wherein the length of the plurality of DNA sequences is 1200 bp and the method achieves 4 assemblies in a single reaction.
15. The method of claim 10, wherein the length of the plurality of DNA sequences is 1400 bp-1600 bp and the method achieves 3 assemblies in a single reaction.
16. The method of claim 10, wherein the length of the plurality of DNA sequences is 1800 bp- 2000 bp and the method achieves 2 assemblies in a single reaction.
17. The method of claim 1, wherein the method is effective in achieving assemblies yielding correct constructs at a fidelity of about 90%.
18. The method of claim 1, wherein the method identifies high-fidelity orthogonal overhang sets that are not self-incompatible or cross-complementary.
19. An in vitro method of scar-less assembly of a plurality of target DNA sequences simultaneously in a single reaction, the method comprising:(a) computationally segmenting each of the plurality of target DNA sequences to identify high-fidelity orthogonal overhang sets of sequences along the length of each of the plurality of target DNA sequences to be assembled in a single reaction, wherein the identified high-fidelity overhang sets of sequences are not self-incompatible or cross-complementary;(b) synthesizing an oligo pool comprising a plurality of single stranded DNA fragments representing each of the plurality of target DNA sequences to be assembled in the single reaction, wherein the plurality of single- stranded DNA fragments is obtained by chemical or enzymatic synthesis, and wherein each of the plurality of single stranded DNA fragments comprise the high-fidelity orthogonal overhang sets of sequences identified in step (a) on each end;(c) adding a type Ils restriction site to each end of each of the plurality of DNA fragments synthesized in step (b);(d) adding an inner and an outer unique adaptor pair for PCR amplification to each end of the plurality of DNA fragments of step (c);(e) amplifying the plurality of DNA fragments;(f) assembling multiple longer sequences per reaction from the amplified DNA fragments using type Ils restriction enzyme cutting and ligation;(g) inserting the amplified plurality of DNA fragments into a plurality of compatible destination vectors; and(h) transforming competent host cells with the destination vectors for assembling the plurality of DNA sequences,wherein the outer unique adaptor pair is used to amplify oligos for the entire high-throughput reaction, while the inner unique adaptor pair is used to amplify oligos for each individual assembly reaction.
20. The method of claim 19, wherein the method is performed simultaneously across a plurality of assembly reactions, wherein each of the plurality of assembly reactions comprises a plurality of target DNA sequences to be assembled in a single reaction.
21. The method of claim 19, wherein the type Ils restriction site is selected from a site specific for a restriction enzyme selected from Bsal, BsmBl, Esp3I, Sapl, PaqCl, Bbslb.
22. The method of claim 19, wherein the computational segmentation comprises:(i) defining a plurality of split windows along the length of each of the plurality of target DNA sequences to be assembled in a single reaction;(ii) identifying potential overhang sets of sequences within each of the split windows; and(iii) using an iterative fidelity evaluation module to optimize selection of high-fidelity orthogonal overhang sets of sequences, wherein a first iteration of the fidelity evaluation module comprises:(a) creating an initial population of overhang sets of sequences comprising the potential overhang sets of sequences identified in step (ii);(b) selecting a parent population of high-fidelity overhang sets of sequences, wherein the selecting comprises weighing fidelity of the initial population of overhang sets of sequences created in step (a) against fidelity data of a reference overhang dataset;(c) randomly selecting overhang sets of sequences in the parent population to cross within the parent population and crossing the overhang sets of sequences to obtain a first generation of overhang sets of sequences;(d) randomly mutating a subset of the first generation of overhang sets of sequences;(e) weighing fidelity of the resulting overhang sets of sequences from step (d) against the reference overhang dataset to identify high-fidelity and low-fidelity overhang sets of sequences; and(f) providing the high-fidelity overhang sets of sequences identified in step (e) as input for subsequent iterations of the fidelity evaluation module to repeat steps (a) to (e),wherein the computational segmentation is performed using a computer system comprising one or more processors; system memory; and one or more computer readable storage media having stored thereon computer-executable instructions that when executed by the one or more processors, cause the computer system to execute steps (a) -(f).
23. The method of claim 22, further comprising randomly shuffling the low-fidelity overhang sets of sequences identified in step (e) and providing as input for subsequent iterations of the fitness evaluation module to repeat steps (a) to (e).
24. The method of claim 22, wherein the defining of the split window for each of the plurality of target sequences to be assembled is based on the length of the target sequence and the length of the single-stranded DNA fragments representing the target DNA sequence to be assembled.
25. The method of claim 24, wherein the length of the plurality of target DNA sequences is selected from 600 bp, 800 bp, 1000 bp, 1200bp, 1400 bp, 1600 bp, 1800 bp, and 2000bp.
26. The method of claim 24, wherein the length of the plurality of target DNA sequences is 600 bp and the method achieves 11 assemblies in a single reaction.
27. The method of claim 24, wherein the length of the plurality of target DNA sequences is 800 bp and the method achieves 7 assemblies in a single reaction.
28. The method of claim 24, wherein the length of the plurality of target DNA sequences is 1000 bp and the method achieves 5 assemblies in a single reaction.
29. The method of claim 24, wherein the length of the plurality of target DNA sequences is 1200 bp and the method achieves 4 assemblies in a single reaction.
30. The method of claim 24, wherein the length of the plurality of target DNA sequences is 1400 bp- 1600 bp and the method achieves 3 assemblies in a single reaction.
31. The method of claim 24, wherein the length of the plurality of target DNA sequences is 1800 bp- 2000 bp and the method achieves 2 assemblies in a single reaction.
32. The method of claim 19, wherein the method is effective in achieving assemblies with about 90% fidelity.
33. A system for assembling a plurality of long DNA sequences in a single reaction comprising:(a) a processor configured to identify candidate high-fidelity overhang sequences for each of the plurality of long DNA sequences: and(b) an iterative fidelity evaluation module configured to identify and select high-fidelity overhang sequences for each of the plurality of long DNA sequences to be assembled in the single reaction,wherein the high-fidelity overhang sequences achieve an assembly fidelity of about 90%, and wherein the system is configured to generate assembly designs for a plurality of reactions simultaneously.
34. A non-transitory computer-readable medium storing instructions that, when executed by a processor, causes the processor to perform the method of claim 1.