Biopolymer synthesis

JP2025516106A5Pending Publication Date: 2026-04-06RIBBON BIOLABS GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2026-04-06

AI Technical Summary

Technical Problem

Current methods for synthesizing large biopolymers, such as genomic-scale polynucleotide molecules, are inefficient and prone to errors due to the complexity of assembling oligos into long sequences.

Method used

The method employs an in silico graph structure to direct the linear ligation of oligos, automatically parallelizing the assembly process to rapidly and reliably produce genome-scale DNA molecules.

Benefits of technology

This approach significantly reduces the time required for synthesizing long polynucleotides from days to hours, while ensuring high fidelity and avoiding unintended sequence structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention provides a method for synthesizing large biopolymers, such as genome-scale polynucleotide molecules. When genome-scale polynucleotides are made by linking several oligos together, the assembly of a desired polynucleotide from the oligos is represented as a tree, where the branches represent nucleotide sequences and the nodes represent the bonds between pairs of oligos. Multiple assembly trees are each generated and scored in silico according to how successful the tree is in directing the assembly of the desired polynucleotide, taking into account the biochemical properties of the oligos and the proposed ligations. The system and method of the present invention selects an appropriate tree, operates the liquid handling system, and performs the operations indicated by the selected tree to make the desired polynucleotide.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Sequence Listing The "Sequence Listing XML" is submitted together with this specification in XML file format, where (i) the file name is RBIO-001-01WO.xml; (ii) the creation date is March 31, 2023; (iii) the file size is 290,939 bytes, and the materials in the XML file are incorporated by reference.

[0002] Technical Field The present invention relates to the synthesis of biopolymers.

Background Art

[0003] Background The Human Genome Project (HGP) was an international effort aimed at determining the sequence of the human genome. Deciphering the sequence of the human genome relied on techniques such as gene mapping by restriction fragment length polymorphism and DNA sequencing by dideoxy chain terminator sequencing, often called Sanger sequencing. Using such techniques, a rough outline of the human genome was created in 2000, and the complete human genome was published on April 13, 2003.

[0004] Considering the demands envisioned for the future of genomics arising from the ability to decipher entire genomes, this project could be referred to as "HGP-read". Currently, "GP-write", a new international consortium, has been formed with the aim of using molecular synthesis, gene editing, and other technologies to manipulate, create, and test biological systems, with the overarching goal of understanding the blueprint of life provided by HGP-read. One of the aims of GP-write is to accelerate synthesis technologies and reduce the synthesis cost to 1 / 1000 within 10 years. The GP-write project takes the position that in order to truly understand the blueprint of genes, it is necessary to "write" DNA and construct human and other genomes from scratch.

Summary of the Invention

Means for Solving the Problem

[0005] Abstract The present invention provides a method for synthesizing large biopolymers such as genomic-scale polynucleotide molecules. In one aspect, the present invention enables the synthesis of genomic-scale polynucleotides by using an in silico graph structure to direct the linear ligation of a number of oligos. The present invention provides a method for automatically parallelizing the joining of smaller oligo sub-parts for the rapid and controlled ordering of sub-parts while avoiding unintended sequence structures. The method is operable using oligos provided in a container such as a multi-well plate and a final desired polynucleotide sequence specified, for example, by a computer file.

[0006] The present invention provides a graph structure for directing the assembly of large polymers. Although the present invention is illustrated below using a tree structure, any suitable graph structure is intended to be useful in the implementation of the methods described herein. Generally, the system of the present invention provides the assembly of a desired polynucleotide sequence from constituent oligos by assembling one or more trees written as computer tree files, in which branches and nodes represent combinations of liquid transfer steps between wells of, for example, a multi-well plate, to be performed by a liquid handling system. The systems and methods of the present invention automatically generate and score an assembly tree according to how successfully a polynucleotide is made when a given tree directs the assembly by a liquid handling system. The systems and methods of the present invention select a suitable tree and operate a liquid handling system to produce the desired polynucleotide.

[0007] If the desired polynucleotide sequence is at least about 100 kb in length and the oligos are short, e.g., less than about several tens of bases in length, the number of potential trees showing the steps for assembling the oligos into the polynucleotide is intractable for the human mind. The system of the present invention automatically generates the trees and applies algorithms to score the trees. The trees receive a score regarding the efficiency and fidelity with which the polynucleotide can be assembled. For example, the score can take into account how well the attachment ends of the oligos anneal faithfully and ligate as intended, and how likely the attachment ends are to misligate and create unintended junctions. The system of the present invention can recursively manipulate very large trees by creating subtrees and joining them together. The system of the present invention can be adapted to the equipment, consumables, and reagents used in a liquid handling system. For example, the software module that generates, scores, and selects the trees can prefer subtrees that can be completely withdrawn from one 96-well plate or 384-well plate in order to minimize plate changes during assembly. Similarly, the tree selection module can prefer wider and shallower trees (over deeper and narrower trees). This is because shallower trees give the most opportunities for parallel processing and enable the synthesis of multiple segments of the polynucleotide in parallel. The result shows that tree selection according to the methods herein using the resulting parallel synthesis can create a molecule that would take 5 days to synthesize using a linear "part A, then part B, then part C" approach in about 5 hours.

[0008] An important insight of the present disclosure is that, for example, a tree selection for controlling a liquid handling system can advance molecular synthesis beyond the concept of superparallel processing. In contrast to a simple linear assembly that sequentially attaches oligos from the 5' end to the 3' end, the present invention provides an increase in efficiency by parallelizing the assembly. For example, the desired sequence is divided in the middle, each half is divided in the middle, etc., and each of the resulting subparts is made in parallel and then they are joined together. The present invention recognizes that a symmetric assembly tree may require many ligation events that may not function biochemically well and that the reaction components may assemble in an undesirable manner. Selecting a tree that avoids undesirable ligation can result in a highly asymmetric tree.

[0009] Avoiding undesirable ligation can also be addressed by appropriate in silico partitioning of the desired sequence to identify the constituent oligos and their attachment ends. Using the method of the invention to optimize the partitioning of the desired sequence into the constructed oligo sequences, ligation that is poorly performed can be avoided. Each partition has its own set of oligos, attachment ends, and an assembly tree. The method of the invention can be used to evaluate a proposed partition by generating, scoring, and selecting a partition that gives an oligo, attachment end, and tree that avoid oligo pairs that do not ligate as intended. Alternatively, a method for selecting a tree that includes a highly asymmetric tree that, by design, avoids the need for reactions of sets of oligos and attachment ends that do not execute as intended for a given partition and associated set of oligos can also be used. By way of example, if segments A and B are intended to form AB (such that C can be added to form ABC), but tend to form BA, the system first selects an assembly tree that forms BC and then mixes it with A to form ABC. Such an assembly tree is biased away from symmetry so that a computer system can search for and write a tree that avoids undesirable ligation. In fact, the asymmetry is too simplistic. The selected tree for optimal assembly can have a topology that deviates from symmetry such that the shape of the tree cannot be understood at once and cannot be given a simple verbal description to explain the shape of the tree. Since the selected tree is designed with such specificity for the operations performed by a liquid handling system, there is no single feature that can explain the tree. The selected trees tend to be shallow and asymmetric and they may tend to include subtrees that contain, for example, 8 or 12 terminal taxa, but any given selected tree can have any topology.

[0010] As described above, undesirable ligation can be avoided by splitting a desired polynucleotide sequence into constituent oligo sequences in a manner that optimizes the ends of the oligos according to the laboratory assembly / synthesis being performed. In the simplest case, the desired sequence is divided in half repeatedly to identify the constituent oligos (the "split"). Any given split gives a set of oligo sequences with specified ends. These ends (sticky or blunt) may not function well when joining the oligos to make a molecule. For example, sticky-ended oligos may have sticky ends that form GC-rich hairpins and are difficult to ligate together. Or the oligos may have two complementary sticky ends, which means the oligos tend to ligate to copies of themselves.

[0011] The present invention includes a method for in silico splitting a desired polynucleotide sequence into constituent oligo sequences, creating an assembly tree for those sequences, and scoring the tree using a set of scores that predict how the ends of those oligos will function. In such an embodiment, each split of the desired polynucleotide sequence into constituent oligos has its own set of assembly trees. Scoring and selection of the trees automate the process of selecting a split that identifies the oligos to be used in making the polynucleotide. By doing so, a liquid handling system can provide oligos that function as intended when making the polynucleotide. The in silico splitting allows the laboratory system to reliably produce the intended product without producing unintended products.

[0012] By using the in silico methods and systems described herein for making polynucleotides, these molecules can be rapidly made with a high degree of parallelization in a liquid handling system. Since poor performing constructs are avoided, these molecules are made with certainty. The methods and systems of the invention are used to make polynucleotides that are greater than 100,000 base pairs in length. A typical building block oligo contains a pair of single DNA strands that are several base pairs in overlap, e.g., up to 50 to 100 bases in length, leaving attachment ends at each terminus. Each such oligo contributes approximately 8 bases to the final molecule. When they are assembled together, there are trillions of potential "trees" that describe the order and parallelization of steps for making one molecule from a set of oligos. The systems and methods of the invention can rapidly create many such trees, score the trees, select one tree, and then operate a liquid handling system according to the selected tree to make genome-scale molecules. Thus, the invention provides for rapid and reliable writing of genome-scale DNA.

[0013] In certain embodiments, the invention provides a method of polymer synthesis. The method includes inputting a desired polynucleotide sequence into a computer system, e.g., as an in silico graph structure including a plurality of branches, each branch specifying a linear order of nucleic acids. The in silico graph structure is programmed to select an optimal combination of branches for specifying the desired sequence and to direct the assembly of a polynucleotide molecule having the desired sequence. Preferably, the branches correspond to subsets of polynucleotides or oligos provided in separate compartments, such as wells of a multi-well plate. The in silico graph structure is capable of residing in a system including program instructions executable to cause the computer system to generate a plurality of trees, each tree representing an ordered combination of branches that results in the desired polynucleotide sequence; to select a tree that provides an optimal combination of branches; and to direct the assembly of the desired polynucleotide sequence using the selected tree structure. Generally, a tree includes branches connected by nodes representing bonds between oligos or sub-polynucleotides. The selecting step may include calculating a score for each tree, the score representing a measure of success that the tree structure results in the assembly of a molecule having the desired polynucleotide sequence. The measure of success may be based on stored scores or probabilities of success in assembling partial constructs and / or probabilities of mis-assembly stored.

[0014] Certain embodiments use recursive assembly and the selecting step includes identifying a subgroup (e.g., of from about 20 to less than 100 oligos); generating, scoring, and selecting an optimal subtree for each subgroup; and combining the optimal subtrees from the subgroups to form an assembly tree having an optimal generation score.

[0015] A preferred method may include the step of directing a fluid handling system to produce a desired polynucleotide. The fluid handling system can be a multi-channel robotic system, an acoustic liquid handling system, or a microfluidic system.

[0016] Embodiments of a computer system can use a software package to search a tree space, identify local maxima therein, and thereby select one or more trees from the local maxima. The method can include the step of joining a subgroup of branches together to form a branch path that includes a desired polynucleotide sequence. Preferably, the selection algorithm selects an ordered nesting of combinations of oligos that successfully produce the desired polynucleotide. The computer system can execute a machine learning algorithm to select the optimal combination of branches. The machine learning algorithm can use one or a combination of a decision tree algorithm, a random forest algorithm, an extreme gradient boosting algorithm, an adaptive boosting algorithm, and a deep learning algorithm. The computer system can perform the search and optimization using, for example, Bayesian algorithms, genetic algorithms, and Monte Carlo analysis.

[0017] In a related aspect, the present invention provides a system for polymer synthesis. The system includes a computer system operable to receive a desired polynucleotide sequence, for example, within an in silico graph structure that includes a plurality of branches, each branch specifying a linear order of nucleic acids. The computer system is programmed to select an optimal combination of branches that specify the desired sequence. The system includes a liquid handling system, and the computer system is programmed to direct the liquid handling system to assemble polynucleotide molecules having the desired sequence.

[0018] In another aspect, the present invention provides a method of synthesizing a polymer. The method includes receiving information describing a desired polynucleotide; generating a plurality of trees, each tree providing an order of ligation between oligos to form the polynucleotide; selecting one of the trees having an optimal generation score; and creating a polynucleotide by ligating the oligos as provided by the selected tree. Optionally, the leaves of the tree represent oligos or their overhangs, and the nodes of the tree represent ligation between oligos. The selecting step may include calculating a score for each tree, the score representing the probability that a polynucleotide will be successfully created using that tree. The method may include scoring the trees using (i) a stored probability of success in ligating the overhangs of the oligos, (ii) a stored estimate of the risk of misligation between unintended pairs of oligos, or (iii) both.

[0019] There can be more than about 10^30 trees that give the order of the bonds between oligos to form a polynucleotide, and in certain recursive assembly embodiments, this method includes performing steps of generating and selecting for a first subset of oligos to form an intermediate tree; and storing the intermediate tree for use as a leaf of a final tree having an optimal generation score (i.e., by performing steps on nested sub-parts of the final tree, not scoring all of the 10^30 trees). For example, this method can include identifying subsets of oligos (e.g., each containing from about 20 to less than 100 oligos); generating, scoring, and selecting a sub-tree for each subset; and joining together one optimally scored tree from each subset to obtain a selected tree having an optimal generation score. For each subset, the software package can generate all possible trees; store some or all of the trees in memory; apply a scoring matrix to each of the trees in memory to score all of the subset's trees; and select the highest scoring tree for the subset.

[0020] Some methods include the step of making a molecule, i.e., a polynucleotide. The step of making a polynucleotide can include, for example, making at least one substantially complete expression vector or biological genome of about 100 kb or more. In some embodiments, this method includes performing the above steps to generate a library of variants.

[0021] In certain embodiments, the step of creating the polynucleotide comprises executing a software package that converts the selected tree into instructions to be executed by a liquid handling system to transfer the oligos between storage containers (e.g., tubes or wells of a multi-well plate) in the order given by the selected tree for creating the polynucleotide. Optionally, the mapping software package creates a map of the selected tree onto the hardware capabilities of the liquid handling system, e.g., creates a description of the placement of the oligos within a multi-well plate, and the described placement minimizes the steps or operations of the liquid handling system for creating the polynucleotide when joining the oligos in the order given by the selected tree.

[0022] The selected tree can be asymmetric (the first leaf is connected to the root of the tree selected by the first number of nodes and edges, and the second leaf is connected to the root of the tree selected by a second number different from the first number). The step of selecting the tree can be operated by searching the tree space, and for at least one subgroup of oligos, a tree search software package searches the tree space to identify local maxima in the tree space and selects the highest scoring tree. The tree search software package can use an exhaustive search strategy, heuristic search, genetic algorithm, Monte Carlo search strategy, Bayesian search strategy, maximum likelihood calculation, or machine learning algorithm. This step can be performed to create assembly trees for multiple subgroups of oligos. The step of selecting the tree with the optimal generation score can include calculating the probability of ligation success and also promoting a tree that facilitates parallelization in the liquid handling system (i.e., generally preferring wide, shallow trees).

[0023] Another aspect provides a system for synthesizing a polymer. The system includes a computer system operable to receive information describing a desired polynucleotide and generate a plurality of trees, each tree giving an order of ligation between oligos (preferably provided in a container such as a well) to form the polynucleotide. The computer selects one of the trees having an optimal generation score and directs the system's liquid handling system to create a polynucleotide by joining the oligos according to the structure of the selected tree. BRIEF DESCRIPTION OF THE DRAWINGS

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7-1

Figure 7-2

Figure 8

Figure 9-1

Figure 9-2

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16-1

Figure 16-2

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

[0025] Detailed Description The present invention provides a method for making a biopolymer having a desired sequence of interest. For example, it may be desirable to make one continuous DNA molecule of a predetermined sequence that is arbitrarily long, such as where the molecule is more than 100,000 base pairs in length. In other examples, it may be desirable to make one or a family of variants that share certain defined sequence characteristics and thus a certain sequence similarity, but where the variants do not exactly match each other. For example, it may be desirable to make multiple variants of a vector for delivering a gene or operon, where the variants have different regulatory elements, or different positions or orders of regulatory elements, so that the variants can be tested and evaluated for gene delivery and expression. In that sense, the present invention provides a method for making one or any number of desired polynucleotide sequences, any or all of which may have a completely defined sequence or may have as variants that embody certain defined parameters.

[0026] The problems that arise when making arbitrarily long biopolymers relate to the actual methods of making the molecules. A common theme among the methods of making molecules is to obtain multiple subunits of the molecule and then to couple those subunits pairwise to each other until a long molecule is made. If the subunits are much shorter than the desired biopolymer, for example, when making oligonucleotides or "oligos" that range in length from a few bases to several tens of bases, to long (longer than tens of thousands of bases) nucleic acids, different approaches to joining the subunits together may exist.

[0027] For example, in what could be called a linear approach, the constituent oligos can be identified in the 5' to 3' direction. In step 1, the second oligo is introduced to and ligated to the most 5'-side oligo (which may have its 5' end blocked, for example, by chemical capping or solid-phase binding). In step 2, the third oligo is introduced to and ligated to the ligation product of step 1. In each subsequent step, the next oligo is ligated to the newly emerged polynucleotide. This linear approach is attractive in its simplicity and seems to have fewer variables that need to be controlled to avoid misassembly. However, the linear approach is time-consuming. Trying to make a 100,000-base polynucleotide from 8-mer oligos would involve 12,500 steps, each of which requires pipetting from one container of oligo to another, introducing ligase, factors, incubating, washing away excess reagents, resuspending in saline or buffer, and then moving on to the next step. It seems to take quite a long time to make a 100k polynucleotide.

[0028] Another approach, called hierarchical assembly, involves joining different adjacent pairs of oligos from distal positions along the final desired polynucleotide sequence to create intermediate layers of two-part oligos (e.g., joining 8-mers to form 16-mers), or to create sub-polynucleotides. Next, the two-part oligos can be joined in pairs (e.g., to form 32-mers). These joining steps can proceed in a pairwise fashion to create larger sub-parts of the desired polynucleotide sequence in each layer of the assembly. Hierarchical assembly is attractive because it allows many of the steps to be parallelized. For example, if the starting oligos are 8-mers, step 1 generates 16-mers, step 2 generates 32-mers, and step 3 generates 64-mers. The length generated by each step can be twice the length generated by the previous step, and theoretically, the 15th step generates a polynucleotide with a length of 131,072 bases. Compared to the 12,500 consecutive steps required for a linear approach, a 100,000-base polynucleotide can be assembled in 15 consecutive steps by a hierarchical approach. However, successful assembly by a hierarchical approach can pose potential problems. For example, if a pair of 8-mer oligos are pipetted together from two reaction vessels to accurately form the desired polynucleotide sub-sequence, the pair of oligos must join in the correct order and do not join on their own. Thus, the possible combinations of steps available under hierarchical assembly are limited to pairwise combinations that facilitate the biochemical joining of the free ends of the pair of oligos intended to be joined and discourage the biochemical joining of unintended ends.

[0029] Hierarchical assemblies can be achieved using oligos that can be biochemically joined in the desired order at the intended termini. For example, if the first and second 8-mers A and B are intended to form the 16-mer AB, it is preferable to use oligos that do not biochemically react to form any unintended products such as BA, AAAAA, ABBB, ABAB, etc. This can be achieved by designing and providing oligos with appropriate attachment termini. For example, if A, B, C, D, E, and F are to join to form ADEBCDEF (these are not IUPAC nucleotide codes or amino acid codes; these are just characters representing virtual oligos), A must have an attachment terminus that anneals to the attachment terminus of D. Similarly, D must have an attachment terminus that anneals to the attachment termini of A and C in the intended order.

[0030] For a given desired polynucleotide sequence, it is possible to identify, using bioinformatics, the constituent oligos having annealing sticky ends. Creating one complete split of the desired polynucleotide sequence into the constituent oligos having the identified sticky ends can be referred to as splitting the desired sequence. When the constituent oligos are 16-mers, although there appear to be more than 4 billion possible 16-mers, a polynucleotide of 100,000 bases in length has only about 6,250 unique 16-mers. That number of oligos can be made, stored in a manageable number, about 66 96-well plates, and then dispensed from there. Similarly, that number of oligos can be provided in about 17 384-well plates and dispensed from there. Thus, using a computer system, it may be possible to split the sequence from a desired polynucleotide sequence of about 100,000 bases in length and identify about 6,000 unique oligos having sticky ends that anneal specifically only to the intended sticky ends among the oligos. These oligos can be synthesized or ordered and provided in reaction vessels such as within the wells of about 17 different 384-well plates. Once the constituent oligos are provided in this way, the assembly of the desired polynucleotide sequence can seemingly proceed in a straightforward manner. In the first step, the first oligo is pipetted second (identifying the oligos by their position along the desired polynucleotide sequence), while the third is pipetted fourth, the fifth is pipetted sixth, and so on. This pair is ligated (e.g., by ligation) to form a subpolynucleotide. At the completion of the first step, there are subpolynucleotides that are twice the length of the oligos, and the process is iteratively repeated while treating the subpolynucleotides in the same way as the oligos. In fact, when the process proceeds to completion in a fully iterative fashion, all maps or graphs of the assembly steps are symmetric.

[0031] The pairwise assembly of oligonucleotides into a desired polynucleotide sequence can be described using a map or graph having a tree structure. A tree is a structure composed of branches and nodes. By using the branches of the tree that represent nucleotide sequences and the nodes that represent the connections between pairs of nucleotide sequences, a tree can be used to describe the assembly of the desired polynucleotide sequence from oligonucleotides. In that representation, the root of the tree is a branch representing the desired sequence that connects at a root node representing the final joining operation (e.g., ligation) that actually forms the desired polynucleotide sequence. The terminal branches (or "terminal taxa") represent the starting constituent oligonucleotides from which the desired polynucleotide sequence is ultimately made. The internal branches represent sub-polynucleotides that are shorter than the desired polynucleotide sequence but are made by joining the starting oligonucleotides.

[0032] An advantage of the tree in the context of making a desired polynucleotide sequence from oligonucleotides is that the tree can be represented as a digital file. The ability to represent the tree as a digital file enables several things. First, a software system can automatically generate the tree using logical instructions embodied in software program instructions. Second, a software system can automatically evaluate, compare, or score the tree. Finally, a software system can read the tree and execute program instructions to direct the operation of automated hardware such as a liquid handling system in a laboratory.

[0033] Any suitable format can be used to store and represent the tree, for example, in a graph database. A graph database is an example of a software system that stores and represents relationships between nodes and edges using either node and edge objects, or an index-free adjacency or adjacency list to represent nodes and edges. In another example, the tree can be stored and represented using the directory structure of a file system. For example, in the case of four 8-mers (e.g., A, B, C, and D that assemble into ABCD) that are assembled into a desired polynucleotide sequence of 32-mers, each oligo can be a FASTA file within a Linux®-type file system with nested directories representing a tree structure. The path can be

Number

[0034] The tree can also be graphically displayed in a figure or drawing.

[0035] FIG. 1 is a drawing of a tree 101 for joining oligos named A, B, C, D, and E. In tree 101, terminal branch 105 represents oligo A, terminal branch 107 represents oligo B, terminal branch 111 represents oligo C, terminal branch 113 represents oligo D, and terminal branch 115 represents oligo E. Node 119 represents the step, or action, where the 3' end of oligo D binds or joins to the 5' end of oligo C to form subpolynucleotide DC. Branch 117 represents subpolynucleotide DC. Node 139 represents the step where the 3' end of oligo B joins to the 5' end of oligo A to form subpolynucleotide AB. Branch 135 represents subpolynucleotide AB. Node 131 represents the binding of subpolynucleotide DC to AB, forming subpolynucleotide DCBA, represented by branch 127. Branch 115 represents oligo E. Node 121 represents the binding of oligo E to subpolynucleotide DCBA. Thus, root branch 125 represents the product of the binding of E to DCBA, which is the desired polynucleotide sequence EDCBA.

[0036] Trees such as tree 101 have several important features. First, they represent the order or combination of operations for forming a polynucleotide sequence described from constituent oligos. Referring to tree 101, each of oligos A, B, C, etc. can be provided within a container such as a well of a multi-well plate. Each well can contain any number (e.g., millions) of clonal copies of one oligo. The tree 101 shown represents that D is joined to C and, in parallel, B is joined to A before DC is joined to BA. Thus, in particular, tree 101 notifies a polymer synthesis method by showing that D is joined to C while B is joined to A and only then is DC joined to BA. Considering the desired polynucleotide sequence EDCBA, tree 101 notifies that the method should not be carried out in the simple order of joining D to E, then C to ED, then B to EDC, and finally A to EDCB. A second important feature of trees is that they can be evaluated or compared with each other, which is further explained below. Finally, another important feature of trees is that they can be represented in a computer file such as a plain text file that represents the tree as a text string using only ASCII characters, enabling the tree to be stored in a compact form and easily searched, manipulated, or evaluated.

[0037] For example, in some embodiments, the tree is stored in a text - formatted file, such as by using a modified implementation of the New Hampshire / Newick format as needed. The Newick file format, also called the New Hampshire format, relies on text strings to encode tree representations. Sub - parts of the tree are represented by codes or text strings, and a nested set of parentheses indicates the tree structure. The Newick format has additional features such as encoding of branch lengths, but those features are not necessary at this point in the description. For an explanation of the Newick format, see Pavlopoulos, 2010, A reference guide for tree analysis and visualization, Bio Data Min 3:1, which is incorporated by reference. The Newick Standard for representing trees in computer - readable form makes use of the correspondence between the tree and the nested parentheses. Considering the tree 101 shown graphically here, the Newick standard presents the tree using the following array of printable characters: (E,((D,C),(B,A)));

[0038] A Newick tree ends with a semicolon. Since the desired polynucleotide sequence is pre-determined to be EDCBA, edge 125 is the root rather than the tip, and node 121 is an internal node. Tips and terminal edges, or terminal taxa, are synonymous. Internal nodes are represented by a pair of matching parentheses. Between them, nodes that branch immediately from a node separated by commas are represented. As used herein, a tree can be described in an ordered tree file that is a modified version of the Newick Standard. The Newick Standard does not require that the left - to - right order of the descendants of a node (including, for example, terminal tips) have any meaning. In the true Newick Standard, (A,(B,C),D); is the same tree as (A,(C,B),D);. Also, strictly speaking, the Newick Standard does not require a root. Here, the present disclosure allows the use of a modified tree file where the left - to - right order of the nodes has meaning and the tree necessarily has a root (where the root may be reading the terminal taxa from left to right). The root of the tree is the desired polynucleotide sequence. Thus, (E,((D,C),(B,A))); is the correct modified Newick tree file for tree 101, and the desired polynucleotide sequence EDCBA is the root and can be read from the modified tree file (modified in the sense that the tree file is a particular implementation of the Newick format).

[0039] The advantage of storing the tree in such a text format is that the tree file can be used to control the operation of a liquid handling robot or a robot liquid handler, sometimes called a liquid handling system that includes them. Any suitable type of liquid handling system can be used. Some versions of liquid handling robots dispense a volume of liquid assigned from an electric pipette or syringe, while some such systems manipulate the position of dispensers and containers (often Cartesian robots such as the XYZ Triton Robot made by TriContinent Scientific), and / or integrate additional experimental devices such as centrifuges, microplate readers, heat sealers, heater / shakers, barcode readers, spectrophotometer devices, storage devices, and incubators. Exemplary liquid handling systems include bench-top 8-channel DNA processing robots, and automated liquid handling systems customized for processes such as the TECAN Freedom EVO, automated liquid handlers sold under the trade name PRIME by HighRes Biosolution, or automated liquid handlers sold under the trade name JANUS by PerkinElmer.

[0040] Other versions of liquid handling systems that include "pickers" or "cherry pickers" may be used, which include, for example, liquid handlers that mimic human operation by using orthogonal three-axis movement implemented on a larger workstation by an arm to perform liquid transfer. Such cherry pickers include pipetting robots sold under the trade name ANDREW+ by Waters Corporation (Milford, MA), or the TOMTEC QUADRA 3 cherry picker automated liquid handling system made by Tomtec, Inc. (Hamden, CT).

[0041] Certain embodiments utilize acoustic systems such as acoustic liquid handlers sold under the trade name ECHO 650 by Beckman Coulter Life Sciences (Indianapolis, IN). In an acoustic liquid handler, a multi-well source plate containing reagents is placed in a source plate gripper. A destination plate is placed in a destination plate gripper and the destination plate is inverted and placed over the source plate. A mechanical motor within the handler positions the source well beneath the destination well. A transducer is placed beneath the source well and is typically hydrophilically bonded to the bottom of the source well in water. The liquid handler uses the transducer to measure the distance to the bottom and meniscus of the well in the source plate and then delivers a burst of ultrasonic energy calculated by the handler to eject a droplet of liquid from the source well in an orbit that impacts the bottom surface of the upper inverted destination well. A liquid handler such as the ECHO 650 is reliably operated to transfer liquid volumes on the order of 2.5 nL. Due to surface tension, the transferred liquid is retained at the bottom of the well of the inverted destination plate. Larger volumes can be transferred by a series of acoustic bursts.

[0042] Regardless of whether the liquid handling system uses a multi-channel processing robot, a cherry picker, or an acoustic liquid handler, what is common to such liquid handling systems is that they can typically be operated under computer control. The computer control system can store information regarding reagents and products, as well as instructions for operating the liquid handling system. The computer control system can include a computer incorporated into the liquid handling system; a computer workstation operably connected to the liquid handling system; a server or cloud computing resources; and one or any combination of, for example, network computers connected over a LAN, or via WiFi or the Internet. The computer control system can read from a tree file and direct the operation of the liquid handling system to synthesize (at least a portion of) a desired polynucleotide sequence.

[0043] The computer control system achieves the synthesis of (at least a portion of) a desired polynucleotide sequence by instructing the liquid handling system to operate so as to execute assembly steps represented by an assembly tree.

[0044] When the desired polynucleotide sequence is made by hierarchical assembly, perhaps the simplest example of a perfect iterative assembly is that described by a symmetric tree. An important insight of the present invention is that a symmetric tree may not describe the best way to make the desired polynucleotide in the laboratory and usually is likely not to. One problem is that a symmetric tree is simple with respect to a number of factors such as the biochemical properties of the reagents, the opportunity to utilize the sequence content of the desired polynucleotide sequence, or the practical considerations of the actual laboratory hardware. For example, multi-well plates often have reagents in rows or columns containing wells that are multiples of 8 or 12, and various laboratory instruments are designed to pipette from or into 8, 12, 16, 24, etc. wells in a single operation. Depending on the experimental apparatus used, a symmetric assembly tree may need to be divided into parts that match the rows or columns of those plates, and it may be required to leave wells empty.

[0045] Perhaps even more importantly, even uniquely matched sticky ends do not always anneal with perfect fidelity, and non-matching sticky ends may anneal to each other. Also, importantly, certain sticky ends have properties that prevent the successful synthesis of the desired polynucleotide sequence even if they result in a unique match. For example, if a sticky end has a self-complementary internal region, that region may form a hairpin. The 5’ GC of a sticky end containing GCAAAACG may anneal with 3’ CG, and that end becomes unavailable for use with another oligo. Or, if a sticky end is extremely AT-rich or GC-rich, it may or may not anneal under conditions where other paired oligos in a multi-well plate are intended to do the opposite.

[0046] The insight of the present disclosure is that one assembly tree represents one set of an ordered combination of steps through which a desired polynucleotide sequence can be created via pair-wise ligation of constituent oligos. Also, there are numerous sets of ordered combinations of steps through which a desired polynucleotide sequence can be created via pair-wise ligation of constituent oligos. This means that there are numerous assembly trees for a given desired polynucleotide sequence.

[0047] Considering that one objective of the present disclosure is to provide a method for successfully creating a desired polynucleotide sequence, the present disclosure provides a method for selecting an appropriate assembly tree. Briefly stated, these methods include generating a number of trees representing the splitting of the sequence into oligos and representing the creation of the desired polynucleotide from the set of oligos, scoring those trees based on practical and biochemical criteria applied to the characteristics of the trees, selecting a tree that yields a score indicating that the polynucleotide is successfully created as per the specification (i.e., selecting a tree that includes an optimal split and an optimal combination of branches specifying the desired sequence), providing the set of oligos indicated by the split, and executing program instructions that direct a liquid handling system to assemble a polynucleotide molecule having the desired sequence.

[0048] Figure 2 shows a polymer synthesis method 201. Method 201 includes a step 205 of inputting a desired polynucleotide sequence into an in silico computer system. The desired sequence can be input in any suitable format such as a simple text string or a FASTA or FASTQ file. The sequence can be input into the computer system, for example, by receiving the sequence via the Internet (205). That is, the polynucleotide can be synthesized or produced, for example, in a facility or laboratory that receives the desired polynucleotide sequence as an order via a web form or email. The next step is to input or instantiate an in silico graph structure at 207. There are various steps or approaches that can be used. For example, the nucleic acid components of the desired polynucleotide sequence can be represented using a graph database such as Neo4J. In certain embodiments, a graph structure composed of branches is generated such that each graph structure defines a linear order of the nucleic acids. In a preferred embodiment, the computer system creates or stores the sequence of each of a plurality of constituent oligos that can be used to construct the desired polynucleotide sequence. The programming logic operates to select an optimal combination of branches that define the desired sequence, which is used to direct the assembly of a polynucleotide molecule having the desired sequence.

[0049] In a preferred embodiment, step 207 of instantiating the graph structure may include generating a plurality of assembly trees (e.g., tree 101, etc.). Each tree represents one set of an ordered combination of steps through which a desired polynucleotide sequence can be made via the joining of pairs of constituent oligos and the subpolynucleotides resulting from the first round of joining. Trees representing assembly products that conflict with the desired polynucleotide sequence can be deleted at the software level, so the trees can be generated indiscriminately with respect to the assembly products they represent, as needed. Importantly, some of all the assembly trees generated will, if executed, yield the desired polynucleotide sequence.

[0050] Method 201 includes step 211 of selecting a tree structure having an optimal combination of branches. Step 211 of selecting a tree by a method conducive to automation is discussed below. Once a tree is selected (211), method 201 includes step 215 of assembling the polynucleotide, typically by executing computer program instructions that read the selected tree and direct laboratory equipment to manipulate reagents to form the desired polynucleotide sequence.

[0051] Step 211 of selecting a tree includes selecting (from among a number of such trees) an assembly tree for which the laboratory equipment has a high probability of successfully forming the desired polynucleotide sequence. Because the specific joining reactions represented by the nodes of a tree are prone to failure or to generating non-specific reaction products for biochemical reasons, some trees may not be associated with a high probability of success.

[0052] Each of the trees represents a series of operations that are performed using tangible hardware and prepared biological reagents. Different trees represent different ordered combinations of steps for creating the same desired polynucleotide sequence. Some trees may include operations that, when performed, result in poor performance. For example, the subpolynucleotide AB can be created by joining together one of two pairs that involve either joining A to B' or A' to B, and AB will be obtained regardless of the combination. However, the difference between A and B' compared to A' and B can be at each attachment end. As a fictional example, the 3' attachment end of A' forms a GC-rich hairpin that inhibits joining A' to B. In this example, attempting to join A' to B is likely to fail. Here, it is preferable (211) to select the tree in which A is joined to B' rather than the tree in which A' is joined to B.

[0053] In another example, the step 211 of selecting a good tree can help avoid non-specific or unintended reaction products. For example, there are different combinations of steps to create ABCD from A, B, C, and D.

[0054] Figure 3 is a graphic representation of a symmetric tree 301 for creating ABCD.

[0055] Figure 4 is a graphic representation of an asymmetric tree 401 for creating ABCD. However, when A and B are mixed, they may tend to form not only AB but also BA. In such a case, the symmetric tree 301 does not necessarily ensure the creation of only ABCD. However, the tree 401 always avoids combining free A and free B in such a way that the 3' end of B is exposed to the 5' end of A. Assuming that all other joins are performed reliably, the tree 401 will then reliably generate ABCD.

[0056] The present invention provides a method useful for automatically selecting tree 401 over tree 301. One suitable approach is, for example, to write all of trees 301 and 401 in Newick format in a single file. Such a file can include (among other things) the following lines: 301((A,B),(C,D)) 401((A,(B,(C,D)) 501(A,((B,C),D))

[0057] In such a file, the first line is the Newick representation of tree 301, and the second line is the Newick representation of tree 401. The system can proceed through the file and assign a score to each tree. The score can be assigned by using prior knowledge about the array problem or probability of success. Such information can be calculated and stored in a suitable data structure such as a matrix or variable structure (an array of hashes is well-suited, where the 5' oligo of the ligation is an index to an array entry that points to the hash unique to that oligo; in the hash unique to that 5' oligo, the 3' oligo of the ligation is a key whose paired value is the ligation score). The score can represent the probability of successfully producing only the intended reaction product at any node of the tree (note that a node represents a combination of pairs of sub-parts). Since oligos have different 5' and 3' ends, the score can be represented as an asymmetric matrix. For example, if the score ranges from 0 to 1 and it is known that mixing B with A forms BA rather than AB, the score for combining A and B to form AB can be set to 0.

[0058] FIG. 5 shows a matrix M501 that gives a score for the probability that the 3'-end of the i-th entry binds to the 5'-end of the j-th entry. Matrix 501 is not empirical and is shown only to illustrate that when B and A are mixed, they form AB and BA, while when C and B are mixed, they form only BC and not CB. Other encodings and examples are within the scope of the present invention. Matrix 501 is one exemplary approach. Matrix 501 is also informative in that, for example, when C is mixed with D, it shows that CD rather than DC is formed. Matrix 501 shows that A, B, and C are not self-polymerases, but reveals that D has a moderate probability of binding to itself. The computer system can automatically apply matrix 501 to the Newick format versions of trees 301 and 401. For example, for each pair of Newick trees, the computer system can calculate the probability that the indicated products are formed. In the first row, the entry "(A,B)" indicates that the intended product is AB when A and B are mixed. The computer can read the probability that the 3'-end of B joins the 5'-end of A and score the entry for (A,B) as zero. Because the mixing of A and B has a high probability of producing an unwanted reaction product. The score for tree 301 can be calculated by multiplying all the scores in that row together. In such a case, the score row for the Newick tree of 301 includes a zero for the entry "(A,B)", and multiplying the entire row results in zero. The computer system scored tree 301 as 0.

[0059] In Newick format, tree 401 shows that A joins to BCD. The computer program examines the probability that the 3' end of A joins to the 5' end of B (1), and the probability that the 3' end of D joins to the 5' end of A (0). The entry for that row is scored as 1 (because the matrix does not show the probability of creating the undesirable product BCDA). In the virtual example, the computer system scored tree 401 as 1.

[0060] The important insight here is that arbitrarily many trees (e.g., millions) can exist. The computer system can write each one to a tree file, but in practice, there are heuristics that can ignore trees. The matrix can be obtained empirically. For example, the matrix can be obtained by some combination of laboratory test results and thermodynamic predictions performed in silico. The computer system can apply this matrix to score each tree. The in silico activity of scoring trees can be parallelized. For example, write half of the trees to one file and send them to one processor or one node of a Beowulf cluster, while sending the other half to another processor in a different file. Interestingly, the process of scoring trees does not necessarily need to save or record all scores. This process can simply hold the last calculated score in a variable while scoring the next tree, then compare the newly calculated score to the last calculated score and keep only the better score. Of course, it may be suitable to write all scores to a log.

[0061] Importantly, the disclosed process scales up arbitrarily. Matrix 501 represents joining four oligonucleotides together. The systems and methods of the present invention can be used to create a desired polynucleotide sequence on the order of 100,000 bases or longer.

[0062] Figures 6-10 are used to illustrate certain embodiments for creating a scoring matrix useful both in determining the partitioning of a desired polynucleotide sequence and in selecting the final assembly tree.

[0063] Figure 6 shows a plurality of partial ds oligos that can be used in the assembly described herein. As shown, each molecule completed by covalent bonding is 8 bases long and they overlap by 4 bases each. Sixteen oligos are shown. After the partitioning step, such oligos are provided to the laboratory compartments. They can be synthesized based on the results of the partitioning, or the laboratory may have a large enough library such that after partitioning, an appropriate library plate can be withdrawn for use. For example, each oligo can be provided to a well of a multi-well plate, and each well contains, for example, millions of clone copies of the oligos shown. The long desired polynucleotide sequence can be made by ligating the oligos together in a pairwise fashion to form longer sub-polynucleotides. The method of the present invention partitions the desired sequence to identify the oligos to be used and predicts the success of creating the desired polynucleotide sequence when ligating the oligos together in a pairwise fashion. This method can include (i) a step of partitioning the desired sequence into a set of oligo sequences; a step of assigning a score (i.e., a positive score indicating the probability of obtaining the intended result) indicating the probability of successfully ligating pairs of oligos together, and (ii) a step of identifying pairs of oligos having overhangs that may misligate (i.e., a negative score indicating the risk of obtaining something other than the intended result).

[0064] FIG. 7 shows a ligation matrix 701 that gives a score representing the probability that two overhangs ligate successfully together. Overall, the matrix represents the scoring of each pair of overhangs of the array. A high score means a high probability of successful ligation. Note that the row / column labels may appear to be duplicated because the matrix covers each pair of overhangs that are proposed to be ligated together when creating one array. Even positions refer to the 5' lag strand and odd positions refer to the 5' lead strand.

[0065] FIG. 8 shows a mapping matrix 801, or a partitioning matrix, which, when a perfectly symmetric assembly tree is assumed, shades blocks that represent any pair of overhangs that could misligate at any point during the assembly process. The mapping matrix 801 is specific to the assembly tree and partitioning, and shades blocks where an oligo could misligate to an unintended oligo, forming something other than the intended result. The mapping matrix provides information used when selecting a partitioning.

[0066] An important part of the splitting is that the computer system separates, in silico, the desired polynucleotide sequence into constituent oligo sequences with specified attachment ends. The computer system can repeatedly propose a plurality (e.g., dozens, hundreds, hundreds of thousands) of splits and execute a tree search process for each split. The mapping matrix is used to score the tree for pairs of overhangs that may misligate, i.e., for undesirable ligation. An example of undesirable ligation and how it can be resolved by splitting and tree searching is self-ligation. When a split proposes an oligo with self-complementary attachment ends (e.g., when automatically generated by the programming logic of the computer system), the mapping matrix penalizes the problem of that oligo binding to itself. The computer system can adjust the split by making one of the attachment ends of the oligo shorter (and 2, 3, etc.) or longer by 1 base and repeat the split. The mapping matrix encodes the biochemical possibilities of annealing between attachment ends. Making the attachment ends 1 or 2 bases longer or shorter allows the mapping matrix (when applied to the assembly tree) to score it so that it can be used for that positioning. The positioning is updated or elevated, which means that the oligo with the attachment ends indicated by that split is directed to be used. From that in silico direction, the indicated oligo can be synthesized or retrieved from a library for use.

[0067] Figure 9 shows a scoring matrix 901 resulting from the combination of a ligation matrix 701 and a mapping matrix 801. Combining the ligation matrix 701 and the mapping matrix 801 gives an overall score (e.g., the negative value of the sum of all shaded scores in this display). For a given assembly tree, the scoring matrix 901 is used to score the tree. Since the tree can be represented as a linear string in text form, a computer software module can read along the tree, score each element, and construct the final score for that tree. By such a process, each tree has its own score. In these embodiments, note that the tree represents both the proposed splitting of a desired sequence into constituent oligo arrays having specified attachment ends and the steps for assembling the desired polynucleotide on a liquid handling system. Thus, selecting a tree can guide both which set of oligos to provide and the operation of the assembly device.

[0068] Figure 10 shows a symmetric assembly tree 1003 and an asymmetric assembly tree 1004. It is important to note that (i) typically there are millions of trees and (ii) the trees can be stored in a plain text file, e.g., in Newick format. That is, the symmetric assembly tree 1003 and the asymmetric assembly tree 1004 can be hundreds, thousands, or millions, or more, that are generated. Scoring each tree using the scoring matrix 901 reveals that the symmetric assembly tree 1003 has some array problems, which means it is difficult to synthesize the desired nucleotide sequence on experimental equipment. However, the shown asymmetric assembly tree 1004 gets a good score. Using the labels from tree 1004, this suggests that d2 should be coupled to d3 and the resulting product should be coupled to d4 (steps not found in tree 1003).

[0069] FIG. 11 shows that different assembly trees have different corresponding mapping matrices. FIG. 10 shows the use of the scoring matrix 901 for assembly tree optimization. The tree shown on the left is the proposed tree before scoring. From these, a mapping matrix is generated, and then this mapping matrix is used to score the tree. Different trees represent different ways of combining components by changing the order of the ligation reactions. Thus, the method of the present invention optimizes the assembly tree by scoring all possible trees, which is done by combining the mapping matrix and the split matrix, and then selecting one tree with an appropriate score. In a simple case, the appropriate score may be the highest numerical score, but for various reasons, other trees may be selected (e.g., second highest, or nearly the highest, or top 10 percent, or good enough, or top quarter). For example, since the full width of a 96-well plate can be used to proceed to assembly 225, some monophyletic groups can select a tree consisting of 8 terminal taxa.

[0070] Note that some of the generated tree structures may not be "good trees", i.e., they may not represent an ordered combination of branches that results in the desired polynucleotide. However, these "bad trees" do not necessarily prevent the selection 211 of appropriate trees. In fact, part of the logic of the selecting step 211 in a computer system may involve selecting a tree that produces the desired polynucleotide sequence. That is, the presence of any bad trees in the tree file may not be a problem since they may simply not be selected. This approach allows for the generation of trees in a quasi-random manner and thus admits one strategy for tree generation and scoring. Trees can be generated regardless of the products they represent, and the selection of the products themselves can be included in the selection of the probability of successfully producing those products using laboratory equipment and reagents. This note is particularly applicable in embodiments of Bayesian tree search that can generate one or a limited number of starting trees, randomly modify the starting trees in a Markov chain style, evaluate the modified trees against the previous trees, and repeat the modification and evaluation from the better of the two (a process that can be parallelized and used to cover a larger tree space by periodically swapping parts of the trees between chains being processed in parallel). Regardless of the specific tree search algorithm used, preferably, the selecting step 211 involves calculating the score for each of a plurality of trees. The calculated score represents a measure of success in being able to successfully generate a molecule having the desired polynucleotide sequence using the tree structure. The measure of success can be a predictor of any relevant factor (e.g., correlated with these) in producing the product, including, for example, not only success or failure but also, optionally, any other physicochemical measure correlated with molecule concentration, relative concentration, or assembly yield.

[0071] Although a matrix was shown in the above embodiments, the probabilities for each ligation step or the probabilities associated with each oligo or sub - polynucleotide can be stored in any suitable format or structure. As shown above, the probabilities were applied to tree nodes in order to evaluate the probability of ligation success. Oligos or sub - polynucleotides can also be scored (compared to the action of joining them). For example, a score can include a score representing a risk that a stretch of bases includes risks such as a problematic restriction site, an extreme GC value, a run of CpG repeats, a risk of thymine dimers, etc. The score can represent the probability of success when assembling a partial construct and / or the stored probability of mis - assembly.

[0072] As described, scores are assigned to multiple assembly trees. Preferably, the computer system selects a tree structure that provides an optimal combination of branches (211). Note that the score may be a numerical value, but it is important to note that it is not necessary or required to select "the" optimal tree. Millions of trees can be scored from 0 to 1, for example. It may be suitable to simply select trees with scores greater than 0.9. For example, these systems can score trees until some (e.g., 15) trees obtain scores greater than 0.90, then select one of those with the highest score and proceed to direct the assembly of the desired polynucleotide sequence (even if millions of trees have not yet been scored). Similarly, the system can output several top-scoring trees (e.g., hundreds), including one with the highest numerical value. However, the highest-scoring tree may not be used. For example, problems may become apparent through human review. Or, a wrapper script for liquid handling equipment can include data regarding the equipment and select a tree with certain characteristics particularly suitable for the liquid handling equipment (e.g., a number of single-line blocks with 96 terminal crades because the liquid handling equipment can efficiently construct them using a 96-well plate). Thus, for example, in some embodiments, the computer system of the present invention searches the tree space to identify local maxima and algorithmically selects one tree from among them, or pushes out several trees (e.g., 10) for human review or for evaluation by a laboratory equipment software wrapper.

[0073] Other strategies for selecting an assembly tree (211) are also within the scope of the present invention. For example, a computer system can execute a machine learning algorithm to select a tree having an optimal combination of branches for assembling (215) molecules. The advantage of using a machine learning algorithm is that, if desired, it is not necessary to provide a matrix of scores for the nodes (and / or branches) of the tree using a machine learning system. Such scores can be used during step 211 of selection. However, these scores can be used as part of the training data when training the machine learning system. Further, it is not necessary to use such scores. For example, a tree file, an output polynucleotide, and a success metric can be used as training data to train the machine learning system. The machine learning system can be trained using training data including, for each of a plurality of polynucleotides produced, the tree used when producing the polynucleotide, one or more alternative trees that were not used if desired, the sequence of the polynucleotide if desired (since that information is specific to the tree and is thus partial if desired), and success metrics such as production time, production cost, number of failures, cost of discarded reagents, etc. Interestingly, a machine learning algorithm can be used to generate a matrix of scores, and that matrix of scores can be used later in a non-machine learning approach for tree selection if desired. Using machine learning to generate the score matrix is attractive because it should work with even a smaller training data set but generalize up to a large data set. The machine learning system can be trained with trees, oligos, and a score matrix, where each tree has, for example, 12 terminal taxa and 12 oligos (plus their attached ends) are provided. After such training, the machine learning system is provided with a test input including 96 oligos (and one or more trees if desired). The trained machine learning system can then output a scoring matrix for a tree having 96 terminal taxa.Any suitable machine learning system may be used, whether the machine learning system is used to generate a score matrix, to select a tree, or both. Suitable machine learning systems may include one or more of a decision tree algorithm, a random forest algorithm, an extreme gradient boosting algorithm, an adaptive boosting algorithm, or a deep learning algorithm such as a deep neural network.

[0074] A particular machine learning algorithm may be selected based on a particular problem. For example, a random forest algorithm is well-suited to evaluate text inputs and compare between a large number of entries having text strings in a comparable format. Thus, a random forest algorithm can be used to select from a very large number of Newick trees.

[0075] In another example, a convolutional neural network (CNN) transforms the dimensionality of the input in the feature representation. As such, a CNN can function well with multiple inputs of different scales or dimensionalities. Thus, a CNN can function well even when trained on small trees but used to score much larger trees. In fact, a CNN can function particularly well when scoring large trees when trained by a multiple instance learning (MIL) model on a bag of small trees that are sub-parts of much larger trees. In MIL, each bag typically has one score attached to the entire bag. Thus, training a CNN by MIL can include presenting training data that includes bags of small trees (e.g., 12 terminal taxa), all the trees within a bag taken from one large tree (e.g., 96 terminal taxa), and a score representing the success of polynucleotide production by that bag. An option for training a machine learning algorithm is to train the ML algorithm in the background while executing Method 201 of the present invention using non-machine learning programming to produce some desired polynucleotide sequences over time. For example, write a tree scorer that multiplies matrix scores by Newick tree elements (e.g., in Python), and then use that system to produce long (>100 kb) polynucleotides over a period of time while passing bags of data to the CNN for MIL training. After a period of time, the CNN may be trained and may be able to step in, score, and select the tree.

[0076] Once the tree is selected (211), method 201 includes step 215 of assembling the desired polynucleotide sequence. Any suitable tools, reagents, or equipment can be used to create the polynucleotide. Preferably, the reagents are provided in some type of container or compartment and include oligos that are stored. In some embodiments, the oligos are provided within the wells of a multi-well plate. In other embodiments, the oligos are each provided within a tube such as a microcentrifuge tube. The oligos may be in an aqueous solution or suspension, or may be provided within a hydrogel bead or matrix as needed. The oligos can be dried, for example, lyophilized, and resuspended or dissolved in step 215 of the assembly. In some embodiments, a microfluidic device is used in the assembly step, and the microfluidic device handles droplets (e.g., aqueous droplets surrounded by immiscible phases) that include a portion of the desired polynucleotide and, optionally, reagents for assembling the desired polynucleotide.

[0077] Assembly 215 can be accomplished by directing a robotic fluid handling device to create the desired polynucleotide.

[0078] The present invention has various important features that can form important parts of specific embodiments. For example, in different use cases, the method of the present invention is suitable for different applications such as creating one specific large-scale polynucleotide having a specified sequence or creating a variant library. In another example, recursive assembly helps to keep the assembly method of the present invention practical when creating very long molecules. In another example, the method of the present invention maps the assembly tree to the equipment so that the tree can be customized or the instructions for the equipment can be customized in a way specific to experimental equipment such as a liquid handling system.

[0079] Use Cases The present disclosure is applicable to a plurality of use cases in which the desired biopolymer is synthesized.

[0080] A first such use case is genome-scale synthesis. The core of genome-scale synthesis is to receive an order that describes a desired polynucleotide sequence of one genome-scale polynucleotide, for example, a length significantly longer than 10 kb, typically longer than 100 kb. The method includes providing a plurality of oligos from which the genome-scale polynucleotide is made; generating each of a plurality of assembly trees that describe an ordered series of operations for joining the oligos together to form the genome-scale polynucleotide; scoring the trees to give an assembly score to each tree, where the score represents a predicted measure of success that the tree will result in a molecule having the desired polynucleotide sequence; selecting a tree having a score that meets a threshold; and directing laboratory equipment to make the polynucleotide represented by the tree.

[0081] Recursive assembly may be used for genome-scale synthesis. As described in more detail below, recursive assembly involves selecting an appropriate subtree for a subset of oligos and treating that subtree as a terminal taxon when selecting a higher-level tree that includes the subtree. The recursion can be multiple, having multiple identified subtrees used at parallel levels of the higher tree or multiple subtrees nested within each other in the higher tree. Of particular importance in genome-scale synthesis is one of the desired polynucleotide sequences or a provisional of a particular number or amount of cloned copies.

[0082] Another use case for the method of the present invention is provisional for a library of variants. Generally, in such cases, a variant refers to several molecules that have a number of common elements and also a number of differences between them. As an example of a library of variants, among the variants, expression vectors for genes or operons in which coding sequences and / or regulatory elements (e.g., promoters and repressors) are rearranged, duplicated, or omitted can be mentioned. A user may, for example, desire to have dozens, hundreds, thousands, or more such variants that include the open reading frame of a protein of interest but include various arrangements of regulatory elements, exons, or both (e.g., by omitting a particular exon, expression by an expression vector of a particular splice variant in question can be promoted). A library of variants can be used to express a wide variety of proteins such as antibodies or enzymes, and then the library can become a useful tool for screening, for example, antibody binding, enzyme efficiency, etc., and can be useful in metabolic engineering.

[0083] Recursive assembly As a practical matter, tree selection may need to be approached heuristically. For example, when assembling a 100 kb sequence from 8mers, there can be more than 10^50 possible assembly trees. With the computing power available, there may not be room to generate and score all the assembly trees. One available heuristic is to break the tree generation / selection task into smaller parts, perform method 201 of the present disclosure on the smaller parts to generate subtrees, and then concatenate the subtrees together into the selected assembly tree. Interestingly, recursive tree selection in silico can provide efficiency when the molecules are synthesized. Each subtree of the recursive tree can represent a subsegment of the desired polynucleotide that can be made using one “execution” of a robotic handler or one plate of reagents, so the recursive tree (even if the recursive subtrees are the product of heuristics to improve the tree search algorithm in silico) can itself serve to minimize plate exchanges or aid in the modularization of equipment to improve parallelization.

[0084] To illustrate an example of recursive assembly, one can take a group of some number (e.g., 12) of contiguous oligos found together in the desired sequence and perform method 201 on those 12 oligos to select one subtree. That subtree can then be used recursively as one of the terminal taxa when selecting a higher level tree.

[0085] Figure 12 shows recursive optimization. Considering combinatorially a large number of possible trees, the partition is divided into subgroups, and each subgroup is optimized independently. Each subgroup is reduced to a "leaf" of a higher-level tree, and this process is repeated. This figure shows two levels of recursion where the subgroups form subtrees that are used as leaves of a higher-level tree. This method can include the step of joining subgroups of branches together to form a branch path containing the desired polynucleotide sequence. The oligos can typically be between about 8 and 40 nucleotides in length (generally having single-stranded attachment ends and double-stranded middle segments). Without recursion, there could be more than 10^50 assembly trees that would need to be made otherwise, which can be computationally difficult for some computing resources. By performing recursive assembly, a fully optimized tree can be identified for any arbitrarily large (e.g., genome-scale) polynucleotide.

[0086] Figure 13 shows a fully optimized tree obtained from recursive assembly. The disclosed method provides for recursively optimizing a complete assembly tree without overly demanding computational resources. As a positive side effect, the tree remains shallow enough to allow for a high level of parallelization during automated assembly.

[0087] Mapping of the tree to the device In all cases, the present disclosure relates to the operation of experimental equipment for making synthetic biopolymers. Different experimental equipment may be used, including cherry pickers, robotic liquid handlers, and acoustic liquid handlers, but the processes share the feature of combining short oligos to make very long (e.g., genome-scale) polynucleotides that are (preferably) covalently contiguous. Such devices generally operate using standardized multiwell plates or arrays, or similarly, strips or racks of tubes having standardized dimensions or numbers. When an assembly tree is selected, the tree is mapped to the experimental equipment.

[0088] In one exemplary embodiment, a computer system that directs a liquid handling system parses an assembly tree into a single clade having the maximum number of terminal taxa without exceeding a number specific to the laboratory consumables, and at the same time also seeks to minimize the number of such clades. For example, the liquid handling system may operate with a 384-well plate. The computer system may include a wrapper script that maps the tree to instructions recognized by the liquid handling system. The wrapper script can receive the number 384 as a flag ("-384") when given a selected tree file as input. Next, the wrapper script can greedily identify groups within the tree and attempt to identify the largest contiguous (i.e., single-clade) subtree having no more than 384 members. The wrapper script can then take each such subtree out and issue pipetting or transfer commands to the liquid handling system. For example, if the subtree includes (D,Q)(P,N)... and the wells are referenced accordingly, the wrapper script can issue commands to transfer 5 nL from source well D and 5 nL from source well Q to destination well 1, while transferring 5 nL from source well P and 5 nL from source well N to destination well 2. In subsequent steps, the destination wells can become source wells and the wrapper script can issue commands to mix components from them. With such an algorithm, the wrapper script maps the assembly tree to the liquid handling system.

[0089] Other examples and embodiments are within the scope of the present disclosure. For example, a liquid handling system can operate with a 384-well plate. The system can have a number (e.g., 17) of jobs in queue. A wrapper script can scan the jobs to find jobs that can be fully accomplished from 192 wells that make up half of the 384-well plate. The script can identify that jobs 2 and 11 are such jobs and can start these jobs to be executed simultaneously from the shared 384-well plate.

[0090] Examples of the use of a library of variants lead, in particular, to benefits from an intelligent mapping of an assembly tree to a liquid handling system. In the case of common variants, a number of variants have a large common segment, although at different positions along the final variant molecule. The script can recognize those common segments and can direct them to be made early and in excess, can store those segments in their own containers (e.g., wells), and then when each variant is made, the system can process the segments as oligos (this is not simultaneous and is also an implementation of recursive assembly).

[0091] In yet another example of mapping a tree to a device, when using a liquid handling robot equipped with a multi-channel pipette, the system can enforce an all-or-non rule for rows or columns of the plate such that the pipette is either (i) used in all wells of that row or (ii) not used at all as the robot passes over the plate. This can lead to a certain apparent inefficiency, for example, when planning to join 94 oligos, the system uses only 88 wells from a 96-well plate but then uses a new plate for the remaining 6 oligos, but this can have the benefit of avoiding cross-contamination.

[0092] The system and method of the present invention are executed using experimental equipment under the control of an operating system.

[0093] Figure 14 shows the components that may be included in the system 1401 of the present invention. Typically, at least one liquid handling system 1415 is under the control of at least one computer system 1409. The system 1401 receives a desired polynucleotide sequence 1427 (e.g., sent by email in FASTA format from a client computer 1405 as needed). The client computer 1405, the computer system 1409, and the liquid handling system 1417 preferably communicate via a communication network 1417 that may include any combination of local network hardware, the Internet, and cellular communication networks. The system may preferably have access to storage 1413. Either or both of the computing system 1409 and the storable 1413 may be provided by one or more computers, servers, or cloud computing resources (e.g., Amazon Web Services (AWS) on-demand server computers). The storage 1413 may preferably contain information regarding a plurality of oligos provided within the experimental device 1419 operably coupled to the liquid handling system 1415. For example, the experimental device 1419 may be a robotic liquid handler having a plurality of multi-well plates housed therein, and each well of the plate may contain a large number (millions) of clone copies of an oligo. The sequences of those oligos and their positions within the plate may be stored in the storage 1413. Upon receiving the desired polynucleotide sequence 1417, the computer system 1409 executes a method 201 for selecting an assembly tree (211). Optionally, this method includes steps of identifying oligos that divide and use the desired polynucleotide sequence 1417, and using an experimental apparatus to create or provide the identified oligos. The system 1401 directs the assembly of the polynucleotide 1451 by the liquid handling system 1415 according to the selected tree.This method generates the polynucleotide 1451, a novel synthetic molecule, which can be provided in a suitable container or tube, such as a microcentrifuge tube in solution, or in a dried (e.g., lyophilized) or matrix form, e.g., embedded within a hydrogel bead, frozen (e.g., at -80 °C) for long-term (over 30 days) storage if needed. The polynucleotide 1451 within this tube can be shipped, optionally packaged (e.g., in dry ice).

[0094] In system 1401, computer 1409 is operable to receive a desired polynucleotide sequence, generate a plurality of trees, each tree giving an order of ligation between oligos 601 to form a polynucleotide; select one of the trees having an optimal generation score; and create polynucleotide 1451 by ligating oligos 601 in the order given by the selected tree 1004. Preferably, the leaves of the tree represent oligos and the nodes of the tree represent ligations between oligos. However, it is important to note that these can be reversed, and systems that store oligos in nodes and represent ligation by edges are typically equivalent semantic variants in all functions. System 1409 selects a tree by calculating a score (211) for each tree that represents the probability that a nucleic acid can be successfully created using that tree. Preferably, the tree is scored using (i) a stored matrix 701 of the probability of success of ligating overhangs of oligos; (ii) a stored estimate matrix 801 of the risk of misligation between unintended pairs of oligos, or (iii) both of these.

[0095] As shown in FIG. 12, the steps of generating and selecting can be performed on a first subset of the oligos to form an intermediate tree; the intermediate tree can be used as a leaf in selecting the final tree having an optimal generation score. A recursive version of this method includes identifying sub-groups, each including a computationally tractable number of oligos (e.g., less than about 20-100 oligos); generating, scoring, and selecting a sub-tree for each sub-group; and joining together one optimal scoring sub-tree from each sub-group to obtain a selected tree having an optimal generation score. Within computer system 1409, a software package can generate all possible rooted branched trees, store all the trees in memory, apply a scoring matrix to each of the trees in memory to score all the trees, and select the highest scoring tree for a sub-group. Computer system 1409 can use the software package to convert the selected tree into instructions to be executed by liquid handling system 1415 and transfer oligo 601 between storage containers including reaction vessels having reagents such as ligase as needed in the order provided by the selected tree to create polynucleotide 1451. In an embodiment of mapping the tree to the instrument, computer system 1409 can create a description of the placement of the oligos within a multi-well plate, and the described placement minimizes the steps or operation time of liquid handling system 1415 that creates polynucleotide 1451 by joining oligo 601 in the order provided by the selected tree 1004. Preferably, the selected tree is asymmetric (the first leaf is connected to the root of the tree selected by a first number of nodes and edges, and the second leaf is connected to the root of the tree selected by a second number that is not equal to the first). Computing system 1409 can perform step 211 of selecting in a biased manner that prefers shallower trees over deeper trees because shallower trees maximize parallelization and minimize the time to create polynucleotide 1451.

[0096] The present invention provides a method for synthesizing large biopolymers such as genome-scale polynucleotide molecules. When a genome-scale polynucleotide is created by joining several oligos together, the assembly of the desired polynucleotide from the oligos is represented as a tree where the branches represent nucleotide sequences and the nodes represent the bonds between pairs of oligos. Each of a number of assembly trees is generated and scored in silico according to how successful that tree is in directing the assembly of the desired polynucleotide, taking into account the biochemical properties of the oligos and the proposed ligations. The systems and methods of the present invention select an appropriate tree, operate a liquid handling system to perform the operations indicated by the selected tree, and create the desired polynucleotide.

[0097] Embodiments can use a microfluidic handler that can transfer the contents of all or part of one or several compartments continuously or in parallel to other pre-specified compartments that may or may not be empty. The reaction product can be purified, for example, by gel electrophoresis, hybrid capture with biotinylated probes, chromatography, or affinity separation methods. The assembly method preferably includes connecting oligos by a ligation reaction including enzymatic, chemical, or adapter ligation. Some embodiments use a ligase such as any one of T3, T4 or T7 DNA ligase, or RNA ligase, polymerase or ribozyme. Preferably, T4 DNA ligase, T7 DNA ligase, T3 DNA ligase, Taq DNA ligase, DNA polymerase, or engineered enzymes are used. Preferably, the following ligation reaction is used: T4 DNA ligase at a concentration of 10 cohesive end units per μL supplemented with 1 mM ATP (Sambrook and Russel, 2014, Chapter 1, Protocol 17). Assembly can include hybridizing matching overhangs of ds oligos or hybridizing appropriate ss oligo linkers. A solid support can be used to immobilize one or more of the oligos, target polynucleotides, or one or more intermediates of the assembly. Immobilization can be performed using beads coated with avidin by modification of the oligo / polynucleotide such as biotinylation or amino modification. The modification can use a surface treated with aminosilane to bind to 3-aminopropyltrimethoxysilane or 3’ glycidoxypropyltrimethoxysilane. Surface chemistries suitable for covalent bonding with nucleic acids include, among others, carboxylic acid, aliphatic amine, aromatic amine, chloromethyl (vinylbenzyl chloride), amide, hydrazide, aldehyde, hydroxyl, thiol, or epoxy.Immobilization / binding to the solid support can be achieved using any of the chemicals described in "Strategies for attaching oligonucleotides to solid supports" by Integrated DNA Technologies, 2014 (22 pages), which is incorporated by reference.

[0098] The oligo can be ordered, provided from a library, or generated by appropriate methods such as chemical polynucleotide (or oligonucleotide) synthesis methods including H-phosphonate, phosphodiester, phosphotriester or phosphite triester synthesis methods, or large-scale parallel oligonucleotide synthesis methods, e.g., oligonucleotide synthesis based on microarrays or microfluidics. The oligo can be generated by enzymatic polynucleotide (or oligonucleotide) synthesis methods, e.g., ssDNA synthesis by DNA polymerase protein or reverse transcriptase protein that generates hybrid RNA-ssDNA molecules. In particular, the enzymatic polynucleotide synthesis reaction is carried out in vitro. The synthesis reaction can be carried out to generate RNA, DNA, xeno nucleic acids (XNA) (generally including 1,5-anhydrohexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycol nucleic acid (GNA), locked nucleic acid, also known as bridged nucleic acid (LNA), peptide nucleic acid (PNA), fluoroarabino nucleic acid (FANA)), or hybrids or any combination of the foregoing.

[0099] Oligos can be modified by any one or more of phosphorylation, methylation, biotinylation, or conjugation to a fluorophore or quencher. Oligos can be capped or blocked (e.g., until deblocked) to prevent ligation / polymerization to additional nucleotides. Suitable blocking or capping chemicals can include those discussed in U.S. Patent 10,041,110; WO2018 / 152323; WO2021 / 058438; WO2021 / 213903; and WO2021 / 116270, which are incorporated by reference. The libraries described herein can include library members that are oligos that can be any or all of the following: unmodified ss; phosphorylated ss; methylated ss; biotinylated ss; phosphorylated, biotinylated, and methylated ss; unmodified ds; phosphorylated ds; methylated ds; biotinylated and phosphorylated ds; biotinylated and methylated ds. Preferably, library members include 5'-phosphorylation. In particular, the libraries described herein include ss oligos containing a fluorophore or quencher, and ds oligos containing a fluorophore or quencher.

[0100] Oligos can be provided in a storage-stable form, preferably a form that is storage-stable for at least 6 months at room temperature. Oligos can be stored in a storage container in a dry state. The dry state can be achieved, for example, by lyophilization, freeze-drying, evaporation, crystallization, etc. Enzymes that catalyze the degradation of nucleic acids are typically active at room temperature in a fluid biomolecule preparation. Since such enzymes generally become inactive upon dehydration, and since the degradation chemical reactions catalyzed by the enzymes typically involve the addition of water (i.e., hydrolysis) to protein or nucleic acid molecules, and thus backbone cleavage of the protein or nucleic acid occurs, storage in a dry state inhibits such enzyme activity. In the dry state, there is little or no water (e.g., less than 5%, 4%, 3%, 2%, or 1% (w / w)) as a chemical reactant to support such enzyme catalysis. In addition, since water is generally not available for such reactions, any non-enzymatic hydrolysis of the protein or nucleic acid is similarly inhibited.

[0101] An oligo can contain any of the following bases: "A" representing deoxyadenosine, "T" representing deoxythymidine, "G" representing deoxyguanosine, or "C" representing deoxycytidine, "U" representing uracil, or other natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine), nucleotide analogs such as inosine and 2'-deoxyinosine and their derivatives (e.g., 7'-deaza-2'-deoxyinosine, 2'-deaza-2'-deoxyinosine), azole- (e.g., benzimidazole, indole, 5-fluorindole) or nitroazole analogs (e.g., 3-nitropyrrole, 5-nitroindole, 5-nitroimidazole, 4-nitropyrazole, 4-nitrobenzimidazole) and their derivatives, acyclic sugar analogs (e.g., hypoxanthine derivatives or indole derivatives, 3-nitroimidazole, or those derived from imidazole-4,5-dicarboxamide), 5'-triphosphates of universal base analogs (e.g., those derived from indole derivatives), isocarboxy styryl and its derivatives (e.g., methyl isocarboxy styryl, 7-propynyl isocarboxy styryl), hydrogen-bonding universal base analogs (e.g., pyrrolopyrimidine), or other chemically modified bases (e.g., diaminopurine, 5-methylcytosine, isoguanine, 5-methyl-isocytosine, K-2'-deoxyribose, P-2'-deoxyribose). The components are linked by a phosphodiester bond or a peptidyl bond, or by a phosphorothioate bond, or by any of other types of nucleotide bonds.

[0102] Specifically, the target polynucleotide has a length of at least 100 base pairs (bp). Specifically, the target polynucleotide may have a length of at least 150, 1,000, 10,000 or 100,000 bp, or even longer lengths may be generated. The method of the present invention may include a finalization step of adding, for example, one or more nucleotides corresponding to those previously removed from the 3' end and 5' end, respectively, to prepare a template for such a target ds polynucleotide for the purpose of assembly of the target ds polynucleotide according to the template sequence, such as generating blunt ends. Specifically, one or more oligos may be selected to generate blunt ends that are complementary to any overhangs of the intermediate polynucleotide prior to the final step, i.e., complementary to the attached ends of the polynucleotide. Specifically, each oligo can be used as a primer in a PCR reaction to amplify the final product, and the remaining oligos can be added to each strand to synthesize a complete target polynucleotide with blunt ends.

[0103] Specifically, the finalization step involves removing residual oligos, oligos, enzymes, and reagents, thereby obtaining a purified DNA product ready for further use, including a purification step of the PCR product generated using a standard kit such as the Monarch PCR & DNA clean up kit (product number T1030) from New England Biolabs to obtain the target ds polynucleotide. The method may include a step of concentrating the target polynucleotide or an intermediate of one or more assemblies by polymerase chain reaction (PCR). The molecule can be purified by immobilizing it on a solid phase using a tag, such as a biotin tag, and concentrating it, for example, using PCR amplification. According to a preferred embodiment, two sets of primers are used for target-specific enrichment and simultaneous removal of the tag. Specifically, a set of primers specific to the 5' end of the leading strand of the polynucleotide to be enriched, and a set of primers specific to the 5' end of the lagging strand, each set including a primer complementary to at least the overhang and a primer complementary to the core sequence of the polynucleotide. By using this set of primers, the target polynucleotide is amplified without the tag sequence. This has the significant advantage that no additional steps are required, for example, to remove the tag sequence by enzymatic digestion.

[0104] The method may include verifying the degree of identity with the desired sequence by sequencing the polynucleotide molecule. Any suitable sequencing method can be used, such as pyrosequencing, Illumina sequencing, SOLiD sequencing, semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, single molecule real-time (SMRT) sequencing, or nanopore DNA sequencing. The method may include, for example, restrictions or chemical modifications to facilitate cloning of the target polynucleotide into a vector or plasmid.

[0105] The polynucleotide molecule can be modified by enzymatic modification using any one or more of methyltransferase, kinase, CRISPR / Cas9, multiplex automated genome engineering (MAGE) using λ-red recombination, conjugate assembly genome engineering (CAGE), Argonaute protein family (Ago) or its derivatives, zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), meganuclease, tyrosine / serine site-specific recombinase (Tyr / Ser SSR), hybridizing molecule, sulfurylase, recombinase, nuclease, DNA polymerase, RNA polymerase or TNase.

Example

[0106] (Example 1) Determination of the optimal assembly tree for the assembly of a 600 bp target molecule using an optimal asymmetric workflow by 50-mer ligation In this example, a 600 bp random sequence (SEQ ID NO: 1) is selected as the sequence of interest (SOI). This sequence is first divided into 12 segments of 50 bp each, and then the optimal assembly tree is determined by considering the best according to a predetermined objective function.

[0107] Selection of the objective function First, a weight is defined for all nodes n according to whether the oligo at each node can react in a unique way (w = 1) or cannot react (w = 0), thereby selecting an objective function for evaluating the division. More specifically, any two oligos have a total of four overhangs (O 1 ,..,O 4 ), where the inventors - O 2 and O 3 match according to Watson-Crick pairing, - None of the other pairs of overhangs match is desired.

[0108] Mathematically, this can be represented by a symmetric adjacency matrix A (A ij = A ji ), where 1 indicates a matching pair and 0 indicates a non-matching pair. Next, the score w n at node n is defined as follows.

Number

[0109] For example, in relation to ligation activity, other metrics can be incorporated based on the frequency of clones recovered from a controlled experiment using an appropriate design of overhang pairs.

[0110] Considering all possible oligonucleotide pairings within a partition having N elements, the adjacency matrix for the entire partition has a total of N × N elements. In order to succeed in assembly, only consecutive elements within the partition need to form appropriate ligation, but when the corresponding oligonucleotides are mixed, other pairings form misligations. Next, the score of the assembly tree is obtained by adding all these elements of the matrix corresponding to appropriate ligation and subtracting the elements corresponding to misligations. Formally, the inventors can consider two other matrices: L ij is the ligation matrix, and for this purpose, each element is 0 except for those corresponding to consecutive overhangs that must be ligated; M ij is the misligation matrix, representing the overhangs that are in contact at any point during ligation (this actually depends on the tree), and thus, this is 1 for all pairs of contacting overhangs and 0 otherwise. Next, the score of the tree is calculated by the following formula:

Number

[0111] Partition of the array First, SEQ ID NO:1 was divided in a simple manner, namely, into equal lengths of 50 nt, with an overhang of length = 4. As a result, 24 single-stranded oligonucleotide sequences (SEQ ID NOs:2 to 25) were obtained, which, when assembled into the target double-stranded polynucleotide, would have the same sequence as the SOI.

[0112] Figure 15 shows the adjacency matrix described above for this particular division. Each colored element represents a pair of overhangs that can be ligated successfully. Not all of these overhangs are exposed to each other during the assembly process; rather, it is the assembly tree that determines which overhangs come into contact. A successful assembly is one in which only consecutive double-stranded oligonucleotides ligate properly, and in this matrix representation, these are the elements on two partial diagonals (and thus the non-zero elements of the ligation matrix); all other non-zero elements are avoided (in this example, M 4,10 , M 5,11 , M 8,12 ,...). Next, it becomes possible to correlate a particular assembly tree with the likelihood of a successful assembly, and thus to find the optimal tree.

[0113] Figure 16 shows a graphical representation of the scoring obtained using the equations given above, assuming an arbitrary non-optimal assembly tree in order to calculate the misligation matrix.

[0114] Determination of the Optimal Assembly Tree Any assembly tree representing an assembly workflow can be represented as a bracketed expression that details the "grouping", i.e., which fragments are assembled first. For example, if there are three fragments a, b, and c, an assembly tree representing the ligation of b.c first, followed by the addition of a, would be {a,{b,c}}.

[0115] To determine the optimal assembly tree, the totality of all possible groupings of the fragments corresponding to all possible assembly trees can be generated. For 12 fragments, there are a total of 58,786 possible assembly trees.

[0116] Figure 17 shows that each of these trees has a corresponding score, representing the likelihood of success of the assembly process. The algorithm scores each of these groupings according to the metric defined above and determines which assembly tree is most likely to result in a successful assembly. This is to avoid connecting nodes that are likely to form misligations. If multiple trees have the same optimal score, the one that requires the fewest number of steps in the default setting is selected.

[0117] Figure 18 gives the optimal assembly tree for a particular case of SOI. In the shown tree, the corresponding grouping is {{{{a,b},c},{d,{e,{f,g}}}},{{{{h,i},j},k},l}}.

[0118] Figure 19 shows the overhang adjacency matrix of the assembly defined by the tree in Figure 18. The dark gray regions within the matrix represent the elements of the misligation matrix and thus represent those overhangs that are in contact at any point during the assembly. With this optimal assembly tree, all misligation events are actually avoided.

[0119] Figure 20 shows an alternative assembly tree.

[0120] Figure 21 gives the corresponding adjacency matrix. In this case, as emphasized in the figure, misligations are not avoided in this sub-optimal tree.

[0121] When a clearly defined scoring metric is provided, this method scans all possible trees and, to determine the highest-ranked one, always finds an assembly tree that maximizes this score, thus guaranteeing it.

[0122] (Example 2) Different scoring functions correspond to different optimal assembly trees The choice of objective function is fundamental to the result of the tree optimization process. In this example, the inventors show how different scoring functions ultimately correspond to different assembly trees.

[0123] Definition of multiple scores In this example, the inventors consider three possible scorings. (S1) The first is obtained by considering a perfect match between overhangs: if two overhangs have a perfect Watson-Crick match in all four bases, a ij specific matrix element is 1, otherwise the element is 0, and then the overall scoring is calculated in the same way as in Example 1. (S2) The second scoring assigns a value of 3 to any GC pairing between corresponding overhangs and a value of 1 to any AT pairing. (S3) The third calculates the hybridization energy of the overhangs and assigns a rescaled energy value between 0 and 10 if the energy exceeds a specific threshold.

[0124] Determination of the optimal assembly tree Using these scoring functions, the inventors calculated the optimal assembly tree for a 128bp sequence (SEQ ID NO: 26), which was first simply divided into 16bp fragments (SEQ ID NOs: 27 - 42).

[0125] Figure 22 gives the adjacency matrix with scoring S1.

[0126] Figure 23 shows an assembly tree with S1 scoring.

[0127] Figure 24 shows an adjacency matrix with scoring S2.

[0128] Figure 25 shows an assembly tree with S2 scoring.

[0129] Figure 26 shows an adjacency matrix with scoring S3.

[0130] Figure 27 gives an assembly tree of S3 scoring

[0131] Figures 22 to 27 show three results by comparing the adjacency matrix and the assembly tree.

[0132] In S1 scoring, a completely symmetric assembly tree seems to be the most likely. Because, except for pairs that never touch between symmetric assemblies (element pairs {6,14}, {7,15}) (shown by the gray boxes and corresponding to the elements of the misligation matrix), only the completely matching overhang pairs are part of the ligation matrix. On the other hand, considering scoring S2, the number of matching pairs increases dramatically, so the possible combinations of misligation also increase; for these reasons, the corresponding optimal tree is more asymmetric and minimizes the number of possible combinations of misligation, but they cannot be completely avoided (the dark shaded elements of the adjacency matrix). S3 scoring is an intermediate solution. Because the energy requirements for overhang optimization are more difficult to satisfy for overhang pairs and do not require a perfect match (as in the case of S1); for these reasons, the number of non-zero elements of the adjacency matrix increases, but still the optimal assembly tree can avoid all problematic combinations.

[0133] (Example 3) Assembly of 256 bp target molecules using an asymmetric workflow and comparison with a symmetric workflow by 16-mer ligation The target sequence (SEQ ID NO: 43) was divided and an assembly tree was defined as described in Example 1, and then the oligos were assembled according to both the symmetric assembly and the asymmetric tree.

[0134] Figure 28 shows the symmetric assembly.

[0135] Figure 29 shows the asymmetric assembly. The asymmetric assembly is optimized by using the algorithm procedure described in Example 1.

[0136] All of the oligos from a to p were obtained from Integrated DNA Technologies (IDT) with standard desalting purification and provided normalized to a concentration of 50 μM in IDTE buffer (pH 7.5). The oligos used in the following assembly were single-stranded and pure.

[0137] Preparation and annealing of the annealing solution Some commercially available buffers such as product number B0202S, which is the T4 ligase reaction buffer from New England Biolabs (NEB), are ready to be mixed in H 2 O and already contain the ATP required for ligase activity. 216 μl of ddH 2 O was prepared in a microcentrifuge tube containing the NEB T4 ligase buffer, and the solution was thoroughly mixed by vortexing. 24 μL of this solution mixture was dispensed into 8 reaction tubes labeled a, b,.., p. 3 μL of each oligo was transferred to the 8 predetermined tubes and mixed thoroughly by pipetting: Oligos a- and a+ in tube a Oligos b- and b+ in tube b Oligos c- and c+ in tube c Oligos d- and d+ in tube d Oligos e- and e+ in tube e Oligos f- and f+ in tube f Oligos g- and g+ within tube g Oligos h- and h+ within tube h Oligos i- and i+ within tube i Oligos j- and j+ within tube j Oligos k- and k+ within tube k Oligos l- and l+ within tube l Oligos m- and m+ within tube m Oligos n- and n+ within tube n Oligos o- and o+ within tube o Oligos p- and p+ within tube p

[0138] The tubes were sealed and incubated in a thermocycler at 98 °C for 30 seconds. Subsequently, the temperature was decreased from 95 °C to 24 °C with a ramp function that decreased the temperature by 1 °C per minute, enabling annealing of the matched pairs of ss oligos. Upon completion, the double-stranded oligos were maintained at 4 °C.

[0139] Preparation of the ligation solution. In the following order, 32.5 μL of nuclease-free ddH 2 O, 7.5 μL of ligase buffer, and 5 μL of ATP were mixed on ice to prepare the ligation solution to a final concentration of 10 mM. The ligation solution was thoroughly mixed by vortexing and centrifuged. 2.5 μL of T4 ligase (NEB, product number M0202) and 2.5 μL of T4 polynucleotide kinase (PNK) (NEB, product number M0201B) were added at a total of 10 and 0.25 units per 1 μL of the final solution, respectively, and mixed thoroughly by gently pipetting. This solution was maintained on ice until needed. 10 μL of the ligation solution was transferred to each of four tubes (b, d, f, h) containing 5 μL of the corresponding ds oligo and mixed by pipetting. Subsequently, the tubes were sealed again.

[0140] Symmetric assembly The oligos were arrayed in rows on a 96-well plate and the pairs of oligos or reaction products were transferred stepwise: i. Initial assembly step. Transfer the contents of oligo mixtures a, c, e, g, I, k, m, and o into the mixtures of oligos b, d, f, h, j, l, n, and p (along with the ligation solution) respectively. ii. After gently mixing by pipetting, incubate these two reactants at 24 °C for 30 minutes. iii. Second assembly step. Transfer the contents of oligo mixtures ab,..., mn (tubes b,..., n) into the mixtures of oligos cd,..., op (tubes d,..., p). iv. After gently mixing by pipetting, incubate these two reactants at 24 °C for 30 minutes. v. Third assembly step. Transfer the contents of oligo mixture a.d (tube d) into the mixture of oligos e.h (tube h), and transfer the contents of oligo mixture j.l into the mixture of oligos m.p (tube p). vi. After gently mixing by pipetting, incubate these two reactants at 24 °C for 30 minutes. vii. The final volume containing the 128 bp product was 80 μL. viii. Purify the 128 bp product using MagMax beads technology, following the manufacturer's instructions with the following modifications to the protocol: The ratio of binding solution buffer to the sample used was 1.66. After bead purification, elute the DNA in 17 μl of ligation mix (T7 ligase (75 U / μL), T4 PNK (0.25 U / μl), 10 mM ATP, 10× T4 NEB buffer, and H 2 O). i. Fourth assembly step. Combine the reaction product a.h (tube h) with the mixture i.p (tube p). ii. Gently mix by pipetting and incubate the reactants at 24 °C for 60 minutes. iii. Purify the final 256 bp product using MagMax beads, following the manufacturer's protocol with the following modifications to the protocol: The ratio of binding solution buffer to the sample used was 1, and elute in 17 μl of ddH 2 O.

[0141] Asymmetric assembly by step-by-step manual pipetting In this case, by proceeding according to the tree of Figure 10b, it is sufficient to simply maintain a record of the order in which the oligos and their reaction products are combined. The reaction conditions are the same as for symmetric assembly. It is important to note that the asymmetric assembly follows the same steps as the symmetric assembly, but the order of transfer is crucial. iv. First assembly step. Transfer the contents of oligo mixture a (tube a) to mixture b (tube b), oligo mixture j (tube j) to mixture k (tube k), oligo mixture m (tube m) to mixture l (tube l), and oligo mixture n (tube n) to mixture o (tube o). v. After gently mixing by pipetting, incubate these two reactants at 24 °C for 30 minutes. vi. Second assembly step. Transfer the contents of oligo mixture a.b (tube b) to mixture c (tube c), the contents of oligo mixture d (tube d) to mixture e (tube e), oligo mixture f (tube f) to mixture g (tube g), oligo mixture i (tube i) to mixture j.k (tube k), and oligo mixture m.l (tube l) to mixture n.o (tube o). vii. After gently mixing by pipetting, incubate these two reactants at 24 °C for 30 minutes. viii. Third assembly step. Combine the reaction product a.c (tube c) with mixture d.e (tube e), combine oligo mixture f.g (tube g) within mixture h (tube h), and combine oligo mixture i.k (tube k) within mixture m.o (tube o). ix. Gently mix by pipetting and incubate the reactants at room temperature for 30 minutes. x. Fourth assembly step. Combine the reaction product a.e (tube e) with mixture f.h (tube h) and combine oligo mixture i.o (tube o) within mixture p (tube p). xi. The final volume containing the 128 bp product was 70 μL. xii. Heat-inactivate the ligation reactants by incubating at 65 °C for 10 minutes. xiii. The bead purification of the 128 bp product was carried out using MagMax bead technology according to the manufacturer's instructions with the following modifications to the protocol: The ratio of the binding solution buffer to the sample used was 1.66. After bead purification, the DNA was eluted in 17 μl of ligation mix (T7 ligase (75 U / μl), T4 PNK (0.25 U / μl), 10 mM ATP, 10× T4 NEB buffer, and H 2 O). xiv. The fifth assembly step. Combine the reaction product a.h (tube h) with the mixture i.p (tube p). xv. Gently mix by pipetting and incubate the reaction at 24 °C for 60 minutes. xvi. The bead purification of the final 256 bp product was carried out using MagMax beads according to the manufacturer's protocol with the following modifications to the protocol: The ratio of the binding solution buffer to the sample used was 1, and it was eluted in 17 μl of ddH 2 O.

[0142] Visualization and comparison of results After ligation, the final product was subjected to quality control using capillary electrophoresis (Fragment Analyzer). Before loading the sample onto the FA, the sample was diluted 2-fold with 1× TE buffer. The sample was run on the FA using the corresponding SS Small Fragment method (Agilent, DNF-476-33) according to the manufacturer's instructions using the Small Fragment Kit (Agilent, DNF-476-0500).

[0143] Figure 30 shows the electropherograms of two assemblies obtained after the purification procedure. The symmetric assembly was a failure because there was no distinct peak at the target length and the target molecule could not be recovered even after purification. In contrast, the asymmetric assembly resulted in the recovery of the target molecule after purification such that a distinct peak was visible.

[0144] Optionally, successful constructs can be isolated from the gel using a standard kit (e.g., Zymoclean), amplified by PCR (e.g., Sambrook and Russell, 2014; Chapter 8), and sequenced.

[0145] (Example 4) Pseudocode for a Monte Carlo algorithm for finding optimal sequence partitions and / or assembly graphs. This example describes a Metropolis algorithm for identifying optimal sequence partitions and tree searches. The Metropolis algorithm is described in the literature of probabilistic search, statistical mechanics, heuristics, and many other fields of science, mathematics, and computer science. This example shows the adaptation of the algorithm to appropriately generate an optimal or near-optimal assembly graph consisting of both sequence partitions and an assembly graph. Similar algorithms for constructing phylogenetic trees in evolution can also be found in the literature, and they have a certain similarity to the graph construction in this example. However, the phylogenetic interpretations of trees, scores, and other characteristics are different from those in this application area. For example, the graph does not need to be binary, and the inventors can desire to allow nodes with more than two branches, or in phylogenetics, while a tree is applied independently to all nucleotides in a sequence, in the case of the inventors, the assembly tree is applied to the sequence as a whole. Also, in phylogenetics, all nodes represent nucleotides, while in the case of the inventors, internal nodes represent partial assemblies. Nevertheless, certain algorithms from phylogenetics for generating tree variants (Felsenstein 2004; Gascuel 2005; Gascuel and Steel 2007), for example, can be innovatively applied to the field of DNA synthesis.

[0146] This example of the Metropolis algorithm reads a string representing a target sequence (Target_Seq). When the PARTITION function is called, an initial partitioning of the target sequence into an array of shorter strings is calculated in an order that will recover Target_Seq. For example, PARTITION can return a uniform partitioning into substrings of length equal to s, or non-uniform lengths that create substrings of some range (s min , s max ) in length, either randomly or according to some criterion. This array is initialized to three variables that are used throughout the process to either iterate the search (Curr_Part), propose a new partition (New_Part), or store the best partition (Best_Part). Similarly, an initial graph is generated by the function MAKETREE that takes a partition (array of strings) and returns a graph object, and at the same time, this initial tree is scored by the function SCORETREE that takes the partition and the graph object as arguments and assigns a score to the tree as in Examples 1 and 2. As with the output of PARTITION, three variables for the initial graph and score are also initialized to iterate and store the optimal combination of partition and tree. (Although a graph object can easily contain the partition and score, the inventors have explicitly separated the three for clarity.)

[0147] Figure 31 shows the steps of the method.

[0148] After initialization, the process enters a loop and aims to generate local changes that probabilistically determine the optimal value of the assembly graph including partitioning, resulting in the generation of a partition or a variant of the graph. The generation of variants can be done in multiple ways. In this example, the inventors assume that they randomly select a node of the graph, for example, with a probability inversely proportional to the score of the node, and attempt to generate a variant that improves it. The inventors assume two possible types of variants or "flips". The first (Flip_Type = 1) function REPARTITION modifies a subarray within a node, for example, by moving the 3' nucleotide of an array within the node to the 5' of its consecutive (rightmost) subarray. For example, if a node consists of two subarrays (N...N T ,GN...N), the algorithm gives (N...N, T GN...N). The reverse can also be done, that is, moving the 5' nucleotide of the array within the node to the 3' of its consecutive leftmost subarray within the node. For example, (N...NT, G ...N)->(N...NT G ,N...N). In the second variant generation type (Flip_Type = 2), the topology of the graph is changed by the function REBRANCHTREE that changes the vertices directly below the subtree. For example, if the focus node has a subtree ((s 1 ,s 2 ),(s 3 ,s 4 ))), this becomes (s 1 ,(s 2 ,(s 3 ,s 4 )),((s 1 ,(s 2 ,s 3 ,)),s 4) can be re-branched like this. Note that Flip_Type = 1 affects the scores of all sub-branches, while Flip_Type = 2 affects only the node directly below. For efficiency, the inventors can assign different frequencies to these types of changes and adapt them to the state of the process, but in this example the inventors maintain them at 50%-50%.

[0149] When a variant of the array or graph is generated and the graph is scored, these are assigned to the variables New_Part, New_Tree, and New_Scor respectively. Next, these "proposed" states are compared to the current state by evaluating whether New_Scor improves Curr_Scor (New_Scor > Curr_Scor?). If so, the new state is accepted by assigning it to the current partition, tree, and score configuration. If the new score is not improved, the inventors calculate RFact (the Boltzmann factor, which could be RFact = EXP(-B*(Curr_Scor - New_Scor)), where B > 0 or directly ratio RFact = New_Scor / Curr_Scor), and select a uniform random number in (0,1). If this number is less than RFact, the new configuration is accepted by assigning it to the current partition, tree, and score configuration; otherwise, the inventors maintain the current configuration and discard the proposal.

[0150] In addition, the inventors compare whether the new configuration is better than the best configuration found up to this point. If so, the new configuration is assigned to the best configuration of the partition, tree, and score.

[0151] This process of generating and evaluating variants is iterated comprehensively and repeatedly. This is because an algorithm never ends, so a stopping criterion is necessary. In this embodiment, the inventors add a counter Search_Steps that maintains tracking the number of times from the last update of the optimal value. If the best configuration has not changed in the last LIMSEARCH iteration (for example, LIMSEARCH = 1000000 times from the last update of the best configuration), the loop ends and the best configuration is returned. Alternatively, other criteria can be implemented. [Number] [Number] [Number]

[0152] (Embodiment 5) An evolutionary heuristic search algorithm for finding the optimal tree. In this embodiment, the search for the optimal or nearly optimal assembly graph is performed by a heuristic algorithm that generates a population of graphs and applies (a) a selection criterion based on a score for generating a "breeding population" and (b) a deformation form generation process for generating modifications close to the breeding population. The deformation forms can be performed in a series of ways. In this embodiment, the inventors interpret it as a single vertex modification for each tree at once in the same or similar way as described in the previous embodiments. In this embodiment, the inventors assume that the splitting of the array is fixed and maintained, but the extension to this model may be relaxed. Other modifications can include "recombination" of the tree or other means for redrawing the "breeding" graph.

[0153] In this embodiment, the sequence input containing the target sequence is split by the process PARTITION (see previous embodiments) that gives an array of subsequences. From this split, a population of N graphs or trees is generated. These are initially N copies of the same tree, which can be, for example, the most symmetric tree.

[0154] Next, the process enters the main iterative loop. Within this loop, each tree is (i) modified by a tree rearrangement algorithm (e.g., REBRANCHTREE in previous embodiments) to generate a population that includes a certain diversity of different trees, and (ii) scored after rearrangement to have an objective measure for selection. From this set, the tree with the highest score is stored. Selection criteria are applied to generate the next breeding population of trees. The selection criteria may be directly scored or may be a function that combines and / or transforms the score with other parameters. For example, the breeding population of N trees can be achieved by random selection with replacement from the N existing trees, with a probability proportional to their scores or selection function. The inventors can add constraints such as minimizing the asymmetry of the tree while increasing the score, or any other convenient limitation. After scoring, the inventors record the highest tree generated, and if it has a higher score than the previously recorded putative highest tree, replace the latter with the newly found highest tree in the breeding population. When the breeding population is generated, its maximum value is evaluated and recorded, this procedure is repeated to further generations of variants, from which the inventors select again, etc. This process effectively improves the average score of the population and further reduces the variance to converge to a local optimum.

[0155] This algorithm, like the Metropolis algorithm of previous embodiments, does not end, so additional stopping criteria can be included. The termination criterion may be a simple one such as a counter or a convergence measure that monitors the stability of the evolutionary process.

[0156] When the algorithm stops, the best tree is returned. This is a non-trivial graph that guides the assembly of the target sequence, as in Example 1.

[0157] (Example 6) Optimization of the assembly tree of a 1028 bp GC-rich target molecule using an asymmetric workflow by ligation of oligonucleotides of different lengths using a liquid handler In this example, the advantages of using an asymmetric workflow and an optimization algorithm to split a molecule with a difficult sequence into components (oligos) of unequal length. The target sequence is 1028 bp and is split in two ways: a fixed-length 16-mer and an adapted size between 15 and 17 nt. For reference, these splits are called c13 and c14, respectively. The inventors compare the score of a balanced (symmetric) assembly tree for naive splitting with the optimized assembly trees for both simple and adapted splitting.

[0158] Selection of the objective function To score the assembly tree, an "incorrect ligation matrix" (M i,j ) is generated. Each element of this matrix is 1 if, as a result of the selected assembly tree, the corresponding overhang pair is exposed during the ligation process but is not going to ligate; otherwise it is 0. The ligation matrix (L i,j ) assigns a score (between 0 and 4000) directly related to the likelihood of ligation of the corresponding pair, regardless of whether they are exposed during the assembly process. Pairs with a high GC content have a larger score. Specifically, if the two overhangs have at least 3 out of 4 matching nucleotides according to Watson-Crick pairing, the matrix element increases 10-fold for each A / T pair and 1000-fold for each G / C pair. To calculate the score, the two matrices are multiplied element by element and then the sum of each element is obtained. The overall score is as follows: [Number] In the formula, n is the total number of overhangs, and the minus sign is necessary to ensure that assembly trees with a high tendency for misligation have lower scores.

[0159] Array splitting and selection of assembly trees By using the scores defined in the previous section, the optimal split and assembly tree are calculated under three different conditions.

[0160] First, a uniform and "simple" split is performed by segmenting the SOI and its reverse complement into 16-mers (with a 4-nt offset to allow for a 4-nt 5' overhang). These oligos include SEQ ID NO: 76 to SEQ ID NO: 202. Specifically, SEQ ID NO: 76 to SEQ ID NO: 138 constitute the leading (or "top") strand of the ds oligo, while SEQ ID NO: 139 to SEQ ID NO: 202 constitute the lagging strand. This set is referred to as "c13".

[0161] Symmetric trees are used only for this purpose as the assembly workflow. The scores are analyzed with this split and topology. 128 oligonucleotide arrays were obtained by naive splitting.

[0162] Second, the same naive split as above is used, but an asymmetric tree is calculated to minimize misassembly by maximizing the score.

[0163] Thirdly, adaptive splitting is performed by maximizing the score locally. The oligonucleotide sequences range between 15 and 17 nt, and each dimer is constrained to have a 4 nt 5’ overhang. This splitting resulted in 126 oligonucleotide sequences. These oligos include SEQ ID NO: 203 to SEQ ID NO: 328. Specifically, SEQ ID NO: 203 to SEQ ID NO: 265 constitute the leading strand, while SEQ ID NO: 266 to SEQ ID NO: 328 constitute the lagging strand. This set is referred to as "c14".

[0164] After the locally optimal splitting is completed, the optimal tree is calculated by further maximizing the score.

[0165] Figure 32 shows the scored assembly tree for a naive split with a symmetric tree,

[0166] Figure 33 shows the scored assembly tree for a naive split with an asymmetric tree,

[0167] Figure 34 shows the scored assembly tree for an adaptive split with a symmetric tree.

[0168] Figure 35 gives a graph showing the products of the ligation and misligation matrices for each case: the darker elements correspond to lower scores.

[0169] As is evident from the scoring, the best results are obtained when the adaptive splitting method satisfies an asymmetric assembly tree. In this case, the score increases by more than twice. The choice of an asymmetric tree means that additional steps are required for the assembly, but for sequences such as SOI with a GC content of more than 60%, this is almost compensated for by the very high probability of a successful assembly.

[0170] Provision of Oligonucleotides The synthesis of the oligos was outsourced to a proven genomics company. The oligo supplier was contacted to ensure that the arraying, if necessary, included empty wells.

[0171] For the oligos from the naive partition with a symmetric tree, one plate of 364 microwells was procured, reflecting the systematic order of the oligos of the symmetric assembly.

[0172] The oligos of the naive partition with an asymmetric tree were procured in one plate of 364 microwells. Although these are the same oligos as the symmetric tree, the inventors chose to obtain easily arrayed oligos and avoid potential errors. The arraying was done by recasting the oligos into a larger symmetric tree and adding zeromers to the unassigned terminal branches at the right end of the tree. With the same logic and procedure, the procurement of oligos of the adaptive partition and asymmetric tree was proceeded with.

[0173] Assembly of the sequences When the oligonucleotides were provided, they were annealed in the same manner as in the above examples to form dimers with 5' overhangs. The assembly step was implemented by symmetrically moving the MCA arm on the Tecan Fluent under the same conditions as in the above examples.

[0174] Comparison of the results by capillary electrophoresis showed that the simple assembly using the symmetric tree failed (no detectable peaks). Both the naive partition and the adaptive partition using the asymmetric tree produced detectable peaks in the electropherogram. It was demonstrated that the yield and purity were higher for the assembly based on the adaptive partition using the asymmetric tree than for the naive partition using the asymmetric tree.

[0175] PCR amplification of the target product was performed to concentrate the sample. After amplification, it was washed with the Zymoclean kit, and sequencing was performed with an Oxford Nanopore minION. Analysis of the sequencing results shows that the assembly is accurate.

[0176] Assembly with an acoustic dispenser In different embodiments, the transfer of oligonucleotides can be achieved by acoustic transfer (Echo Liquid Handler 525, Beckman-Coulter). In this embodiment, spotting with zeroomers is not required, and it is possible to maximize the use of microtiter plates. Instead of mapping a tree to an array handled by the symmetric movement of a multi-channel pipette, the mapping is performed as a "worklist", i.e., a set of instructions (e.g., in XML format) indicating the exact order in which the transfers are to be made. This is not strictly parallel, but due to the accelerated speed of acoustic dispensing, it is substantially parallel compared to, for example, the required incubation time of the reaction.

[0177] (Example 7) In this example, Bayesian learning is used to improve the scoring of the binary assembly tree by using the sequencing data from the assembly. Selection of surrogate scores, splitting of sequences, and estimation of assembly trees First, for the sequence of interest (SOI), split and calculate the tree τ as in the above examples. The scores used in this process are just a surrogate set of [Number] It is understood that they are nothing more than. The purpose of this example is to demonstrate how the Bayesian scheme can be used to improve the knowledge about these parameters in order to improve scoring, estimation, and construction of the assembly tree, so the initial choice should have a large impact on the results. (The inventors can similarly randomly select the assembly tree.)

[0178] Array assembly As a second step of the process, once the assembly tree is determined, use a liquid handler and an appropriate enzyme environment to complete the assembly in a manner equivalent to the above-described examples.

[0179] Next-generation sequencing data Thirdly, after assembling the molecules, sequencing is performed using an Oxford Nanopore minION to determine the width of the construct obtained after the assembly process. From the sequencing reads, the data is analyzed to count the correct assemblies and alternative assemblies, as well as the unassembled fragments, and the frequency f of the target molecules with SOI is calculated. Mathematically, f is referred to as the frequency of the target construct.

Number

[0180] Bayesian scheme for calculating parameter distributions After collecting the data, the distribution of the parameters is estimated by the Bayesian method. For example, according to the Bayesian paradigm, the information distribution P of the parameter p when the data f is given is as follows.

Number

Number

[0181] Selection of the score of the likelihood function. In this example, the inventors assume that the topology is fixed and focus on improving the parameter p by using the obtained data. ij In this example, the data consists of array reads at the root-node A, which corresponds to the assembled array. The Bayesian scheme is as follows.

Number

[0182] Since B and C are from separate assemblies, their distributions are independent. Their distributions are such that they are spread out by themselves to their respective nodes.

[0183] Figure 36 shows a minimum example of the likelihood components in the assembly tree. The capital Roman letters are variables (e.g., the number of molecules of the NGS reads), P indicates the likelihood component, p is the (estimated) parameter of the likelihood,

Number

[0184] Illustrated with the simple tree of Figure 35, it is as follows.

Number

Number

Number

[0185] For more complex trees, such as the tree of this embodiment, the sum is performed over all internal nodes and implemented in computer code in an automated way. Thus, in this embodiment, there is no need to analytically represent the mathematical formula.

[0186] Regarding the internal probability P[T|p,ci,cj], consider that the final construct SOI or any target intermediate construct T is composed of two reagents ci and cj. Additionally, consider that there may also be misassemblies m. Thus, each node has a certain probability p of an exactly assembled molecule of n T of n m misassembly, as well as n ci and n cj unassembled reagents ci and cj of n. The probability of any given state is the following multinomial distribution.

Equation

[0187] Prior selection Probability p ij When given, they can be used to calculate a nearly optimal tree in a way equivalent to the above and above embodiments. However, in this Bayesian context, although assembly trees are generated using surrogate parameters

Equation

[0188] Figure 37 shows three alternative prior distributions following a beta distribution. In these possible prior distributions: solid line: uninformative uniform prior distribution (corresponding to beta[1,1]. Dashed line: beta distribution that prioritizes perfectly matching overhangs (beta[1,10]). Dotted line: beta distribution that prioritizes non-matching overhangs (beta[10,1]).

[0189] Bayesian estimation In this case, the estimation of the posterior parameter distribution is computationally performed by combining the above equations and implementing code to numerically sum. The marginal probability P[f] of the data is achieved by integrating over all parameter values.

Number

[0190] Many methods for numerical summation and integration are available in the literature, and especially in mathematical and statistical software such as Mathematica, Matlab®, Maple, Octave, R, etc.

[0191] The output of the calculation may be a function or data structure that depends on the input f, and may be an evaluable object in software that uses a particular interpolation method. In one embodiment, it can be assumed that each parameter follows a normal distribution with a particular mean and variance, and an approximation can be used such that only the latter can be used as the output. Iterative scheme

[0192] In a further embodiment of the method, the iterative scheme can be implemented in the following manner to improve the estimated distribution of the parameters using the cumulative data. 1. Initial surrogate distribution F o Set [p] and P o [p] = F o Set [p] 2. Prior

Number

Number

Number

[0193] Thus, each time a new molecule is assembled, additional information is generated to improve the scoring, thereby causing X o to converge to an accurate predictor. This method can be extended to construct multiple molecules in parallel, generate a larger dataset, and improve parameter estimation therefrom.

[0194] In another embodiment, instead of using surrogate parameters to estimate the optimal tree, a complete search can be performed in the parameter space to identify the best tree in the product space of the tree x parameters.

[0195] In a further embodiment, considering the hyper - distribution of the trees, the co - estimation of the trees and parameters based on the observed data can be further improved.

[0196] Array

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemistry

Chemistry

Chemistry

Chemistry

Chemistry

Chemistry

Chemistry

Chemistry

Chemistry

[0197] Further Embodiments The various described embodiments and claims of the present invention can be used in combination with one or more other embodiments / claims, unless they are technically incompatible.

Claims

1. A method for synthesizing polymers, A step of inputting a desired polynucleotide sequence into an in silico graph structure containing multiple branches, wherein each of the multiple branches identifies a linear order of nucleic acids or xenonucleic acids, A method wherein the in silico graph structure is programmed to select the optimal combination of branches and vertices that identify the desired sequence, and to direct the assembly of polynucleotide molecules having the desired sequence.

2. The method according to claim 1, wherein the branch corresponds to an oligonucleotide or polynucleotide in the isolated compartment.

3. The aforementioned in silico graph structure is used in a computer system. Multiple trees are generated, each tree representing an ordered combination of operations for generating the desired polynucleotide sequence. Select the tree that provides the optimal combination of operations. The method according to claim 1, wherein the system resides with a selected tree and includes executable program instructions to orient the assembly of the desired polynucleotide sequences.

4. The method according to claim 3, wherein the tree includes branches representing arrays connected by nodes representing joins generated between the arrays.

5. The method according to claim 3, wherein the selecting step comprises calculating a score for each tree, the score representing a measure of success in yielding a molecule having the desired polynucleotide sequence of the tree.

6. The method according to claim 5, wherein the measure of success is based on the probability of success when assembling a stored score or partial construct, and / or the probability of a stored misassembly.

7. The method according to claim 1, wherein the selection step includes: identifying subgroups, each containing fewer than approximately 20 oligos; generating, scoring, and selecting branches for each subgroup; and combining the best branches from the subgroups to form an assembly tree having an optimal generation score.

8. The method according to claim 7, further comprising the step of directing a fluid handling system to produce the desired polynucleotide.

9. The method according to claim 8, wherein the fluid handling system directs the assembly of the desired polynucleotides in a microfluidic device.

10. The method according to claim 9, wherein the microfluidic device includes a droplet containing a portion of the desired polynucleotide and, optionally, a reagent for assembling the desired polynucleotide.

11. The method according to claim 10, wherein the droplet comprises an aqueous phase surrounded by an immiscible phase.

12. The method according to claim 1, wherein the software package searches the graph structure to identify local maximums and thereby selects one or more optimal branches.

13. The method according to claim 1, further comprising the step of joining together subgroups of branches to form a branched pathway containing a desired polynucleotide sequence.

14. The method according to claim 1, wherein the selection algorithm selects a nested order of oligos to be combined in such a way as to optimize the pathway to the desired polynucleotide.

15. The method according to claim 1, wherein the in silico graph structure executes a machine learning algorithm to select the optimal combination of branches.

16. The method according to claim 15, wherein the machine learning algorithm uses one selected from the group consisting of a decision tree algorithm, a random forest algorithm, an extreme gradient boosting algorithm, an adaptive boosting algorithm, and a deep learning algorithm.

17. The method according to claim 1, wherein the polymer is a nucleic acid.

18. A method for synthesizing polymers, The steps include receiving sequence information describing a desired polynucleotide, A step of generating multiple trees using a computer system, wherein each tree provides an order of bonding between oligos to form the desired polynucleotide, A step of selecting one of the trees having the optimal generation score, The steps include: directing a liquid handling system by a computing system to produce the desired polynucleotide by joining the oligos as given by a selected tree; Methods that include...

19. The method according to claim 18, wherein the leaves of the tree represent the oligos and the nodes of the tree represent the connections between the oligos.

20. The method according to claim 18, wherein the selection step includes calculating a score for each tree, the score representing the probability that nucleic acids are successfully produced using that tree.

21. The method of claim 18, further comprising the step of (i) the stored probability of success when ligating the overhang of the oligo; (ii) an estimate of the stored risk of misligation between unintended pairs of the oligos; or (iii) scoring the tree using both.

22. There are more than 10^30 trees that give the order of bonding between the oligos to form the nucleic acid, and the method, The steps include: performing a generation and selection step for a first subset of the oligos to form an intermediate tree; and storing the intermediate tree to be used as a leaf when selecting a final tree having the optimal generation score. The method according to claim 18, including the method described in claim 18.

23. The method according to claim 18, further comprising the steps of: identifying subgroups, each containing fewer than approximately 20 of the oligos; generating, scoring, and selecting subtrees for each subgroup; and joining together the best scoring trees from each subgroup to obtain the selected tree having the best generated score.

24. The method according to claim 23, wherein for each subgroup, the software package generates all possible trees; stores all of the trees in memory; scores all of the trees by applying a scoring matrix to each of the trees in memory; and selects the best scoring tree for the subgroup.

25. The method according to claim 18, wherein the desired polynucleotide comprises a substantially complete expression vector or biological genome.

26. The method according to claim 18, further comprising repeating the above steps to generate a library of variants.

27. The method according to claim 18, further comprising performing the steps described above in parallel to simultaneously generate a wide variety of polynucleotides or their variants.

28. The method according to claim 18, wherein the step of producing the desired polynucleotide comprises running a software package to translate the selected tree into instructions to be executed by a liquid handling system to transfer the oligos between reaction vessels in the order given by the selected tree for producing the desired polynucleotide.

29. The method of claim 18, further comprising the step of running a software package that creates a description of the arrangement of the oligos in a multiwell plate, wherein the described arrangement minimizes the steps or operating time of a liquid handling system that produces the nucleic acid by conjugating the oligos in an order given by the selected tree.

30. The method according to claim 18, wherein a first leaf is connected to the root of the selected tree by a first number of nodes and edges, and a second leaf is connected to the root of the selected tree by a second number not equal to the first.

31. The method according to claim 18, wherein for at least one subgroup of the oligos, the software package searches the tree space to identify the local maximum in the tree space and selects the best scoring tree.

32. The method according to claim 30, wherein the software package uses Bayesian estimation, maximum likelihood calculation, or artificial intelligence.

33. The method according to claim 18, further comprising the step of creating a tree for the subgroup of the oligos.

34. The method according to claim 18, wherein selecting the tree having the optimal generation score includes calculating the probability of ligation success and promoting a tree that facilitates parallelization on the liquid handling system.