Method for synthesizing nucleic acid molecules

By combining multiple overlapping oligonucleotides and performing annealing, ligation and amplification steps, the problem of difficulty in efficiently synthesizing growth DNA molecules in the prior art is solved, and efficient and high-fidelity DNA synthesis is achieved.

CN120112653APending Publication Date: 2025-06-06TELESIS BIO INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280101379.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to synthesize any possible DNA molecules efficiently and with high fidelity, especially in applications where ultra-long sequences and high sequence fidelity are required.

Method used

Any possible DNA sequences are synthesized by combining multiple overlapping oligonucleotides in the reaction cell and performing annealing, ligation and amplification steps. This method uses overlapping oligonucleotide libraries to assemble DNA molecules in a hierarchical manner and realize the synthesis of longer DNA molecules.

Benefits of technology

The efficient synthesis of any possible DNA molecule from a limited oligonucleotide library is achieved, enabling the generation of DNA molecules of thousands of base pairs in length, and maintaining a low error rate of less than 1 error per 2,000 nucleotides in 2,000 nucleotides.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112653A_ABST
    Figure CN120112653A_ABST
Patent Text Reader

Abstract

The present invention provides methods for synthesizing a product DNA molecule of any possible DNA sequence from a universal library of overlapping oligonucleotides. The method involves combining a plurality of the overlapping oligonucleotides in a reaction cell, wherein the sequence of the plurality of oligonucleotides comprises at least a sub-sequence of the product DNA molecule. The method also involves annealing the plurality of oligonucleotides, performing a ligation step, and performing an amplification step, thereby synthesizing a subsequence of the product DNA molecule. The invention can be used to synthesize DNA molecules of any possible sequence from the universal library, which can be accomplished by a hierarchical assembly scheme. In one embodiment, the universal library comprises less than 10,000 prefabricated oligonucleotides that can be synthesized into any possible DNA sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides compositions, methods and kits for synthesizing any possible DNA molecule from a limited oligonucleotide library. Background Art

[0002] There is a continuing and growing demand for diverse and known sequence oligonucleotides in the fields of synthetic biology and gene editing and therapeutics. Existing methods for synthesizing small oligonucleotides involve chemical synthesis via solid phase, sequential coupling of nucleotides to generate oligonucleotides of desired length and sequence. The resulting oligonucleotides are then released from the solid phase, deprotected, and collected by other methods to be assembled into larger oligonucleotides. Although automated, these processes are susceptible to side reactions and base errors, limiting the length of the resulting oligonucleotides. For applications requiring ultra-high sequence fidelity, these methods have additional limitations.

[0003] Enzymatic methods for synthesizing oligonucleotides also exist and involve the use of enzymes such as terminal deoxynucleotidyl transferase (TdT), a template-independent polymerase that catalyzes the incorporation of deoxyribonucleotides into the 3'-hydroxyl end of a DNA template. However, this enzyme exhibits a strong preference for specific nucleotide bases and cannot reliably add nucleotides in the desired order and length.

[0004] There remains a need for efficient and high-fidelity methods for synthesizing oligonucleotides so that users can generate oligonucleotides of any desired length and sequence. Summary of the invention

[0005] The present invention provides a method for synthesizing a product DNA molecule of any possible DNA sequence from a library of overlapping oligonucleotides (which can be a universal library). The method involves combining a plurality of overlapping oligonucleotides in a reaction pool, wherein the sequence of a plurality of oligonucleotides at least comprises a subsequence of a product DNA molecule. The method further involves annealing a plurality of oligonucleotides, performing a connection step and performing an amplification step, thereby synthesizing a subsequence of a product DNA molecule. The present invention can be used to synthesize a DNA molecule of any possible sequence from a library, which can be accomplished by a hierarchical assembly scheme. In one embodiment, the library is a universal library having less than 10,000 prefabricated oligonucleotides, and the use method can synthesize the prefabricated oligonucleotides into any possible DNA sequence. The length of the product DNA molecule can be at least 100 base pairs or at least 150 base pairs. When subsequent DNA assembly techniques are adopted, DNA molecules of thousands of base pairs can be synthesized. In any embodiment, the product DNA molecule can have an error rate of less than 1 error per 2,000 nucleotides.

[0006] In the first aspect, the invention provides a method for synthesizing a DNA molecule with a desired sequence. The method relates to annealing at least two oligonucleotides to an anchor chain so that at least two oligonucleotides annealed to the anchor chain are adjacent to each other on the anchor chain. In any embodiment, the oligonucleotide can be adjacent to its variable sequence on the anchor chain. At least two oligonucleotides can each be included in a primer binding site on a 3' or 5' end, and a variable sequence on a relative 5' or 3' end, and a conservative flanking sequence between the primer binding site and the variable sequence. The anchor chain can have a conservative flanking sequence complementary to the conservative flanking sequence on at least two oligonucleotides, and can further have at least one variable sequence. At least a portion of at least one variable sequence on the anchor chain is complementary to at least a portion of a variable sequence on each of at least two oligonucleotides. The present invention involves the steps of ligating at least two oligonucleotides annealed to an anchor strand to produce a first dsDNA molecule, subjecting the first dsDNA molecule to an amplification step, the first dsDNA molecule having a desired sequence and comprising primer binding sites at the 3' and 5' ends, conserved flanking sequences within each of the 3' and 5' ends, and variable sequences within the conserved flanking sequences. In one embodiment, the first dsDNA molecule has a variable sequence of about 8 or about 10 nucleotides in length. In either embodiment, one or more (or all) of the primer binding sites may be universal primer binding sites.

[0007] The method may also involve contacting the first dsDNA molecule with a restriction endonuclease to produce a first dsDNA fragment having a 3' and / or 5' overhang sequence comprising a portion of the variable sequence from the first dsDNA molecule, providing at least one additional dsDNA fragment comprising a 3' and / or 5' overhang sequence that is at least partially complementary to the overhang sequence of at least one of the first dsDNA fragments. The 3' and / or 5' overhang sequence may contain at least a portion of the variable sequence. The method also involves annealing the first dsDNA fragment and at least one additional dsDNA fragment through the 3' and / or 5' overhang sequence, and connecting the annealed dsDNA fragments to produce a second dsDNA molecule having a conserved flanking sequence within each of the 3' and 5' ends, and a variable sequence within the 3' and 5' conserved flanking sequence that is longer than the variable sequence on the first dsDNA molecule (and an optional primer binding site on the 3' and / or 5' end). In one embodiment, the length of the variable sequence of the second dsDNA molecule is about 16 base pairs. The method may also involve amplifying the second dsDNA molecule. In any embodiment, the restriction endonuclease may be a type II restriction endonuclease (e.g., type IIS or type IIT). In any embodiment and any step of the method, at least one additional dsDNA fragment may be a product of a parallel DNA synthesis reaction. The first dsDNA molecule may have a restriction endonuclease recognition site on the 5' or 3' side of the molecule, and the first additional dsDNA fragment may be derived from restriction cleavage of a dsDNA molecule having a restriction endonuclease recognition site on the opposite 3' or 5' side of the molecule.

[0008] The method may further involve contacting at least one second dsDNA molecule with a restriction endonuclease to produce a plurality of second dsDNA fragments comprising 3' and / or 5' overhang sequences and conserved flanking sequences within each of the 3' or 5' ends (and optional primer binding sites on the 3' and / or 5' ends). The 3' and / or 5' overhang sequences may contain at least a portion of the variable sequence. The method may further involve the steps of providing at least one (second) additional dsDNA fragment, the first dsDNA molecule comprising a 3' and / or 5' overhang sequence that is at least partially complementary to the overhang sequence of at least one of the second dsDNA fragments, annealing the plurality of second dsDNA fragments to one or more (second) additional dsDNA fragments via the 3' and / or 5' overhang sequence; and performing a ligation step to produce a third dsDNA molecule having conserved flanking sequences on the 3' and 5' ends, and a variable sequence within the conserved flanking sequence that is longer than the variable sequence of the second dsDNA molecule (and an optional primer binding site on the 3' and / or 5' end). In one embodiment, the variable sequence is about 28 base pairs in length. The method may include performing an amplification step on the third dsDNA molecule. The at least one (second) additional dsDNA fragment may be a product of a parallel DNA synthesis reaction. The second dsDNA molecule may have a recognition site for a restriction endonuclease on the 5' or 3' side of the molecule, and the (second) additional dsDNA fragment may be derived from restriction cleavage of a dsDNA molecule having a recognition site for a restriction endonuclease on the opposite 3' or 5' side of the molecule.

[0009] The method may further involve contacting at least one third dsDNA molecule with a restriction endonuclease to produce a plurality of third dsDNA fragments, the third dsDNA fragments comprising 3' and / or 5' overhang sequences and conserved flanking sequences within each of the 3' or 5' ends (and optional primer binding sites on the 3' and / or 5' ends); the fragments may contain at least a portion of the variable sequence on the 3' and / or 5' overhangs. The method may also involve providing at least one (third) additional dsDNA fragment comprising a 3' and / or 5' overhang sequence that is at least partially complementary to the overhang sequence of at least one of the third dsDNA fragments. The overhang sequence may contain at least a portion of the variable sequence. The method may also involve annealing the plurality of third dsDNA fragments to one or more (third) additional dsDNA fragments via the 3' and / or 5' overhang sequence. The method may further involve a ligation step to produce a fourth dsDNA molecule having conservative flanking sequences on the 3' and 5' ends, and a variable sequence within the conservative flanking sequence that is longer than the variable sequence of the third dsDNA molecule (and an optional primer binding site on the 3' and / or 5' ends). The (third) additional dsDNA fragment may be the product of a parallel DNA synthesis reaction. The third dsDNA molecule may have a recognition site for a restriction endonuclease on the 5' or 3' side of the molecule, and the (third) additional dsDNA fragment may be derived from restricted cleavage of a dsDNA molecule having a recognition site for a restriction endonuclease on the relative 3' or 5' side of the molecule. The method may also involve an amplification step to the fourth dsDNA molecule. In one embodiment, the length of the variable sequence is about 100 base pairs.

[0010] In another embodiment, step a) further relates to at least two paired oligonucleotides annealed to the paired anchor strand, so that at least two paired oligonucleotides combined with the paired anchor strand are adjacent to each other on the paired anchor strand, which can occur at its variable sequence. At least two paired oligonucleotides can have a primer binding site on 3' or 5' end, and a variable sequence on relative 5' or 3' end, and a conservative flanking sequence between the primer binding site and the variable sequence. The paired anchor strand can have a conservative flanking sequence complementary to the conservative flanking sequence on the oligonucleotides of at least two pairs, and can further have at least one variable sequence. A part of the variable sequence on the paired anchor strand can overlap with a part of the variable sequence on the first anchor strand. The variable sequence can be located between two sequences complementary to the conservative flanking sequence. Between the first and second anchor strands, at least a portion of the variable sequence is different. In any embodiment, one or more (or all) in the primer binding site can be a universal primer binding site.

[0011] The method may further involve ligating at least two paired oligonucleotides annealed to the anchor strand, performing an amplification step to produce a paired dsDNA molecule having a desired sequence and comprising a primer binding site at the 3' and 5' ends, a conserved flanking sequence within each of the 3' and 5' ends, and a variable sequence within the conserved flanking sequence that partially overlaps with the variable sequence of the first dsDNA molecule. In one embodiment, at least two oligonucleotides and the first anchor strand, and at least two paired oligonucleotides and the paired anchor strand may be annealed in a simultaneous reaction in the same pool. The method may further involve contacting the first dsDNA molecule and the paired dsDNA molecule with a restriction endonuclease to produce at least one dsDNA fragment and at least one paired dsDNA fragment, each fragment comprising at least one 3' and / or 5' overhang sequence; and at least a portion of the 3' or 5' overhang sequence from the first dsDNA fragment may be complementary to at least a portion of the 5' or 3' overhang sequence from the paired dsDNA fragment (and an optional primer binding site on the 3' or 5' end). The method may further involve annealing the first dsDNA fragment and the paired dsDNA fragment through their complementary overhang sequences, and performing a ligation step to produce a second dsDNA molecule having a conserved flanking sequence within each of the 3' and 5' ends, and a variable sequence within the 3' and 5' conserved flanking sequences that is longer than the variable sequence on the first dsDNA molecule (and an optional primer binding site on the 3' and / or 5' ends). The method may also involve performing an amplification step on the second dsDNA molecule.

[0012] In further embodiments, the method may further involve contacting at least one second dsDNA molecule and at least one paired second dsDNA molecule with a restriction endonuclease to produce a plurality of second dsDNA fragments and paired second dsDNA fragments, each fragment comprising a 3' and / or 5' overhang sequence. At least two of the plurality comprise a conserved flanking sequence within each of the 3' or 5' ends. At least a portion of the 3' or 5' overhang sequence from the second dsDNA fragment may be complementary to at least a portion of the 5' or 3' overhang sequence from the paired second dsDNA fragment. The method may further involve annealing the second dsDNA fragment and the paired second dsDNA fragment through their complementary overhang sequences, and performing a ligation step to produce a third dsDNA molecule, the third dsDNA molecule comprising a conserved flanking sequence within each of the 3' and 5' ends, and a variable sequence within the 3' and 5' conserved flanking sequences that is longer than the variable sequence on the second dsDNA molecule (and an optional primer binding site on the 3' and / or 5' ends). At least a portion of the variable sequence on the third dsDNA molecule may overlap with at least a portion of the variable sequence on the paired third dsDNA molecule.The method may also involve subjecting the third dsDNA molecule to an amplification step.

[0013] In another embodiment, the method further involves contacting at least one third dsDNA molecule and at least one paired third dsDNA molecule with a restriction endonuclease to produce a plurality of third dsDNA fragments and paired third dsDNA fragments, each fragment comprising a 3' and / or 5' overhang sequence; the third dsDNA fragments may comprise at least a portion of a variable sequence on the 3' and / or 5' overhang. At least two of the plurality of fragments may have conserved flanking sequences within the 3' or 5' end. At least a portion of the 3' or 5' overhang sequence from the third dsDNA fragment may be complementary to at least a portion of the 5' or 3' overhang sequence from the paired third dsDNA fragment. The method may further involve the steps of annealing the third dsDNA fragment and the paired third dsDNA fragment through their complementary overhang sequences, and performing a ligation step to produce a fourth dsDNA molecule having a conserved flanking sequence within each of the 3' and 5' ends, and a variable sequence within the 3' and 5' conserved flanking sequences that is longer than the variable sequence on the third dsDNA molecule (and an optional primer binding site on the 3' and / or 5' ends). The method may also involve performing an amplification step on the fourth dsDNA molecule.

[0014] In any embodiment, the first dsDNA molecule can have a variable region of 8 to 12 base pairs. In any embodiment, the paired dsDNA molecule can have a variable region of 8 to 12 base pairs. In any embodiment, the second dsDNA molecule can have a variable sequence of 14 to 18 base pairs. In any embodiment, the third dsDNA molecule can have a variable sequence of 24 to 32 base pairs. In any embodiment, the fourth dsDNA molecule can have a variable sequence of 90 to 110 base pairs. In any embodiment, at least two oligonucleotides can have a variable sequence of 4 to 6 nucleotides. In any embodiment, the anchor chain can have a sequence complementary to the conservative flanking sequence on 3' and 5' ends on at least two oligonucleotides. In any embodiment, the anchor chain can have a sequence complementary to the conservative flanking sequence on 3' and 5' ends on at least two oligonucleotides. In any embodiment, the amplification step can be carried out by polymerase chain reaction (PCR). In any embodiment, the variable sequence of anchor chain or dsDNA product molecule (such as first, second, etc.) can be equal to the length of the variable sequence on at least two oligonucleotides.In any embodiment, anchor chain can have a variable sequence between two sequences complementary to the conservative flanking sequences on at least two oligonucleotides.In any embodiment, anchor chain can have a variable sequence between two sequences complementary to the conservative flanking sequences on at least two oligonucleotides.In any embodiment, at least two oligonucleotides combined with anchor chain can be adjacent to each other at their variable sequence on anchor chain.In any embodiment, a part of the variable sequence on anchor chain complementary to the conservative flanking sequences on at least two oligonucleotides can be 2 to 6 nucleotides, or 14 to 18 nucleotides, or 26 to 30 nucleotides, or 90 to 110 nucleotides.In any embodiment, at least two oligonucleotides and anchor chain are programmed so that dsDNA molecules have at least one recognition site of restriction endonuclease.In any embodiment, restriction endonuclease can be IIS type endonuclease.In any embodiment, anchor chain can have 4 to 6 degenerate nucleotides. In any embodiment, at least one additional dsDNA fragment can be from a parallel synthesis reaction. In any embodiment, the 3' and / or 5' overhang sequence can have a portion of a variable sequence from a first dsDNA molecule. In any embodiment, the ligation step can occur spontaneously. In any embodiment, at least one additional dsDNA fragment can have a variable sequence that is at least partially complementary to a variable sequence from a first dsDNA molecule. In any embodiment, the method can further involve the step of ligating at least two oligonucleotides bound to an anchor strand.

[0015] In yet another aspect, the invention provides the composition of at least two oligonucleotides, each of which is included in a primer binding site (for example universal primer binding site) on 3' or 5' end, and a variable sequence on relative 5' or 3' end, and a conservative flanking sequence between this primer binding site and this variable sequence. In this composition, the anchor chain can have a sequence complementary to the conservative flanking sequence on at least two oligonucleotides, and can further have at least one variable sequence between two sequences complementary to the conservative flanking sequence. At least a portion of the variable sequence on the anchor chain can be complementary to at least a portion of the variable sequence on at least two oligonucleotides. In one embodiment, the anchor chain can have a sequence complementary to the conservative flanking sequence at its 3' and 5' end on at least two oligonucleotides. The anchor chain can have a variable sequence between two sequences complementary to the conservative flanking sequence.

[0016] In another aspect, the present invention provides a method for storing data in a DNA sequence. The method involves determining a DNA sequence encoding non-genetic information according to a coding scheme that translates the non-genetic information from a reference language into a DNA sequence, or vice versa; synthesizing a DNA sequence encoding non-genetic information according to any method disclosed herein; and thereby storing data in a DNA sequence.

[0017] In another aspect, the present invention provides a method for synthesizing a DNA sequence encoding a guide RNA. The method involves determining a DNA sequence encoding a guide RNA; synthesizing a DNA sequence encoding a guide RNA according to any method disclosed herein.

[0018] In another aspect, the invention provides an oligonucleotide library comprising 1,536 different positions. The library may have 1,024 positions having oligonucleotides with unique variable sequences of non-degenerate nucleotides; and an additional 512 different positions, each of the 512 positions having an anchor having a variable sequence of at least three non-degenerate nucleotides and at least four degenerate nucleotides.

[0019] In one embodiment, the oligonucleotide has a primer binding site on the 3' or 5' end, and a variable sequence on the relative 5' or 3' end, and a conservative flanking sequence between the primer binding site and the variable sequence. The anchor chain can have a conservative flanking sequence complementary to the conservative flanking sequence on at least two oligonucleotides, and also have at least one variable sequence. At least a portion of at least one variable sequence on the anchor chain can be complementary to at least a portion of the variable sequence on at least two oligonucleotides.

[0020] In another embodiment, the oligonucleotide library has 4,608 different positions. 4,096 positions have oligonucleotides whose unique variable sequences are non-degenerate nucleotides. The library can also have another 512 different positions, each of which has an anchor chain, and each of the 512 positions has at least three non-degenerate nucleotides and at least five degenerate nucleotides. The oligonucleotides in the library can be any oligonucleotides described herein.

[0021] In any embodiment of the library, some oligonucleotides may contain 5' phosphate groups (e.g., deoxyribose 5' phosphates). In various embodiments, at least 25% or at least 30% or about one-third of the oligonucleotides in the library may have 5' phosphates. In any embodiment, the oligonucleotides containing 5' phosphates are not anchor chains. 5' phosphates may be present on 5' nucleotides. In any embodiment, at least 40% or at least 50% or at least 75% of the oligonucleotides that are not anchor chains in the library may contain 5' phosphates. In one embodiment, all "O2" oligonucleotides have 5' phosphates. The 5' phosphate on one of the two oligonucleotides to be connected may contribute to the action of the ligase. The position containing the anchor chain in the library may contain the anchor chain oligonucleotides of each possible sequence with a variable sequence, i.e., multiple oligonucleotide sequences may be present at this position. The variable sequence of the anchor chain may have five or six nucleotides, or may be other situations of the anchor chain as described herein. At other positions in the library (e.g., containing non-anchor chain oligonucleotides), there may be oligonucleotides of unique sequences, i.e., single sequences at this position. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figures 1A to 1B . Figure 1A A schematic diagram of synthesizing a DNA molecule of a desired sequence according to one embodiment of the present invention is provided. Figure 1B Further reactions of the products from the schematic diagram for synthesizing a DNA molecule of the desired sequence are provided. Degenerate nucleotides are labeled "N".

[0023] Figures 2A to 2B . Figure 2A A schematic diagram containing details of an embodiment of a DNA synthesis reaction of the present invention is provided. Figure 2B Another portion of the reaction is shown. These figures show the complementary 3' and 5' overhang sequences between the dsDNA molecules and the fragments. Degenerate nucleotides are marked with "N".

[0024] Figure 3A schematic diagram of an embodiment of a general scheme for hierarchical assembly using the methods of the present invention is provided. Hierarchical assembly can be utilized by synthesizing dsDNA molecules with overlapping variable sequences. Assembly can also include the addition of dsDNA fragments with 5' and 3' overhangs, which can be added to the final assembly (or any assembly step). The length of the variable sequence is for illustrative purposes only.

[0025] Figures 4A to 4D . Figure 4A A gel image of the PCR1 product after the first ligation step (L0) is provided, in which three oligonucleotides from the library are combined, ligated, and PCR amplified using a single universal primer pair. Each of the 98 bp PCR products contains 10 bp of synthetic DNA (variable sequence) that is utilized for downstream assembly. Figure 4B Provided is a gel image showing the PCR2 product after the first digestion and connection step (DL1) of the PCR1 product. These PCR2 products are produced by combining two PCR1 products, removing the flanking sequences on each product (by enzymatic digestion), and then connecting the products. Although the total length of the PCR2 product is shorter than the PCR1 product in Figure 1 (56bp), the variable sequence of dsDNA has increased to 16bp, which indicates that the variable sequence length increases as the workflow progresses. Figure 4C Provided is a gel image showing the PCR3 product after the second digestion and ligation step (DL2). These PCR3 products are produced by combining two PCR2 products, digesting a flanking sequence on each product, and then connecting. Note that the 1st and 4th PCR products contain 5' thiophosphate capped ends, which help to reduce downstream misconnection events, and in addition, these products also provide universal priming sequences for subsequent PCR4 amplification. Among all the sequences on this gel, the length of the variable sequence of the PCR3 dsDNA molecule is 28bp. Figure 4D Provided is a gel image showing PCR4 products after the third digestion and ligation step (DL3). These PCR4 products are produced by combining four PCR3 products, digesting one or two flanking sequences on each product, and then ligating them together. Of all the sequences on the gel, the variable sequence of PCR4 dsDNA is 100 bp in length, and they have 40 bp flanking sequences on both sides, which can be digested away using BsmBI, thereby enabling further assembly into even larger DNA fragments.

[0026] Figure 5A schematic diagram of an embodiment of the present invention for storing digital information in DNA is provided. A 16 bp product DNA molecule encoding four bytes of information is generated. This example shows how to use the method of the present invention to encode non-genetic information (here "the cat in the hat") into DNA.

[0027] Figure 6 Schematic diagram of an embodiment of the present invention applied to synthesize 120bp product DNA, which is the initial guide structure of the transcription element with promoter, guide RNA, Cas9 handle and terminator. In this embodiment, the first cycle of PCR uses two primers with two variable bases at their 3' ends. This converts the other 16bp product DNA molecules into 20bp products. Subsequent steps of PCR incorporate the transcription element.

[0028] Figure 7 is a graphic representation of a gel image showing the fully synthesized assembled 3,942 bp spike protein gene. The full length was confirmed by cloning and DNA sequencing analysis and determined to have an error rate of approximately 1 error per 5,400 bp before applying the enzymatic error correction step. DETAILED DESCRIPTION

[0029] The present invention provides a method for assembling the DNA molecule of any sequence with high fidelity using a universal oligonucleotide library. The method involves using an oligonucleotide library with a DNA molecule member, so that all possible DNA sequences can be assembled from the library using the method. In one embodiment, the oligonucleotide library has less than 10,000 members. In order to realize the method of assembling any possible DNA sequence from a library with a limited number of members, many efforts have been made. The inventors have found that any possible DNA sequence can be easily assembled using materials and methods disclosed herein. Therefore, the present invention enables the creation of a library less than 10,000 oligonucleotides, from which all possible oligonucleotide sequences can be assembled. The library less than 10,000 oligonucleotides can be conveniently provided on a small device (e.g., DNA chip), and devices and instruments are provided to selectively assemble any DNA sequence using only members of the oligonucleotide library.

[0030] Oligonucleotide library

[0031] In various embodiments, the oligonucleotide members in the library can be DNA of various lengths. In various embodiments, the oligonucleotide library can have less than 20,000 members, or less than 15,000 members, or less than 12,000 members, or less than 10,000 members, or less than 9,000 members, or less than 8,000 members, or less than 7,000 members, or less than 6,000 members.

[0032] In various embodiments, the library can contain at least 2,000 members, or at least 3,000 members, or at least 5,000 members, yet also contain less than 20,000 members, or less than 15,000, or less than 12,000 members, or less than 10,000 members, or less than 9,000 members, or less than 7,000 members, or less than 6,000 members, with all possible combinations and sub-combinations. The method can use the oligonucleotide members in the library to synthesize all possible polynucleotide sequences. In various embodiments, the present invention allows only to start assembling different sequences (such as variable sequences) from the oligonucleotides in the library More than 4 billion (for 16 aggressiveness) and up to more than 1 trillion (for 20 aggressiveness) polynucleotides. In various embodiments, each oligonucleotide in the library can be used 100 to 10,000 times in the synthesis of product DNA molecules.

[0033] The assembled product DNA molecule can be of any size, for example, it can be more than 100 bp, or more than 250 bp, or more than 500 bp, or more than 750 bp, or more than 1 kbp, or more than 1.5 kbp, or more than 2 kbp, or more than 5 kbp, or more than 10 kbp, or 100 bp to 500 bp, or 80 bp to 500 bp, or 80 bp to 750 bp, or 80 bp to 1000 bp, or 1 kbp to 4 kbp, or 1 kbp to 5 kbp, or 2 kbp to 10 kbp, or 5 kbp to 15 kbp, or 5 kbp to 20 kbp. The term "oligonucleotide" refers to a polynucleotide having a length of 100 bp to 100 bp, or a length of 500 bp to 100 bp, or a length of 4000 bp to 500 bp, or a length of 1000 bp to 150 bp, or a length of 500 bp to 100 bp, or a length of 200 bp to 250 bp, or a length of 500 kbp to 100 bp, or a length of 1000 bp to 150 bp, or a length of 1000 bp to 150 bp, or a length of 2000 bp to 250 bp, or a length of 500 kbp to 10 ...2000 bp to 250 bp, or a length of 500 kbp to 100 bp, or a length of 1000 bp to 150 bp, or a length of 1000 bp to 150 bp, or a length of 1000 bp to 15

[0034] In other embodiments, the method can also be used together with even smaller libraries to assemble a large number of sequences that may be needed, such as assembling more limited and more targeted sequences in the definition category of such sequences. The example of the definition category can include the gene set related to a specific biological function, or the gene from a specific organism. In any embodiment, the product DNA molecule synthesized in the method can be completely by and only use the oligonucleotides from the oligonucleotide library to synthesize." Universal library " is a polynucleotide molecule library from which any possible DNA sequence can be assembled. On a broad level, the universal library can contain the polynucleotides that can be assembled into any possible DNA sequence. However, within the DNA sequence category of a specific definition, a smaller universal library (or library of definition category) containing a sequence of interest can be used, such as, for RNA metabolism or for being related to transcription or for regulating, RNA metabolism, translation, protein folding, protein export, RNA (rRNA, tRNA, microRNA), ribosome biogenesis, rRNA modification, DNA replication, DNA repair, DNA topology, DNA metabolism, chromosome segregation, cell division and tRNA modified gene or sequence of the library of the DNA sequence. Thus, in some embodiments, any of these and any other categories can be considered as a library of defined categories of sequences of interest for more specific purposes. The definition of the DNA sequences included in the library of defined categories or the library of sequences of interest may be subject to certain user manipulations depending on the needs of the application. Thus, the methods disclosed herein can assemble all possible sequences of interest, which are actually a subset of all possible sequences.

[0035] Oligonucleotides

[0036] In one embodiment, at least two oligonucleotides can be the DNA of any convenient length.For example, the length of at least two oligonucleotides can be greater than 12 nucleotides, or is not limited to about 20 to 65 nucleotides, or 20 to 35 nucleotides, or 35 to 55 nucleotides, or 25 to 65 nucleotides, or 30 to 60 nucleotides, or 40 to 50 nucleotides, or 40 to 60 nucleotides, or about 42 to 48 nucleotides, or about 44 or about 45 nucleotides.The anchor chain used in the method can be 20 to 60 nucleotides, or 20 to 70 nucleotides, or 30 to 60 nucleotides, or 30 to 70 nucleotides, or 35 to 65 nucleotides, or 40 to 60 nucleotides, or 40 to 50 nucleotides.In one embodiment, at least two oligonucleotides are 40 to 50 nucleotides, and the anchor chain is 35 to 45 or 45 to 55 nucleotides.Primer binding sites can be added to or included in these oligonucleotide lengths.Oligonucleotides can exist with any combination or subcombination of length provided herein. In any embodiment, the oligonucleotide can have only nucleotides without non-standard bases. In any embodiment, the oligonucleotide can have only nucleotides with standard bases, that is, all nucleotides in the oligonucleotide have standard bases A (adenine), T (thymine), C (cytosine) or G (guanine). In other embodiments, any one of the oligonucleotides can contain one or more non-standard bases. The oligonucleotide and / or anchor chain can have a sequence for binding primers, which can be used for PCR or another DNA amplification program. When determining the nucleotide size of the library, the ordinary technician of the present disclosure will realize the optimal size of the oligonucleotide used in the method by considering the ability of the oligonucleotide length to anneal to other oligonucleotides. Any one of the oligonucleotides can be programmed and synthesized into a recognition site with a restriction endonuclease in the resulting dsDNA molecule. Restriction enzymes can be restriction enzymes such as IIS type restriction enzymes that recognize asymmetric DNA sequences and cut many nucleotides (e.g., within 1 to 5 or 1 to 10 or 1 to 20 nucleotides) outside its recognition sequence. Any oligonucleotide described herein can be a member of a library of oligonucleotides, including all combinations and subcombinations of described oligonucleotides.

[0037] The method of the present invention synthesizes a product DNA molecule with a "desired sequence", which can be a predetermined sequence, i.e., a sequence determined by a user before starting the method. The product DNA molecule of the desired sequence can be any molecule produced by the method, including but not limited to the first dsDNA molecule, the second dsDNA molecule, the third dsDNA molecule, the fourth dsDNA molecule, and another dsDNA molecule. The dsDNA fragment or another dsDNA fragment can be derived from the restriction enzyme digestion of any product dsDNA molecule. In various embodiments, the DNA molecule of the desired sequence can be at least 16bp, or at least 20bp, or at least 30bp, or at least 40bp, or at least 60bp, or at least 80bp, or at least 90bp, or at least 100bp, or at least 175bp, or at least 200bp, including or not including conservative flanking sequences or primer binding sites. The oligonucleotides (and / or anchor chains) used in the method can contain variable sequences, which can correspond to a part of the variable sequence of the product dsDNA molecule. Therefore, in any embodiment, it can be considered that the product dsDNA molecule has or does not have primer binding sites added to the 3' and 5' ends. In any embodiment, the variable sequence or subsequence of the oligonucleotide and / or anchor chain can be at least 4bp, or at least 5bp, or at least 8bp, or at least 16bp, or at least 28bp, or at least 100bp. The length of the variable sequence of the product DNA molecule can depend on the steps in the method, and can be provided by the combination of oligonucleotides and dsDNA fragments. In various embodiments, the length of the variable sequence can be at least 8% of the oligonucleotide or anchor chain length, or at least 10%, or at least 15%, or at least 20%, or at least 25%. In other embodiments, the length of the variable sequence of the dsDNA molecule can be at least 8% of the product dsDNA molecule, or at least 10%, or at least 25%, or at least 50%, or at least 60%, or at least 75%.

[0038] method

[0039] Figure 1ADescribe the method for synthesizing product dsDNA molecules according to the method of the present invention. O1 and O2 are at least two oligonucleotides, and O3 is an anchor chain. In this embodiment, O1 to O2 each have a primer binding site 101 (which can be a universal primer binding site) on the 5' or 3' end, and a variable sequence 105 on the relative 3' or 5' end. O1 to O2 also each have a conservative flanking sequence 110, which is depicted between the universal primer binding site 101 and the variable sequence 105 in this embodiment. In this embodiment, the total length of O1 and O2 is 40 to 50 or about 45 nucleotides. "Conservative flanking sequence" (CFS) is used to help the oligonucleotide anneal to the target oligonucleotide with complementary CFS, and can also have a primer binding site (e.g., for use later in the method). In various embodiments, CFS can be at least 8 nucleotides, or at least 12 nucleotides, or at least 15 nucleotides, or 15 to 25 nucleotides, 18 to 22 nucleotides, or 8 to 15 nucleotides, or 10 to 18 nucleotides, or 12 to 25 nucleotides, or 12 to 30 nucleotides, or 15 to 20 or 15 to 30 nucleotides, 18 to 22 nucleotides, or 18 to 30 nucleotides, or 18 to 60 nucleotides, but can use any convenient length that can help annealing and provide primer binding site. In a specific embodiment, CFS is 18 to 22 or about 20 nucleotides. In various embodiments, the CFS on the anchor chain can be complementary to the CFS on at least two oligonucleotides, and they are combined (for example, in combination with the concentrated oligonucleotides) thus. In various embodiments, conservative flanking sequences can be present in each molecule in combination with concentration. Conservative flanking sequences can be between primer binding site and variable sequence on at least two oligonucleotides. In either embodiment, the conserved flanking sequences may include a 5' cap to prevent degradation of the dsDNA molecule, and / or a primer binding sequence to aid in amplification.

[0040] In this embodiment, anchor chain O3 has a conservative flanking sequence 110 complementary to the conservative flanking sequence on at least two oligonucleotides (or complementary to at least a portion of the conservative flanking sequence on at least two oligonucleotides that is sufficient to anneal the oligonucleotides) at the 3' and 5' ends. In this embodiment, the length of O3 is about 50 nucleotides, wherein each CFS is about 20 nucleotides. O3 also has at least one variable sequence 105, which is located between two conservative flanking sequences 110 in this embodiment (of the anchor chain) and is depicted as being about 10 nucleotides in length. In other embodiments, the variable sequence can be moved to another position on the oligonucleotide, as long as enough space can be left so that CFS can promote annealing and / or provide a primer binding site (if used). In any embodiment, at least 10, or at least 15, or at least 18 nucleotides of the conservative flanking sequence can be present on both sides of the anchor chain variable sequence. In this embodiment, the variable sequence 105 on anchor chain O3 includes a degenerate nucleotide N, which is as follows here: Figure 1A Six degenerate nucleotides depicted. At least a portion of at least one variable sequence 105 on the anchor chain is complementary to at least a portion of the variable sequence 105 on at least two oligonucleotides, and in one embodiment can be complementary across the entire variable sequence. One or more of the at least two oligonucleotides can be further synthesized to have a recognition site for a restriction endonuclease when assembled into a dsDNA molecule; the anchor chain can also be synthesized to contain a recognition site for a restriction endonuclease so that when combined with at least two oligonucleotides, the recognition site exists and is active. In one embodiment, the restriction endonuclease can be a IIS type restriction endonuclease. The recognition site on at least two oligonucleotides and the anchor chain can be programmed to be located outside the variable sequence on the assembled molecule, but the restriction endonuclease can be cut inside the variable sequence. In one embodiment, the recognition site is contained in a conserved flanking sequence; and in one embodiment, the restriction endonuclease cuts within the variable sequence of the dsDNA molecule. In any embodiment, the nucleotide sequences of any two nucleotide sequences that are complementary or overlapping can have at least 90% sequence identity, or at least 95% sequence identity, or at least 98% sequence identity, or 100% sequence identity. The product DNA molecules can be optionally assembled with conserved flanking sequences for use in continuous procedures (e.g., PCR or other DNA amplification).

[0041] The method involves the step of annealing at least two oligonucleotides O1 to O2 to the anchor chain O3, so that at least two oligonucleotides bound to the anchor chain are adjacent to each other on the anchor chain. In one embodiment, at least two oligonucleotides can be adjacent to their variable sequences. In one embodiment, the variable sequence 105 of at least two oligonucleotides is annealed to the variable sequence of the anchor chain 105. It is worth noting that in the embodiment shown, each of O1 to O2 has a variable sequence of 5 nucleotides, and O3 has a variable sequence of ten nucleotides. After annealing, the corresponding variable sequence anneals and forms a base pair. In one embodiment, the variable sequence forms a continuous sequence after connection. In the present invention, "binding" or "annealing" is used interchangeably with respect to polynucleotides, and refers to the formation of double-stranded DNA molecules by standard Watson-Crick base pairing. In various embodiments, annealing can occur when at least 25%, or at least 50%, or at least 75%, or at least 90% of the nucleotides are paired with complementary oligonucleotide bases. When the first oligonucleotide contains nucleotides adjacent to the nucleotides on the second oligonucleotide, when the two oligonucleotides are combined with the identical (third) complementary oligonucleotide (e.g., anchor chain) at least at its variable sequence by Watson-Crick base pairing, i.e., nucleotides and complementary (e.g., anchor) chains adjacent to each other, the two oligonucleotides are "adjacent" to each other. In any embodiment, the anchor chain can be fully bonded to the first and second oligonucleotides. Fig. 1 and Fig. 2 illustrate this concept. Therefore, when adjacent, each nucleotide of the anchor chain can be annealed with the nucleotides of one of the at least two oligonucleotides. After the amplification reaction, each nucleotide in the at least two oligonucleotides can be annealed to the nucleotides on the anchor chain. In any embodiment, the anchor chain can be annealed to two and no more than two oligonucleotides.

[0042] The method also involves a step of ligating at least two oligonucleotides annealed to an anchor strand to produce a ("first") dsDNA molecule. The step of ligation or "ligating" can mean contacting the annealed dsDNA fragments or dsDNA molecules with a ligase, or allowing the ligation to occur spontaneously. A ligase is an enzyme that catalyzes the joining of two polynucleotide molecules by forming a new chemical bond. In one embodiment, a ligase can connect adjacent (or contiguous) polynucleotides that are bound to the same complementary polynucleotide strand. In any of the methods, any DNA ligase can be used, for example, T4 DNA ligase and E. coli DNA ligase are just two examples, but another DNA ligase can also be used.

[0043] The method may involve the step of performing an amplification step to produce a product dsDNA molecule. In any of the steps in any of the methods, the amplification step may involve, for example, PCR, isothermal amplification, rolling circle amplification, loop-mediated isothermal amplification, or another DNA amplification method for dsDNA molecules (e.g., O1 to O3 and O4 to O6 (when present)) to produce a first dsDNA molecule O7 (and / or O8). In the embodiment of FIG. 1 , as an example, the variable sequences of O7 and O8 are 10-mers, but one of ordinary skill in the art with the benefit of this disclosure will appreciate that any variable sequence of appropriate length may be used in the method, depending on the step in the method. For example, in any embodiment, a variable sequence of 6 to 20, or 6 to 12, or 6 to 14, or 8 to 12, or 10 to 16, or 15 to 25, or 20 to 30, or 30 to 50, or 40 to 100, or 60 to 120 nucleotides, or other number of nucleotides, can be used on the anchor strand, and optionally can correspond in length to the sum of the variable sequences on at least two oligonucleotides or dsDNA fragments of the synthetic dsDNA molecule (minus overlapping nucleotides). In any embodiment, at least two oligonucleotides can have variable sequences of different lengths. For example, one of the two oligonucleotides can have a variable sequence of 4 nucleotides, and the second oligonucleotide can have a variable sequence of 6 nucleotides or other combinations to form a first dsDNA molecule. In the embodiment depicted in Figure 1, a first dsDNA molecule (e.g., O7 or O8) is synthesized that has universal primer binding sites 101 at the 3' and 5' ends, conserved flanking sequences 110 within each of the 3' and 5' ends, and variable sequences 105 within the conserved flanking sequences 110. This arrangement can be utilized in any embodiment. With respect to a DNA molecule, "within" refers to a feature that exists closer to the center of the DNA molecule (and further away from the 5' or 3' end) than a reference feature.

[0044] In some embodiments, multiple product DNA molecules can be "multiplexed," i.e., synthesized in the same reaction pool. In other embodiments where there is a need, the DNA molecules can be synthesized individually (e.g., "in parallel") in their own reaction pool (and subsequently combined). Reactions can be multiplexed using two or more binding sets of at least two oligonucleotides and at least one anchor chain. As with any of the methods, the method depicted in Figure 1 can be performed as a single synthesis of one dsDNA molecule, or as a multiplexed synthesis of at least two dsDNA molecules in the same pool. The synthesized DNA molecules can later (optionally) be connected by methods disclosed herein, such as illustrated in Figures 1 to 2. When multiplexing is utilized, more than one DNA molecule or more than two DNA molecules are synthesized in simultaneous reactions in the same pool. In Figure 2AIn the embodiment, multiplexing is depicted as paired oligonucleotides O4 to O5 and paired anchor strand O6, which form a single paired dsDNA molecule O8, which is paired with the first dsDNA molecule O7. However, the paired dsDNA molecules (e.g., O8) can be synthesized in parallel and separate reaction pools, and the first dsDNA molecule and the paired dsDNA molecule (or fragments thereof) combined in a subsequent step. In either embodiment, dsDNA molecules (or dsDNA fragments) are "paired" when they have overlapping sequences at the variable sequence. In this embodiment, a 10 bp variable sequence in the first dsDNA molecule and the paired dsDNA molecule and a 4 bp overlap in the variable sequence between the dsDNA molecules are depicted (e.g., as in Figure 2B In any embodiment of the method, however, the overlap may be at least 1 bp, or at least 2 bp, or at least 3 bp, or at least 4 bp, or at least 5 bp, or at least 6 bp, or at least 8 bp, or at least 12 bp or more. Any two dsDNA molecules (or dsDNA fragments), whether the first dsDNA molecule, the second dsDNA molecule, the third dsDNA molecule, the fourth dsDNA molecule, etc., or another dsDNA molecule may be a paired dsDNA molecule (or dsDNA fragment).

[0045] dsDNA fragments can be produced (e.g., by restriction enzymes acting on dsDNA molecules, or in any embodiment, by separate synthesis) to produce paired dsDNA fragments with prominent 3' and / or 5' sequences, which can be at their variable sequences and can overlap at least in part. Therefore, such dsDNA fragments can be annealed at 3' and / or 5' overhangs to form larger dsDNA molecules. "Overlapping" (or complementary) sequences are sequences that contain complementary sequences for a series of nucleotides that are sufficient to anneal by Watson-Crick base pairing under standard reaction conditions. In any embodiment, the method can utilize overlapping 1, 2, 3, 4, 5, 6, 7, or 8 nucleotides (which can be continuous nucleotides), or at least 1, or at least 2, or at least 3, or at least 4, or at least 5, or at least 6, or at least 7, or at least 8, or at least 12, or at least 15 nucleotides (any of which can be continuous nucleotides) of dsDNA fragments or polynucleotides. In any embodiment, the overlap can be at their variable sequences. "Overhang" or "overhang" sequence refers to a 3' or 5' single-stranded DNA sequence extending from a double-stranded DNA sequence. In various embodiments, the overhang can be at least 2, or at least 3, or at least 4 nucleotides, or at least 6, or at least 8, or at least 10 nucleotides. At least two paired oligonucleotides can anneal or be bound to their corresponding paired anchor strands. "Paired anchor strands" are only such anchor strands: they have sequences complementary to at least two paired oligonucleotides so that they can fully anneal to work in the method.

[0046] The method may involve a further step of synthesizing a larger product DNA molecule. The method may include the step of contacting the first dsDNA molecule and the paired dsDNA molecule (e.g., O7 to O8) with a restriction enzyme to produce a first dsDNA fragment and a paired dsDNA fragment, the first dsDNA fragment and the paired dsDNA fragment having 3' and / or 5' complementary overhang sequences and a portion of the variable sequence from the first dsDNA molecule and the paired dsDNA molecule, respectively (in Figure 2B In any embodiment, the first dsDNA molecule and the paired dsDNA molecule can have complementary and overlapping variable sequences and can have conserved flanking sequences containing recognition sites for restriction endonucleases (e.g., Type IIS restriction endonucleases). The restriction enzyme can cut within the variable sequence of the first dsDNA molecule (and the paired dsDNA molecule, when present) to produce complementary overhanging 3' and / or 5' sequences on the first dsDNA fragment and the paired dsDNA fragment.

[0047] The method may involve the step of providing at least one additional dsDNA fragment having a 3' and / or 5' overhang sequence complementary to the overhang sequence of at least one other dsDNA fragment (e.g., the first dsDNA fragment or the paired dsDNA fragment) in the synthesis method. The overhang sequence of the additional dsDNA fragment may contain at least a portion of the variable sequence (e.g., as depicted in Figure 2). The overhang sequence may contain at least a portion of the variable sequence of the first dsDNA molecule (and the paired dsDNA molecule (when present)), and thus the fragment may have an overlapping or complementary sequence at the variable sequence. At least one additional dsDNA fragment may be derived from a restriction endonuclease reaction on the dsDNA molecule or its paired dsDNA (e.g., O7 and O8), or may be a DNA fragment produced in another parallel reaction, and may also be synthesized separately. Each of these dsDNA fragments has a 3' or 5' overhang sequence complementary to at least one other dsDNA fragment in the method. However, in other embodiments, multiple additional dsDNA fragments may be input simultaneously and assembled into product dsDNA molecules. Some dsDNA fragments may have 3' and 5' overhang sequences that are complementary to two other dsDNA fragments in the reaction (e.g., Figure 1B ) and can therefore be inserted between the two fragments to extend the product DNA molecule, these dsDNA fragments will therefore have 3' and 5' overhangs, each of which is complementary to the other 3' or 5' overhang of the dsDNA fragment being synthesized.

[0048] Additional dsDNA fragments can be used in any of the embodiments. "Additional dsDNA fragments" (or additional dsDNA molecules) are general terms and are not necessarily specific to any particular step in the method. Such additional dsDNA fragments can be annealed with another dsDNA fragment having a complementary sequence at any step in the method, such as at 3' and / or 5' overhangs. The additional dsDNA fragments can have a sequence that is at least partially complementary to the 3' and / or 5' overhangs on at least one other dsDNA fragment in the method, and the other dsDNA fragments can be the first dsDNA fragment, or the second dsDNA fragment, or the third dsDNA fragment, or the fourth dsDNA fragment, or another additional dsDNA fragment. When a dsDNA molecule is cut with a restriction endonuclease, it can leave 3' and 5' overhangs. Therefore, if two dsDNA molecules are cut with a restriction endonuclease, the resulting two dsDNA fragments can be annealed and connected by their complementary 3' and 5' overhangs (e.g., as Figure 2BIn various embodiments, the complementarity or overlap can be at least 5, or at least 6, or at least 8, or at least 10, or at least 12, or at least 15 nucleotides, which can be continuous nucleotides. In any embodiment, the overhang of the additional dsDNA fragment or any dsDNA fragment can be in the variable sequence. In any embodiment, the additional dsDNA fragment can be derived from another dsDNA molecule, for example, by cutting with a restriction endonuclease to produce another dsDNA fragment. Therefore, the additional dsDNA fragment can have a next series of nucleotides to be synthesized into the final product dsDNA molecule at the 3' and / or 5' overhang to form the desired (predetermined sequence) dsDNA molecule. In any embodiment, any additional dsDNA molecule can have the same structure as the first dsDNA molecule or the second dsDNA molecule, etc. (but with changes in the variable sequence). And any additional dsDNA fragment can have the same structure as the first dsDNA fragment, the second dsDNA fragment, etc. For example, in any embodiment, the additional dsDNA molecules can have primer binding sites at the 3' and / or 5' ends, variable sequences, and conserved flanking sequences on either side of the variable sequences, identical to, for example, O7 or O8. The additional dsDNA fragments can (but need not) be derived from restriction enzyme digestion of the additional dsDNA molecules.

[0049] The method may involve a step of annealing the first dsDNA fragment to at least one paired or additional dsDNA fragment via their complementary overhang sequences (the overlap may be at their variable sequences). The method may also involve a ligation step to produce a second dsDNA molecule (e.g., O9, Figure 1A ) (depicted here as having a 16-mer variable sequence) having a conserved flanking sequence (CFS) 110 within each of the 3' and 5' ends, and a variable sequence 105 within the 3' and 5' conserved flanking sequences that is longer than the variable sequence on the first dsDNA molecule.

[0050] Thus, the method may further involve the step of contacting at least one second dsDNA molecule with a restriction enzyme to produce a plurality of second dsDNA fragments comprising 3' and / or 5' overhang sequences. At least two of the plurality of fragments may have a conserved flanking sequence within each of the 3' and / or 5' overhangs. The method may further involve the steps of annealing at least one of the second dsDNA fragments to one or more paired or additional dsDNA fragments having a complementary 3' or 5' overhang sequence (which may be at a variable sequence), and performing a ligation step to produce at least one third dsDNA molecule having a conserved flanking sequence at the 3' and 5' ends (or, optionally, within the primer binding sites at the 3' and 5' ends) and a variable sequence within the conserved flanking sequence that is longer than the variable sequence of the second dsDNA molecule. In the embodiment depicted in Figures 1 and 2, at least one third dsDNA molecule has a variable sequence of 28 bp and overlaps 4 bp with at least one paired third dsDNA molecule.

[0051] The method may further involve the steps of reacting at least one third dsDNA molecule with a restriction enzyme to produce at least one third dsDNA fragment (e.g., O11 having a 3' and / or 5' overhang sequence), optionally annealing at least one third dsDNA fragment to one or more paired or additional dsDNA fragments (e.g., O12 having a complementary 5' or 3' overhang sequence), and performing a ligation step to produce a fourth dsDNA molecule. At least two of the dsDNA fragments in the mixture may have a conserved flanking sequence within the variable sequence and a variable sequence at the 3' or 5' overhang. Thus, the fourth dsDNA molecule may have a conserved flanking sequence within the 3' and 5' ends, an optional primer or universal primer binding sequence 120, and a variable sequence that is longer than the variable sequence of the third dsDNA molecule (optionally between CFSs). In the embodiments depicted in Figures 1 and 2, at least one fourth dsDNA molecule has a variable sequence of 100 bp O14. As with the other steps, one or more additional dsDNA fragments may be included in the reaction to further extend the variable sequence of the product dsDNA molecule. At least one additional dsDNA fragment may be derived from a parallel (or multiplexed) reaction. Figure 1BAs depicted as fragments 125 and 130 in the reaction, one or more of the dsDNA fragments can have overhangs at both the 3' and 5' ends, and these overhangs have sequences complementary to the overhang sequences of two other dsDNA fragments in the reaction. These additional dsDNA fragments can be produced by including two restriction recognition sites on the dsDNA molecule contacted with a restriction endonuclease (optionally within the CFS), and then the restriction endonuclease can cut the dsDNA molecule into at least three fragments. Therefore, the dsDNA fragments can be connected in the annealing and ligation reactions to form longer product dsDNA molecules. In any of the embodiments or steps of the method, two or more dsDNA fragments can be included in the reaction, and the dsDNA fragments can have 3' and 5' overhang sequences and do not have a CFS, that is, these dsDNA fragments that have both 5' and 5' overhang sequences can all be variable sequences.

[0052] These methods provide great versatility in the synthetic product dsDNA molecule of the desired sequence. In any embodiment, the dsDNA molecule can be synthesized by multiple dsDNA fragments (for example, two, or three, or four, or five, or six, or more than six fragments). Multiple fragments can each contain a part of the product dsDNA molecule to be synthesized. In any embodiment, the step in the synthesis or the final step in the synthesis can include at least one dsDNA fragment in the annealing reaction, and the at least one dsDNA fragment includes at least a portion of the desired sequence so that the desired sequence is present on the product dsDNA molecule. For example, in any embodiment, 5' caps and / or primer binding sites 120 can be added to the 3' and / or 5' ends of the product dsDNA molecule. In the embodiment of using multiple fragments in the synthesis, the first fragment on the 5' end of the molecule assembled and the last fragment on the 3' end can have a 5' cap. The 5' cap can help prevent the degradation of the DNA molecule end, and the initiator sequence is convenient for amplification when needed. In various embodiments, the 5' cap can be any suitable cap that protects the oligonucleotide from degradation, such as a phosphorothioate bond between at least one of the last 2 nucleotides, or the last 3 or last 4 or last 5 nucleotides at the 5' and / or 3' ends.

[0053] Addition reactions can also be carried out. Figure 1BIn the depicted embodiment, the fourth dsDNA molecule O14 has a variable sequence of 100bp. Parallel reactions can produce multiple other dsDNA molecules, which have complementary and / or overlapping sequences with the third dsDNA fragment that will form the fourth dsDNA molecule. Other dsDNA can also have a variable sequence of, for example, 100bp or any suitable length. Any one of the dsDNA molecules can be cut with one or more restriction endonucleases to produce multiple dsDNA fragments with 3' and / or 5' overhangs, which contain complementary sequences of 3' and / or 5' overhangs of one or two other dsDNA fragments. Therefore, in any embodiment and in any step, synthesis can include one or more dsDNA fragments, which have 3' and 5' overhang sequences complementary to the overhang sequences on at least one other dsDNA fragment in the mixture. Overhangs can include at least a portion of the variable sequence of each dsDNA molecule. The dsDNA fragments can be combined to synthesize much longer variable sequences in the product dsDNA molecule.

[0054] A more detailed illustration of the method of the invention is provided in Figure 2. At least two oligonucleotides O1 to O2 and an anchor strand O3 are depicted, as well as paired oligonucleotides O4 to O5 and a paired anchor strand O6. Figure 2B Complementary overlapping 3' and 5' overhang sequences that appear after restriction endonuclease digestion of a first dsDNA molecule and a paired dsDNA molecule are shown. The variable sequence is depicted as a 10-mer and forms part of the 3' or 5' overhang sequence in the oligonucleotide. A second dsDNA molecule (O9) (illustrated with a variable sequence of 16 nucleotides) is also depicted, which is synthesized after annealing and amplification (PCR2) of the first dsDNA fragment and the paired dsDNA fragment. In the embodiment depicted in Figure 2, annealing, ligation 0 (L0) and PCR1 occur between at least two oligonucleotides and an anchor strand, depicted here in the binding sets O1 to O3 and O4 to O6, to form a first dsDNA molecule (O7) and a paired dsDNA molecule (O8). Digestion with a restriction endonuclease is then performed to produce a first dsDNA fragment and a paired dsDNA fragment, which are then ligated 1 (L1) with another dsDNA fragment (here, the paired dsDNA fragment), and then PCR2 is performed to form a second dsDNA molecule, depicted here as having a variable sequence (O9) that is a 16-mer.

[0055] The second dsDNA molecule can then be digested with a restriction endonuclease to form a second dsDNA fragment and ligated with an additional dsDNA fragment (L2), followed by PCR3 to form a third dsDNA molecule, which is depicted as having a 28-mer variable sequence (O10) ( Figure 1B). The third dsDNA molecule can then be digested with a restriction endonuclease, and in one embodiment forms a third dsDNA fragment (e.g., O11), which can be combined and connected (L3) with another dsDNA fragment (e.g., O12). One or more dsDNA fragments with 3' and 5' overhangs (e.g., 125, 130) can be included, and PCR3 is performed to produce a fourth dsDNA molecule (O14), which is depicted as having a 100-mer variable sequence. One or more additional dsDNA fragments (e.g., 125, 130, O12) can be included in the reaction, which can be derived from multiplexed or parallel synthesis reactions. In this embodiment, the additional dsDNA fragments are derived from additional dsDNA molecules digested with restriction endonucleases. Any dsDNA molecule can be synthesized to have two restriction sites to produce dsDNA fragments with overhangs at both the 5' and 3' ends. "Parallel DNA synthesis reactions" can be reactions for synthesizing DNA molecules of different desired sequences (i.e., sequences different from the primary reactions parallel to them). The parallel reactions can be performed as separate reactions in separate pools, but can also be multiple reactions in the same pool. The DNA molecules of the desired sequence from the parallel synthesis reactions may contain overlaps at variable sequences with the DNA molecules of the desired sequence in the primary reaction.

[0056] The terms "first dsDNA molecule", "second dsDNA molecule", "third dsDNA molecule", "fourth dsDNA molecule", "dsDNA fragment", "additional dsDNA molecule", and "paired dsDNA molecule", "first anchor", and the like, "paired anchor" are relative terms provided to help track the molecule through any step in the method, and do not necessarily refer to any absolute point or DNA molecule or fragment in the reaction. The "paired" dsDNA molecule or fragment contains a variable sequence that overlaps with and is at least partially complementary to (e.g., at least 3, or at least 4, or at least 5 consecutive bp) the variable sequence of a reference dsDNA molecule or fragment. In one embodiment, the "paired" dsDNA molecule is multiplexed with a reference dsDNA molecule and an "additional dsDNA molecule" synthesized in parallel synthesis. For example, a first dsDNA molecule contains a variable sequence, and its paired dsDNA molecule may contain a variable sequence that at least partially overlaps with the variable sequence of the first dsDNA molecule, thereby enabling them to be synthesized into a single larger dsDNA molecule. In another embodiment, the variable sequence of the first dsDNA molecule will at least partially overlap with the variable sequence of at least one additional dsDNA molecule. The second dsDNA molecule contains the variable sequences of the first and paired (or additional) dsDNA molecules, and in turn may at least partially overlap with the paired dsDNA fragment or additional dsDNA fragment having a variable sequence that is at least partially complementary. The third dsDNA molecule contains a portion of the variable sequence from at least one second dsDNA molecule, and may further contain a portion of the variable sequence from the first dsDNA molecule, and may also have the variable sequence of one or more additional dsDNA molecules. The fourth dsDNA molecule may contain variable sequences from the first dsDNA molecule (and its paired molecule), the second dsDNA molecule (and its paired molecule), and the third dsDNA molecule (and its paired molecule); in some embodiments, the fourth dsDNA molecule contains the variable sequences of multiple third dsDNA molecules. This can continue, and five to ten or more dsDNA molecules can be synthesized in a hierarchical manner, such as Figure 3 When digested by a Type IIS restriction endonuclease, the dsDNA molecules will produce dsDNA fragments having 3' and / or 5' overhang sequences that are complementary to at least one other dsDNA fragment in the mixture at their variable sequences (or produced by parallel synthesis reactions). In any embodiment, any of the dsDNA molecules can be formed without the use of blunt-end ligation. Thus, whether the first, second, third, fourth, paired, additional, etc. dsDNA molecule, any of the steps or methods can involve the use of annealing rather than utilizing blunt-end ligation.

[0057] Therefore, method allows to produce the product DNA molecule with variable sequence of any length, and does not need traditional oligonucleotide synthesizer, and traditional oligonucleotide synthesizer usually relies on chemical synthesis (such as phosphoramidite chemistry).On the contrary, method can only rely on the synthesis based on enzymatic as described herein, and therefore can produce DNA molecule or polynucleotide as required.DNA molecule can refer to single-stranded polynucleotide or double-stranded DNA combined by Watson-Crick base pairing.The method can also relate to carrying out multiple PCR circulation or another DNA amplification program to any product DNA molecule.In certain embodiments, can only use enzyme and the buffer supporting enzyme to carry out method to polynucleotide.

[0058] In any embodiment, the method or any step of the method can be performed without cloning or without the need for cloning, such as without the use of host cells at any time point of the method. In any embodiment, the method or any step of the method can be performed entirely in vitro, or can be performed without the use of living cells for any purpose in the method. In any embodiment, the method or any step of the method can be performed without the use of terminal deoxynucleotidyl transferase (TdT) or without the use of non-template-dependent DNA polymerase. In any embodiment, the method or any step of the method can produce a traceless product DNA molecule. Traceless DNA means such DNA, which does not have any nucleotides introduced by or from the process of preparing its synthetic DNA or nucleotides, at least with respect to the variable sequence of the product DNA molecule (e.g., residual nucleotides from a joint or adapter or flanking sequence). In any embodiment, the method or any step of the method can produce a product DNA molecule without a barcode or a nucleotide sequence placed for the original purpose of identification. A barcode can be a sequence that is not otherwise required but has a specific sequence and is used to identify a DNA sequence. In various examples and embodiments, the barcode sequence is 6 to 8 nucleotides in length, or 4 to 10 nucleotides in length. In any embodiment, the method or any step of the method can be performed without any portion of any oligonucleotide used in the method being immobilized, i.e., bound to a solid phase or solid support (e.g., beads, DNA chips, microfluidic surfaces, etc.). In any embodiment, the oligonucleotides can be annealed in solution and can be ligated in solution, i.e., no oligonucleotide is bound or partially bound to a solid phase or solid support (e.g., a DNA chip, beads, surface, or other solid phase) during the step or method.

[0059] In any embodiment, any step of the method or method can synthesize product DNA molecules without using and without carrying out chemical assembly technology (such as phosphoramidite chemistry).In any embodiment, the method can synthesize the DNA molecules of the desired sequence without using joints, adapters or spacer oligonucleotides or sequences."Joints", "adapters" or "spacer" molecules can be short oligonucleotides that can be connected to the ends of other DNA or oligonucleotide molecules.Joints, adapters or spacers can also be used to provide the release of polynucleotides from solid supports, or connect or tether polynucleotides to solid supports.These molecules or sequences can provide sticky ends and / or overhangs that allow connection.Joints, adapters or spacer DNA sequences can also include, for example, recognition sites (such as for endonucleases), primer binding sites, polyU sequences, or can be sequences with one or more uracil residues.Joints, adapters or spacer DNA as mentioned herein can be such sequences, which do not contain the nucleotide sequence of a part of the DNA molecules of the desired sequence synthesized in the method. In any embodiment of the method, oligonucleotide or anchor chain for synthesizing the DNA molecule of desired sequence can have one or more of the above-mentioned structures, but the structure is not provided on a separate joint, adapter or spacer molecule. In any embodiment, the method can be carried out without using a joint, adapter or spacer DNA or sequence. In any embodiment of the method, only oligonucleotide or anchor chain are used to synthesize the DNA molecule of desired sequence, and oligonucleotide or anchor chain contain at least a portion of the nucleotide sequence that will be present in the synthetic DNA molecule of desired sequence. In any embodiment, the part can be at least 6 or at least 8 or at least 10 or at least 16 or at least 28 or at least 50 or at least 100 continuous nucleotides. This is another advantage of the method, and makes the method more suitable for automation.

[0060] In any embodiment, the method or any step of the method can only use the enzymatic assembly of oligonucleotides to assemble product DNA molecules. In any embodiment, the method or any step of the method can be carried out by extracting at least two oligonucleotides and anchor chains from a library comprising less than 20,000 members or from any library described herein. In any embodiment, at least two oligonucleotides and anchor chains can be selected from an oligonucleotide library with less than 10,000 members, or selected from any oligonucleotide library described herein. In any embodiment, the method or any step of the method does not utilize or need not use a carrier in the method.

[0061] The product DNA molecule can be optionally formed to have conservative flanking sequences and / or optionally have universal primer binding sites on the 3' and 5' ends of the product DNA molecule. At least two oligonucleotides can be formed to have one or more primer binding sites, and the primer binding site can provide binding sites for primers in an amplification program (e.g., by PCR). Once the anchor chain (e.g., sufficiently long product DNA molecules have been synthesized) is no longer needed, the primer combined with the conservative flanking sequences can be used to increase, and the universal primer binding site is not needed.

[0062] The method can be facilitated by using a recognition site for a restriction endonuclease that can be effectively activated or inactivated. For example, with reference to Figure 1, when combined with anchor chain O3, O1 can have an inactive recognition site, while O2 can have an active site programmed into the sequence. The recognition site can be used to digest the formed dsDNA molecule O7 / O8, and digest the dsDNA molecule at one position, leaving 5' and 3' overhangs. When multiplexed or in a single synthesis reaction, restriction sites can be formulated within the sequence so that the restriction site is active on one side (e.g., 3' or 5' side) of the variable sequence of the first dsDNA molecule, and inactive on the opposite side, and vice versa for the second dsDNA molecule of the pair that will be combined in the synthesis step of the present invention. Therefore, when digested, each dsDNA will produce two dsDNA fragments, which can then be annealed. However, in some embodiments where larger dsDNA is obtained (e.g., dsDNA having at least a 20-mer or 28-mer or similar variable sequence), and when additional dsDNA fragments having complementary 3' and 5' overhangs 130 from parallel reactions are contemplated, the dsDNA molecule can be formulated so that it has active restriction recognition sites on both sides of the dsDNA molecule. Thus, when digested, it will be cut into at least three fragments, at least one of which has both a 3' and 5' overhang sequence, which can be at the variable sequence. Then, according to Figure 1B, this additional dsDNA fragment can be included in the annealing and ligation reaction using at least one 3' end and at least one 5' end of the dsDNA molecule. This makes it possible to greatly increase the length of the variable sequence in the product dsDNA molecule. By using primers with nucleotide mismatches, the recognition (and restriction) site can be "opened" or "closed" so that the product dsDNA molecule no longer has an active recognition site (or has an active recognition site that was not previously available). For example, the restriction site for BsaI is 5'-GGTCTC(N1)-3' (SEQ ID NO: 17). By changing one nucleotide in the sequence, the restriction recognition site can be activated or inactivated (e.g., "opened" or "closed"). This can be achieved by using primers with a single mismatch, thereby changing the resulting sequence. This can be used for any restriction endonuclease and can be used to place or remove recognition sites at either or both ends of the DNA.

[0063] In any embodiment, the method may include removing the conservative flanking sequences and / or primer binding sites on one or both sides of the DNA molecule after amplification to produce the step of product DNA molecules. The method of removing flanking sequences is known in the art. In some embodiments, conservative flanking sequences and / or primer binding sites can be used to increase the length of the product DNA molecule, or the product DNA molecule can be surrounded by transcription elements (which can be on the 5' and / or 3' side of the variable sequence) or other beneficial sequences to be used in the final desired sequence. For example, flanking sequences can be set to provide a promoter before the variable sequence (e.g., 5'), and / or a terminator (i.e., a regulatory sequence) can be provided after the variable sequence (or 3'). In one embodiment, the product DNA molecule is a gRNA sequence (e.g., 16 to 20bp). The flanking sequence can be optionally set to provide a promoter in front of the gRNA sequence, and a Cas9 handle and a terminator are provided thereafter. Therefore, in some embodiments, the product DNA molecule can be extended to cover primer binding sites and / or flanking sequences and / or one or more regulatory sequences and / or Cas9 handles, any of which can provide more utility than just binding primer sites. In any of these embodiments of the methods, the primer binding site can be a universal primer binding site.

[0064] Any method disclosed herein can be performed in an automated method, such as by an automated instrument. An automated method is a method that does not require human intervention after the method is started-the method does not require any human operation from this point to completion. The automated instrument may contain a component for selecting an oligonucleotide member from an oligonucleotide library. The DNA sequence to be assembled can be uploaded, recorded or stored on a non-transient computer-readable medium. The non-transient computer-readable medium may be programmed to perform an automated step when inserted into a processor attached to or contained in an automated instrument or otherwise communicated electronically with the processor. The automated step may be any automated step disclosed herein for performing any method disclosed herein. Therefore, the present invention also provides a non-transient computer-readable medium programmed with the position of each member of the oligonucleotide library described herein, wherein the oligonucleotide library is present on a suitable support structure of the oligonucleotide library. In one embodiment, the non-transient computer-readable medium is programmed with at least 6,000 or at least 9,000 positions of oligonucleotide library members. The medium may also be programmed with instructions to combine 4 to 6 members of the binding set from the library and assemble the members of the binding set into product DNA molecules according to the method described herein. A "member" of a library is one or more polynucleotides at a certain position. The oligonucleotide library may be contained on any type of medium, such as a multiwell plate or a plurality of plates.

[0065] The present invention also provides a test kit with an oligonucleotide library described herein on a medium. Medium can be any suitable medium, for example, DNA chip, one or more beads, microtubules, one or more 96 orifice plates, one or more 384 orifice plates, one or more 1536 orifice plates, one or more microfluidic reaction supports, one or more microtiter plates, one or more nanotiter plates, one or more picotiter plates or other solid supports or solid phase surfaces of the oligonucleotide members of the library that can be retained or more. When utilizing more than one medium, medium can exist with the quantity that is enough to hold the oligonucleotide library. The medium containing the oligonucleotide library can contain the member of any suitable volume, and example includes 1nl until 100ul or 10nl until the volume of 100ul. DNA chip (or DNA microarray) is the solid surface with microscopic position set, and oligonucleotide can be attached and / or stored thereon.

[0066] The method of the present invention can synthesize product DNA molecules with extremely low error rates. In various embodiments, the method can produce any product DNA molecule described herein, and its error rate is less than 1 in 1,000 base pairs, or less than 1 in 2,000 base pairs, or less than 1 in 2,400 base pairs, or less than 1 error in 2,500 base pairs, or less than 1 error per 3,000 bases, or less than 1 error per 5,000 base pairs, or less than 1 error per 5,300 base pairs, or less than 1 error per 6,000 base pairs, or less than 1 error per 8,000 base pairs, or less than 1 error per 12,000 base pairs, or less than 1 error per 14,000 base pairs.

[0067] General steps

[0068] In either embodiment, the method can begin by pooling at least two oligonucleotides, for example from an oligonucleotide library, and an anchor. Figure 1A to Figure 1B General examples are described in . In the examples described here, multiplexing will be utilized and at least two oligonucleotides and anchor strands have been prepared with a BsaI restriction site, but any restriction enzyme or type IIS restriction enzyme may be utilized.

[0069] The oligonucleotide pool may undergo an annealing step and a ligation step (e.g., LO and PCR1). The ligation step may be performed by contacting the oligonucleotide pool with a ligase, such as T4 DNA ligase. However, any ligase may be used in any step of the present invention. Ligation may be performed by annealing complementary 5' and 3' overhang sequences on dsDNA fragments produced by restriction endonuclease digestion. Ligation may also involve contacting the oligonucleotides with a ligase, and forming covalent bonds between adjacent nucleotides. The polymerase chain reaction (PCR) is a common reaction in biology known to those of ordinary skill. PCR may be used in the present invention according to normal procedures and well-known techniques. PCR (PCR1) causes the amplification of the oligonucleotides, Figure 1A In the example of , the oligonucleotides are depicted as O7 and O8 and have a variable sequence that is a 10-mer. The method can involve the steps of digestion with a restriction endonuclease and annealing and ligation with additional dsDNA fragments, followed by PCR (D1 and PCR2) steps. The oligonucleotide set can be digested with a restriction enzyme (e.g., a type II restriction endonuclease). Digestion produces dsDNA fragments with 3' and 5' overhangs, which can then be annealed to other fragments with complementary overhangs, ligated and amplified in the PCR2 step to form dsDNA molecules with variable sequences, in Figure 1AIn the embodiment of the invention, it is depicted as a 16-mer (O9). The product can be subjected to another digestion (DL2) to generate dsDNA fragments, followed by annealing, ligation and PCR3 steps to form dsDNA molecules. Figure 1A is depicted as having a 28-mer variable sequence (O10, Figure 1B ). The steps of digestion with restriction endonucleases, annealing with additional dsDNA fragments and ligation (DL3) can then be utilized. The ligation can optionally involve dsDNA fragments from parallel reactions or synthesized otherwise, which have 3' and / or 5' overhangs that are complementary to at least one other dsDNA fragment in the mixture. Therefore, optionally, the annealing step can involve adding or using additional dsDNA fragments, which have 3' and 5' overhang sequences that have complementary sequences to two other dsDNA fragments in the mixture. Annealing of the dsDNA fragment mixture produces longer dsDNA molecules. In this way, the length of the variable sequence can be rapidly increased. In this example, the product dsDNA molecule (O14) comes from the combination of two dsDNA fragments with two additional dsDNA fragments (all from 28-mers). The product dsDNA molecule of the desired sequence in this embodiment is Figure 1B 100-mer variable sequence depicted in .

[0070] Primer binding site

[0071] In any embodiment of the method, the primer binding site can be present on some DNA molecules. In certain embodiments, the site can be present on at least two oligonucleotides and on product dsDNA molecules (e.g., the first dsDNA molecule). The primer binding site can be a part of a conservative flanking sequence, or be different from a conservative flanking sequence. However, in certain embodiments, different primer binding sites can be eliminated in any step after no longer utilizing the anchor chain. For example, these sites can be eliminated after forming at least one first dsDNA molecule or the second dsDNA molecule, and thereafter a part of the conservative flanking sequence is used as the primer binding site. Therefore, at least two oligonucleotides can have primer binding sites, which are then present in at least one first dsDNA molecule, but any one or more of the second dsDNA molecule, the third dsDNA molecule, and the fourth dsDNA molecule can lack (or can have) primer binding sites. The primer binding site can also be added to the dsDNA molecule in any convenient step in the method, such as when forming the final product dsDNA molecule, it may be found that a convenient method for having an amplified product is required. The length of primer binding site and / or the length of complementary part between primer and primer binding site can be at least 4, or 5, or 6 nucleotides, or at least 10, or at least 15, or at least 18, or at least 20 nucleotides, or at least 25 nucleotides, or less than 15 nucleotides, or less than 12 nucleotides, or less than 10 nucleotides, or less than 8 nucleotides, which can be continuous nucleotides in any embodiment.But specific length is not required, only primer binding site is required to allow the combination of primers and the amplification of molecules.In any embodiment, primer binding site can be universal primer binding site, and can have identical sequence on all molecules in the mixture with primer binding site, so as to enable amplification mixture from a single primer set.In any step of amplification, all dsDNA molecules to be amplified can have universal primer binding site of identical sequence.In any embodiment, at least two oligonucleotides or DNA molecules of desired sequence can have single (i.e. only one) primer binding site on 3' and / or 5' end.

[0072] In any step or embodiment of any method, one or more or all primer binding sites on the polynucleotide used in the method can be universal primer binding sites.Universal primers are complementary and can be combined with universal primer binding sites.Universal primers are used to allow a primer set or a small primer set to be amplified and assembled on the whole mixture of oligonucleotide or pool or on a subset of oligonucleotide."Universal primer binding site" can be a primer sequence common to a specific oligonucleotide or DNA molecule set.For example, in various embodiments, the universal primer binding site can be present in at least 25% or at least 50% or at least 60% or about two-thirds or at least 70% or at least 80% or at least 90% or at least 95% or at least 98% or 100% polynucleotide in a specific mixture or pool.In certain embodiments, primer binding sites (including universal primer binding sites) can be located only on the terminal 50nt at either end or both ends of a DNA molecule, or only on terminal 30nt or 25nt or 20nt or 15nt. In other embodiments, primer binding sites (including universal primer binding sites) may be located on conservative flanking sequences. In any embodiment, primer binding sites (including universal primer binding sites) may not be located on variable sequences. In other embodiments of the method, it may be desirable to utilize common primers or non-universal primers. Although this will require the use of other primers and complementary primer binding sites to amplify and assemble, it can also provide flexibility for the method when needed. In certain embodiments, primer binding sites can only be found on a sequence (or its complementary sequence) of oligonucleotide. Common primers and primer binding sites differ from universals only in that they cannot be found on most of the sequences in the method. "Pool" of term oligonucleotide is used herein according to common meaning, and it indicates the oligonucleotides in different and independent reaction pools.

[0073] Mutable Sequence

[0074] As the method proceeds, whether in multiplex mode or in parallel, the variable sequence in the dsDNA molecule can grow longer as the method proceeds, due to the gradual or continuous combination of more DNA and / or oligonucleotides containing variable sequences that will become part of the product dsDNA molecule. In either embodiment, the length of the variable sequence in the first dsDNA molecule can be equal to the length of the variable sequence from at least two oligonucleotides combined and annealed on the anchor strand. In either embodiment, the variable sequence in the first dsDNA molecule can be 6 to 14 base pairs, or 7 to 13 base pairs, or 8 to 12 base pairs, or about 10 base pairs, which can be adjusted according to the dsDNA molecule to be synthesized. In either embodiment, the second dsDNA molecule can have a variable sequence of 8 to 24 or 10 to 22 or 12 to 20, or 14 to 18 or 15 to 17 base pairs (or, as in either step, the length of the variable sequence in the dsDNA fragment synthesizing it minus overlapping nucleotides). In any embodiment, the third dsDNA molecule comprises a variable sequence of 18 to 38 or 20 to 36 or 24 to 32 or 26 to 30 or 27 to 29 base pairs. In any embodiment, the fourth dsDNA molecule can have a variable sequence of 70 to 130, or 80 to 120, or 90 to 110, or 70 to 200 base pairs. However, the length of the variable sequence in any step is not fixed and can be changed to any length that is convenient or required in the application.

[0075] The variable sequence can be such a sequence, which will be present in the product dsDNA molecule, or form the "desired sequence" of the DNA molecule of the desired sequence, and does not form a part for a primer binding site or a conservative flanking sequence. Therefore, the variable sequence is an important component of the DNA molecule of the desired sequence. Therefore, the variable sequence in each construct will change according to the part of the final product DNA molecule carried by it and the product dsDNA molecule being synthesized. The product DNA molecule can be a DNA molecule with a desired sequence. In one embodiment, all variable sequences in at least two oligonucleotides will be present in the product dsDNA molecule produced when carrying out any method.

[0076] The variable sequence in at least two oligonucleotides (or in any step of the method or in the molecule) can be at least 4 nucleotides, or at least 5 nucleotides, or at least 6 nucleotides, or at least 10 or at least 12 or at least 15 or at least 18 or at least 20 nucleotides, or 3 to 7 nucleotides, or 4 to 6 nucleotides, or 4 to 8 nucleotides, or 6 to 10 nucleotides, or 6 to 12 nucleotides, or 12 to 16 nucleotides, or 14 to 18 nucleotides. The variable sequence of anchor chain can be equal to the length of the variable sequence in at least two oligonucleotides. In any embodiment, the variable sequence on anchor chain can be annealed completely with the variable sequence on at least two oligonucleotides.

[0077] In any embodiment, the variable sequence can exist as a continuous sequence. In other embodiments, the nucleotides of the variable sequence can be separated to include variable regions individually or in the form of a group of two or three or four or more continuous nucleotides in the entire oligonucleotide sequence. The variable sequence can be at least a portion of the desired sequence or product dsDNA molecule to be synthesized in the method. In one embodiment, the library can contain different oligonucleotides for each possible variable sequence of the oligonucleotide, and each different sequence can be present in different positions in the oligonucleotide library. Therefore, each oligonucleotide with different variable sequences can be located at different positions in the oligonucleotide library. For example, O1 in at least two oligonucleotides has a variable sequence. When the variable sequence is five nucleotides, O1 can have 1024 possible nucleotide sequences, i.e., 4x4x4x4x4 equals 1024 variable sequences of O1, and each variable sequence can be present in 1024 different positions in the library. This is also true for O2 to O6, as depicted in the embodiment in Figure 1. In various embodiments, the variable sequence can be a sequence that limits each of "at least two oligonucleotides" in different positions in the library. In various embodiments, half (or only a portion) of the variable sequence is passed to the next step in the method, with the remainder of the variable sequence provided by the dsDNA fragment that is combined or annealed with the fragment of the invention.

[0078] In any embodiment, the variable sequences of two dsDNA molecules (including product dsDNA molecules) can overlap, i.e., have complementary sequences of two or more nucleotides. In some embodiments, any two dsDNA molecules can contain variable sequences that overlap by at least 1, or at least 2, or at least 3, or at least 4, or at least 5, or at least 6, or at least 7, or at least 8 nucleotides, or 1 to 6 nucleotides, or 2 to 4 or 2 to 5 nucleotides, or 3 to 10 nucleotides, or about 4 nucleotides, or at least 10 nucleotides, or more than 8 nucleotides. In various embodiments, the first dsDNA molecule or the second dsDNA molecule or the third dsDNA molecule or additional dsDNA molecules described herein can have variable sequences that overlap with other dsDNA molecules as described. For example, a first dsDNA molecule may have a variable sequence that overlaps with a variable sequence of a dsDNA molecule it is paired with or another dsDNA molecule, and a second dsDNA molecule, a third dsDNA molecule, or another dsDNA molecule may all similarly have a variable sequence that overlaps with a variable sequence of a dsDNA molecule it is paired with or another dsDNA molecule. However, a dsDNA molecule may also have a variable sequence that overlaps with a variable sequence of any other dsDNA molecule (e.g., a second dsDNA may overlap with a third dsDNA from a parallel synthesis reaction).

[0079] In any embodiment, the dsDNA fragments may also have 3' and / or 5' overhang sequences containing variable sequences that overlap with variable sequence overhangs of other (paired) dsDNA fragments by at least 3, or at least 4, or at least 5, or at least 6, or at least 7, or at least 8, or at least 10, or more than 8 nucleotides. Thus, a first dsDNA fragment may have a variable sequence on a 3' and / or 5' overhang that overlaps with a variable sequence on a paired dsDNA fragment at its 5' or 3' overhang. A second dsDNA fragment may have a 3' and / or 5' overhang sequence containing a variable sequence that overlaps with a variable sequence of a paired dsDNA fragment or another dsDNA fragment. The 3' and / or 5' overhangs may be generated by restriction endonucleases acting on dsDNA molecules, and may also be synthesized separately and provided to any one of the reactions. In any step of the method, the dsDNA fragments may have 3' and / or 5' overhang sequences as part of the variable sequence and may be used to anneal and combine them with one or more other dsDNA fragments at their variable sequences.

[0080] Degenerate nucleotides

[0081] One or more of at least two oligonucleotides and / or anchor chains used in the method may optionally have one or more degenerate nucleotides. In one embodiment, only anchor chains contain degenerate nucleotides. Degenerate nucleotides refer to nucleotides present at degenerate positions in oligonucleotide sequences. In any embodiment, degenerate nucleotides may be present in the variable sequence of anchor chains or other oligonucleotides in the method and as a part of the variable sequence. Degenerate nucleotides in oligonucleotides are nucleotides that can be any one of A, C, T or G (i.e., nucleotide positions that have been randomized in library members). Randomization can be performed by simply providing a sequence with all four bases during oligonucleotide synthesis, thereby producing oligonucleotides with randomized positions. However, in some embodiments, degenerate nucleotides may be universal bases, which may be base paired with all four standard bases. Examples include deoxyinosine, 2-deoxyinosine, nitroindole, 5-nitroindole, 2'-deoxyhygromycin, 3-nitropyrrole, dP, dK or other universal bases that may be used to reduce degeneracy. 3-nitropyrrole 2'-deoxynucleoside and 5-nitroindole 2'-deoxynucleoside can also be used as degenerate bases.Oligonucleotides with one or more degenerate nucleotides are degenerate oligonucleotides.Degenerate oligonucleotides can be co-located in the same (degenerate oligonucleotide) position in the oligonucleotide library.Therefore, degenerate oligonucleotides can be present in a certain position as a group of slightly different sequences, wherein each degenerate oligonucleotide has different sequences due to degenerate nucleotides, but all co-located in the same position.In certain embodiments, the degenerate nucleotides on an oligonucleotide can anneal to the nucleotides of the variable sequence on another (target) oligonucleotide, such as depicted in Figures 1 to 2.In various embodiments, any one of the oligonucleotides in the concentration can have the degenerate nucleotides in its variable sequence.In certain embodiments, at least one oligonucleotide in the concentration has degenerate nucleotides.In certain embodiments, only the anchor chain has the variable sequence containing degenerate nucleotides.With reference to Figures 1 to 2, the anchor chain is depicted as having the degenerate oligonucleotides (designated as "N") in the variable sequence. A "binding set" is a group of oligonucleotides that bind to each other in a step of the methods disclosed herein and form a dsDNA product in the method. Thus, O1 to O3 are as follows Figure 1A, O4 to O6. In some embodiments, the oligonucleotides of the binding set can substantially bind to each other (i.e., not just a small amount of binding). In another embodiment, the oligonucleotides of the binding set bind to each other without mismatched base pairs. In some embodiments, the binding set includes at least one oligonucleotide that fully binds to one or two or more other oligonucleotides in the binding set. In various embodiments, at least one oligonucleotide of the binding set binds to one or two or more other oligonucleotides of the binding set by at least 80%, or at least 90%, or at least 95%, or 100% (i.e., without mismatched bases). A "target" oligonucleotide is a second oligonucleotide that is expected to bind to a first oligonucleotide in the method.

[0082] In any embodiment, the method may involve the step of annealing two (and optionally only two) oligonucleotides to a third oligonucleotide (e.g., an anchor strand). Two (and only two) oligonucleotides may be combined with the same third oligonucleotide. The method may involve the step of performing PCR on the annealed oligonucleotides to form a dsDNA molecule. The two oligonucleotides may have a variable sequence at the 3' and / or 5' end, a primer binding site at the relative 5' and / or 3' end, and a conserved flanking sequence between the variable sequence and the primer binding site.

[0083] One or more anchor chains in the method can have 3 or 4 or 5 or 6 or 7 or 8 or 3 to 5 or 3 to 6 or 3 to 7 or 3 to 8 or 4 to 5 or 4 to 6 or 4 to 7 or 4 to 8 or 6 to 10 or more than 8 or more than 10 or more than 12 degenerate nucleotides in its variable sequence.One or more degenerate nucleotides in oligonucleotide can exist as a continuous sequence to comprise degenerate sequence, or degenerate nucleotide can separate individually or with the form of the group of two or more continuous degenerate nucleotides in whole oligonucleotide (for example anchor chain).In certain embodiments, degenerate nucleotide is only present in the variable sequence of oligonucleotide, or is only present in the variable sequence of anchor chain.In one embodiment, 60% or less or 70% or less nucleotide is degenerate oligonucleotide in the variable sequence of anchor chain.

[0084] The degenerate oligonucleotide present at a certain position in an oligonucleotide library has multiple sequences at the position, and can be grouped together and considered as a member of a library. For example, an anchor oligonucleotide (or other oligonucleotide) with, for example, five degenerate nucleotides can have 1024 possible sequences (4x4x4x4x4=1024), but all 1024 sequences can be co-located at a single limited position in the library. The position containing multiple degenerate oligonucleotide sequences in an oligonucleotide library is referred to as a degenerate oligonucleotide position. Multiple degenerate oligonucleotides (each of slightly different sequences) can be co-located at a single position in an oligonucleotide library. Although in some embodiments, all possible sequences of degenerate oligonucleotides provide (for example, all 1024 possible sequences of degenerate oligonucleotides with 5 degenerate nucleotides) at the same position, in other embodiments, multiple degenerate oligonucleotides can be located at multiple different positions in an oligonucleotide library with a convenient number of groups. Therefore, degenerate nucleotides allow users to greatly reduce the number of positions in an oligonucleotide library. However, in some embodiments, a degenerate oligonucleotide at a position may contain a universal base, and therefore may have a "single" or less sequence at that position, even if that position is a degenerate oligonucleotide position. In some embodiments, because many degenerate nucleotides are universal nucleotides, the number of degenerate oligonucleotides at that position may be reduced.

[0085] Therefore, although the oligonucleotide with one or more degenerate nucleotides can be located together at the single limited position in the library, the oligonucleotide with the variable sequence that does not contain degenerate nucleotides can have its own limited position in the library separately, i.e. the individual position of each sequence. The oligonucleotide with one or more degenerate nucleotides can be located together at this single position with all possible sequences of the oligonucleotide at each degenerate position that is present in a single position. In one embodiment, only anchor chain has degenerate nucleotides, and " at least two oligonucleotides " do not have degenerate nucleotides.

[0086] For illustration, consider that the anchor chain O3 in Fig. 1 has ten variable nucleotides, including six degenerate nucleotide positions.Ten nucleotides of variable sequence usually need to exceed 1 million positions, but O3 has six degenerate nucleotides.Therefore, O3 with 4 non-degenerate nucleotides can be present in the library for L1 ... L256 position of O3 (4x4x4x4), and each position contains the specific sequence of the non-degenerate part of variable sequence.And 256 positions can have a group of oligonucleotides, which is degenerate nucleotide D1 ... D4096 provides all possible sequences (4x4x4x4x4x4x4 or 4096).Therefore, degenerate sequence D1-D4096 can all be present in each of the variable positions L1 to L256 for O3, and each position has the different sequences for the non-degenerate position on the sequence, and the oligonucleotides of all possible sequences at the degenerate position. Therefore, at the position L1 of example O3, there can be a variable sequence SEQ ID NO:3NNNACTCNNN (V1), which has the non-degenerate part of the variable sequence and the oligonucleotides of all possible sequences (D1-D4096) of the degenerate nucleotides in each position. At the position L2 for O3, the degenerate sequence D1 to D4096 will all have the non-degenerate sequence V2. At the position L3 for O3, the degenerate sequence D1 to D4096 will all have the non-degenerate sequence V3, and so on. Therefore, the degenerate sequence D1 to D4096 for O3 is all present at the position L1 to L256 for O3, wherein each degenerate sequence has the non-degenerate part of the variable sequence. Therefore, all anchor chains will contain the same sequence set part, but the sequences of all anchor chains at the degenerate nucleotides will be different. Therefore, in this example, for O3, the library can have 256 positions.

[0087] Oligonucleotide library

[0088] The present invention also provides a method for synthesizing product DNA molecules from an oligonucleotide member library according to the methods disclosed herein. The oligonucleotide member library can have less than 10,000 or less than 5,000 oligonucleotide members (or positions), and the oligonucleotide members in the library are sufficient to assemble any possible polynucleotide sequence. The method involves assembling oligonucleotide members from the library to obtain product DNA molecules.

[0089] Reference Figure 1A, all oligonucleotides O1-O6 are members in the library. O1 to O6 can each have a variable sequence. The oligonucleotide library can be included in any one or more of a DNA chip, a solid support, a solid phase, a bead, a microfluidic surface, a cell culture plate (e.g., 96 holes, 384 holes, or 1536 holes), etc., or wherein the oligonucleotide can be stored in a limited position and can be used for other structures retrieved and used in the method. In certain embodiments, the library will be directed to different positions of each of the possible variable sequences containing O1 to O6. In other embodiments, degenerate nucleotides are used on one or more polynucleotides.

[0090] When O1 to O2 and O4 to O5 have variable regions with 5 variable nucleotides, the number of positions that accommodate possible oligonucleotide sequences is 4 to the fifth power, so 4x4x4x4x4 equals 1,024. Thus, in some embodiments, there are defined oligonucleotide sequences at 4,096 defined positions, wherein at each of these positions there is a single or unique defined variable sequence for O1 to O2 and O4 to O5. Thus, the O1 oligonucleotide can have five variable nucleotides and therefore has 1024 possible sequences that can be present at 1024 defined positions for O1, each with a single defined variable sequence, and similarly for O2 and O4 to 5.

[0091] In this example, anchor chains O3 and O6 are added, and each anchor chain has four non-degenerate nucleotides and six degenerate nucleotides. Therefore, the library can also have 256 positions for each of O3 and O6 to accommodate oligonucleotides with non-degenerate nucleotides, wherein each position has different sequences for non-degenerate nucleotides. In addition, each of the 256 positions can also have all possible degenerate sequences, so for the nucleotide set of variable sequence, 4,096 degenerate oligonucleotide sequences are present together at each of the 256 positions. Therefore, this example provides a total of only 4,608 different positions (4x1024+2x256=4,608) in the entire library, from which all possible DNA sequences can be assembled. Even if the library size is doubled to accommodate parallel synthesis, only 9,216 members can be obtained.

[0092] A location in an oligonucleotide library can be a well of a plate, a tube, or any other structure or force that isolates an oligonucleotide member at a distinct location, sufficiently spatially separated from other members of the library to allow it to be accessed individually and as a species at that distinct location.

[0093] Oligonucleotide can be maintained at its different positions as a single molecule (which can be amplified) or as multiple copies of the same molecule (a small volume can be taken out from it and used for synthesis procedures). The different positions of each sequence are recognizable for a software program, which can be configured with a mechanical support or device in the method of the present invention for retrieving specific library members from different positions for oligonucleotide library members that need to be limited. In one embodiment, the oligonucleotide library can be located in a set of assay plates or tubules, each assay plate or tubule containing a member of the oligonucleotide library, and an instrument assembly can go to the assay plate or tubule and retrieve the oligonucleotide library member according to a software instruction, which can be located on a non-transient computer-readable medium. Non-transient computer-readable media can also contain programmed instructions and / or steps for synthesizing product DNA molecules according to any one of the methods disclosed herein, and programmed instructions and / or steps can be provided to an instrument communicating with a computer-readable medium. Programmed instructions or steps can guide the instrument to assemble the DNA molecules of a predefined sequence according to any method disclosed herein, or perform any one of the methods provided herein.

[0094] The member of oligonucleotide library is present in different positions, is separated in space from other members of library.Therefore, the member of library can be the specific sequence (single or multiple copies) present in its position.Non-degenerate oligonucleotide can have a sequence present in its library position.When using degenerate sequence, considering the quantity of degenerate nucleotides on oligonucleotide, the library member containing degenerate sequence can contain all possible degenerate sequences (or the subset of all possible sequences in some embodiments) of oligonucleotide member, and is present in different positions.Therefore, in the library position of non-degenerate nucleotide sequence, the member of library can contain a sequence.In the library position of degenerate nucleotide sequence, member can contain multiple sequences, in view of the degenerate nucleotides in oligonucleotide sequence, comprise the sequence in all possible sequences of oligonucleotide.

[0095] In certain embodiments, there may be many uninterested sequences in all possible sequences.Therefore, only the subset of all possible degenerate sequences needs to be present in the different positions in the definition category library to assemble all possible sequences of interest.In any embodiment, different positions can be limited by any suitable technology, for example, the reference point in the micrograph or the grid of the solid support containing the oligonucleotide library.In certain embodiments, the different positions of any or all oligonucleotide sequences can be stored on the non-transient computer readable medium and / or communicated by the non-transient computer readable medium.

[0096] DNA with overhangs

[0097] If desired, the product DNA molecules of any synthesis method disclosed herein can be assembled into larger product dsDNA molecules. In certain embodiments, the product dsDNA molecules of any one of the methods can be double-stranded blunt-end DNA. DNA molecules can be synthesized so that the variable sequences between the product dsDNA molecules contain overlapping sequences. The product dsDNA molecules can be digested with restriction endonucleases, which cut in the variable sequence and leave 3' and / or 5' overhang sequences or "sticky ends" in the resulting dsDNA fragments. These overhang sequences can overlap (and be complementary) with the nucleotides in the overhang sequence of another digested dsDNA fragment. These overhang sequences can then be used to assemble dsDNA fragments into larger DNA molecules by annealing of complementary 3' and / or 5' sequences. In other embodiments, product dsDNA molecules can be synthesized, which have single-stranded overhang sequences of one or more nucleotides, or 4 nucleotides, or 5 nucleotides, or 6 nucleotides, or 7 nucleotides, or 8 nucleotides, or more nucleotides, or 9 or 10 nucleotides, or more than 10 nucleotides, and are provided as additional dsDNA fragments. Overlapping or complementary nucleotides within these overhangs can then be used to anneal and ligate oligonucleotides or dsDNA fragments.

[0098] Restriction recognition site

[0099] IIS type restriction enzyme cuts DNA at a defined distance from its recognition site, and leaves 5' and / or 3' single-stranded overhangs. The recognition site can be provided as being located outside the variable sequence, and the cleavage site can be provided as being located within the variable sequence, thereby leaving 3' and / or 5' overhangs on the resulting dsDNA fragment. IIS type restriction endonucleases also find applications in the present invention for producing other dsDNA fragments with single-stranded overhangs. The single-stranded overhangs can be present at 3' and / or 5' ends, depending on the position of the dsDNA fragment relative to other fragments in the molecule. The dsDNA molecule can be programmed or synthesized to have an active recognition site on the 3' and / or 5' side of the dsDNA molecule and on one or both sides of the variable sequence. The dsDNA molecule can also be programmed to have a cleavage site within the variable sequence or towards 5' and / or 3' ends. The dsDNA fragments can be connected by annealing and connecting the dsDNA fragments with complementary protruding 3' and / or 5' sequences to form longer DNA molecules. Multiple additional dsDNA fragments with 3' and / or 5' overhangs (e.g. Figure 1B125 and 130 in the above step) can be annealed to dsDNA fragments with complementary 5' and / or 3' overhangs (e.g., O11 and O12). Thus, the dsDNA fragments of any step can be annealed to one or more additional dsDNA fragments to more rapidly increase the size of the variable sequence of the product dsDNA molecule. In this hierarchical manner, dsDNA molecules with variable sequences of about 100 bp or greater can be synthesized (e.g., Figure 1B and Figure 3 ).

[0100] In any embodiment, the restriction enzyme utilized in the present invention can be an IIS type restriction enzyme. In one embodiment, an IIS type restriction enzyme is an enzyme that only cuts dsDNA. In one embodiment, an IIS type restriction site can be encoded into a conserved flanking sequence, as shown in Figure 1. Any IIS type restriction enzyme can be utilized. In various embodiments, the restriction site can be a BsmBI site, or a BsmBi site, or an EciI site, or a BspMI site, or a FauI site, etc. BsmBI recognition sequence 5'-CGTCTC (N) -3' (SEQ ID NO: 16). The enzyme typically cuts the 3' side of N. BsaI is another IIS type restriction enzyme, which recognizes the sequence 5'-GGTCTC (N1) -3' (SEQ ID NO: 17), and typically cuts to the 3' side of N. With the aid of the present disclosure, a person of ordinary skill will recognize that many other IIS type restriction enzymes can be used in the present invention. These people can also encode recognition sites for specific restriction endonucleases in CFS so that they will cut within the variable sequence and provide protruding nucleotides. In any embodiment of the method disclosed herein, any one of the DNA molecules utilized in the method or produced by the method can contain one or more IIS type restriction endonuclease recognition sites. In various embodiments, the restriction enzyme can leave an overhang of at least 2bp, or at least 3bp, or at least 4bp.

[0101] In some embodiments of the method, the restriction recognition site on the dsDNA molecule can be "opened" or "closed" as needed. Although the dsDNA molecule can include a restriction recognition site (for example, on a conservative flanking sequence (CFS)), the CFS sequence can be changed in the method, thereby replacing the restriction recognition site with a sequence that is not recognized by the enzyme. This can be achieved by utilizing a primer with at least one base mismatch with the site on the CFS in the amplification step. During amplification, the use of mismatched primers leads to a change in the sequence during amplification, and thus the recognition site is inactivated. Therefore, at the time point when the site is no longer useful in the method, the sequence of the recognition site can be modified using a primer mismatched at the recognition site during amplification, thereby modifying the sequence in the amplified product and effectively "closing" the recognition site. Primers can be mismatched with a sequence on the recognition site or a sequence near the recognition site that covers at least a portion of the recognition site. Primers can be mismatched with the recognition site at least one nucleotide, or at one nucleotide, or at two nucleotides, or at three nucleotides.

[0102] DNA Data Storage

[0103] DNA is stable even after thousands of years and even in many extreme environments, which makes it have a huge advantage in storing information. Any of the methods disclosed herein can be applied to digital data encoding into DNA. One or more product DNA molecules can have a sequence containing the non-genetic information of coding. One or more product DNA molecules can have a sequence corresponding to the information bytes of the non-genetic information of coding. The bytes of information can be decoded with reference to the coding scheme or key that assigns one or more letters, words, characters or numbers to each coded byte of information. Non-genetic information can be, for example, the content of a word, a phrase, an identification watermark, text information, a book or a library, or any other information that can be provided with a reference language.

[0104] For example, Figure 5As described, it is possible to synthesize a DNA molecule with a 16bp variable sequence, and to easily accommodate four bytes of information on the variable sequence, wherein each byte is encoded by the nucleotide sequence of distribution. In this example, the sequence of four nucleotides represents an information byte, which can correspond to a character or symbol (such as a letter, a number or other symbol). Therefore, in this example, 256 characters (4x4x4x4) can be encoded in each information byte. Therefore, the alphabet of any language in the world can be easily accommodated in the information of these 256 bytes and a sufficient number of digits and other digits or other characters that are also used for communication. In various embodiments, non-genetic information can be encoded in a reference language, such as English, French, German, Italian, Spanish, Latin, Japanese, Indian, Chinese, Russian or any language. The reference language can also include numbers and special characters, even if they are not formal components of the reference language. But any information can be encoded in the DNA sequence in any language. In various embodiments, the information can be at least 100 characters long, or at least 500 characters long, or at least 1000 characters long, or at least 10,000 characters long. Nucleotides with non-standard bases can also be used, which can expand the number of characters available.

[0105] The product DNA may also encode characters (e.g., letters, words, numbers, punctuation marks, word characters, or other characters used for communication) that indicate where in the sequence the information encoded by the DNA molecule is to be placed. Figure 5Depict each 16bp product DNA molecule of four bytes with four nucleotides. The last byte in each product DNA sequence indicates the position of the first three bytes in the information; This is conveniently a number, but can be any character that can be placed in a qualifiable sequence. Although 4 nucleotide bytes provide up to 256 identifiers, the byte can be any convenient nucleotide length. For example, a byte can be composed of 3 nucleotides, or 5 nucleotides, or 6 nucleotides (allowing 4,096 identifiers), or 7 or even 8 nucleotides, or more than 8 nucleotides, thereby allowing more identifiers to be included. A limited number of identifiers can also be expanded by placing DNA molecules into a single hole until the number of identifiers, and then assembling the information from DNA in the order of the hole sequence. Using this method, with only 4 nucleotide bytes, even a single 384-well plate can contain more than 98,000 DNA molecules (256 molecules x384 holes), these molecules can be assembled to provide nearly 300,000 bytes of information (except identifiers), providing more than 153 million words in a single 384-well plate, or more than 550,000 pages of text (using standard pages and 512 words / kb). When five nucleotide bytes are used to exceed 1,024 molecules, 384 holes can be identified separately, i.e. 393,000 molecules in a single plate, or more than 1 million bytes of information. Multiple plates can be used to accommodate much more information. Therefore, according to the method, unlimited encoding and storage of unlimited information can be performed. Due to the high stability and small size of DNA, the entire information library can be encoded according to the method.

[0106] Thus, the present invention provides a method for storing data in a DNA sequence, which may involve determining a DNA sequence encoding non-genetic information according to a coding scheme that can translate the non-genetic information from a reference language into a DNA sequence, or vice versa; synthesizing a DNA sequence encoding non-genetic information according to the methods disclosed herein; and thereby storing the data in the DNA sequence. The method may optionally be repeated until the non-genetic information is recorded in the sequence.

[0107] An encoding scheme is a set of codes that assigns specific characters of a reference language to specific codons (e.g., 4- or 3-nucleotide codons, e.g., Figure 5). For example, the standard DNA codon table is a coding scheme, but it may be advantageous to use a coding scheme that is not easily transcribed. Examples of coding schemes are known to those of ordinary skill in the art, such as any of those disclosed in the following: U.S. Patent No. 10,818,378, which is hereby incorporated by reference in its entirety (including all tables, drawings, and claims); or Marillonnet et al., Nature Biotech., Vol. 21, pp. 224 to 226 (2003). However, many such coding schemes are known, and new coding schemes can be easily designed by those skilled in the art. After synthesizing dsDNA molecules according to any of the methods described herein, additional DNA joining techniques known in the art can be used to combine dsDNA molecules to construct larger dsDNA molecules that contain coding information and can be stored indefinitely. Therefore, in one embodiment, non-genetic information can be provided according to a coding scheme and translated into a DNA sequence that can be synthesized, thereby storing the data in the DNA sequence.

[0108] CRISPR guide RNA

[0109] The present invention can also be applied to the synthesis of guide RNA (gRNA) for CRISPR-Cas9 methods. Any sequence of gRNA can be quickly constructed using these methods. Guide RNA constructs can also be constructed from oligonucleotides in an oligonucleotide library. Product DNA molecules can be synthesized in a method with a DNA sequence encoding an initial guide structure. The initial guide RNA structure can encode a gRNA having prokaryotic or eukaryotic transcription elements necessary for in vitro transcription in an appropriate order, such as any one or more of a promoter, a gRNA sequence, and a terminator. In some embodiments, gRNA can encode Cas9 binding hairpins (Cas9 handles). In some embodiments, transcription elements include promoters and / or terminators. In some embodiments, product DNA molecules can encode 20 bases of gRNA. Figure 6 An embodiment is depicted in which a dsDNA molecule is synthesized as an initial guide structure with a transcription element. In any of the methods disclosed herein, the product dsDNA molecule can encode a guide structure or a gRNA or other RNA molecule. Since all possible polynucleotide sequences can be assembled from an oligonucleotide library, any initial guide structure or gRNA or RNA can be assembled in this method.

[0110] Example

[0111] In one embodiment, the method involves annealing at least two oligonucleotides of about 30 to 60 nucleotides in length to an anchor strand of about 30 to 70 nucleotides in length according to the methods disclosed herein.

[0112] In another embodiment, the method involves annealing at least two oligonucleotides of about 40 to 50 nucleotides in length to an anchor strand of about 40 to 50 nucleotides in length according to the methods disclosed herein.

[0113] In another embodiment, the method involves annealing at least two oligonucleotides of about or about 40 to 50 nucleotides in length to an anchor strand of about 40 to 60 nucleotides in length. In various embodiments, the anchor strand may utilize 4 to 6 or 6 degenerate oligonucleotides.

[0114] In another embodiment, the method involves annealing at least two oligonucleotides of about or about 40 to 50 nucleotides in length to an anchor strand of about 45 to 55 nucleotides in length. In various embodiments, the anchor strand may utilize 4 to 6 or 6 degenerate oligonucleotides.

[0115] In any of these embodiments, the method can produce dsDNA molecules having an error rate of less than one error per 5,300 base pairs.

[0116] Example 1 – Hierarchical synthesis

[0117] This example shows the synthesis of dsDNA molecules of desired sequence with 100 base pair variable regions in a hierarchical approach.

[0118] The "L0" ligation reaction includes two oligonucleotides O1 and O2 (45 nucleotides each), each having a variable sequence of 5 nucleotides, a conserved flanking sequence of about 20 nucleotides, and a primer binding site of about 20 nucleotides. Anchor chain O3 is programmed to have a variable sequence of 10 nucleotides and a length of 50 nucleotides. The oligonucleotides are selected so that the sequence generated by the synthesis (L0) of O1 to O3 will contain a 10 nucleotide variable sequence, which will be part of the 100 nucleotide variable sequence of the predetermined total dsDNA molecule; and will have a variable sequence of about 10 nucleotides. The oligonucleotides are also selected to encode a restriction site for BsaI (a type IIS nuclease) on the 5' side of the DNA molecule (for subsequent connection with a paired dsDNA molecule having an active recognition site on the 3' side of the DNA molecule).

[0119] Prepare a solution (2ul of 100pM pool) containing oligonucleotides O1 to O2 (two oligonucleotides) and O3 (anchor). Place the oligonucleotides in a well containing T4 DNA ligase buffer (0.5ul), water (2.4ul) and T4 DNA ligase (0.1ul). Incubate the solution at 16°C for 1 hour and then at 65°C for 10 minutes.

[0120] After the ligation step (L0) with T4 DNA ligase, water (2ul), tailed 5' and 3' primers targeting conserved primer binding sites (1ul, 1uM), high-fidelity thermostable DNA polymerase (5ul) were used. (New England Biolabs, Ipswich, MA) and the L0 reaction product were subjected to a PCR amplification step (PCR1). The PCR protocol was as follows: 98°C for 30 seconds, followed by 30 cycles of 98°C (10 seconds), 50°C (10 seconds), and 65°C (15 seconds). Enzymatic purification was performed by adding 2uL of a 10-fold diluted calf intestinal phosphatase (CIP) + exonuclease I ("CE") stock solution and incubating at 37°C for 10 minutes. A 10-fold diluted proteinase K (2uL) was added and then incubated at 37°C for 15 minutes, followed by incubation at 95°C for 10 minutes. 4% EXE- The purified 98 bp product was confirmed on a PCR gel (ThermoFisher Corp., Waltham, MA). The product had a variable sequence of 10 nucleotides.

[0121] Then the digestion and ligation step (DL1) is performed. Water (2.3ul), T4 ligation buffer (0.5ul), BsaI enzyme (0.1ul), T4 DNA ligase (0.1ul), and the PCR1 products are mixed together. Additional dsDNA fragments with variable sequence overhangs and 4bp overlaps with the variable sequence of the first dsDNA molecule are added by parallel PCR1 synthesis reactions. Additional dsDNA fragments can be derived from, for example, dsDNA molecules with recognition sites on the opposite sides of the dsDNA molecules. The mixture is incubated at 37°C for 1 minute, then at 16°C for 1 minute, and circulated 10 times. Finally, the mixture is kept at 80°C for 20 minutes. Then the DL1 product is subjected to a PCR step (PCR2) in a mixture of water (2ul), 5' and 3' primers (1uM) and DNA polymerase (5ul), and then diluted 150 times. PCR cycles and CIP+CE and proteinase K are performed as described above. The dsDNA molecules produced have variable sequences of 16 nucleotides.

[0122] Another digestion and ligation step (DL2) was performed using 2.3ul water, 10x T4 ligation buffer (0.5ul), BsaI (0.1ul), T4 DNA ligase (0.1ul) and 2ul PCR2 product. Additional dsDNA fragments with variable sequence overhangs and 4bp overlap with the first dsDNA molecule were added by parallel PCR2 synthesis reactions. The mixture was incubated at 37°C for 1 minute, then at 16°C for 1 minute, and cycled 10 times. Finally, the mixture was kept at 80°C for 20 minutes. The DL2 product was then subjected to a PCR step (PCR3) in a mixture of water (2ul), 5' and 3' primers (1uM), and the above DNA polymerase (5ul), and then diluted 150 times. PCR cycles and calf intestinal phosphatase (CE) and proteinase K digestions were performed as above. The resulting dsDNA molecules had a variable sequence of 28 nucleotides.

[0123] The digestion reaction is performed and the resulting dsDNA fragments are combined with dsDNA fragments from two additional parallel reactions, one of which is a reaction that produces two dsDNA fragments, both of which are variable sequences and are derived from the digestion of a dsDNA molecule with three restriction recognition sites, thereby generating two variable sequences without flanking sequences (e.g. Figure 1B 125, 130 in ). A third additional parallel reaction is performed to maintain conserved flanking sequences from the opposite (3') end to allow efficient ligation and enable universal primers to be used for downstream PCR (e.g., Figure 1B The amplified products were verified on gel, showing the presence of 88, 68, 68 and 88 bp products.

[0124] The ligation step was performed on the dsDNA fragment (DL3) using 16.5ul of water, 10x T4 ligation buffer (2.5ul), BsaI (0.5ul), T4 DNA ligase (0.5ul) and 5ul of the pooled PCR3 product. The mixture was incubated at 37°C for 1 minute, then at 16°C for 1 minute and cycled 25 times. Finally, the mixture was kept at 80°C for 20 minutes. The product was then subjected to a PCR step (PCR4) in water (6ul), 5' and 3' primers (2ul of 1uM), the above DNA polymerase (10ul) and 2ul of the digested and ligated product. PCR cycles as well as CE and proteinase K were performed as before. The amplified product was verified on a gel, showing the presence of a 180bp product. The molecule was sequenced and found to have the correct sequence, including an error-free 100 nucleotide variable sequence.

[0125] Example 2 - Oligonucleotide Library

[0126] This example shows the construction of a universal oligonucleotide library. Considerations when selecting a library include whether the flanking sequences act as robust universal priming sequences and ensuring that the 5' and 3' flanking sequences are different enough so that the PCR primer sequences do not cross-react during the PCR step. The common feature of all flanks is a type IIS site, and this is maintained in the flanking sequences and designed around it. These sequences are generated by computational design, but can also be generated manually.

[0127] Different flanking sequences were selected empirically by using approximately eight sequences and testing them directly in PCR. The best performing flanking sequence set based on the empirical data was then selected. The "flanking sets" were tested using 5' and 3' primer pairs, 5' only, and 3' only to ensure that the expected PCR product was generated.

[0128] After the flanking sequences are selected, variable sequences are added to the sequence. Note that all possible permutations of the variable bases are required to construct a library that can synthesize any possible DNA sequence. For example, if five variable bases are added to the 3' end of O1, there will be 4 to the power of 5 or 1,024 different O1 sequences in separate microtiter wells, where 4 is the number of available DNA bases and 5 is the number of variable bases used for the O1 oligonucleotide. These variable sequences are generated by available computational design programs, but can also be generated manually.

[0129] In the case of O1, five variable bases are added to the 3' end. In the case of O2, five variable bases are added to the 5' end. In the case of O3, a variable sequence containing four non-degenerate bases is added to the central portion of the oligonucleotide to support the connection of O1 and O2 at their adjacent interfaces, and then it is surrounded by degenerate N bases on each side because these bases prevent unnecessary expansion of the library. Degenerate N bases are synthesized on an oligonucleotide synthesizer by combining all four DNA bases at position N, so the O3 anchor oligonucleotide is a mixture of sequences. The O3 anchor oligonucleotide has a total of 6 N positions, and therefore in a single library well, there are a total of 4 to the power of 6 or 4,096 different molecules. Not all molecules in the library are viable O3 anchors for O1+O2 connection, but only a portion of the 4,096 molecules are needed to support a robust L0 connection.

[0130] The oligonucleotides that make up the library are then synthesized in a microtiter plate format so that all oligonucleotide members have discrete well positions within the library. The wells are in single microtube or 96-well and 384-well microtiter plate formats, but they can be any format that allows physical separation of the library oligonucleotide members. The position of each member is precisely known and can be accessed when the oligonucleotide components are brought together manually or by laboratory liquid handling automation.

[0131] When synthesizing a sequence (for example, a 100 bp sequence that is part of a specific gene), the following steps are followed:

[0132] Three oligonucleotides (O1, O2 and O3) are pooled into a single well and correspond to the first 10 bp (bases 1 to 10) of the 100 bp variable sequence to be synthesized in this example.

[0133] Then another three oligonucleotides (i.e., the next set of O1, O2 and O3) are pooled into adjacent wells. These oligonucleotides constitute another 10 bp variable sequence, but overlap 4 bp with the first oligonucleotide set above, thus constituting bases 6 to 14 of the 100 bp sequence in this example.

[0134] This process is repeated until there are enough starting pools to make the entire DNA molecule with a 100 bp variable sequence. In this example, there are 16 starting pools, and the sequence of each pool overlaps with the next by 4 bp. After all pools are established in the reaction wells, the synthesis process begins.

[0135] Table 1: This table shows the number of oligonucleotide members in the entire library set required to construct any DNA molecule with a variable sequence of 10→16→28→100 bp. The total number of library members required is 9,216.

[0136] Library ID O1 Assembly 1 O2 Assembly 1 O3 Assembly 1 O1 Assembly 2 O2 Assembly 2 O3 Assembly 2 total 1 1024 1024 256 1024 1024 256 4608 2 1024 1024 256 1024 1024 256 4608 Total --> 9216

[0137] Table 2: This table shows the nucleotide length of each of the oligonucleotide members in the library set. The length of non-degenerate nucleotides of the variable sequence is shown in parentheses.

[0138] Library ID O1 Assembly 1 O2 Assembly 1 O3 Assembly 1 O1 Assembly 2 O2 Assembly 2 O3 Assembly 2 1 45(5) 45(5) 50(4) 45(5) 45(5) 50(4) 2 45(5) 45(5) 50(4) 45(5) 45(5) 50(4)

[0139] Example 3 - Preparation for SARS CoV-2 Spike Protein Gene Assembly

[0140] This example shows the assembly of 72 dsDNA molecules with overlapping 100 base pair variable sequences in a hierarchical method for synthesizing approximately 4 kb of SARS-CoV-2 spike protein. The 100 bp variable sequence in each dsDNA molecule is a subsequence of the SARS-CoV-2 spike protein gene. The dsDNA molecules containing the subsequences were synthesized as in Example 1 to produce seventy-two 180 bp product sequences with 100 bp variable sequences overlapping by approximately 4 bp. In the PCR4 step, the dsDNA molecules were biotinylated using biotinylated primers and standard methods and then combined into a single pool.

[0141] DNA capture and 100 bp fragment release (using flank removal)

[0142] According to the manufacturer's instructions ( DNA microbeads (spherical particles with a silica core covered with a layer of paramagnetic material) were prepared and used by Biotechnology (Life Technologies, Oslo, Norway).

[0143] Resuspend the new microbeads in a vial and vortex for about 30 seconds or tilt and rotate for 5 minutes. Transfer the microbeads (50ul beads / sample) to a centrifuge tube containing the pooled PCR4 product. Add 1ml of 1x Bind and Wash (B&W) buffer to the tube and vortex the tube for 5 seconds. Place the tube on a magnet for 1 minute to bind the DNA, and discard the supernatant. Remove the tube from the magnet and resuspend the washed beads in at least 1ml of wash buffer, or resuspend in the microbeads of the initial volume taken from the initial vial. Repeat this wash twice in total. Resuspend the microbeads in 2x B&W buffer with 2x volume of the original beads (e.g., 100ul bead stock to 200ul 2x B&W).

[0144] Add an equal volume of biotinylated DNA (such as 50ul PCR4 pool + 50ul pre-washed beads) to dilute the NaCl concentration in 2x B&W buffer from 2M to 1M for optimal binding and immobilization. Incubate the sample at room temperature for about 30 minutes with gentle rotation. Then capture the beads on a magnet for 2 minutes and wash with 1x B&W buffer. Repeat this wash three times in total.

[0145] The captured beads were resuspended in 1x NEB3 buffer (1x BsmBI buffer, 10x r3.1 buffer diluted with ddH2O) in the same volume as the PCR pool used. 2ul of type II enzyme (BsmBI) was added per 50ul volume, the beads were resuspended and incubated at 55°C for 60 minutes, then cooled to room temperature. The beads were captured on a magnet for 2 to 3 minutes, and the liquid digest was transferred to a new tube or well containing the released 100bp fragment pool.

[0146] The PCR4 pool digest was used as a template for polymerase chain assembly (PCA) for 30 cycles to assemble the final dsDNA molecules. Although various cycles can be used, the PCA cycle parameters are as follows: 1. 98°C, 1 minute; 2. 98°C, 30 seconds; 3. 72°C, 30 seconds (increase of 15 seconds / cycle); 4. 65°C for 1 minute (increase of 15 seconds / cycle); 5. 60°C for 1 minute; 6. 55°C for 45 seconds, 7. cycle back to 2. perform 29 more cycles, and finally 8. 72°C for 5 minutes, and 9. store at 10°C.

[0147] Add 5ul of PCA reaction product to 20ul of PCR master mix containing primers matching the 5' and 3' ends of the spike protein gene sequence, and perform 30 cycles of PCR according to the following cycle parameters: 1. 98℃ for 1 minute; 2. 98℃ for 30 seconds, 3. 72℃ for 45 seconds (increase by 15 seconds / cycle); 4. 60℃ for 1 minute (increase by 15 seconds / cycle); 5. Cycle back to 2. Nine times; 6. Then 98℃ for 30 seconds; 7. 72℃ for 3 minutes; 8. 65℃ for 4 minutes; 9. Cycle back to step 6, perform nineteen more times; then 10. 72℃ for 5 minutes; 11. Store at 10℃.

[0148] like Figure 7 As shown, a 3,942 bp product was obtained. Before applying the enzymatic error correction step, the identity of the full-length gene was confirmed as the SARS-Cov2 spike protein gene by cloning and DNA sequencing analysis and was determined to have an error rate of approximately one error per 5,400 bp.

[0149] sequence

[0150] SEQ ID NO: 1, DNA, artificial sequence

[0151] AGGGA

[0152] SEQ ID NO: 2, DNA, artificial sequence

[0153] CGTTG

[0154] SEQ ID NO: 3, DNA, artificial sequence

[0155] NNNACTCNNN

[0156] SEQ ID NO: 4, DNA, artificial sequence

[0157] TTGCG

[0158] SEQ ID NO: 5, DNA, artificial sequence

[0159] TAGCG

[0160] SEQ ID NO: 6, DNA, artificial sequence

[0161] NNNTACGNNN

[0162] SEQ ID NO: 7, DNA, artificial sequence

[0163] AGGGAGTTGC

[0164] SEQ ID NO: 8, DNA, artificial sequence

[0165] TTGCGTAGCG

[0166] SEQ ID NO: 9, DNA, artificial sequence

[0167] AGGGAG

[0168] SEQ ID NO: 10, DNA, artificial sequence

[0169] TTGC

[0170] SEQ ID NO: 11, DNA, artificial sequence

[0171] GCAACTCCCT

[0172] SEQ ID NO: 12, DNA, artificial sequence

[0173] TTGCGTAGCG

[0174] SEQ ID NO: 13, DNA, artificial sequence

[0175] CGCTAC

[0176] SEQ ID NO: 14, DNA, artificial sequence

[0177] GCAA

[0178] SEQ ID NO: 15, DNA, artificial sequence

[0179] AGGGAGTTGCGTAGCG

[0180] SEQ ID NO: 16, DNA, BsmBI recognition site, Bacillus stearothermophilus

[0181] CGTCTC(N)

[0182] SEQ ID NO: 17, DNA, Bsal recognition site, Bacillus thermophilus

[0183] GGTCTC(N)

[0184] Although the present invention has been described with reference to presently preferred embodiments, it should be understood that various modifications can be made without departing from the spirit of the invention. Accordingly, the present invention is limited only by the appended claims.

Claims

1. A method for synthesizing a DNA molecule having a desired sequence, wherein include: a) annealing at least two oligonucleotides to an anchor strand such that the at least two oligonucleotides annealed to the anchor strand are adjacent to each other on the anchor strand; wherein each of the at least two oligonucleotides comprises a primer binding site at the 3' or 5' end, and a variable sequence at the opposite 5' or 3' end, and a conserved flanking sequence between the primer binding site and the variable sequence; and wherein the anchor strand comprises a conserved flanking sequence complementary to the conserved flanking sequences on the at least two oligonucleotides, and further comprises at least one variable sequence, wherein at least a portion of the at least one variable sequence on the anchor strand is complementary to at least a portion of the variable sequence on the at least two oligonucleotides; b) ligating the at least two oligonucleotides annealed to the anchor strand to produce a first dsDNA molecule; c) subjecting said first dsDNA molecule having a desired sequence and comprising a conserved flanking sequence within each of said 3' and 5' ends and a variable sequence within said conserved flanking sequence to an amplification step.

2. The method of claim 1, further comprising contacting the first dsDNA molecule with a restriction endonuclease to generate a first dsDNA fragment comprising a 3' and / or 5' overhang sequence comprising a portion of the variable sequence from the first dsDNA molecule, providing at least one first additional dsDNA fragment comprising a 3' and / or 5' overhang sequence that is at least partially complementary to an overhang sequence of at least one of the first dsDNA fragments; annealing the first dsDNA fragment and at least one first additional dsDNA fragment via the 3' and / or 5' overhang sequences; and The annealed dsDNA fragments are ligated to generate a second dsDNA molecule comprising a conserved flanking sequence within each of the 3' and 5' ends, and a variable sequence within the 3' and 5' conserved flanking sequences that is longer than the variable sequence on the first dsDNA molecule.

3. The method of claim 2, wherein the at least one first additional dsDNA fragment is the product of a parallel DNA synthesis reaction, further wherein the first dsDNA molecule has a recognition site for a restriction endonuclease on the 5' or 3' side of the molecule, and the first additional dsDNA fragment is derived from restriction cleavage of a dsDNA molecule having a recognition site for a restriction endonuclease on the opposite 3' or 5' side of the molecule.

4. The method of any one of claims 2 to 3, further comprising contacting at least one second dsDNA molecule with a restriction endonuclease to produce a plurality of second dsDNA fragments, the plurality of second dsDNA fragments comprising 3' and / or 5' overhang sequences and conserved flanking sequences within each of the 3' or 5' ends; providing at least one second additional dsDNA fragment comprising a 3' and / or 5' overhang sequence that is at least partially complementary to an overhang sequence of at least one of the second dsDNA fragments; annealing the plurality of second dsDNA fragments to the at least one second additional dsDNA fragment via the 3' and / or 5' overhang sequences; and The ligation step is performed to generate a third dsDNA molecule comprising conserved flanking sequences on the 3' and 5' ends, and a variable sequence within the conserved flanking sequences that is longer than the variable sequence of the second dsDNA molecule.

5. The method of claim 4, wherein the at least one second additional dsDNA fragment is the product of a parallel DNA synthesis reaction, further wherein the second dsDNA molecule has a recognition site for a restriction endonuclease on the 5' or 3' side of the molecule, and the second additional dsDNA fragment is derived from restriction cleavage of a dsDNA molecule having a recognition site for a restriction endonuclease on the opposite 3' or 5' side of the molecule.

6. The method according to any one of claims 4 to 5, further comprising reacting the at least one third dsDNA molecule with a restriction endonuclease to produce a plurality of third dsDNA fragments, the third dsDNA fragments comprising 3' and / or 5' overhang sequences and conserved flanking sequences within each of the 3' or 5' ends; providing at least one third additional dsDNA fragment comprising a 3' and / or 5' overhang sequence that is at least partially complementary to an overhang sequence of at least one of the third dsDNA fragments; annealing the plurality of third dsDNA fragments to the at least one third additional dsDNA fragment via the 3' and / or 5' overhang sequences; and A ligation step is performed to generate a fourth dsDNA molecule comprising conserved flanking sequences on the 3' and 5' ends, and a variable sequence within the conserved flanking sequences that is longer than the variable sequence of the third dsDNA molecule.

7. The method of claim 6, wherein the at least one third additional dsDNA fragment is the product of a parallel DNA synthesis reaction, further wherein the third dsDNA molecule has a recognition site for a restriction endonuclease on the 5' or 3' side of the molecule, and the first additional dsDNA fragment is derived from restriction cleavage of a dsDNA molecule having a recognition site for a restriction endonuclease on the opposite 3' or 5' side of the molecule.

8. The method according to claim 1, wherein step a) further comprises annealing at least two paired oligonucleotides to the paired anchor strands so that the at least two paired oligonucleotides bound to the paired anchor strands are adjacent to each other on the paired anchor strands, wherein the at least two paired oligonucleotides comprise a primer binding site on the 3' or 5' end, and a variable sequence on the opposite 5' or 3' end, and a conserved flanking sequence between the primer binding site and the variable sequence; and wherein the paired anchor strands comprise a conserved flanking sequence complementary to the conserved flanking sequence on the at least two paired oligonucleotides, and further comprise at least one variable sequence, and wherein a portion of the variable sequence on the paired anchor strands overlaps with a portion of the variable sequence on the first anchor strand, d) ligating the at least two paired oligonucleotides annealed to the anchor strand; e) performing an amplification step to generate a paired dsDNA molecule having a desired sequence and comprising primer binding sites at the 3' and 5' ends, a conserved flanking sequence within each of said 3' and / or 5' ends, and a variable sequence within said conserved flanking sequence that partially overlaps with said variable sequence of said first dsDNA molecule.

9. The method of claim 8, wherein the at least two oligonucleotides and the first anchor strand, and the at least two paired oligonucleotides and the paired anchor strand are annealed in simultaneous reactions in the same pool.

10. The method of claim 8, further comprising contacting the first dsDNA molecule and the paired dsDNA molecule with a restriction endonuclease to produce at least one dsDNA fragment and at least one paired dsDNA fragment, each fragment comprising at least one 3' and / or 5' overhang sequence; and wherein at least a portion of the 3' or 5' overhang sequence from the first dsDNA fragment is complementary to at least a portion of the 5' or 3' overhang sequence from the paired dsDNA fragment, The at least one first dsDNA fragment and the paired dsDNA fragment are annealed through their complementary overhang sequences and subjected to a ligation step to produce a second dsDNA molecule comprising a conserved flanking sequence within each of the 3' and 5' ends, and a variable sequence within the 3' and 5' conserved flanking sequences that is longer than the variable sequence on the corresponding first dsDNA molecule.

11. The method of claim 10, further comprising contacting the at least one second dsDNA molecule and the at least one paired second dsDNA molecule with a restriction endonuclease to produce a plurality of second dsDNA fragments and paired second dsDNA fragments, each fragment comprising a 3' and / or 5' overhang sequence, wherein at least two of the plurality comprise a conserved flanking sequence within each of the 3' or 5' ends; and wherein at least a portion of the 3' or 5' overhang sequence from the second dsDNA fragment is complementary to at least a portion of the 5' or 3' overhang sequence from the paired second dsDNA fragment, annealing the second dsDNA fragment and the paired second dsDNA fragment through their complementary overhang sequences; and The ligating step is performed to produce a third dsDNA molecule comprising a conserved flanking sequence within each of the 3' and 5' ends, and a variable sequence within the 3' and 5' conserved flanking sequences that is longer than the variable sequence on the second dsDNA molecule.

12. The method of claim 11, further comprising contacting the at least one third dsDNA molecule and the at least one paired third dsDNA molecule with a restriction endonuclease to produce a plurality of third dsDNA fragments and paired third dsDNA fragments, each fragment comprising a 3' and / or 5' overhang sequence, wherein at least two of the plurality comprise a conserved flanking sequence within the 3' or 5' end; and wherein at least a portion of the 3' or 5' overhang sequence from the third dsDNA fragment is complementary to at least a portion of the 5' or 3' overhang sequence from the paired third dsDNA fragment, annealing the third and paired third dsDNA fragments through their complementary overhang sequences; and A ligation step is performed to generate a fourth dsDNA molecule comprising a conserved flanking sequence within each of the 3' and 5' ends, and a variable sequence within the 3' and 5' conserved flanking sequences that is longer than the variable sequence on the third dsDNA molecule.

13. The method of any one of claims 1 to 12, wherein the first dsDNA molecule comprises a variable sequence of 8 to 12 base pairs.

14. The method of claims 8 to 12, wherein the paired dsDNA molecules comprise a variable sequence of 8 to 12 base pairs.

15. The method of any one of claims 2 to 7 or 10 to 12, wherein the second dsDNA molecule comprises a variable sequence of 14 to 18 base pairs.

16. The method of any one of claims 4 to 7 or 11 to 12, wherein the third dsDNA molecule comprises a variable sequence of 24 to 32 base pairs.

17. The method of any one of claims 6 to 7 or 12, wherein the fourth dsDNA molecule comprises a variable sequence of 90 to 110 base pairs.

18. The method according to any one of claims 1 to 12, wherein the at least two oligonucleotides have a variable sequence of 4 to 6 nucleotides.

19. The method according to any one of claims 1 to 12, wherein the amplification step is performed by polymerase chain reaction (PCR).

20. The method of any one of claims 1 to 12, wherein the variable sequence of the anchor strand is equal in length to the length of the variable sequences on the at least two oligonucleotides.

21. The method according to any one of claims 1 to 12, wherein the anchor strand comprises a variable sequence present between two sequences complementary to the conserved flanking sequences on the at least two oligonucleotides.

22. The method according to any one of claims 1 to 12, wherein the at least two oligonucleotides bound to the anchor strand are adjacent to each other on the anchor strand at their variable sequences.

23. The method according to any one of claims 1 to 12, wherein the portion of the variable sequence on the anchor strand that is complementary to the conserved flanking sequences on the at least two oligonucleotides comprises 2 to 6 nucleotides, or 14 to 18 nucleotides, or 26 to 30 nucleotides, or 90 to 110 nucleotides.

24. The method according to any one of claims 1 to 12, wherein the at least two oligonucleotides and the anchor further comprise a recognition site for a restriction endonuclease.

25. The method according to any one of claims 1 to 24, wherein the restriction endonuclease is a Type IIS endonuclease.

26. The method of any one of claims 1 to 12, wherein the ligating step occurs spontaneously.

27. The method of any one of claims 1 to 12, wherein the anchor strand comprises 4 to 6 degenerate nucleotides.

28. The method of claim 28, wherein the degenerate nucleotides comprise universal or randomized bases.

29. The method of any one of claims 1 to 29, wherein the DNA molecule of the desired sequence has an error rate of less than 1 base pair per 2,000 base pairs relative to the desired sequence.

30. The method of any one of claims 1 to 29, wherein the DNA molecule of the desired sequence has an error rate of less than 1 base pair per 14,000 base pairs relative to the desired sequence.

31. A method according to any one of claims 1 to 31, wherein DNA molecules of desired sequence are assembled from a library of less than 20,000 or 10,000 members.

32. The method of any one of claims 1 to 31, wherein the primer binding site comprises a universal primer binding site.

33. The method of any one of claims 1 to 32, wherein the product DNA molecule is at most 4,000 bp or at most 5,000 bp in length.

34. A composition comprising at least two oligonucleotides, each oligonucleotide comprising a primer binding site at the 3' or 5' end, and a variable sequence at the opposite 5' or 3' end, and a conserved flanking sequence between the primer binding site and the variable sequence; and wherein the anchor strand comprises a sequence complementary to the conserved flanking sequences on the at least two oligonucleotides, and Further comprising at least one variable sequence, wherein at least a portion of the at least one variable sequence on the anchor strand is complementary to at least a portion of the variable sequences on the at least two oligonucleotides.

35. The composition of claim 34, wherein the anchor comprises a sequence complementary to the conserved flanking sequences on the at least two oligonucleotides at their 3' and 5' ends.

36. A composition according to any one of claims 34 to 35, wherein the anchor comprises a variable sequence between two sequences complementary to the conserved flanking sequences.

37. The composition of any one of claims 34 to 36, wherein the primer binding site comprises a universal primer binding site.

38. A method for storing data in a DNA sequence, wherein include: determining the sequence of the DNA encoding the non-genetic information according to an encoding scheme that translates the non-genetic information from a reference language into a DNA sequence and vice versa; synthesizing the sequence of the DNA encoding the non-genetic information according to the method of any one of claims 1 to 33; as well as The data is thus stored in the DNA sequence.

39. The method of claim 37, wherein the DNA is encoded in bytes of 4 or 5 or 6 nucleotides.

40. A method for synthesizing a DNA sequence encoding a guide RNA, wherein include: determining the sequence of the DNA encoding the guide RNA; The sequence of the DNA encoding the guide RNA is synthesized according to any one of claims 1 to 33.

41. An oligonucleotide library comprising 1,536 different positions, the different positions comprising a. 1,024 positions comprising oligonucleotides having unique variable sequences of non-degenerate nucleotides; and b. an additional 512 different positions, each of the 512 positions comprising an anchor having a variable sequence comprising at least three non-degenerate nucleotides and at least four degenerate nucleotides.

42. The oligonucleotide library of claim 41, wherein the oligonucleotide comprises a primer binding site on the 3' or 5' end, and a variable sequence on the opposite 5' or 3' end, and a conserved flanking sequence between the primer binding site and the variable sequence; and The anchor strand comprises a conserved flanking sequence complementary to the conserved flanking sequences on the at least two oligonucleotides, and further comprises at least one variable sequence, wherein at least a portion of the at least one variable sequence on the anchor strand is complementary to at least a portion of the variable sequence on the at least two oligonucleotides.

43. The oligonucleotide library of claim 42, wherein the 512 different positions comprise anchor oligonucleotides having every possible sequence of the variable sequence of the anchor.

44. The oligonucleotide library of claim 43, wherein the variable sequence of the anchor strand comprises five or six nucleotides.

45. The oligonucleotide library of claim 41 comprising 4,608 different positions comprising a. 4,096 positions comprising oligonucleotides having unique variable sequences comprising non-degenerate nucleotides; b. an additional 512 different positions, each of at least the 512 positions comprising an anchor having a variable sequence comprising at least three non-degenerate nucleotides and at least five degenerate nucleotides.

46. ​​The oligonucleotide library of claim 45, wherein the oligonucleotide comprises a primer binding site on the 3' or 5' end, and a variable sequence on the opposite 5' or 3' end, and a conserved flanking sequence between the primer binding site and the variable sequence; and The anchor strand comprises a conserved flanking sequence complementary to the conserved flanking sequences on the at least two oligonucleotides, and further comprises at least one variable sequence, wherein at least a portion of the at least one variable sequence on the anchor strand is complementary to at least a portion of the variable sequence on the at least two oligonucleotides.

47. The oligonucleotide library of claim 46, wherein the 512 different positions comprise anchor oligonucleotides having every possible sequence of the variable sequence of the anchor.

48. The oligonucleotide library of claim 47, wherein the variable sequence of the anchor strand comprises five or six nucleotides.

Citation Information

Patent Citations

  • Encoding text into nucleic acid sequences

    US10818378B2