High-throughput long gene synthesis method and application thereof

By designing non-spaced and non-overlapping oligonucleotide fragments in a single reaction system and utilizing a ligase-mediated unwinding-annealing method, the complexity and high cost of existing gene synthesis methods are solved, enabling efficient and low-cost synthesis of multi-gene libraries.

CN121975792APending Publication Date: 2026-05-05WESTLAKE UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WESTLAKE UNIV
Filing Date
2025-04-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing gene synthesis methods suffer from problems such as complex operation, high cost, short synthesis length, high error rate, and difficulty in assembling complex systems. In particular, it is difficult to achieve high efficiency, low cost, and high accuracy in large-scale multi-gene library synthesis.

Method used

A three-step method was used to synthesize multiple genes in a single reaction system. By designing oligonucleotide fragments without gaps or overlaps, ligase-mediated unwinding-annealing was used for seamless splicing, reducing the number of steps and lowering the mutation and error rates.

Benefits of technology

It simplifies operations, reduces costs, improves the accuracy and efficiency of gene synthesis, reduces reagent consumption, and is suitable for high-throughput synthesis of multiple gene libraries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121975792A_ABST
    Figure CN121975792A_ABST
Patent Text Reader

Abstract

The invention discloses a high-throughput long gene synthesis method and application thereof. Methods for assembling (e.g., synthesizing) nucleic acid sequences of a variety of target genes, pools of oligonucleotide fragments, and uses thereof are provided. The method and the oligonucleotide fragment pool provided by the invention can be used for simultaneously assembling full-length nucleic acid sequences of various target genes in a single or multiple reaction systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of gene synthesis. Specifically, this invention relates to a method for synthesizing multiple genes in one or more reaction systems, including the design of oligonucleotide fragments and their applications. Background Technology

[0002] Gene synthesis is a widely used technique in life science research. Although the price of synthesizing 1kb of DNA using commercial gene synthesis methods has dropped to $100, large-scale DNA synthesis remains very expensive for many research and applications.

[0003] In 2010, Sriram Kosuri developed a method for high-throughput parallel gene synthesis using oligonucleotide libraries. This method involves splitting a single gene into overlapping fragments and assigning them to separate assembly sub-libraries. These fragments are then assembled into a full-length gene using a Golden Gate based on IIS restriction endonucleases (BtsI, BspQI, BsrDI, EarI, BsaI, BsmBI, SapI, BbsI), or by generating single-stranded DNA using DpnII / USER / λ exonucleases. Sriram Kosuri's gene synthesis method still involves single-gene, single-well synthesis, with each gene assigned to a single reaction for splicing. This method is complex and involves numerous steps for synthesizing gene libraries containing thousands of genes.

[0004] Sriram Kosuri's research group developed DropSynth, a method for multiplex gene synthesis in emulsions, in 2018 and 2020, along with its upgraded version, DropSynth 2.0. This method divides each gene into multiple oligonucleotides, each containing a corresponding IIS restriction enzyme site and a flanking sequence of a barcode tag. Different magnetic beads are used to adsorb oligonucleotides of different genes into different droplet compartments, with each droplet containing oligonucleotides of the same gene. Then, overlapping PCR using a high-fidelity polymerase is used to obtain the full-length gene library. The DropSynth method involves cumbersome droplet preparation and is only suitable for synthesizing genes <800 bp in length. Furthermore, the fidelity of high-fidelity polymerase assembly is only about 20%.

[0005] WO / 2023 / 096890 discloses a method for assembling oligonucleotides into polynucleotides based on a Zip sequence. This method divides each gene into n fragments, uses methods such as exonucleases to cut the Zip region into sticky ends, and then assembles the first and second fragments of multiple genes through ligation, circularization, and Zip region excision, followed by the assembly of the third, fourth, and so on, up to the nth fragment. However, when using the Zip sequence-based oligonucleotide assembly method to synthesize long genes, the steps of enzyme digestion and circularization need to be repeated, making the operation cumbersome and time-consuming. Furthermore, its accuracy is only 4% when synthesizing hundreds of 500 bp genes.

[0006] Currently used single-gene synthesis methods are mostly costly. A few large-scale multi-gene synthesis methods have been reported, but they still face challenges such as complex reaction systems, short synthesis lengths, high error rates, low yields, and difficulties in assembling complex systems. Therefore, there is a need to develop new high-throughput gene synthesis methods to provide simplified, more accurate, and lower-cost gene synthesis approaches. Summary of the Invention

[0007] The method of this invention enables the simultaneous synthesis of one or more genes in a single reaction system, reducing the reaction steps in multi-gene library synthesis and reducing reagent consumption. For example, the three-step gene synthesis method of this invention is simple to operate, involves fewer steps, requires no special equipment, and is easier to implement; it uses ligase-mediated unwinding-annealing for seamless splicing, minimizing mutations and errors caused by traditional overlapping PCR methods.

[0008] In some aspects, the present invention provides a method for synthesizing multiple target genes in a single reaction system, comprising: designing corresponding oligonucleotide fragments for each of the multiple target genes, wherein the full-length sequence of each target gene is sequentially split without gaps or overlaps to obtain sequences of multiple oligonucleotide fragments; synthesizing an oligonucleotide fragment pool based on the designed sequences of the multiple oligonucleotide fragments; phosphorylating the 5' end of the oligonucleotide fragments; and assembling the 5' phosphorylated oligonucleotide fragments into the full-length nucleic acid sequence of the target gene by ligase chain reaction (LCR). Optionally, the oligonucleotide fragments are designed with tag sequences at both ends, and the method includes removing the tag sequences before LCR.

[0009] In some implementations, the method for designing a corresponding oligonucleotide fragment pool for a target gene includes: sequentially splitting the full-length sequence without gaps or overlaps from two different splitting starting points using two splitting methods, resulting in two different sets of oligonucleotide fragments. The oligonucleotide fragments obtained from the same splitting starting point have no gaps or overlaps, while the oligonucleotide fragments obtained from different splitting methods have overlapping or complementary regions of 6-5000 nt (e.g., 10-2000 nt, 10-1000 nt, 1000-500 nt, 20-500 nt, 20-250 nt, 6-500 nt). Optionally, the two splitting methods include: (a) splitting the first strand of the target gene at a first splitting starting point and splitting the second strand at a second splitting starting point; or (b) splitting either the first or second strand of the target gene at a first splitting starting point and then splitting again at a second splitting starting point. Optionally, the target gene is sequence-optimized using a computer program before splitting, for example, through codon optimization.

[0010] In some embodiments, the oligonucleotide fragments of the multiple target genes constitute an oligonucleotide fragment pool, which may contain two or more oligonucleotide fragment sub-libraries; optionally, each sub-library contains oligonucleotide fragments from 1 to 1000 target genes, preferably from 1 to 100 target genes.

[0011] In some implementations, the method for designing corresponding oligonucleotide fragments for a target gene includes: for each target gene, designing the full-length sequence of the target gene to have a first pair of tag sequences at both ends. Optionally, the first pair of tag sequences includes an enzyme recognition site and / or a barcode sequence, wherein the enzyme recognition site is an enzyme recognition site of a restriction endonuclease, and the barcode sequence is gene type-specific, sub-library-specific, or gene-specific.

[0012] In some implementations, a method for designing a pool of oligonucleotide fragments for a target gene includes: each oligonucleotide fragment sequence is designed to have a second pair of tag sequences at both ends, optionally the second pair of tag sequences comprising a barcode sequence and one or more restriction enzyme recognition sites, the barcode sequence being subpool specific.

[0013] In some embodiments, the oligonucleotide fragment having the second tag sequence is amplified by polymerase chain reaction (PCR) using primers corresponding to the second tag sequence prior to phosphorylation of the 5' end of the oligonucleotide fragment. In some embodiments, the number of cycles of the polymerase chain reaction is 2 to 50, more preferably 8 to 20; preferably, the polymerase is a high-fidelity polymerase.

[0014] In some embodiments, the second pair of tag sequences is removed and phosphorylated at its 5' end by contacting the oligonucleotide fragment having the second pair of tag sequences with an enzyme capable of hydrolyzing phosphodiester bonds and generating a 5' phosphate group. In some embodiments, the enzyme is a restriction endonuclease, and the second pair of tag sequences contains the enzyme recognition site of the restriction endonuclease; preferably, the restriction endonuclease is a restriction endonuclease capable of cleaving double strands and generating blunt ends, more preferably, a MlyI enzyme. In some embodiments, the enzyme is a restriction endonuclease and an S1 nuclease, and the second pair of tag sequences contains the enzyme recognition site of the restriction endonuclease; preferably, the restriction endonuclease is an IIS type restriction endonuclease. In some embodiments, the enzyme is a different nicking enzyme, and the second pair of tag sequences contains the enzyme recognition sites of the different nicking enzymes. In some embodiments, the enzyme is a restriction endonuclease and a nicking enzyme, and the second pair of tag sequences contains the enzyme recognition sites of both the restriction endonuclease and the nicking enzyme; preferably, the restriction endonuclease is an IIS type restriction endonuclease.

[0015] In some embodiments, the primers corresponding to the second pair of tag sequences include a forward primer and a reverse primer, the forward primer comprising a 5' end thiophosphorylation modification and optionally a 3' uracil modification, the reverse primer comprising a reverse complementary sequence to the tag sequence, and the enzyme recognition site being selected from the enzyme recognition sites of restriction endonucleases, nicking enzymes, and USER enzymes, or combinations thereof; optionally, the reverse primer comprises a 5' phosphorylation modification.

[0016] In some embodiments, for reverse primers that do not contain 5' phosphorylation modification, the method further includes the step of preparing a 5' phosphorylated single-stranded oligonucleotide fragment after amplification, wherein: the amplified oligonucleotide fragment is contacted with a phosphokinase to phosphorylate one side of its 5' end, or the amplified oligonucleotide fragment is contacted with an enzyme corresponding to the enzyme recognition site; one side of the tag in the second tag sequence is removed and the 5' end of the second tag sequence is phosphorylated; the 5' phosphorylated oligonucleotide fragment is treated with an exonuclease (e.g., λ exonuclease) to obtain a single-stranded oligonucleotide fragment; and the second tag sequence is removed and the 5' end of the single-stranded oligonucleotide fragment is phosphorylated.

[0017] In some embodiments, for reverse primers containing 5' phosphorylation modification, the method further includes the step of preparing a 5' phosphorylated single-stranded oligonucleotide fragment after amplification, wherein: the oligonucleotide fragment is contacted with an exonuclease (e.g., λ exonuclease) to obtain a single-stranded oligonucleotide fragment, and then a second pair of tag sequences is removed and the 5' end of the single-stranded oligonucleotide fragment is phosphorylated.

[0018] In some embodiments, the second pair of tag sequences is removed by contacting the single-stranded oligonucleotide fragment with an enzyme corresponding to the enzyme recognition site (with primers added if necessary), wherein the primers contain the second pair of tag sequences or their reverse complementary sequences.

[0019] In some embodiments, the nicking enzyme is selected from Nt.BspQI, Nt.CviPII, Nt.BstNBI, Nb.BsrDI, Nb.BtsI, Nt.AlwI, Nb.BbvCI, Nt.BbvCI, Nb.BsmI, Nb.BssSI, Nt.BsmAI, and Cas9 nicking enzyme, etc.

[0020] In some implementations, the length of the first pair of tag sequences is 4 nt to 60 nt, for example, 10 nt to 60 nt, 10 nt to 50 nt, 20 nt to 50 nt, 20 nt to 40 nt, 30 nt to 50 nt, or 30 nt to 40 nt; preferably, the length of the first pair of tags is 10 nt to 30 nt.

[0021] In some implementations, the length of the second pair of tag sequences is 4 nt to 60 nt, for example, 10 nt to 60 nt, 10 nt to 50 nt, 20 nt to 50 nt, 20 nt to 40 nt, 30 nt to 50 nt, or 30 nt to 40 nt; preferably, the length of the second pair of tags is 10 nt to 30 nt.

[0022] In some embodiments, for oligonucleotide fragments that do not have a second pair of tag sequences at both ends, the 5' phosphorylation is performed by a phosphokinase-mediated reaction or a chemical phosphorylation method (such as phosphoramide or chemical modification reagents); preferably, the phosphokinase is a T4 polynucleotide kinase.

[0023] In some embodiments, the length of the oligonucleotide fragment is about 20 nt to about 10,000 nt, about 50 nt to about 5,000 nt, about 50 nt to about 2,500 nt, about 50 nt to about 2,000 nt, about 50 nt to about 1,500 nt, about 50 nt to about 1,000 nt, about 20 nt to about 500 nt, about 70 nt to about 500 nt, about 90 nt to about 500 nt, about 20 nt to about 400 nt, about 50 nt to about 400 nt, about 70 nt to about 400 nt, about 90 nt to about 400 nt, about 20 nt to about 300 nt, about 50 nt to about 300 nt, about 70 nt to about 300 nt, about 90 nt to about 300 nt, or about 200 nt to about 300 nt.

[0024] In some implementations, the distance between different splitting starting points is approximately 6 nt to approximately 5000 nt, for example, approximately 10 nt to approximately 2000 nt, approximately 10 nt to approximately 1500 nt, approximately 10 nt to approximately 1000 nt, approximately 10 nt to approximately 500 nt, approximately 30 nt to approximately 500 nt, approximately 30 nt to approximately 400 nt, approximately 30 nt to approximately 200 nt, approximately 40 nt to approximately 150 nt, approximately 40 nt to approximately 100 nt, approximately 50 nt to approximately 200 nt, approximately 60 nt to approximately 150 nt, approximately 70 nt to approximately 200 nt, approximately 70 nt to approximately 150 nt, approximately 80 nt to approximately 200 nt, approximately 80 nt to approximately 150 nt, approximately 90 nt to approximately 200 nt, approximately 90 nt to approximately 150 nt, approximately 100 nt to approximately 200 nt, or approximately 100 nt to approximately 150 nt.

[0025] In some embodiments, 5' phosphorylated oligonucleotide fragments are assembled into the full-length nucleic acid sequence of the target gene by ligase chain reaction (LCR); preferably, an oligonucleotide fragment pool or oligonucleotide fragment sublibrary assembles oligonucleotide fragments of multiple target genes into the corresponding full-length nucleic acid sequences of the target genes in a separate reaction system. In some embodiments, the LCR is performed at a fixed temperature or under high-low temperature cycling. Optionally, in the high-low temperature cycling, the high temperature is 90-98°C, for example 90°C, 91°C, 92°C, 93°C, 94°C, 95°C, 96°C, 97°C, or 98°C, and the low temperature is 37-75°C, for example 37°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, or 75°C. Optionally, the fixed temperature is 4-75°C, for example 4°C, 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C or 75°C.

[0026] In some implementations, the number of cycles of the LCR is 2-100 cycles, for example, 2 cycles, 3 cycles, 4 cycles, 5 cycles, 10 cycles, 15 cycles, 20 cycles, 25 cycles, 30 cycles, 35 cycles, 40 cycles, 45 cycles, 50 cycles, 55 cycles, 60 cycles, 65 cycles, 70 cycles, 75 cycles, 80 cycles, 85 cycles, 90 cycles, 95 cycles, or 100 cycles.

[0027] In some embodiments, the LCR is performed under high-temperature-medium-low-temperature cycling, and the ligase is a thermostable ligase; preferably, the ligase is Hi-T4. TM Thermostable DNA ligase, HiFi Taq DNA ligase, Taq DNA ligase or DNA ligase.

[0028] In some embodiments, the LCR is performed at a fixed temperature, and the ligase is selected from HiFiTaq DNA ligase, 9°NTM DNA ligase, Pfu DNA ligase, Taq DNA ligase, T4 DNA ligase, T3 DNA ligase, and T7 DNA ligase.

[0029] In some embodiments, the method further includes the step of amplifying the nucleic acid sequences of the multiple target genes by means of a first pair of tag sequences (e.g., by polymerase chain reaction) after assembly into a full-length nucleic acid sequence.

[0030] On the other hand, the present invention provides an oligonucleotide fragment pool for synthesizing multiple target genes, comprising corresponding oligonucleotide fragments designed and synthesized for each of the multiple target genes, wherein the sequences of the oligonucleotide fragments are obtained by sequentially splitting the full-length sequence of each target gene in two splitting methods without gaps or overlap. Optionally, the oligonucleotide fragment pool may comprise two or more oligonucleotide fragment sublibraries, wherein each oligonucleotide sublibrary contains oligonucleotide fragments of one or more target genes from the multiple target genes.

[0031] In some embodiments, the two splitting methods involve sequentially splitting the full-length sequence without gaps or overlaps from two different splitting starting points, resulting in two different sets of oligonucleotide fragments. The oligonucleotide fragments obtained from splitting at the same starting point have no gaps or overlaps, while the oligonucleotide fragments obtained from different splitting methods have overlapping or complementary regions of 6-5000 nt (e.g., 10-2000 nt, 10-1000 nt, 1000-500 nt, 20-500 nt, 20-250 nt, 6-500 nt). Optionally, the two splitting methods include: (a) splitting the first strand of the target gene at a first splitting starting point and splitting the second strand at a second splitting starting point; or (b) splitting either the first or second strand of the target gene at a first splitting starting point and then splitting again at a second splitting starting point. Optionally, the sequence of the target gene has undergone sequence optimization, for example, codon optimization.

[0032] In some implementations, for each target gene, the full-length sequence of the target gene has a first pair of tag sequences at both ends. Optionally, the first pair of tag sequences includes an enzyme recognition site and / or a barcode sequence, wherein the enzyme recognition site is an enzyme recognition site of a restriction endonuclease, and the barcode sequence is gene type-specific, sub-library-specific, or gene-specific.

[0033] In some embodiments, each oligonucleotide fragment sequence has a second pair of tag sequences at both ends. Optionally, the second pair of tag sequences comprises a barcode sequence and one or more enzyme recognition sites. The barcode sequence is sub-library specific. Optionally, the enzyme recognition sites are selected from the enzyme recognition sites of restriction endonucleases, nicking enzymes, and USER enzymes, or combinations thereof.

[0034] In some embodiments, the length of the oligonucleotide fragment is about 20 nt to about 10,000 nt, about 50 nt to about 5,000 nt, about 50 nt to about 2,500 nt, about 50 nt to about 2,000 nt, about 50 nt to about 1,500 nt, about 50 nt to about 1,000 nt, about 20 nt to about 500 nt, about 70 nt to about 500 nt, about 90 nt to about 500 nt, about 50 nt to about 400 nt, about 70 nt to about 400 nt, about 90 nt to about 400 nt, about 20 nt to about 300 nt, about 50 nt to about 300 nt, about 70 nt to about 300 nt, about 90 nt to about 300 nt, or about 200 nt to about 300 nt.

[0035] In some implementations, the distance between different splitting starting points is approximately 6 nt to approximately 5000 nt, for example, approximately 10 nt to approximately 2000 nt, approximately 10 nt to approximately 1500 nt, approximately 10 nt to approximately 1000 nt, approximately 10 nt to approximately 500 nt, approximately 30 nt to approximately 500 nt, approximately 30 nt to approximately 400 nt, approximately 30 nt to approximately 200 nt, approximately 40 nt to approximately 150 nt, approximately 40 nt to approximately 100 nt, approximately 50 nt to approximately 200 nt, approximately 60 nt to approximately 150 nt, approximately 70 nt to approximately 200 nt, approximately 70 nt to approximately 150 nt, approximately 80 nt to approximately 200 nt, approximately 80 nt to approximately 150 nt, approximately 90 nt to approximately 200 nt, approximately 90 nt to approximately 150 nt, approximately 100 nt to approximately 200 nt, or approximately 100 nt to approximately 150 nt. Attached Figure Description

[0036] Figure 1 This is a schematic diagram of the oligonucleotide fragment to be synthesized. BCO1 and BCO2 represent the tag sequences (containing barcodes) at both ends, and the black bars represent the oligonucleotide fragments obtained from gene splitting.

[0037] Figure 2This diagram illustrates gene splitting and LCR assembly. Each full-length gene has tag sequences (BCG1 and BCG2) at both ends of its two strands. As shown, each gene is designed for splitting in two ways, each starting from a different splitting point to divide the gene into multiple oligonucleotide fragments. The oligonucleotide fragments obtained from splitting at different starting points partially overlap or are complementary. The oligonucleotide fragments obtained from splitting multiple genes can form a sub-library, and these sub-libraries constitute an oligonucleotide fragment library.

[0038] Figure 3 An exemplary assembly flow of the gene synthesis method of the present invention is shown, including "generating an oligonucleotide fragment library", "PCR amplification to enrich the oligonucleotide fragments in each sub-library", "removing BCO1 and BCO2" and "assembling into a gene by LCR".

[0039] Figure 4 The results of 2% agarose gel electrophoresis show the extraction and enrichment of oligonucleotide fragments from different sub-libraries BCO1 and BCO2.

[0040] Figure 5 The results of the analysis of the assembled products using capillary electrophoresis are shown.

[0041] Figure 6 The gene coverage obtained by genome assembly using the method of the present invention is shown.

[0042] Figure 7 The results of 2% agarose gel electrophoresis (left) and 8% PAGE electrophoresis (right) of the products prepared by the method of the present invention are shown. Lanes A and 1: Markers; Lanes B and C: Fragments with unilateral 5' phosphorylation and unilateral 5' thiophosphorylation after PCR with thiophosphorylation primers and phosphorylation primers and then treated with phosphokinase; Lanes 2 and 4: Fragments with unilateral 5' phosphorylation and unilateral 5' thiophosphorylation after phosphokinase treatment; Lanes 3 and 5: Single-stranded fragments after digestion with λ exonuclease; Lanes 6-9: Enzyme digestion products after adding BCO and the corresponding enzyme. Detailed Implementation

[0043] In this invention, unless otherwise stated, all scientific and technical terms used herein have the same meaning as commonly understood by those skilled in the art. All patents, patent applications, and other publications cited herein are incorporated herein by reference in their entirety. Furthermore, the terms and laboratory procedures used herein, including those related to molecular biology, biochemistry, nucleic acid chemistry, organic chemistry, genomics, genetics, and synthetic biology, are widely used terms and routine procedures in their respective fields. If any definition presented herein conflicts with a definition presented in a patent, patent application, or other publication incorporated herein by reference, the definition presented herein shall prevail.

[0044] In this disclosure, the use of the singular includes the plural unless otherwise specified. Furthermore, unless otherwise specified, the use of “or” means “and / or”. Similarly, “comprising” and “including” are not intended to be limiting. The terms “about” or “approximately” refer to the acceptable error range of a particular value as determined by a person skilled in the art, depending in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, “about” may refer to within one or more standard deviations. Optionally, “about” may refer to a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value.

[0045] The term "gene," as used herein, refers to any naturally occurring or artificially synthesized DNA molecule, including coding regions, non-coding regions, and functional segments thereof, as well as combinations containing regulatory elements. A gene can be a DNA sequence with a specific function, which may include, but is not limited to: encoding proteins or polypeptides (such as structural genes); regulating gene expression (such as regulatory elements like promoters, enhancers, and terminators); and participating in the storage, replication, or transmission of genetic information. Genes encompass naturally occurring sequences or their variants (e.g., optimized gene sequences, such as codon optimization, GC content adjustment, etc.).

[0046] As used herein, “nucleic acid” means any natural or modified DNA or RNA molecule, including single-stranded, double-stranded or hybridized forms thereof, as well as analogs containing chemical modifications.

[0047] As used herein, the term "sequence" refers to the order of nucleotides in a gene or nucleic acid. Gene sequences are typically deoxyribonucleic acid (DNA) sequences and are usually double-stranded. Nucleic acid sequences can be DNA or ribonucleic acid (RNA) sequences, and can be single-stranded or double-stranded. Sequences can be mutated to differ from a reference sequence (e.g., a wild-type sequence). Any given gene sequence may contain sequence information of the given gene sequence and its inverse complementary sequence. In some cases, a DNA sequence may contain sequence information of the corresponding RNA sequence transcribed from that DNA. The sequence may be a letter representation of a polynucleotide. The sequence may be a piece of information that a computer processor can use.

[0048] As used in this article, the term "split" refers to dividing the sequence of a target gene into several oligonucleotide fragments without gaps or overlap, starting from a specific position on the sequence. These oligonucleotide fragments are adjacent to each other and are called a group of oligonucleotide fragments. "Sequential" means in the order from the 5' end to the 3' end or from the 3' end to the 5' end of the target gene sequence. Such splits are usually pre-designed by computer.

[0049] As used in this article, the term "split point" refers to the dividing point or position between two adjacent oligonucleotide fragments when splitting a target gene sequence. The positions of these split points are usually pre-designed by a computer. For example, the first split point might be located at the 100th nucleotide from the 5' end to the 3' end, the second split point at the 300th nucleotide from the 5' end to the 3' end, and so on. The intervals between split points can be the same or different. The position of the first split point is also called the split start point. Starting from the split start point, the sequence can continue to be split at one or more split points. The number of split points is usually determined by the sequence length of the target gene, the position of the split start point, and the split interval. For example, a gene with a sequence length of 900 bp, split at the 100th nucleotide from the 5' end to the 3' end as the first split point, using an equal split interval of 150 nt, would have 6 split points.

[0050] As used herein, the term “split interval” refers to the distance between adjacent split points, such as the distance between the first and second split points (e.g., 200 nt), which is equivalent to the length of the oligonucleotide fragment obtained from splitting from two adjacent split points.

[0051] As used in this article, "oligonucleotide fragment" refers to a sequence fragment of approximately 20-10,000 nucleotides in length obtained by splitting a target gene, which can be single-stranded or double-stranded. Oligonucleotide fragments also include fragments obtained by adding tag sequences to both ends of the split sequence fragment.

[0052] As used herein, the terms "oligonucleotide fragment pool," "oligonucleotide fragment library," "oligonucleotide pool," and "oligonucleotide library" are used interchangeably and refer to a combination of multiple oligonucleotide fragments used in a reaction system. These oligonucleotide fragments can be combinations of fragments synthesized based on gene split sequences (optionally, tagged sequences are added to both ends of the sequence). For the oligonucleotide fragments obtained after gene splitting, their sequences are identical or complementary to a portion of the sequence of the corresponding target gene, and the complementary regions are adjacent to each other. The oligonucleotide fragment library of the present invention contains oligonucleotide fragments obtained from the splitting of one or more target genes. In some embodiments, the oligonucleotide fragment library may contain two or more subsets, namely oligonucleotide sub-libraries (or simply sub-libraries), each containing several oligonucleotide fragments obtained from the splitting of target genes (e.g., genes from the same family).

[0053] In this article, when used in relation to oligonucleotide fragments, "adjacent" means that the sequences of two oligonucleotide fragments are complementary or substantially complementary to the sequences of two consecutive regions of the same template DNA.

[0054] As used herein, the terms "flat-ended" and "short-ended" refer to the ends of a double-stranded nucleic acid molecule, wherein substantially all nucleotides at the end of one strand of the nucleic acid molecule are paired with corresponding nucleotide bases in the other strand of the same nucleic acid molecule. If the end of a nucleic acid molecule includes a single-stranded portion of at least one nucleotide length, the nucleic acid molecule is not flat-ended and is referred to herein as a "protruding end" or "sticky end." In some embodiments, when nicking enzymes and / or restriction endonucleases with different enzyme recognition sites are used in combination with a double-stranded nucleic acid molecule, they can cleave the double or single strand at different locations in the nucleic acid sequence and produce sticky ends of a specific length.

[0055] The terms "barcode sequence" or "barcode" are used interchangeably herein and refer to a DNA sequence designed to be detected and identified. Barcode sequences can have a length of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50 or more nucleotides. Barcode sequences should also possess at least the following characteristics: (a) balanced GC content, preferably 35%-65%, more preferably 40%-60%; (b) no consecutive repeating bases or fewer than 5 consecutive repeating bases; (c) avoidance of homology with the target sequence or primers, with a similarity to the target gene sequence of less than 80%, preferably less than 50%; and (d) avoidance of secondary structure formation. Barcode sequences designed for methods such as polymerase chain reaction, single-cell sequencing, etc., are suitable for use in the methods of this invention, provided that these sequences achieve the objectives of this invention. The methods for obtaining the barcode sequence are well known to those skilled in the art. For example, barcode sequences can be designed and generated using the Python toolkit BioPython, the R toolkit DNABarcodeCompatibility, online tools (such as Barcode Generator), or software such as Barcoding Toolbox, DNABarcoder, Primer3, etc., or sequences can be obtained using existing barcode libraries (such as Illumina TruSeq Index, NexteraIndex).

[0056] As used herein, the term "tag sequence" refers to a nucleic acid sequence of 4-60 nt in length located at both ends of the full-length sequence of a target gene or at both ends of an oligonucleotide fragment sequence, which may be designed to contain a barcode. Preferably, the tag sequence is 10-30 nt in length. The tag sequences mentioned in this invention are typically used in pairs, consisting of a 5' tag sequence and a 3' tag sequence. In some embodiments, the tag sequences used in this invention contain enzyme recognition sites, such as enzyme recognition sites for restriction endonucleases and / or enzyme recognition sites for nicking enzymes and / or USER enzyme recognition sites, etc.

[0057] As used herein, the term "polymerase" refers to an enzyme that performs template-guided polynucleotide synthesis. DNA polymerases add a free nucleotide only to the 3' end of the newly formed strand. This results in the newly formed strand elongating in the 5' to 3' direction. DNA polymerases can only add nucleotides to a 3'-OH group that is already present; therefore, they require a primer to add the first nucleotide. Non-restrictive examples of polymerases include prokaryotic DNA polymerases (e.g., PolI, PolII, PolIII, PolIV, and PolV), eukaryotic DNA polymerases, archaeal DNA polymerases, telomerases, reverse transcriptases, and RNA polymerases.

[0058] As used herein, the term "ligase chain reaction (LCR)" has the meaning commonly understood by those skilled in the art, referring to a reaction in which, in one cycle, denaturation and melting at high temperature, followed by annealing, hybridize two oligonucleotide fragments in a sequence-complementary manner to a single-stranded template DNA, forming a nicked double-stranded fragment, wherein the 5' end of the oligonucleotide contains a phosphate group, and then, under appropriate conditions and in the action of a ligase, the two oligonucleotide fragments are joined to form a complete nucleotide chain. The reaction involving multiple cycles of denaturation, annealing, and ligation to assemble oligonucleotide fragments into a double-stranded nucleic acid molecule is called a ligase chain reaction. For example, as... Figure 2 As shown, two adjacent oligonucleotide fragments from the same cleavage origin can partially overlap or complement another oligonucleotide fragment from a different cleavage origin, forming a gap between the two adjacent oligonucleotide fragments. After 5' phosphorylation, a phosphodiester bond is formed at the gap under the action of a ligase, thereby connecting the adjacent oligonucleotide fragments. Ligases that can be used for LCR can be thermostable ligases, including but not limited to HiFiTaq DNA ligase, 9°N DNA ligase, etc. DNA ligase, Hi-T4 TM Thermostable DNA ligases, Taq DNA ligases; or other ligases that can catalyze the formation of phosphodiester bonds between adjacent 5' phosphate ends and 3' hydroxyl ends on double-stranded DNA, such as T4 DNA ligase, T7 DNA ligase, E. coli DNA ligase, and T3 DNA ligase.

[0059] The term "linkage" refers to the formation of a covalent bond or linkage between the ends of two or more nucleic acids (i.e., oligonucleotides and / or polynucleotides) in a template-driven reaction. The nature of such bonds and linkages can vary considerably, and linkages can be achieved through enzymatic or chemical processes. As used herein, linkages are typically achieved enzymatically to form a phosphodiester bond between the 5' carbon of the terminal nucleotide of one oligonucleotide and the 3' carbon of another oligonucleotide.

[0060] Sequence identity, as a percentage (%) relative to a reference nucleic acid sequence (or peptide sequence), refers to the percentage of nucleotides (or amino acid residues) in the candidate sequence that are identical to those in the reference nucleic acid sequence after alignment and the introduction of gaps (if necessary) to achieve maximum sequence identity, and without considering any conserved substitutions as part of sequence identity. Alignment used to determine the percentage of sequence identity can be performed in various ways within the scope of those skilled in the art, for example, using publicly available computer software such as BLAST, BLAST-2, CLUSTALW, ALIGN, or Megalign (DNASTAR) software. Those skilled in the art can determine appropriate parameters for aligning sequences, including any algorithms required to achieve maximum alignment across the full length of the sequences being compared.

[0061] The term "substantially identical" applied to nucleic acid or amino acid sequences means that, using the procedures described above (e.g., BLAST) and standard parameters, the nucleic acid or amino acid sequence includes sequences that have at least 90% or more, at least 95%, at least 98%, or at least 99% sequence identity compared to a reference sequence. For example, the BLASTN procedure (for nucleotide sequences) uses a word length (W) of 11, an expected value (E) of 10, M=5, N=-4, and a comparison of the two strands as default values ​​(see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1992)).

[0062] As used herein, the term "complementary" refers to hybridization or base pairing or double-stranded fragment formation between nucleotides or nucleic acids, such as between the two strands of a double-stranded DNA molecule, or between primer binding sites on an oligonucleotide primer and a single-stranded nucleic acid. Typically, complementary nucleotides are A and T (or A and U) or C and G.

[0063] As used herein, the terms “assembly,” “synthesis,” “assembly process,” or “synthetic process” refer to a reaction or series of reactions in which product components of two or more polynucleotide molecules (e.g., oligonucleotide fragments of the present invention) are linked to form a continuous (e.g., replicated by DNA or RNA polymerase) and longer product component (e.g., the full-length gene of the present invention). Each individual reaction used to complete the assembly process may be referred to as an assembly reaction. The assembly process in the present invention includes ligation reactions, such as ligase-mediated ligation reactions. The assembly process may further include amplification reactions.

[0064] As used herein, a target gene refers to a desired assembly product having a known nucleic acid sequence (i.e., a target nucleic acid sequence). The known nucleic acid sequence can be obtained using sequence design strategies known in the art or from existing genomic databases. The known nucleic acid sequence can be a sequence that has undergone further sequence optimization (e.g., codon optimization). The target gene can also be a nucleic acid sequence designed based on a known amino acid sequence, which can be achieved using conventional techniques in the art.

[0065] As used herein, the term "primer" in relation to polymerase chain reaction includes both forward and reverse primers. The terms "upstream primer" and "forward primer" are used interchangeably herein and refer to oligonucleotide molecules designed to be complementary to the sense strand (5' to 3' end) of a double-stranded target DNA sequence, thereby guiding DNA polymerase to extend in the 3' to 5' direction. The terms "downstream primer" and "reverse primer" are used interchangeably herein and refer to oligonucleotide molecules designed to be complementary to the antisense strand (3' to 5' end) of a double-stranded target DNA sequence, thereby guiding DNA polymerase to extend in the 5' to 3' direction. In some embodiments, the primers used in this invention are designed for tag sequences at both ends of the oligonucleotide fragment. In some embodiments, the primers used in this invention are designed with a 5' phosphorylation modification for the forward primer and a 3' uracil modification and / or a 5' thiophosphorylation modification for the reverse primer.

[0066] Methods for synthesizing multiple genes

[0067] In some aspects, the present invention provides a method for synthesizing multiple target genes. The multiple target genes can be synthesized in parallel in one or more reaction systems. Specifically, the method includes:

[0068] (a) For each of the multiple target genes, a corresponding oligonucleotide fragment is designed, wherein the full-length sequence of each target gene is sequentially split without gaps or overlaps to obtain the sequences of multiple oligonucleotide fragments;

[0069] (b) Based on the sequences of multiple oligonucleotide fragments, an oligonucleotide fragment pool is synthesized, and the synthesized oligonucleotide fragments can be single-stranded or double-stranded;

[0070] (c) Optionally, the desired oligonucleotide fragment is amplified by PCR using primers corresponding to specific tag sequences at both ends of the oligonucleotide fragment;

[0071] (d) Phosphorylate the 5' end of the oligonucleotide fragment (if the oligonucleotide fragment has tag sequences at both ends, the tag sequences can be removed simultaneously with 5' phosphorylation); and

[0072] (e) The 5' phosphorylated oligonucleotide fragments are assembled into a full-length nucleic acid sequence by a ligation reaction to obtain one or more desired target genes.

[0073] In some implementations, the full-length sequence of the target gene is designed with tag sequences (referred to as the first pair of tag sequences) at both ends. Further, each oligonucleotide fragment that is split can also be designed with tag sequences at both ends (referred to as the second pair of tag sequences). Alternatively, tag sequences may not be designed at the ends of the oligonucleotide fragments.

[0074] In some implementations, the method for designing corresponding oligonucleotide fragments for a target gene includes: for each target gene, sequentially splitting its full-length sequence without gaps or overlaps from two different splitting origins to obtain two different sets of oligonucleotide fragments. There are no gaps or overlapping regions between oligonucleotide fragments split from the same splitting origin, and the two oligonucleotide fragments from different splitting origins may have overlapping or complementary regions of 6-5000 nt (see [link to implementation details]). Figure 2 ).

[0075] When splitting a full-length gene sequence, the first strand can be split according to the first splitting point, and the second strand according to the second splitting point. Alternatively, either strand of the double helix can be split using either of the two splitting methods.

[0076] Preferably, referring to the target gene sequence, each splitting point under the first splitting start method corresponds to the interior of the oligonucleotide fragment sequence or its reverse complementary sequence under the second splitting start method, and preferably is at least 30 nt away from both ends of the oligonucleotide fragment sequence or its reverse complementary sequence under the second splitting start method. For example, for a target gene with a nucleic acid sequence of 1000 bp, the first splitting start is located at the 100th nucleotide position in the direction from the 5' end to the 3' end of the target gene sequence. Assuming that splitting is performed at 200 nt intervals, oligonucleotide fragments are obtained, respectively composed of nucleotides 1-100, 101-300, 301-500, 501-700, 701-900, and 901-1000 of the target gene sequence. The second splitting start is located at the target gene sequence... At the 150th nucleotide position from the 5' to the 3' end of the sequence, assuming a 200nt split interval, oligonucleotide fragments are obtained, consisting of nucleotides 1-150, 151-350, 351-550, 551-750, 751-950, and 951-1000 of the target gene sequence, respectively. The oligonucleotide fragments obtained from the first splitting point and the second splitting point have an overlap of 150nt and / or 50nt. Using two staggered splitting methods ensures that the splitting point of the first splitting method is not split in the second splitting method, and vice versa. This allows the oligonucleotide fragments obtained from the first and second splitting methods, or their reverse complementary sequences, to serve as templates for correct matching and ligation of adjacent oligonucleotide fragments in subsequent ligation reactions. It should be understood that the length of the oligonucleotide fragments does not have to be the same; that is, different splitting intervals can be used when sequentially splitting the target gene sequence, resulting in oligonucleotide fragments of different lengths.

[0077] Preferably, the method of the present invention for designing corresponding oligonucleotide fragments for target genes is implemented by a computer program, for example, by a script written based on programming / analysis software (such as R, Python, etc.). Under the guidance of the method of the present invention for designing oligonucleotide fragment pools, oligonucleotide fragments designed for target genes can be easily obtained by computer programs.

[0078] In some implementations, the target gene is sequentially split into two or more oligonucleotide fragments at the same splitting interval under each splitting method. The oligonucleotide fragments obtained by the same splitting method are adjacent to each other, while the oligonucleotide fragments obtained by the two splitting methods have an overlap or complementary region of 6-5000 nt.

[0079] Optionally, the target gene is sequence optimized prior to splitting. This sequence optimization can be performed using computer programs employing methods conventional in the art, such as codon optimization, GC content adjustment, removal of restriction enzyme sites, removal of repetitive sequences, and epigenetic optimization.

[0080] In some embodiments, each oligonucleotide fragment sequence is designed to also have tag sequences at both ends (referred to as a second pair of tag sequences). Optionally, the second pair of tag sequences contains enzyme recognition sites, such as enzyme recognition sites of restriction endonucleases and / or nicking enzymes and / or USER enzymes.

[0081] In some implementations, with the same splitting start point, the interval between each splitting point (i.e., the length of the resulting oligonucleotide fragment sequence) is approximately 20 nt to approximately 10,000 nt. Preferably, the splitting interval is 20 nt to approximately 5,000 nt, approximately 50 nt to approximately 5,000 nt, approximately 50 nt to approximately 2,500 nt, approximately 50 nt to approximately 2,000 nt, approximately 50 nt to approximately 1,500 nt, approximately 50 nt to approximately 1,000 nt, 20 nt to approximately 500 nt, approximately 50 nt to approximately 500 nt, approximately 60 nt to approximately 500 nt, approximately 70 nt to approximately 500 nt, approximately 80 nt to approximately 500 nt, approximately 90 nt to approximately 500 nt, approximately 100 nt to approximately 500 nt, 20 nt to approximately 500 nt. Approximately 400nt, approximately 50nt to approximately 400nt, approximately 60nt to approximately 400nt, approximately 70nt to approximately 400nt, approximately 80nt to approximately 400nt, approximately 90nt to approximately 400nt, approximately 100nt to approximately 400nt, approximately 20nt to approximately 300nt, approximately 50nt to approximately 300nt, approximately 60nt to approximately 300nt, approximately 70nt to approximately 300nt, approximately 80nt to approximately 300nt, approximately 90nt to approximately 300nt, approximately 100nt to approximately 300nt, or approximately 200nt to approximately 300nt.

[0082] In some implementations, the distance between different splitting start points is approximately 6 nt to approximately 5000 nt, for example, approximately 10 nt to approximately 2000 nt, approximately 10 nt to approximately 1500 nt, approximately 10 nt to approximately 1000 nt, approximately 10 nt to approximately 500 nt, approximately 10 nt to approximately 300 nt, approximately 10 nt to approximately 240 nt, approximately 10 nt to approximately 230 nt, approximately 20 nt to approximately 500 nt, approximately 20 nt to approximately 220 nt, approximately 20 nt to approximately 210 nt, approximately 20 nt to approximately 200 nt, approximately 20 nt to approximately 150 nt, approximately 20 nt to approximately 100 nt, approximately 20 nt to approximately 90 nt, approximately 20 nt to approximately 50 nt, approximately 30 nt to approximately 200 nt, approximately 30 nt to approximately 150 nt, approximately 30 nt to approximately 100 nt, approximately 30 nt to approximately 90 nt. Approximately 30nt to approximately 50nt, approximately 40nt to approximately 200nt, approximately 40nt to approximately 150nt, approximately 40nt to approximately 100nt, approximately 40nt to approximately 90nt, approximately 40nt to approximately 50nt, approximately 50nt to approximately 200nt, approximately 50nt to approximately 150nt, approximately 50nt to approximately 100nt, approximately 60nt to approximately 200nt, approximately 60nt to approximately 150nt, approximately 60nt to approximately 100nt, approximately 70nt to approximately 200nt, approximately 70nt to approximately 150nt, approximately 70nt to approximately 100nt, approximately 80nt to approximately 200nt, approximately 80nt to approximately 150nt, approximately 80nt to approximately 100nt, approximately 90nt to approximately 200nt, approximately 90nt to approximately 150nt, approximately 100nt to approximately 200nt, or approximately 100nt to approximately 150nt.

[0083] After designing oligonucleotide fragments for each target gene, the resulting fragments are synthesized for the next ligation reaction. These short oligonucleotide fragments can be synthesized by commercial gene synthesis companies or in laboratories, and can be single-stranded or double-stranded.

[0084] In some embodiments, 5' phosphorylated oligonucleotide fragments are assembled into the full-length nucleic acid sequence of the corresponding target gene using a ligase chain reaction (LCR). In some embodiments, the oligonucleotide fragments obtained from the dissection of the various target genes constitute an oligonucleotide fragment library for parallel synthesis of the full-length nucleic acid sequence of the target gene in a single reaction system. In some embodiments, the oligonucleotide fragments obtained from the dissection of the various target genes constitute multiple oligonucleotide fragment sublibraries, wherein each sublibrary consists of 2 to 1000 oligonucleotide fragments obtained from the dissection of the target genes, and each sublibrary is used for parallel synthesis of the full-length nucleic acid sequence of the corresponding target gene in a single reaction system.

[0085] In some embodiments, the method of the present invention for synthesizing multiple target genes further includes the step of amplifying the nucleic acid sequences of the multiple target genes by means of a first pair of tag sequences (e.g., by polymerase chain reaction) after assembly into nucleic acid sequences.

[0086] In some embodiments, the plurality of target genes synthesized in a single reaction system by the method of the present invention may be 1 to 1000 genes or more, for example 1-200, 200-400, 400-600, 600-800, 800-1000 or more genes.

[0087] In some implementations, the number of oligonucleotide fragments in an oligonucleotide fragment library or sublibrary used to synthesize multiple target genes in a single reaction system can be from 2 to 10,000 (inclusive).

[0088] Design of label sequences

[0089] In some embodiments, the method of the present invention includes, when designing the gene sequence splits and oligonucleotide fragment sequences, further designing tag sequences to be attached to both ends of the full-length gene sequence and to both ends of the split oligonucleotide fragments. In some embodiments, the design includes attaching a first pair of tag sequences (which may be gene type-specific, oligonucleotide fragment sub-library-specific, or gene-specific to distinguish different gene classifications or sub-libraries or single genes) to both ends of the full-length gene sequence and a second pair of tag sequences (preferably sub-library-specific) to both ends of the split oligonucleotide fragments. As used herein, with respect to tag sequences, "gene type-specific" means that genes of the same type are designed to have the same tag sequence, "oligonucleotide fragment sub-library-specific" means that oligonucleotide fragments in the same sub-library have the same tag sequence, and "gene-specific" means that each of a variety of target genes has a different tag sequence.

[0090] The tag sequence can be designed as a pair of sequences containing a gene, sub-library, or gene type-specific barcode. Optionally, the tag sequence of the present invention is also designed to contain an enzyme recognition site. The enzyme recognition site can be designed in or outside the barcode. The design of gene, sub-library, or gene type-specific barcode sequences is well known in the art. Typically, the barcode sequence in the tag sequence has the following characteristics: (a) balanced GC content, preferably 35%-65% GC content; (b) preferably no more than 5 consecutive repeating bases; (c) avoidance of homology with the target gene sequence, with a similarity of less than 80%, preferably less than 50%; and (d) avoidance of secondary structure formation. The methods for obtaining the barcode sequence are well known to those skilled in the art. For example, the barcode sequence can be generated using the Python toolkit BioPython, the R toolkit DNABarcodeCompatibility, online tools (e.g., Barcode Generator), or software such as BarcodingToolbox, DNABarcoder, Primer3, etc. Alternatively, suitable sequences can be selected from existing barcode libraries (e.g., Illumina TruSeq Index, Nextera Index) for use as the barcode sequence of this invention. In some embodiments, the tag sequence can be designed such that the enzyme recognition site is located at the 5' end, 3' end, or within the sequence of the barcode.

[0091] In some embodiments, for a target gene designed to have a first pair of tag sequences at both ends of its full-length sequence, the first pair of tag sequences includes an enzyme recognition site and optionally a barcode. Preferably, the enzyme recognition site is an enzyme recognition site of a restriction endonuclease. In some embodiments, multiple target genes are synthesized by the method of the present invention, wherein the first pair of tag sequences for each target gene contains a different barcode. In some embodiments, the length of the first pair of tag sequences is 4 nt to 60 nt, for example, 10 nt to 60 nt, 10 nt to 50 nt, 20 nt to 50 nt, 20 nt to 40 nt, 30 nt to 50 nt, or 30 nt to 40 nt. Preferably, the length of the first pair of tags is 10 nt to 30 nt.

[0092] In some implementations, for oligonucleotide fragments designed to have an additional second pair of tag sequences at both ends of their sequence, the second pair of tag sequences includes a barcode and, optionally, an enzyme recognition site. Optionally, oligonucleotide fragments of the same target gene may have the same second pair of tag sequences.

[0093] In some embodiments, the second pair of tag sequences contains one or more enzyme recognition sites. As used herein, the terms "enzyme recognition site" or "enzyme cleavage site" refer to a specific base pair sequence or specific base (e.g., a uracil base) at a location in a double-stranded nucleic acid molecule or a complementary double-stranded segment, which is capable of being recognized and bound by a corresponding enzyme, optionally cleaving the double or single strand at a specific location in the sequence, or creating a single nucleotide gap. For restriction endonucleases, they are capable of binding to the DNA molecule at the recognition site and cleaving the DNA strand inside or outside the recognition site. Specific recognition sites of restriction endonucleases are known to those skilled in the art.

[0094] In some embodiments, the enzyme recognition site is the recognition site of a restriction endonuclease (e.g., type II and type II). Restriction endonucleases are generally classified into four main classes—type I, type II, type III, and type IV—based on differences in subunit composition, cleavage location, recognition site, and cofactors. Type II restriction endonucleases recognize specific nucleotide sequences and cleave double-stranded DNA at a fixed position inside or outside the recognized sequence, producing a DNA product with blunt or sticky ends bearing 3'-OH and 5'-phosphate groups. In some embodiments, the restriction endonuclease is a type II restriction endonuclease. In some embodiments, the restriction endonuclease is an IIS-type restriction endonuclease that recognizes consecutive non-palindromic sequences and cleaves the DNA double strand outside the recognition site. Examples include, but are not limited to, AcuI, AlwI, BaeI, BbsI, BbvI, BccI, BceAI, BcgI, BciVI, BcoDI, BsmAI, BfuAI, BspMI, BmrI, BpmI, BpuEI, BsaI, and BsaXI. BseRI, BsgI, BsmAI, BsmFI, BsmI, BspCNI, BspMI, BsrDI, BsrI, BtsCI, BtsIMutI, CspCI, EarI, EciI, Esp3I, FauI, FokI, HgaI, HphI, HpyAV, MboII, MlyI, MmeI, MnlI, NmeAIII, PaqCI, PleI, SapI, SfaNI. In a preferred embodiment, the restriction endonuclease is MlyI, which recognizes the GAGTC sequence and cleaves the DNA double strand outside the recognition site to produce blunt-ended fragments.

[0095] In some embodiments, the enzyme recognition site is the enzyme recognition site of a nicking enzyme. As used herein, the term "nicking enzyme" refers to an enzyme capable of recognizing a specific sequence (i.e., the recognition site) and cleaving one strand of a DNA double helix to create a gap (also known as a nick). Nicking enzymes as used herein encompass naturally occurring nicking enzymes or variants thereof, provided they retain nicking enzyme activity. Examples of nicking enzymes include, but are not limited to, nicking endonucleases such as Nt.BspQI, Nt.CviPII, Nt.BstNBI, Nb.BsrDI, Nb.BtsI, Nt.AlwI, Nb.BbvCI, Nt.BbvCI, Nb.BsmI, Nb.BssSI, Nt.BsmAI, and Cas9 nicking enzymes, etc. Those skilled in the art will understand how to select specific nicking enzymes (e.g., in combination with specific restriction endonucleases or different nicking enzymes) to achieve the objectives of this invention, and the recognition sites and mechanisms of action of nicking enzymes in the prior art are also known.

[0096] In some embodiments, the enzyme recognition site is the enzyme recognition site of a restriction endonuclease. In some embodiments, the enzyme recognition site is the enzyme recognition site of both a restriction endonuclease and a nicking enzyme. In some embodiments, the enzyme recognition site is the enzyme recognition site of two different nicking enzymes.

[0097] In some embodiments, the lengths of the second pair of tag sequences are 4 nt to 60 nt, for example, 10 nt to 60 nt, 10 nt to 50 nt, 20 nt to 50 nt, 20 nt to 40 nt, 30 nt to 50 nt, and 30 nt to 40 nt. Preferably, the lengths of the second pair of tags are 10 nt to 30 nt.

[0098] PCR amplification

[0099] Optionally, the oligonucleotide fragment having a second tag sequence is amplified after the oligonucleotide fragment is synthesized. In some embodiments, the oligonucleotide fragment having a second tag sequence is amplified by polymerase chain reaction (PCR) using primers containing the second tag sequence or its inverse complementary sequence.

[0100] In some embodiments, the number of cycles in the polymerase chain reaction is 2 to 50 cycles, for example, 3 to 40 cycles, 4 to 30 cycles, 5 to 30 cycles, 6 to 30 cycles, or 7 to 30 cycles; more preferably 8 to 20 cycles. The polymerase can be a high-fidelity polymerase. High-fidelity DNA polymerases have a polymerization center and a digestion center. The polymerization center has 5'→3' DNA polymerase activity, capable of catalyzing the synthesis of DNA chains along the 5'→3' direction; the digestion center has 3'→5' exonuclease activity, capable of repairing base mismatches. Common high-fidelity DNA polymerases include Pfu, KOD, Phusion, Q5, PrimeSTAR, Vent, and other DNA polymerases.

[0101] 5' phosphorylation of oligonucleotide fragments

[0102] As used herein, the term "5' phosphorylation" refers to the process by which the 5' terminal nucleotide of an oligonucleotide fragment carries a phosphate group through appropriate methods. Such appropriate methods include directly adding a phosphate group to the 5' terminal nucleotide of the oligonucleotide fragment via, for example, a phosphokinase-mediated reaction (such as T4 polynucleotide kinase) or a chemical phosphorylation method (such as the phosphoramide method or chemical modification reagents); performing a PCR reaction on the oligonucleotide fragment using 5' phosphorylated primers; and contacting and reacting the oligonucleotide fragment with an enzyme capable of hydrolyzing phosphodiester bonds and generating a 5' phosphate group, thereby phosphorylating the 5' terminal of the reaction product. The 5' phosphorylation in the methods of synthesizing various target genes of the present invention can be achieved in a variety of ways.

[0103] In some embodiments, the 5' end of the oligonucleotide fragments in the oligonucleotide fragment pool is phosphorylated by a phosphokinase-mediated reaction, preferably, the phosphokinase being a T4 polynucleotide kinase. In this case, it is preferable that the oligonucleotide fragments do not have tag sequences at both ends.

[0104] In some embodiments, the 5' phosphorylation is performed by a chemical phosphorylation method. Preferably, oligonucleotide fragments of multiple target genes are phosphorylated in a single reaction system.

[0105] For oligonucleotide fragments with a second tag sequence, after amplification and enrichment, the second tag sequence is removed and its 5' end is phosphorylated by contacting the oligonucleotide fragment with the second tag sequence with an enzyme capable of hydrolyzing phosphodiester bonds and generating 5' phosphate groups.

[0106] In some embodiments, an oligonucleotide fragment having a second pair of tag sequences is contacted with a restriction endonuclease to remove the second pair of tag sequences and phosphorylate its 5' end, wherein the second pair of tag sequences contains an enzyme recognition site, and the restriction endonuclease is capable of recognizing this enzyme recognition site. Preferably, the restriction endonuclease is a restriction endonuclease that cleaves the DNA double strand at a specific location outside its enzyme recognition site to produce blunt ends, such as MlyI enzyme.

[0107] In some embodiments, the oligonucleotide fragment having a second pair of tag sequences is contacted with a restriction endonuclease and an S1 nuclease to remove the second pair of tag sequences and phosphorylate their 5' ends, wherein the second pair of tag sequences contains an enzyme recognition site, and the restriction endonuclease is capable of recognizing this enzyme recognition site. Preferably, the restriction endonuclease is an IIS type restriction endonuclease.

[0108] In some embodiments, an oligonucleotide fragment having a second pair of tag sequences is contacted with different nicking enzymes to remove the second pair of tag sequences and phosphorylate their 5' ends, wherein the second pair of tag sequences contains different enzyme recognition sites, and the different nicking enzymes are capable of recognizing the different enzyme recognition sites respectively.

[0109] In some embodiments, the oligonucleotide fragment having a second pair of tag sequences is contacted with a restriction endonuclease and a nicking enzyme to remove the second pair of tag sequences and phosphorylate their 5' ends, wherein the second pair of tag sequences contains an enzyme recognition site, and the restriction endonuclease and nicking enzyme are capable of recognizing the enzyme recognition site, respectively. Preferably, the restriction endonuclease is an IIS type restriction endonuclease.

[0110] Ligase chain reaction (LCR) process

[0111] As described above, the method for synthesizing multiple target genes of the present invention further includes the step of assembling oligonucleotide fragments into a full-length nucleic acid sequence via a ligase chain reaction, thereby synthesizing multiple target genes. Specifically, under conditions that allow for complementary base pairing, any two adjacent oligonucleotide fragments at the same splitting point in the oligonucleotide fragment pool can be used as "parts" to hybridize complementaryly with another oligonucleotide fragment at a different splitting point, which serves as a "template," forming a double-stranded region with a gap in one strand. The adjacent oligonucleotide fragments are then linked by a ligase-mediated ligation reaction. Subsequently, the ligation product and its adjacent oligonucleotide fragments hybridize complementaryly with the "template" oligonucleotide fragment, or hybridize complementaryly with two other adjacent oligonucleotide fragments as a "template" alone, undergoing a similar ligase-mediated ligation reaction, thereby assembling multiple oligonucleotide fragments into a full-length nucleic acid sequence.

[0112] In some embodiments, the oligonucleotide fragment is a double-stranded oligonucleotide fragment, and the ligase chain reaction can be carried out under high-temperature to medium-low-temperature cycling. For example, the temperature cycling consists of high-temperature denaturation at 90-98°C followed by medium-low-temperature annealing and ligation at 37-75°C.

[0113] In some embodiments, the ligase used in the ligase chain reaction is a heat-stable ligase, such as HiFiTaq DNA ligase, 9°N DNA ligase, etc. DNA ligase, Hi-T4 TM Thermostable DNA ligase, Taq DNA ligase.

[0114] In some embodiments, the number of cycles of the ligase chain reaction can be 2-100 cycles, for example, 2 cycles, 3 cycles, 4 cycles, 5 cycles, 10 cycles, 15 cycles, 20 cycles, 25 cycles, 30 cycles, 35 cycles, 40 cycles, 45 cycles, 50 cycles, 55 cycles, 60 cycles, 65 cycles, 70 cycles, 75 cycles, 80 cycles, 85 cycles, 90 cycles, 95 cycles, or 100 cycles.

[0115] Preparation and assembly of 5' phosphorylated single-stranded oligonucleotide fragments

[0116] In some embodiments, the method of the present invention for synthesizing multiple target genes uses a 5' phosphorylated single-stranded oligonucleotide fragment as a substrate for a ligation reaction. The 5' phosphorylated single-stranded oligonucleotide fragment can be obtained by further preparing a 5' single-stranded oligonucleotide fragment from a double-stranded oligonucleotide fragment after amplification and enrichment.

[0117] In some embodiments, the amplification and enrichment of the oligonucleotide fragment is performed using amplification primers corresponding to the second pair of tag sequences, including forward amplification primers and reverse amplification primers, wherein the forward amplification primers contain a 5' end thiophosphorylation modification and optionally a 3' uracil modification, and the reverse amplification primers contain the reverse complementary sequence of the tag sequence. Optionally, the reverse amplification primers may also contain a 5' phosphorylation modification.

[0118] The second pair of tag sequences contains enzyme recognition sites selected from restriction endonucleases, nicking enzymes, and USER enzymes, or combinations thereof.

[0119] In some implementations, the steps for preparing 5' phosphorylated single-stranded oligonucleotide fragments include:

[0120] (a) Contact the amplified oligonucleotide fragment with phosphokinase to phosphorylate its 5' end on one side;

[0121] (b) Contacting a 5'-phosphorylated oligonucleotide fragment with an exonuclease (e.g., λ exonuclease) to obtain a single-stranded oligonucleotide fragment; and

[0122] (c) Remove the second tag sequence and phosphorylate the 5' end of the single-stranded oligonucleotide fragment.

[0123] In some embodiments, the second pair of tag sequences is removed by contacting the single-stranded oligonucleotide fragment with an enzyme corresponding to the enzyme recognition site and a primer, wherein the primer contains the second pair of tag sequences or their reverse complementary sequence. Optionally, the second pair of tag sequences contains an enzyme recognition site for a restriction endonuclease and / or a nicking enzyme, wherein the restriction endonuclease may be a type II or type II restriction endonuclease.

[0124] In some implementations, the steps for preparing 5' phosphorylated single-stranded oligonucleotide fragments include:

[0125] (a) Contact the amplified oligonucleotide fragment with the enzyme corresponding to the enzyme recognition site, remove the second pair of tag sequences on one side and phosphorylate the 5' end on one side;

[0126] (b) Contacting a 5'-phosphorylated oligonucleotide fragment with an exonuclease (e.g., λ exonuclease) to obtain a single-stranded oligonucleotide fragment; and

[0127] (c) Remove the second tag sequence and phosphorylate the 5' end of the single-stranded oligonucleotide fragment.

[0128] In some embodiments, the second pair of tag sequences is removed by contacting the single-stranded oligonucleotide fragment with an enzyme corresponding to the enzyme recognition site and optionally a primer, wherein the primer contains the second pair of tag sequences or their reverse complementary sequence.

[0129] In some embodiments, the enzyme recognition site is the uracil-modified site of the USER enzyme. The USER (uracil-specific excision reagent) enzyme is a mixture of uracil-DNA glycosylase (UDG) and endonuclease VIII, capable of creating a single nucleotide nick at the uracil position. During the oligonucleotide fragment amplification and enrichment step, the USER enzyme recognition site can be introduced into the amplification product using primers with a 3' uracil-modified tag sequence. After preparing the amplification product into a single-stranded oligonucleotide fragment, it is contacted with the primers and the USER enzyme, creating a nick at the 3' uracil base of the tag sequence, thereby removing the tag sequence and phosphorylating its 5' end.

[0130] In some embodiments, the ligase chain reaction can be carried out under high-temperature-medium-low-temperature cycling. For example, the temperature cycling is from a high temperature of 90-98°C to a medium-low temperature of 37-75°C, and the ligase is preferably a thermostable ligase. In some embodiments, the ligase chain reaction for assembly using single-stranded oligonucleotide fragments as substrates can be carried out at a constant temperature of 4-75°C (i.e., a fixed temperature).

[0131] In some embodiments, the number of cycles of the ligase chain reaction can be 2-100 cycles, for example, 2 cycles, 3 cycles, 4 cycles, 5 cycles, 10 cycles, 15 cycles, 20 cycles, 25 cycles, 30 cycles, 35 cycles, 40 cycles, 45 cycles, 50 cycles, 55 cycles, 60 cycles, 65 cycles, 70 cycles, 75 cycles, 80 cycles, 85 cycles, 90 cycles, 95 cycles, or 100 cycles.

[0132] In some embodiments, the ligase is selected from T4 DNA ligase, T7 DNA ligase, E. coli DNA ligase, T3 DNA ligase, HiFi Taq DNA ligase, 9°N DNA ligase, etc. DNA ligase, Hi-T4 TM Thermostable DNA ligase, Taq DNA ligase.

[0133] Specific embodiments of the present invention

[0134] The tag sequences at both ends of the oligonucleotide fragments were removed using restriction endonucleases and then phosphorylated at 5'. In some embodiments, the method of the present invention for synthesizing multiple target genes includes:

[0135] 1) Using the methods disclosed herein, corresponding oligonucleotide fragments are designed for each of the multiple target genes. The oligonucleotide fragments of all target genes together constitute an oligonucleotide library. Optionally, the oligonucleotide library can be further divided into two or more oligonucleotide sublibraries, wherein the nucleic acid sequences of the multiple target genes have the same or different first pair of tag sequences BCG1 and BCG2 at both ends, and any oligonucleotide fragment in the same sublibrary of the oligonucleotide library has the same second pair of tag sequences BCO1 and BCO2 at both ends.

[0136] 2) Phosphorylate the 5' end of the oligonucleotide fragment and remove the second pair of tag sequences BCO1 and BCO2;

[0137] 3) The 5' phosphorylated oligonucleotide fragments are assembled into nucleic acid sequences to obtain a variety of target genes.

[0138] Optionally, the lengths of the first pair of tag sequences and the second pair of tag sequences are 4nt-60nt, and preferably 10nt-30nt.

[0139] Optionally, prior to 5' phosphorylation, the corresponding oligonucleotide fragments are enriched by PCR using a polymerase and primers containing a second tag sequence or its reverse complementary sequence. Optionally, the polymerase is a high-fidelity polymerase, and the number of PCR cycles is 2-50, preferably 8-20 cycles. In a specific embodiment, the oligonucleotide library or sublibrary is amplified and enriched by PCR, for example, in a 20 μl PCR system containing MegaFi TM Pro Fidelity 2X PCR 10 μl, primers containing BCO1 and BCO2 or their reverse complementary sequences 1 μl each, oligonucleotide library or sub-library 1-10 ng, add nuclease-free water to 20 μl; amplification reaction: 8-20 cycles of 95℃ for 2 min, 95℃ for 10 s, 50℃ for 15 s, 72℃ for 20 s, 72℃ for 1 min, recovery and purification.

[0140] In one specific implementation, the oligonucleotide fragment having a second pair of tag sequences is contacted with a restriction endonuclease to remove the second pair of tag sequences and phosphorylate its 5' end, wherein the second pair of tag sequences contains the enzyme recognition site of the restriction endonuclease. Preferably, the restriction endonuclease is a restriction endonuclease that cleaves the double strand at a specific location outside its recognition site to produce blunt ends, such as MlyI. For example, in a 50 μl system: 5 μl of 10X rCutSmartBuffer, 1 μl of MlyI, the enriched oligonucleotide pool or sub-pool, and nuclease-free water added to a final volume of 50 μl, the mixture is incubated at 37°C for 1-2 h and then purified.

[0141] In another specific embodiment, the oligonucleotide fragment having the second pair of tag sequences is contacted with a restriction endonuclease and / or S1 nuclease to remove the second pair of tag sequences and phosphorylate their 5' ends, wherein the second pair of tag sequences contains the enzyme recognition site of the restriction endonuclease. Preferably, the restriction endonuclease is an IIS type restriction endonuclease.

[0142] In one specific implementation, 5' phosphorylated oligonucleotide fragments are assembled into nucleic acid sequences using LCR (Lipolysis Chromatography), which can be performed using a cycle of high-temperature denaturation followed by medium-low-temperature ligation. The ligase used in the LCR is a thermostable ligase, preferably a ligase without blunt-end activity, such as HiFi Taq DNA ligase or 9°N ligase. TMDNA ligase, Pfu DNA ligase, and Taq DNA ligase. For example, in the following 20 μl LCR system: 10X HiFi Taq DNA Ligase Buffer 2 μl, HiFi Taq DNA ligase, 5' phosphorylated oligonucleotide library or sublibrary, add nuclease-free water to 20 μl, and then perform 1-100 cycles of incubation at 90-98°C for 10 seconds and 37-75°C for 1-30 min.

[0143] In one specific implementation, 5' phosphorylated oligonucleotide fragments are assembled into nucleic acid sequences using LCR (Lipolysis Chromatography), wherein the LCR reaction is performed at an isothermal temperature of 4-75°C, i.e., an isothermal ligation reaction. The ligase used in the ligation is any DNA ligase, such as HiFiTaq DNA ligase or 9°N ligase. TM DNA ligase, Pfu DNA ligase, Taq DNA ligase, T4 DNA ligase. For example, in the following 20 μl LCR system: 10X HiFi Taq DNA Ligase Buffer 2 μl, HiFiTaq DNA ligase, 5' phosphorylated single-stranded oligonucleotides, add nuclease-free water to 20 μl, and then incubate at 4-75°C for 15 min-16 h.

[0144] A specific embodiment of the present invention is as follows:

[0145] 1) Design and synthesize oligonucleotide pools containing tag sequences;

[0146] 2) Add primers and polymerase to amplify and enrich oligonucleotide chains, such as a 20 μl PCR system: MegaFi TM ProFidelity 2X PCR 10 μl, primers containing BCO1 and BCO2 or their reverse complementary sequences 1 μl each, oligonucleotides 1-10 ng, and nuclease-free water added to 20 μl;

[0147] 3) Amplification reaction: 95℃ for 2 min, 95℃ for 10 s, 50℃ for 15 s, 72℃ for 20 s, 8-20 cycles, 72℃ for 1 min, recovery and purification;

[0148] 4) Add a restriction endonuclease corresponding to the enzyme recognition site in the tag sequence to remove the tag sequence of the oligonucleotide chain and phosphorylate it. For example, in a 50 μl digestion system, add 5 μl of 10X rCutSmart Buffer, 1 μl of MlyI, and the amplified double-stranded fragment. Add nuclease-free water to a final volume of 50 μl, incubate at 37°C for 1-2 h, and then recover and purify. The restriction endonuclease is capable of cleaving the double strand at a specific location outside its recognition site to produce blunt-ended double-stranded fragments.

[0149] 5) Step 4) above can also be performed using a combination of other IIS restriction enzymes and S1 nucleases. In this case, the IIS restriction enzyme cuts the double strand at a specific location outside its recognition site and produces a double-stranded fragment with sticky ends. Subsequently, the S1 nuclease degrades the sticky ends, thereby removing the tag sequence.

[0150] 6) Add ligase for ligation, such as a 20μl LCR system: 2μl of 10X HiFi Taq DNA Ligase Buffer, HiFiTaq DNA ligase, the digested double-stranded fragment, add nuclease-free water to 20μl, incubate at 90-98℃ for 10 seconds, then at 37-75℃ for 1-30 min, for 2-100 cycles.

[0151] 7) For sequences longer than 1kb, it is preferable to divide them into several gene fragments, perform LCR on each fragment separately, and then perform LCR / PCR to assemble the full-length sequence.

[0152] 8) If necessary, enrich the full-length gene using BCG1 and BCG2 and purify it; to avoid bias in PCR enrichment of different genes, it is preferable to use kits such as the Micellula DNA Emulsion & Purification Kit (EURx, Poland) for emulsion PCR enrichment.

[0153] 9) The sample is cleaved by restriction endonuclease, purified, and ligated into a vector.

[0154] The tag sequences at both ends of the oligonucleotide fragment were removed using a nicking enzyme and then 5' phosphorylated.

[0155] In some embodiments, the method of the present invention for synthesizing multiple target genes includes:

[0156] 1) Using the methods disclosed herein, corresponding oligonucleotide fragments are designed for each of the multiple target genes. The oligonucleotide fragments of all target genes together constitute an oligonucleotide library. Optionally, the oligonucleotide library can be further divided into two or more oligonucleotide sublibraries, wherein the nucleic acid sequences of the multiple target genes have the same or different first pair of tag sequences BCG1 and BCG2 at both ends, and any oligonucleotide fragment in the same sublibrary of the oligonucleotide library has the same second pair of tag sequences BCO1 and BCO2 at both ends.

[0157] 2) Phosphorylate the 5' end of the oligonucleotide fragment and remove the second pair of tag sequences BCO1 and BCO2;

[0158] 3) The 5' phosphorylated oligonucleotide fragments are assembled into nucleic acid sequences to obtain a variety of target genes.

[0159] Optionally, the lengths of the first pair of tag sequences and the second pair of tag sequences are 4nt-60nt, and preferably 10nt-30nt.

[0160] Optionally, prior to 5' phosphorylation, the corresponding oligonucleotide fragment is enriched by PCR using a polymerase and primers containing a second tag sequence or its reverse complementary sequence. Optionally, the polymerase is a high-fidelity polymerase, and the PCR cycle number is 2-50 cycles, preferably 8-20 cycles. In a specific embodiment, the oligonucleotide fragment is amplified and enriched by PCR, for example, in a 20 μl PCR system containing MegaFi TM Pro Fidelity 2X PCR 10 μl, primers containing BCO1 and BCO2 or their reverse complementary sequences 1 μl each, oligonucleotide pool 1-10 ng, add nuclease-free water to 20 μl; amplification reaction: 8-20 cycles of 95℃ for 2 min, 95℃ for 10 s, 50℃ for 15 s, 72℃ for 20 s, 72℃ for 1 min, recovery and purification.

[0161] In one specific implementation, an oligonucleotide fragment having a second pair of tag sequences is contacted with different nicking enzymes to remove the second pair of tag sequences and phosphorylate their 5' ends, wherein the second pair of tag sequences contains the corresponding enzyme recognition sites of the different nicking enzymes.

[0162] In another specific embodiment, the oligonucleotide fragment having the second pair of tag sequences is contacted with a restriction endonuclease and a nicking enzyme to remove the second pair of tag sequences and phosphorylate their 5' ends, simultaneously forming sticky ends of 4-90 nt. The second pair of tag sequences contains the enzyme recognition sites of the restriction endonuclease and the nicking enzyme. Preferably, the restriction endonuclease is an IIS type restriction endonuclease. For example, in the following 50 μl system: 5 μl of 10X rCutSmartBuffer, 1 μl of nicking enzyme, 1 μl of IIS restriction endonuclease, enriched oligonucleotide sublibrary, and nuclease-free water added to a final volume of 50 μl, incubated at 37°C for 1-2 h, and then purified.

[0163] In one specific implementation, 5' phosphorylated oligonucleotide fragments are assembled into nucleic acid sequences using LCR (Lipolysis), which can be performed using high-temperature-low-temperature cycling. The ligase used in the LCR is a thermostable ligase, preferably a ligase without blunt-end activity, such as HiFi Taq DNA ligase or 9°N ligase. TMDNA ligase, Pfu DNA ligase, and Taq DNA ligase. For example, in the following 20 μl LCR system: 10X HiFiTaq DNA Ligase Buffer 2 μl, HiFiTaq DNA ligase, 5' phosphorylated oligonucleotide sublibrary, nuclease-free water to 20 μl, then perform 1-100 cycles of incubation at 90-98°C for 10 seconds and 37-75°C for 1-30 min.

[0164] In one specific implementation, a 5' phosphorylated oligonucleotide fragment is assembled into a nucleic acid sequence using LCR. Preferably, the oligonucleotide fragment has sticky ends, and the LCR reaction is performed at an isothermal temperature of 4-75°C, i.e., an isothermal ligation reaction. The ligase used in the ligation reaction can be any DNA ligase, such as HiFi Taq DNA ligase, 9°N... TM DNA ligase, Pfu DNA ligase, Taq DNA ligase, T4 DNA ligase. For example, in the following 20 μl LCR system: 10X HiFiTaq DNA Ligase Buffer 2 μl, HiFiTaq DNA ligase, 5' phosphorylated oligonucleotides with sticky ends, add nuclease-free water to 20 μl, and then incubate at 4-75°C for 15 min-16 h.

[0165] Digesting double-stranded oligonucleotide fragments into single-stranded oligonucleotide chains and then linking them together to form a single-stranded genome assembly.

[0166] In some embodiments, the method of the present invention for synthesizing multiple target genes includes:

[0167] 1) Using the methods disclosed herein, corresponding oligonucleotide fragments are designed for each of the various target genes. All oligonucleotide fragments from the target genes collectively constitute an oligonucleotide library. Optionally, the oligonucleotide library can be further divided into two or more oligonucleotide sublibraries. The nucleic acid sequences of the various target genes have different or identical first tag sequences BCG1 and BCG2 at both ends, and any oligonucleotide fragment in the same sublibrary of the oligonucleotide library has identical second tag sequences BCO1 and BCO2 at both ends. This scheme requires the preparation of single-stranded oligonucleotide fragments, i.e., a single-strand assembly scheme.

[0168] 2) Prepare 5' phosphorylated single-stranded oligonucleotide fragments;

[0169] 3) Assemble 5' phosphorylated single-stranded oligonucleotide fragments into nucleic acid sequences to obtain multiple target genes.

[0170] In some embodiments, the second tag sequence BCO1 has a 5' end phosphorylation modification and a 3' uracil modification, and BCO2 contains an enzyme recognition site, preferably an enzyme recognition site of a restriction endonuclease or nicking enzyme. In some embodiments, the corresponding oligonucleotide fragment is enriched by PCR before preparing the 5' end phosphorylated single-stranded oligonucleotide fragment, wherein the forward primer used in the PCR contains BCO1 and the 5' end phosphorylation modification and / or 3' uracil modification, and the reverse primer contains the reverse complementary sequence of BCO2.

[0171] In some embodiments, the method for preparing the 5' phosphorylated single-stranded oligonucleotide fragment includes:

[0172] 1) Phosphorylation of the 5' end of the oligonucleotide fragment is achieved by contacting the oligonucleotide fragment with a restriction endonuclease or nicking enzyme to remove BCO2, or by contacting it with a phosphokinase (e.g., T4 PNK).

[0173] 2) Digest oligonucleotide chains with 5' end phosphorylation modification to prepare single-stranded oligonucleotide fragments;

[0174] 3) Contact the single-stranded oligonucleotide fragment with an enzyme (e.g., USER enzyme) and / or primers to remove BCO1, thereby phosphorylating the 5' end of the single-stranded oligonucleotide fragment.

[0175] Optionally, the lengths of the first pair of tag sequences and the second pair of tag sequences are 4nt-60nt, and preferably 10nt-30nt.

[0176] In one specific implementation, modified primers are used to enrich the corresponding oligonucleotide fragments via PCR while simultaneously phosphorylating one side of the 5' end of the oligonucleotide fragments. Preferably, the primers used for the PCR are, for example, a forward primer Fs containing a BCO1 sequence with 5' end thiophosphorylation modification, and a reverse primer Rs containing a BCO2 inverse complementary sequence and optionally 5' phosphorylation modification, preferably using a phosphokinase to introduce the 5' phosphorylation modification. The 3' ends of primers Fs and Rs contain enzyme recognition sites selected from the enzyme recognition sites of restriction endonucleases, nicking enzymes, and USER enzymes, or combinations thereof. Then, a polymerase is added to amplify and enrich the oligonucleotides. Optionally, the polymerase is a high-fidelity polymerase, and the number of PCR cycles is 2-50 cycles, preferably 8-20 cycles.

[0177] In one specific implementation, oligonucleotide fragments are amplified and enriched by PCR, wherein the 5' end of primer Fs is phosphorylated, and the 5' end of primer Rs is phosphorylated or unmodified, for example, in the following 20 μl PCR system: MegaFi TMProFidelity 2X PCR 10 μl, modified primers Fs and Rs 1 μl each, oligonucleotide pool 1-10 ng, add nuclease-free water to 20 μl; amplification reaction: 8-20 cycles of 95℃ for 2 min, 95℃ for 10 s, 50℃ for 15 s, 72℃ for 20 s, 72℃ for 1 min, recovery and purification.

[0178] In another specific embodiment, the oligonucleotide fragment having the second pair of tag sequences is contacted with a phosphokinase, retaining the tag sequences and phosphorylating them at one side of the 5' end, wherein the second pair of tag sequences contains the enzyme recognition site of the restriction endonuclease and / or the nicking enzyme or USER enzyme. Alternatively, unilateral 5' phosphorylation can be achieved by contacting the second pair of tag sequences with the restriction endonuclease and / or the nicking enzyme or USER enzyme to remove the unilateral tag sequences.

[0179] Optionally, after obtaining a single-sided 5' phosphorylated oligonucleotide fragment, the oligonucleotide fragment is contacted with an exonuclease to digest the 5' phosphorylated oligonucleotide chain to prepare a single-stranded oligonucleotide fragment.

[0180] In some embodiments, the method for preparing the 5' phosphorylated single-stranded oligonucleotide fragment includes: contacting the single-stranded oligonucleotide fragment having a second tag sequence with a primer and an enzyme containing the reverse complementary sequence of the corresponding tag sequence, and removing the second tag sequence to prepare the 5' phosphorylated single-stranded oligonucleotide fragment. The enzyme may be a restriction endonuclease, a nicking enzyme, a USER enzyme, or a combination thereof.

[0181] In one specific implementation, a single-stranded oligonucleotide fragment having a second pair of tag sequences is contacted with primers Fs and Rs containing reverse complementary sequences of the tag sequences and different restriction endonucleases, the second pair of tag sequences is removed and its 5' end is phosphorylated, wherein the second pair of tag sequences contains the enzyme recognition sites of the different restriction endonucleases.

[0182] In another specific implementation, a single-stranded oligonucleotide fragment having a second pair of tag sequences is contacted with primers Fs and Rs containing reverse complementary sequences of the tag sequences and different nicking enzymes to remove the second pair of tag sequences and phosphorylate their 5' ends, wherein the second pair of tag sequences contains the corresponding enzyme recognition sites of the different nicking enzymes.

[0183] In another specific embodiment, a single-stranded oligonucleotide fragment having a second pair of tag sequences is contacted with the USER enzyme to remove the second pair of tag sequences and phosphorylate its 5' end, wherein the second pair of tag sequences contains 3' and / or 5' uracil modifications.

[0184] In one specific implementation, 5' phosphorylated single-stranded oligonucleotide fragments are assembled into nucleic acid sequences using LCR (Lipolysis), which can employ high-temperature-low-temperature cycling. The ligase used in the LCR is a thermostable ligase, preferably a ligase without blunt-end activity, such as HiFi Taq DNA ligase or 9°N ligase. TM DNA ligase, Pfu DNA ligase, and Taq DNA ligase. For example, in the following 20 μl LCR system: 10X HiFiTaq DNA Ligase Buffer 2 μl, HiFiTaq DNA ligase, 5' phosphorylated single-stranded oligonucleotide fragment, nuclease-free water to 20 μl, then incubate at 90-98°C for 10 seconds, 37-75°C for 1-30 min for 1-100 cycles.

[0185] In one specific implementation, 5' phosphorylated single-stranded oligonucleotide fragments are assembled into nucleic acid sequences using LCR, followed by an isothermal reaction at 4-75°C, i.e., an isothermal ligation reaction. The ligase used in the ligation reaction can be any DNA ligase, such as HiFi Taq DNA ligase, 9°N... TM DNA ligase, Pfu DNA ligase, Taq DNA ligase, T4 DNA ligase. For example, in the following 20 μl LCR system: 10X HiFiTaq DNA Ligase Buffer 2 μl, HiFiTaq DNA ligase, 5' phosphorylated single-stranded oligonucleotide fragment, add nuclease-free water to 20 μl, and then incubate at 4-75°C for 15 min-16 h.

[0186] Compared to the single-strand preparation method described above, the present invention is more preferably a double-strand preparation method that does not include digesting double strands into single strands, because this avoids the cumbersome steps and operational losses of single-strand preparation, and also avoids the steps of preparing sticky ends with exonuclease or nicking enzyme.

[0187] One-pot assembly of oligonucleotide fragments without tagged sequences at both ends and without amplification of the oligonucleotide fragments.

[0188] In some implementations, the cleaved oligonucleotide fragments do not require tag sequences at both ends, and the method does not require a PCR amplification enrichment step. Preferably, the first strand of the target gene is cleaved at a first cleavage starting point, and the second strand is cleaved at a second cleavage starting point. This scheme is also known as a one-pot assembly scheme. This type of method is preferably suitable for cases with 1-1000 target genes, and more preferably for genome assembly with 1-100 genes.

[0189] More specifically, the method includes:

[0190] 1) Design corresponding oligonucleotide fragments for each of the multiple target genes using the methods disclosed herein, wherein the oligonucleotide fragments do not contain a second pair of tag sequences at both ends, and the oligonucleotide fragments of all target genes together form an oligonucleotide pool, wherein the nucleic acid sequences of the multiple target genes have different or the same first pair of tag sequences BCG1 and BCG2 at both ends.

[0191] 2) Phosphorylate the 5' end of the oligonucleotide fragment, preferably by means of a phosphokinase (e.g., T4 polynucleotide kinase);

[0192] 3) The 5' phosphorylated oligonucleotide fragments are assembled into nucleic acid sequences through a ligation reaction, thereby obtaining a variety of target genes.

[0193] In one specific implementation, the 5' end of the oligonucleotide fragment is phosphorylated by T4 polynucleotide kinase. For example, in the following 50 μl phosphorylation system: oligonucleotide pool, 5 μl 10XT4 PNK reaction buffer, 5 μl 10 mM ATP, 10 units of T4 polynucleotide kinase, and nuclease-free water to bring the volume to 50 μl, the mixture is incubated at 37°C for 30 min, and the product is recovered and purified.

[0194] In one specific implementation, 5' phosphorylated oligonucleotide fragments are assembled into nucleic acid sequences using LCR, which can be performed using high-temperature-low-temperature cycling. The ligase used in the LCR is a thermostable ligase. For example, in the following 20 μl LCR system: 2 μl of 10X HiFi Taq DNA Ligase Buffer, HiFi Taq DNA ligase, 5' phosphorylated oligonucleotide fragments, and nuclease-free water to a final volume of 20 μl, followed by 1-100 cycles of incubation at 90-98°C for 10 seconds and 37-75°C for 1-30 minutes.

[0195] In one specific implementation, 5' phosphorylated single-stranded oligonucleotide fragments are assembled into nucleic acid sequences using LCR (Lipolysis Chromatography), wherein the LCR reaction is performed at an isothermal temperature of 4-75°C, i.e., an isothermal ligation reaction. The ligase used in the ligation can be any DNA ligase, such as HiFiTaq DNA ligase, 9°N... TM DNA ligase, Pfu DNA ligase, Taq DNA ligase, T4 DNA ligase. For example, in the following 20 μl LCR system: 10X HiFi Taq DNA Ligase Buffer 2 μl, HiFi Taq DNA ligase, 5' phosphorylated single-stranded oligonucleotide fragment, add nuclease-free water to 20 μl, and then incubate at 4-75℃ for 15 min-16 h.

[0196] Evaluation of gene synthesis efficiency

[0197] The effectiveness of gene synthesis or oligonucleotide fragment assembly can be evaluated using multiple parameters, including coverage, fidelity, and loss rate.

[0198] As used herein, the term "coverage" refers to the ratio of the number of target nucleic acid sequence types actually obtained through assembly to the predetermined number of target nucleic acid sequence types. The assembled nucleic acid sequences are optionally determined using conventional methods in the art (e.g., second- or third-generation sequencing). Coverage can be calculated using sequence alignment methods known in the art. In some embodiments, the nucleic acid sequences obtained by the method of synthesizing one or more target genes according to the present invention have a coverage of at least 50% relative to the target nucleic acid sequence, preferably at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95%.

[0199] As used in this article, the term "fidelity" refers to the median percentage of 100% correct sequences among all types of nucleic acid sequences assembled in the library, i.e., 100% accuracy.

[0200] As used herein, the term "loss rate" refers to the percentage of oligonucleotides that are sequenced at a depth less than or equal to 1%, 5%, or 10% of the median sequencing depth of all target oligonucleotide chains in the total library after Illumina sequencing by retrieving oligonucleotide chains with additional tag sequences (e.g., BCG1 and BCG2 in the examples) designed at both ends of the full-length gene. The percentage of genes corresponding to these oligonucleotides is the gene loss rate.

[0201] Beneficial effects of the present invention

[0202] By using tag sequences and corresponding primers, different genes (target genes) or oligonucleotide fragments can be easily retrieved and separated into different sub-libraries by PCR according to the actual situation, while simultaneously enriching the oligonucleotide fragments.

[0203] High-quality double-stranded DNA fragments (i.e., double-stranded oligonucleotide fragments) can be obtained by using high-fidelity polymerase PCR. Illumina sequencing analysis showed that the single-base error rate was no higher than 1 / 4000.

[0204] Through LCR-based seamless ligation, a 1.8kb gene was obtained with a fidelity of 60.11%.

[0205] The three-step synthesis method of the present invention, utilizing oligonucleotide pools and primers, reduces the cost of gene (nucleic acid sequence) synthesis. The average cost of preparing approximately 500-12,500 1kb genes is $2.95 / 1kb, including $2.24 / 1kb for oligonucleotide pool and primer synthesis and $0.71 / 1kb for enzyme preparation and purification.

[0206] This protocol differs from DropSynth 2.0 in the following ways: 1) It does not require magnetic beads; 2) It does not require separating each oligonucleotide chain of each gene into different droplet compartments for assembly; 3) Different target genes can share the same tag sequence, resulting in less use of barcodes; 4) It uses LCR instead of polymerase chain assembly (PCA), significantly reducing the error rate associated with PCR; 5) It can synthesize longer genes; 6) It offers greater experimental operability, fewer steps, and 3 times higher fidelity.

[0207] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0208] Example

[0209] Example 1. Design, synthesis and amplification of oligonucleotide chains

[0210] 1.1 Synthesis of gene fragments with barcodes

[0211] The synthesized genes included 328 beta-lactamase genes (approximately 0.8 kb), 493 fluorescent protein genes, and 2507 TnpB genes (1.6-2.2 kb, totaling three gene classes). Based on the known sequences of these genes, fragment splitting was designed on a computer: necessary gene / sequence optimization was performed using a self-written Python script. Each gene was then sequentially split into oligonucleotide chains at two different splitting start points. Oligonucleotide chains from the same splitting start point had no gaps or overlaps, while oligonucleotide chains from different splitting start points had overlapping or complementary regions of 6-5000 nt. Figure 2 As shown. Each segment has a pair of barcodes at both ends, namely BCO1 and BCO2; the full-length gene has another pair of barcodes at both ends, namely BCG1 and BCG2.

[0212] Barcode design is as follows: Generate bars of a certain length using a self-written Python script, preferably 10-30 nt. During design, the bars should have the lowest possible similarity to each other, with a similarity to the target gene below 90%, preferably below 80%. They should avoid non-specific amplification during PCR and unnecessary restriction enzyme sites, and have a GC content of 35-65%. Each barcode may or may not carry the corresponding IIS restriction enzyme site. Each sub-library should use a pair of independent bars as primers or for PCR retrieval. For example, for a pool of 48,000 oligonucleotide chains, 1-8,000 pairs of bars can be used for retrieval, preferably 1-400 pairs.

[0213] The assembly method is based on the LCR principle. Oligonucleotides from gene A and gene B, which have significant nucleotide differences, also exhibit specificity and do not share the same ternary linker (i.e., three oligonucleotide chains forming a triangular pair), making them impossible to splice together during assembly. If gene A and gene B splits to produce partially highly similar oligonucleotides, it is preferable to assign gene A and gene B to different sub-libraries.

[0214] Oligonucleotide fragments with pre-designed sequences corresponding to beta-lactamase, fluorescent protein, and TnpB gene were sent for synthesis to form an oligonucleotide fragment pool.

[0215] 1.2 Oligonucleotide chain synthesis and amplification

[0216] according to Figure 1-3 The gene fragments of 0.8-3kb were dissected to obtain a pool of 48,000 oligonucleotide chains of 250-300nt length (synthesized by Twist). The oligonucleotide chains from different subpools of the synthesized oligonucleotide pools were enriched by PCR using their corresponding BCO1 and BCO2 molecules. The PCR experimental conditions were as follows: MegaFi TM Pro Fidelity 2X PCR 10 μl, primers containing BCO1 and BCO2 or their reverse complementary sequences 1 μl each, oligonucleotide 1 ng, and Nuclease-free Water added to 20 μl; Amplification reaction: 12 cycles of 95℃ for 2 min, 95℃ for 10 s, 50℃ for 15 s, and 72℃ for 20 s, followed by 72℃ for 1 min, and then recovery and purification.

[0217] The purified recovered product was analyzed by 2% agarose gel electrophoresis, and the results are as follows: Figure 4As shown, Illumina sequencing analysis was performed. The median error rate of completely correct oligonucleotide single bases was no higher than 1 / 4000, comparable to the error rate of mutation-free reference sequences. When a threshold of 5% of the median sequencing reads was set, the loss rate of oligonucleotide chains before assembly was 0.85%, corresponding to a loss of 7.34% of the gene. Specific loss rates are shown in Table 1.

[0218] Table 1. Percentage of oligonucleotide and gene loss under different threshold settings.

[0219] Set threshold % Oligonucleotide chain loss rate % Corresponding gene loss rate % 1 0.22 2.28 5 0.85 7.34 10 2.77 18.55

[0220] Example 2. Assembling oligonucleotide chains

[0221] The synthesized oligonucleotide chain was digested by enzymes and then... Figure 3 The LCR method was used for assembly. The 50 μl digestion system was as follows: 5 μl 10XrCutSmart Buffer, 1 μl MlyI enzyme, double-stranded fragment, and nuclease-free water to bring the volume to 50 μl. The mixture was incubated at 37°C for 2 h and then purified. Digestion with the IIS-type restriction endonuclease MlyI achieved 5' phosphorylation of the oligonucleotide fragment during the excision of BCO1 and BCO2. This example, compared to alternative one-pot assembly, allows for the assembly of more genes by splitting different sub-libraries; compared to single-strand preparation methods, it eliminates the single-strand preparation and purification steps; and compared to methods combining nicking enzymes and / or restriction endonucleases and / or S1 nucleases, only one enzyme is needed to complete the excision of the second tag sequence and 5' phosphorylation of the oligonucleotide fragment.

[0222] 20μl LCR system: 10X HiFi Taq DNA Ligase Buffer 2μl, HiFi Taq DNA ligase, digested double-stranded fragment, add nuclease-free water to 20μl, incubate at 90-98℃ for 10 seconds, then at 37-75℃ for 1-30 min, for 2-100 cycles.

[0223] The LCR product was purified and detected by capillary electrophoresis, such as... Figure 5 As shown, the results indicate that 18% of the oligonucleotides in the LCR system were spliced ​​together to form the target sequence of approximately 950 bp.

[0224] Example 3. Validating the coverage and fidelity of the genome assembly method

[0225] Genes of varying lengths, ranging from 0.8 to 2.2 kb, were assembled using the method described in this application. The assembled sequences were then ligated to plasmids using restriction endonucleases. The coverage and fidelity of the obtained sequences were assessed using Nanopore sequencing.

[0226] like Figure 6As shown, the coverage of genes shorter than 1kb can reach 87.94%, and the coverage of genes around 1.8kb is 50%. Nanopore sequencing results further demonstrate that assembling genes shorter than 1kb is easier. This invention also achieved ideal coverage when assembling longer 1.8kb genes. Taking a longer gene with an average length of 1.8kb as an example, the percentage of correct sequences in the final obtained sequence was 60.11%. The method of this application has significant advantages over the DropSynth 2.0 technology used in Sriram Kosuri's laboratory, which only achieves a fidelity of 22.6% and 27.6% when assembling genes around 1kb.

[0227] The results of Examples 1-3 show that the gene synthesis method of this application meets the research requirements for large-scale gene characterization experiments in terms of coverage and accuracy (fidelity).

[0228] Example 4. Other genome assembly schemes

[0229] 4.1 The second tag sequence was removed using a nicking enzyme and 5' phosphorylation was performed.

[0230] 1) Design oligonucleotide fragments with tag sequences at both ends, and then synthesize an oligonucleotide fragment pool;

[0231] 2) Add polymerase and primer pairs for oligonucleotide chain amplification and enrichment, such as a 20 μl PCR system: MegaFi TM ProFidelity 2X PCR 10 μl, primers containing BCO1 and BCO2 or their reverse complementary sequences 1 μl each, oligonucleotide chains 1-10 ng, and nuclease-free water added to 20 μl;

[0232] Amplification reaction: 95℃ for 2 min, 95℃ for 10 s, 50℃ for 15 s, 72℃ for 20 s, 8-20 cycles, 72℃ for 1 min, recovery and purification;

[0233] 3) Use two cutting enzymes or a combination of cutting enzyme and IIS enzyme sequentially to remove the tag sequence and form sticky ends. For example, in a 50 μl digestion system: 5 μl 10X rCutSmart Buffer, 1 μl cutting enzyme, 1 μl IIS enzyme, double-stranded fragment, add nuclease-free water to 50 μl, incubate at 37°C for 1-2 h, and then recover and purify. Preferably, use a combination of cutting enzymes that can produce long sticky ends or a combination of cutting enzyme and IIS enzyme. Alternatively, two cutting enzymes or a combination of cutting enzyme and IIS enzyme can be used to remove the tag sequence and form blunt ends.

[0234] 4) Add ligase for ligation, such as a 20μl LCR system: 2μl of 10X HiFi Taq DNA Ligase Buffer, HiFi Taq DNA ligase, the digested double-stranded fragment, add nuclease-free water to 20μl, incubate at 90-98℃ for 10 seconds, then at 37-75℃ for 1-30 min, for 1-100 cycles.

[0235] The two steps described above can be combined into one reaction system;

[0236] 5) If necessary, enrich the full-length gene using BCG1 and BCG2 and purify it; to avoid bias in PCR enrichment of different genes, it is preferable to use kits such as the Micellula DNA Emulsion & Purification Kit (EURx, Poland) for emulsion PCR enrichment.

[0237] 6) The sample is cleaved by restriction endonuclease, purified, and ligated into a vector.

[0238] 4.2 Single-strand assembly for PCR amplification using modified primers

[0239] 1) Design oligonucleotide fragments with tag sequences at both ends, and then synthesize an oligonucleotide fragment pool;

[0240] 2) Synthesize modified primers: for example, forward primer Fs contains the BCO1 sequence and is thiophosphorylated at the 5' end, and reverse primer Rs contains the BCO2 reverse complementary sequence and is phosphorylated at the 5' end.

[0241] 3) Add polymerase to amplify and enrich oligonucleotide chains, such as in a 20 μl LCR system: MegaFi TM Pro Fidelity2X PCR 10 μl, modified primers Fs and Rs 1 μl each, oligonucleotides 1-10 ng, add nuclease-free water to 20 μl;

[0242] Amplification reaction: 95℃ for 2 min, 95℃ for 10 s, 50℃ for 15 s, 72℃ for 20 s, 8-20 cycles, 72℃ for 1 min, recovery and purification, as follows. Figure 7 Lanes B and C are shown;

[0243] 4) In practice, the reverse primer Rs in step 2) may also lack 5' phosphorylation modification, while unilateral 5' phosphorylation is achieved through subsequent phosphokinase treatment. For example, a 50 μl phosphorylation system can be prepared as follows: oligonucleotide, 10X T4 PNK reaction buffer 5 μl, 10 mM ATP 5 μl, T4 polynucleotide kinase 10 units, and nuclease-free water added to a final volume of 50 μl. The mixture is incubated at 37°C for 30 min, then purified. Figure 7 Lanes 2 and 4 are shown;

[0244] 5) Add λ exonuclease to prepare single-stranded oligonucleotide chains, such as a 50 μl LCR system: double-stranded DNA, 5 μl 10X λ Exonuclease Reaction Buffer, 5 units of λ exonuclease, add nuclease-free water to a final volume of 50 μl, incubate at 37°C for 30 min, recover and purify, results as shown below. Figure 7 Lanes 3 and 5;

[0245] 6) Add primers and restriction endonucleases containing the tag sequence or its reverse complementary sequence to the λ exonuclease digestion product. For example, in a 50 μl LCR system: 5 μl 10X rCutSmart Buffer, 1 μl restriction endonuclease, single-stranded fragment, primers containing BCO1 and BCO2 or their reverse complementary sequences, and add nuclease-free water to a final volume of 50 μl. Incubate at 37°C for 1-2 h, then recover and purify. Figure 7 Lanes 7-9, below are the cut BCO1 and BCO2 tag sequences;

[0246] 7) During implementation, the removal of the tag sequence in step 6) can be achieved by using a 3' uracil-modified primer in combination with the uracil-specific removal reagent USER enzyme to achieve bilateral or unilateral cleavage. For example, in a 50 μl LCR system: 5 μl of 10X rCutSmart Buffer, 1 μl of USER enzyme, incubated at 37°C for 1-2 h.

[0247] 8) Add ligase for ligation, such as a 20μl LCR system: 2μl of 10X HiFi Taq DNA Ligase Buffer, HiFi Taq DNA ligase, digested single-stranded oligonucleotide fragments, add nuclease-free water to 20μl, incubate at 90-98℃ for 10 seconds, then at 37-75℃ for 1-30 min, for 1-100 cycles;

[0248] 9) If necessary, enrich the full-length gene using BCG1 and BCG2 and purify it; to avoid bias in PCR enrichment of different genes, it is preferable to use kits such as the Micellula DNA Emulsion & Purification Kit (EURx, Poland) for emulsion PCR enrichment.

[0249] 10) The sample is cleaved by restriction endonuclease, purified, and ligated into a vector.

[0250] 4.3 One-pot cooking method for oligonucleotides without tagged sequences at both ends

[0251] 1) The designed and synthesized oligonucleotide fragments fully cover the forward and reverse sequences of the gene and do not contain any additional tag sequences;

[0252] 2) Add phosphokinase for phosphorylation, such as a 50 μl phosphorylation system: oligonucleotide, 10X T4PNK reaction buffer 5 μl, 10 mM ATP 5 μl, T4 polynucleotide kinase 10 units, add nuclease-free water to 50 μl, incubate at 37°C for 30 min, and then recover and purify.

[0253] 3) Add ligase for ligation, such as a 20μl LCR system: 10X HiFi Taq DNA Ligase Buffer 2μl, HiFi Taq DNA ligase, phosphorylated oligonucleotides, add nuclease-free water to 20μl, incubate at 90-98℃ for 10 seconds, incubate at 37-75℃ for 1-30 min, 1-100 cycles.

[0254] 4) If necessary, enrich the full-length gene using BCG1 and BCG2 and purify it; to avoid bias in PCR enrichment of different genes, it is preferable to use kits such as the Micellula DNA Emulsion & Purification Kit (EURx, Poland) for emulsion PCR enrichment.

[0255] 5) The sample is cleaved by restriction endonuclease, purified, and ligated into a vector.

Claims

1. A method for synthesizing multiple target genes in a single reaction system, comprising: For each of the multiple target genes, corresponding oligonucleotide fragments are designed, wherein the full-length sequence of each target gene is sequentially split without gaps or overlaps to obtain the sequences of multiple oligonucleotide fragments; Based on the designed sequences of multiple oligonucleotide fragments, an oligonucleotide fragment pool is synthesized. Phosphorylation of the 5' end of the oligonucleotide fragment; and The 5' phosphorylated oligonucleotide fragments are assembled into the full-length nucleic acid sequence of the target gene by ligase chain reaction (LCR); Optionally, the oligonucleotide fragment is designed with tag sequences at both ends, and the method includes removing the tag sequences before LCR.

2. The method of claim 1, wherein the method for designing a corresponding oligonucleotide fragment pool for the target gene comprises: The full-length sequence was split sequentially without gaps or overlap from two different splitting starting points using two different splitting methods, resulting in two different sets of oligonucleotide fragments; Oligonucleotide fragments obtained from the same splitting origin do not have gaps or overlapping regions, while oligonucleotide fragments obtained from different splitting methods have overlapping or complementary regions of 6-5000 nt. Optionally, the target gene can be sequence optimized, for example, by a computer program, before splitting, by codon optimization.

3. The method of claim 1 or 2, wherein the oligonucleotide fragments of the plurality of target genes constitute an oligonucleotide fragment pool, optionally, the oligonucleotide fragment pool comprises two or more oligonucleotide fragment sub-libraries; Optionally, each sub-library contains oligonucleotide fragments derived from 1 to 1000 target genes, preferably from 1 to 100 target genes.

4. The method of any one of claims 1-3, wherein the method for designing corresponding oligonucleotide fragments for the target gene comprises: For each target gene, the full-length sequence of the target gene is designed to have a first pair of tag sequences at both ends. Optionally, the first pair of tag sequences includes an enzyme recognition site and / or a barcode sequence, wherein the enzyme recognition site is an enzyme recognition site of a restriction endonuclease, and the barcode sequence is gene type-specific, sub-library-specific, or gene-specific.

5. The method of any one of claims 1-4, wherein the method for designing a corresponding oligonucleotide fragment pool for the target gene comprises: Each oligonucleotide fragment sequence is designed to have a second pair of tag sequences at both ends. Optionally, the second pair of tag sequences comprises a barcode sequence and one or more enzyme digestion recognition sites, wherein the barcode sequence is sub-library specific.

6. The method of claim 5, wherein the oligonucleotide fragment having the second pair of tag sequences is amplified by polymerase chain reaction (PCR) using primers corresponding to the second pair of tag sequences prior to phosphorylation of the 5' end of the oligonucleotide fragment.

7. The method of claim 6, wherein the number of cycles of the polymerase chain reaction is 2 to 50 cycles, more preferably 8 to 20 cycles; Preferably, the polymerase is a high-fidelity polymerase.

8. The method of claim 6, wherein the second pair of tag sequences is removed and its 5' end phosphorylated by contacting the oligonucleotide fragment having the second pair of tag sequences with an enzyme capable of hydrolyzing phosphodiester bonds and generating 5' phosphate groups.

9. The method of claim 8, wherein the enzyme is a restriction endonuclease, and the second pair of tag sequences contains the enzyme recognition site of the restriction endonuclease, preferably, the restriction endonuclease is a restriction endonuclease capable of cleaving double strands and producing blunt ends, more preferably MlyI enzyme.

10. The method of claim 8, wherein the enzyme is a restriction endonuclease and an S1 nuclease, and the second pair of tag sequences contains an enzyme recognition site of the restriction endonuclease, preferably, the restriction endonuclease is an IIS type restriction endonuclease.

11. The method of claim 8, wherein the enzyme is a different nicking enzyme, and the second pair of tag sequences contains the enzyme recognition site of the different nicking enzyme.

12. The method of claim 8, wherein the enzyme is a restriction endonuclease and a nicking enzyme, and the second pair of tag sequences contains enzyme recognition sites of the restriction endonuclease and the nicking enzyme, preferably, the restriction endonuclease is an IIS type restriction endonuclease.

13. The method of claim 6, wherein the primers corresponding to the second pair of tag sequences comprise a forward primer and a reverse primer, the forward primer comprising a 5' end thiophosphorylation modification and optionally a 3' uracil modification, the reverse primer comprising a reverse complementary sequence to the tag sequence, and the enzyme recognition site being selected from the enzyme recognition sites of restriction endonucleases, nicking enzymes, and USER enzymes, or combinations thereof. Optionally, the reverse primer contains a 5' phosphorylation modification.

14. The method of claim 13, for a reverse primer not containing 5' phosphorylation modification, the method further includes the step of preparing a 5' phosphorylated single-stranded oligonucleotide fragment after amplification, wherein: The amplified oligonucleotide fragment is contacted with a phosphokinase to phosphorylate its 5' end on one side, or the amplified oligonucleotide fragment is contacted with the enzyme corresponding to the enzyme recognition site, and the tag on one side of the second pair of tag sequences is removed and its 5' end is phosphorylated on one side. Single-stranded oligonucleotide fragments are obtained by treating oligonucleotide fragments with phosphorylated 5' ends on one side with an exonuclease (e.g., λ exonuclease). The second tag sequence is removed and the 5' end of the single-stranded oligonucleotide fragment is phosphorylated.

15. The method of claim 13, for a reverse primer containing a 5' phosphorylation modification, the method further includes the step of preparing a 5' phosphorylated single-stranded oligonucleotide fragment after amplification, wherein: Single-stranded oligonucleotide fragments are obtained by contacting oligonucleotide fragments with exonucleases (e.g., λ exonuclease). The second tag sequence is removed and the 5' end of the single-stranded oligonucleotide fragment is phosphorylated.

16. The method of claim 14 or 15, wherein, The second pair of tag sequences is removed by contacting the single-stranded oligonucleotide fragment with an enzyme corresponding to the enzyme recognition site and, optionally, a primer containing the second pair of tag sequences or their reverse complementary sequence.

17. The method of any one of claims 11-16, wherein the nicking enzyme is selected from Nt.BspQI, Nt.CviPII, Nt.BstNBI, Nb.BsrDI, Nb.BtsI, Nt.AlwI, Nb.BbvCI, Nt.BbvCI, Nb.BsmI, Nb.BssSI, Nt.BsmAI and Cas9 nicking enzyme or combinations thereof.

18. The method of any one of claims 4-17, wherein the length of the first pair of tag sequences is 4 nt to 60 nt, for example, 10 nt to 60 nt, 10 nt to 50 nt, 20 nt to 50 nt, 20 nt to 40 nt, 30 nt to 50 nt, or 30 nt to 40 nt; Preferably, the length of the first pair of tags is 10 to 30 nt.

19. The method of any one of claims 5-18, wherein the length of the second pair of tag sequences is 4 nt to 60 nt, for example, 10 nt to 60 nt, 10 nt to 50 nt, 20 nt to 50 nt, 20 nt to 40 nt, 30 nt to 50 nt, or 30 nt to 40 nt; Preferably, the length of the second pair of labels is 10 to 30 nt.

20. The method of any one of claims 1-4, wherein the 5' end phosphorylation is performed by a phosphokinase-mediated reaction or a chemical phosphorylation method (such as phosphoramide or a chemically modified reagent); Preferably, the phosphokinase is T4 polynucleotide kinase.

21. The method of any one of claims 1-20, wherein the length of the oligonucleotide fragment is about 20 nt to about 10,000 nt, about 50 nt to about 5,000 nt, about 70 nt to about 500 nt, about 90 nt to about 500 nt, about 50 nt to about 400 nt, about 70 nt to about 400 nt, about 90 nt to about 400 nt, about 50 nt to about 300 nt, about 70 nt to about 300 nt, about 90 nt to about 300 nt, or about 200 nt to about 300 nt.

22. The method of any one of claims 2-21, wherein the distance between different splitting starting points is from about 6nt to about 5000nt, for example, from about 10nt to about 2000nt, from about 30nt to about 200nt, from about 40nt to about 150nt, from about 40nt to about 100nt, from about 50nt to about 200nt, from about 60nt to about 150nt, from about 70nt to about 200nt, from about 70nt to about 150nt, from about 80nt to about 200nt, from about 80nt to about 150nt, from about 90nt to about 200nt, from about 90nt to about 150nt, from about 100nt to about 200nt, or from about 100nt to about 150nt.

23. The method of any one of claims 1-22, wherein the 5' phosphorylated oligonucleotide fragment is assembled into the full-length nucleic acid sequence of the target gene by ligase chain reaction (LCR); Preferably, the oligonucleotide fragment pool or oligonucleotide fragment sublibrary assembles oligonucleotide fragments of multiple target genes into the full-length nucleic acid sequence of the corresponding target gene in a separate reaction system.

24. The method of claim 23, wherein the LCR is performed at a fixed temperature or under high-temperature-medium-low-temperature cycling; Optionally, in the high-temperature-medium-low-temperature cycle, the high temperature is 90-98℃ and the medium-low temperature is 37-75℃; Optionally, the fixed temperature is 4-75°C.

25. The method of claim 23 or 24, wherein the number of cycles of the LCR is 2 to 100 cycles, for example, 2 cycles, 3 cycles, 4 cycles, 5 cycles, 10 cycles, 15 cycles, 20 cycles, 25 cycles, 30 cycles, 35 cycles, 40 cycles, 45 cycles, 50 cycles, 55 cycles, 60 cycles, 65 cycles, 70 cycles, 75 cycles, 80 cycles, 85 cycles, 90 cycles, 95 cycles, or 100 cycles.

26. The method of any one of claims 23-25, wherein the LCR is performed under high-temperature-medium-low-temperature cycling, and the ligase is a thermostable ligase; Preferably, the ligase is Hi-T4. TM Thermostable DNA ligase, HiFi Taq DNA ligase, Taq DNA ligase or DNA ligase.

27. The method of any one of claims 23-25, wherein the LCR is performed at a fixed temperature, and the ligase is selected from HiFi Taq DNA ligase, 9°NTM DNA ligase, Pfu DNA ligase, Taq DNA ligase, T4 DNA ligase, T3 DNA ligase, and T7 DNA ligase.

28. The method of any one of claims 23-27, further comprising the step of amplifying the nucleic acid sequences of the plurality of target genes by means of a first pair of tag sequences (e.g., by polymerase chain reaction) after assembly into a full-length nucleic acid sequence.

29. A pool of oligonucleotide fragments for synthesizing multiple target genes, comprising corresponding oligonucleotide fragments designed and synthesized for each of the multiple target genes, wherein, The sequences of oligonucleotide fragments were obtained by sequentially splitting the full-length sequence of each target gene in two ways without gaps or overlap. Optionally, the oligonucleotide fragment pool comprises two or more oligonucleotide fragment sub-libraries.

30. The oligonucleotide fragment pool of claim 29, wherein the two splitting methods involve sequentially splitting the full-length sequence without gaps or overlaps from two different splitting starting points to obtain two different sets of oligonucleotide fragments; Oligonucleotide fragments obtained from the same splitting origin do not have gaps or overlapping regions, while oligonucleotide fragments obtained from different splitting methods have overlapping or complementary regions of 6-5000 nt. Optionally, the sequence of the target gene is optimized, for example, by codon optimization.

31. The oligonucleotide fragment pool of claim 29 or 30, wherein For each target gene, the full-length sequence of the target gene has a first pair of tag sequences at both ends. Optionally, the first pair of tag sequences includes an enzyme recognition site and / or a barcode sequence, wherein the enzyme recognition site is an enzyme recognition site of a restriction endonuclease, and the barcode sequence is gene type-specific, sub-library-specific, or gene-specific.

32. The oligonucleotide fragment pool of any one of claims 29-31, wherein Each oligonucleotide fragment sequence has a second pair of tag sequences at both ends. Optionally, the second pair of tag sequences comprises a barcode sequence and one or more enzyme restriction sites, wherein the barcode sequence is sub-library specific. Optionally, the enzyme recognition site is selected from the enzyme recognition sites of restriction endonucleases, nicking enzymes, and USER enzymes, or combinations thereof.

Citation Information

Patent Citations

  • Compositions and methods for polynucleotide assembly

    WO2023096890A1