Method for preparing long double-stranded nucleic acid
By annealing short-chain single-stranded nucleic acids to form cracked intermediates and utilizing the organism's repair capabilities, the mismatch problem in long-chain nucleic acid synthesis is solved, achieving high-accuracy and low-cost nucleic acid synthesis, which is suitable for multi-gene synthesis and gene chips.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CODER THERAPEUTICS CO LTD
- Filing Date
- 2025-11-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies suffer from nucleic acid sequence mismatch problems when synthesizing long-chain nucleic acids, resulting in low product accuracy and a high likelihood of assembly failure, especially in the synthesis of highly homologous gene fragments and special sequences.
Short-chain single-stranded nucleic acids are annealed to form a cracked intermediate, and then the crack is repaired to form a complete long-chain double-stranded nucleic acid. The annealing process is separated from the synthesis of phosphodiester bonds, and the nucleic acid repair ability of organisms is used for repair and amplification.
It improves the accuracy of nucleic acid synthesis and reduces the mismatch rate, saves synthesis costs and time, is suitable for multi-gene synthesis and gene chip development, and can assess the host's nucleic acid repair capacity.
Smart Images

Figure PCTCN2025133887-FTAPPB-I100001 
Figure PCTCN2025133887-FTAPPB-I100002 
Figure PCTCN2025133887-FTAPPB-I100003
Abstract
Description
A method for preparing long double-stranded nucleic acids Technical Field
[0001] This application relates to the field of biomedicine, specifically to an improved method for de novo synthesis of nucleic acids, and more specifically to a method for synthesizing long double-stranded nucleic acids from short single-stranded nucleic acids. Background Technology
[0002] De novo nucleic acid synthesis, the process of constructing entirely new DNA or RNA molecules from basic nucleotide units, is of great significance to synthetic biology, gene therapy, and biotechnology. It not only provides a powerful tool for studying gene function and regulatory mechanisms but also enables scientists to design and construct synthetic genes and regulatory elements with specific functions, advancing our understanding of the origin and evolution of life. In the medical field, de novo nucleic acid synthesis helps develop personalized medicines and vaccines, providing new strategies for treating hereditary diseases and cancer. Furthermore, in agricultural biotechnology, this technology improves crops by synthesizing genes for specific traits, enhancing crop resistance and yield. De novo nucleic acid synthesis also provides forensic medicine with precise genetic analysis tools, helping to solve genetic evidence problems in criminal investigations. With decreasing synthesis costs and increasing synthesis efficiency, de novo nucleic acid synthesis technology is gradually becoming an important pillar of biological science research and applications, continuously expanding the boundaries of our understanding of life sciences.
[0003] Regarding specific de novo synthesis methods, there are different solutions for short-chain and long-chain nucleic acids. For short-chain nucleic acids, the main approach is to utilize chemical methods using automated nucleic acid synthesizers. For long-chain nucleic acids, the current method primarily involves synthesizing short-chain nucleic acids first and then assembling them into long-chain nucleic acids. Currently used assembly methods mainly employ enzymes, such as polymerase cycle assembly (PCA) and ligase-based assembly (LBA). However, current synthesis methods suffer from problems such as nucleic acid sequence mismatches, which can be partially addressed by introducing high-fidelity enzymes. Nevertheless, we still need to explore more suitable solutions. Summary of the Invention
[0004] This application provides a novel method for synthesizing long double-stranded nucleic acids from short single-stranded nucleic acids.
[0005] As shown in Figure 1, this application adopts a synthetic pathway of short single-stranded nucleic acid - split intermediate - long double-stranded nucleic acid, which enables the efficient and low mismatch rate assembly of short single-stranded nucleic acid (10-60bp or 10-300bp) into long double-stranded nucleic acid (greater than 200bp or greater than 500bp), and further increases the yield of long double-stranded nucleic acid through nucleic acid amplification and other means.
[0006] In the current long-fragment nucleic acid synthesis process, which mainly uses PCA and LBA to synthesize long-chain double-stranded nucleic acids from short-chain single-stranded nucleic acids, the annealing process of the short-chain single-stranded nucleic acids and the synthesis of phosphodiester bonds are carried out in parallel. That is, when the short-chain single-stranded nucleic acids form local reverse complementary pairing regions, nucleic acid polymerization extension (for PCA) or cleft ligation (for LBA) will occur. In this case, once mismatches occur due to the high homology of the short-chain single-stranded nucleic acid segments (i.e., the homology of the reverse complementary regions is less than 100%), the mismatches will be fixed on the product, reducing the accuracy of the long-chain double-stranded nucleic acid sequence. For short-chain single-stranded nucleic acids containing special sequences (self-pairing, high GC, or repetitive sequences), not only may mismatches occur between the short-chain single-stranded nucleic acids, but stem-loop structures as shown in Figure 5B, Figure 8A, and even Figure 22 may also form. Once such stem-loop structures are fixed on the product, they may not only reduce the accuracy of the final product, but may also cause assembly failure. Unlike existing technologies, the method employed in this application separates the annealing process from the synthesis of phosphodiester bonds. Specifically, the short-chain single-stranded nucleic acid is first annealed to form a cracked intermediate, and then the crack is repaired to form a complete long-chain double-stranded nucleic acid. This intermediate provides some buffer space for the mismatched reverse complementary regions, rather than directly fixing the mismatch. In this process, the short-chain single-stranded nucleic acid does not remain unchanged after the formation of the intermediate but rather undergoes a dynamic equilibrium. Therefore, this method allows for a more accurate sequence of the nucleic acid product obtained.
[0007] In a first aspect, this application provides a method for synthesizing long double-stranded nucleic acids from short single-stranded nucleic acids, characterized in that an intermediate is present during the synthesis process. The intermediate is formed by complementary pairing of the short single-stranded nucleic acids. This complementary pairing can be accomplished through annealing.
[0008] In some embodiments, the intermediate is a defective long double-stranded nucleic acid, specifically a defect referring to a gap in the nucleic acid. The gap refers to the absence of a phosphodiester bond between two adjacent nucleotides in the long double-stranded nucleic acid.
[0009] In some embodiments, the method further includes obtaining the complete long double-stranded nucleic acid by repairing the crack.
[0010] In some embodiments, the short single-stranded nucleic acid contains between 10 and 300 nucleotides. Specifically, the short single-stranded nucleic acid contains between 10 and 60 nucleotides or between 60 and 300 nucleotides.
[0011] In some embodiments, the length of the long double-stranded nucleic acid is more than twice that of the short single-stranded nucleic acid, preferably more than five times, and more preferably more than ten times.
[0012] In some embodiments, the intermediate has at least two slits.
[0013] In a preferred embodiment, the short single-stranded nucleic acid should have 100% homology with a portion of the long double-stranded nucleic acid.
[0014] In some embodiments, the short single-stranded nucleic acid is designed and synthesized by the following methods:
[0015] (1) Design a nucleotide sequence containing the long double-stranded nucleic acid;
[0016] (2) Divide each single strand of the double helix of the nucleotide sequence into several short single-stranded nucleic acids; and
[0017] (3) Synthesize the short single-stranded nucleic acid.
[0018] In step (2), the split points in the two chains are staggered from each other, preferably by at least 10 nucleotides.
[0019] In a specific implementation, the short single-stranded nucleic acids obtained after dividing the long double-stranded nucleic acid are named oligo-1, oligo-2, oligo-3, oligo-4, oligo-5, ..., oligo-(2n-1), oligo-2n, where n is a positive integer. Oligo-1, oligo-3, oligo-5, ..., oligo-(2n-1) are located on single strand -1, and oligo-2, oligo-4, oligo-6, ..., oligo-2n are located on single strand -2. Single strand -1 and single strand -2 are two single strands of the long double-stranded nucleic acid, respectively.
[0020] In a specific implementation, for the short single-stranded nucleic acid, oligo-1 corresponds to bases a1 to b1 of single-stranded oligo-1, oligo-2 corresponds to bases a2 to b2 of single-stranded oligo-2, oligo-3 corresponds to bases a3 to b3 of single-stranded oligo-1, ..., oligo-2n corresponds to base a1 to b3 of single-stranded oligo-2. 2n to b 2n The number of bases.
[0021] In this nucleotide sequence, the nucleotides on single strand-1 and single strand-2 are arranged in opposite directions.
[0022] For example, the nucleotide sequence on single strand-1 is from the 5' end to the 3' end of single strand-1, more specifically, the 5' start position on single strand-1 is the first base; and the nucleotide sequence on single strand-2 is from the 3' end to the 5' end of single strand-2, more specifically, the 3' start position on single strand-2 is the first base.
[0023] For example, the nucleotide sequence on single strand-1 is from the 3' end to the 5' end of single strand-1, more specifically, the 3' start position on single strand-1 is the first base; and the nucleotide sequence on single strand-2 is from the 5' end to the 3' end of single strand-2, more specifically, the 5' start position on single strand-2 is the first base.
[0024] In a specific implementation, the single chain-1 has a total of m1 bases, the single chain-2 has a total of m2 bases, and for a1, a2, ..., a... 2n b1, b2, ..., b 2n a1 = a2 = 1; b (2n-1) =m1;b 2n =m2; a3=b1+1, a4=b2+1,…,a 2n =b (2n-2) +1.
[0025] In a specific implementation, a3 and a4, a5 and a6, ..., a 2n With a (2n-1) b1 and b2, b3 and b4, ..., b (2n-3) With b (2n-2) Or a2 and a3, a4 and a5, ..., a (2n-2) With a (2n-1) b2 and b3, b4 and b5, ..., b (2n-2) With b (2n-1) The difference in value represents the number of nucleotides that are misaligned between adjacent dividing points of single strand-1 and single strand-2. In a preferred embodiment, the number of misaligned nucleotides is 10 or more.
[0026] In a specific implementation, a1 and a4, a3 and a6, ..., a 2n With a (2n-3) b1 and b4, b3 and b6, ..., b 2n With b (2n-3) b 2n With b (2n-3) Or a3 and a4, a5 and a6, ..., a (2n-3) With a (2n-2) b1 and b2, b3 and b4, b5 and b6, ..., b (2n-3) With b (2n-2)The difference in value represents the number of nucleotides that are misaligned between adjacent dividing points of single strand-1 and single strand-2. In a preferred embodiment, the number of misaligned nucleotides is 10 or more.
[0027] In a specific embodiment, the aforementioned staggered sequences are overlapping sequences of the two short single-stranded nucleic acids. The arrangement of the overlapping sequences and the complementarity of the overlapping sequences between the two short single-stranded nucleic acids are crucial to the formation of the intermediate. The length of the overlapping sequences also affects the formation of the intermediate. In a preferred embodiment, the length of the overlapping sequences is at least 10, more preferably at least 12, even more preferably at least 14, and even more preferably at least 16.
[0028] In some embodiments, the method of repairing the crack on the intermediate is to form a phosphate diester bond at the crack.
[0029] In one embodiment, when the intermediate is preferably a linear nucleic acid, the repair method includes the following steps:
[0030] i. Attach the intermediate to the carrier;
[0031] ii. Transfer the vector into the host body.
[0032] The cracks in the intermediate are repaired within the host body by the host's built-in repair system.
[0033] In some embodiments, the vector is an expression vector. For example, the expression vector is a plasmid.
[0034] In another embodiment, where the intermediate is preferably a circular nucleic acid, more preferably a plasmid, the repair method includes transferring the plasmid into the host.
[0035] In some embodiments, the host is an organism capable of accommodating a foreign genome, preferably one that can also provide conditions for the amplification of the foreign genome.
[0036] In a preferred embodiment, the host is a model organism. For example, the model organism is *Escherichia coli*. Another example is yeast.
[0037] In some embodiments, the number of slits is 2-50.
[0038] In some embodiments, the number of slits is 51 or more.
[0039] In some embodiments, the method further includes repairing cracks on the intermediate before the vector or plasmid is transferred into the host. In a preferred embodiment, the repair includes the use of a ligase, preferably the use of T4 ligase.
[0040] For example, the repair occurs before the intermediate is attached to the carrier.
[0041] For example, the repair occurs after the intermediate is attached to the carrier.
[0042] On the other hand, this application also relates to the use of the aforementioned synthesis method in the synthesis of long double-stranded nucleic acids.
[0043] In some embodiments, the number of long double-stranded nucleic acids is 1.
[0044] In some embodiments, the number of long double-stranded nucleic acids is greater than or equal to 2.
[0045] In some implementations, there is no upper limit to the number of long double-stranded nucleic acids.
[0046] In some embodiments, the long double-stranded nucleic acid is a linear nucleic acid.
[0047] In some embodiments, the long double-stranded nucleic acid is a circular nucleic acid, preferably a plasmid.
[0048] In some embodiments, the long double-stranded nucleic acid is a nucleic acid with a complex double-stranded structure, preferably a triangular Y-shaped double-stranded structure or a cross-shaped double-stranded structure.
[0049] Many commonly used methods for synthesizing long double-stranded nucleic acids from short single-stranded nucleic acids involve the complementation of overlapping sequences on the short single-stranded nucleic acids. More specifically, this refers to designing short single-stranded nucleic acids with staggered dividing points on their corresponding long double-stranded nucleic acids. In this case, most short single-stranded nucleic acids can pair with at least two other short single-stranded nucleic acids. However, due to the high homology between some gene fragments, this pairing is not always perfectly correct, thus requiring methods to correct these errors. However, some existing synthetic methods, because they ligate or amplify fragments during overlapping sequence complementation, irreversibly leave these errors on the genome, losing the opportunity for correction. This application differs from these methods in that, during the complementary formation of long double-stranded nucleic acids from overlapping sequences, the fragments are not immediately ligated or amplified. Instead, a long double-stranded nucleic acid with a "crack" is formed, i.e., the intermediate. Unlike the final product, the intermediate has a "crack," which provides an opportunity to correct errors.
[0050] The method for synthesizing long double-stranded nucleic acids disclosed in this application has the following advantages:
[0051] (1) It has higher accuracy and lower fragment mismatch rate in gene synthesis, and has great advantages in the application of multi-gene synthesis, especially gene synthesis with high homology.
[0052] (2) No products such as PCR primers are used in the synthesis process, which reduces the synthesis cost to a certain extent.
[0053] (3) Make full use of the organism's nucleic acid repair capacity and amplify nucleic acid while repairing the crack, which simplifies the steps and saves time to a certain extent.
[0054] (4) When synthesizing multiple long double-stranded nucleic acids, the preference is very small or even non-existent, and the proportion of each synthesized long double-stranded nucleic acid in the final product is not much different.
[0055] (5) It can assess the host’s nucleic acid repair capacity and screen out hosts with stronger repair capacity.
[0056] (6) Applying it to gene chips can further promote the development of gene chips.
[0057] Other aspects and advantages of this application will readily be apparent to those skilled in the art from the detailed description below. Only exemplary embodiments of this application are shown and described in the following detailed description. As will be appreciated by those skilled in the art, the content of this application enables them to make modifications to the disclosed specific embodiments without departing from the spirit and scope of the invention to which this application pertains. Accordingly, the descriptions in the accompanying drawings and specification of this application are merely exemplary and not restrictive. Attached Figure Description
[0058] The specific features of the invention involved in this application are shown in the appended claims. The features and advantages of the invention can be better understood by referring to the exemplary embodiments and drawings described in detail below. A brief description of the drawings is as follows:
[0059] Figure 1 shows a flowchart of the synthesis of long double-stranded nucleic acids from short single-stranded nucleic acids.
[0060] Figure 2 shows the difference in the effectiveness of short single-stranded nucleic acids, whether or not they have been purified by PAGE, when assembling into intermediates.
[0061] Figure 3 illustrates two methods for assembling the intermediate and the carrier according to this application. Figure 3A shows the Gibson assembly method; Figure 3B shows the chain substitution assembly method.
[0062] Figure 4 shows a more detailed schematic diagram of the Gibson assembly, in which the sequence of the intermediate is derived from the coding sequence of green fluorescent protein (GFP).
[0063] Figure 5 illustrates the different behaviors of different types of short single-stranded nucleic acids in the synthesis of long double-stranded nucleic acids. Figure 5A shows the flowchart of the synthesis and verification of long double-stranded nucleic acids from short single-stranded nucleic acids; Figure 5B shows the synthesis process of four different types of short single-stranded nucleic acids; Figure 5C shows the agarose gel electrophoresis results of the intermediates synthesized from the four different types of short single-stranded nucleic acids; Figure 5D shows the length and corresponding synthesis yield of the long double-stranded nucleic acids synthesized from the four different types of short single-stranded nucleic acids.
[0064] Figure 6 shows the results of fluorescence detection of colonies on the plate in Example 5.
[0065] Figure 7 shows the agarose gel electrophoresis results of the intermediates assembled from short single-stranded nucleic acids with different overlapping sequence lengths in Example 6.
[0066] Figure 8 shows a schematic diagram of the assembly of short single-stranded nucleic acids containing special fragments into intermediates. Figure 8A shows the circularity analysis (left) and predicted self-pairing structures (right) of nucleic acids containing self-pairing hairpin structures in Example 7 performed on mfold.org; Figure 8B shows a schematic diagram of the partitioning of tandem repeat sequences; Figure 8C shows the distribution of the number of repeats and the corresponding number of clones in the repeat sequence sequencing results, where 6 repeats is the correct number of repeats.
[0067] Figure 9 shows an agarose gel electrophoresis image of the product obtained by assembling short single-stranded nucleic acids containing a specific sequence using the PCA method in Example 7.
[0068] Figure 10 shows the performance of short single-stranded nucleic acids in Example 8 when synthesizing two long double-stranded nucleic acids of different lengths. Figure 10A shows a schematic diagram of the formation of intermediates from short single-stranded nucleic acids separated from long double-stranded nucleic acids of different lengths (gfp and bacterial rhodopsin); Figure 10B shows the agarose gel electrophoresis results of the intermediates formed after annealing and assembling the short single-stranded nucleic acids of gfp and bacterial rhodopsin in the same reaction; Figure 10C shows a statistical graph of the sequencing results of the assembled and amplified synthesized intermediates and sampled for sequencing.
[0069] Figure 11 shows the performance of short single-stranded nucleic acids in Example 8 when synthesizing two long double-stranded nucleic acids of the same length and high homology. Figure 11A shows a schematic diagram of the formation of intermediates from short single-stranded nucleic acids divided on homologous long double-stranded nucleic acids (nr2f1 and nr2f2); Figure 11B shows the number of identical bases on corresponding homologous fragments in the two fragments nr2f1 and nr2f2, with the homologous fragments being 60 nt in length. Figure 11C shows the agarose gel electrophoresis results of the intermediates formed after annealing and assembling the short single-stranded nucleic acids of nr2f1 and nr2f2 in the same reaction; Figure 11D shows the sequencing results of the synthesized intermediates after assembly, amplification, and sequencing.
[0070] Figure 12 shows the results of assembling homologous genes nr2f1 and nr2f2 using the PCA method in Example 8. Figure 12A shows the agarose gel electrophoresis images of the products obtained from synthesizing homologous genes nr2f1 and nr2f2 in the same reaction; Figure 12B shows the sequence alignment results of the final products with the standard sequences of nr2f1 and nr2f2, respectively.
[0071] Figure 13 shows the length distribution of the 10 genes in Example 9 and the agarose gel electrophoresis diagram of the intermediates assembled in one reaction after filling the 10 genes to 1170-bp (Figure 13A) and 1050-bp (Figure 13B).
[0072] Figure 14 shows the length distribution of 100 genes in Example 9.
[0073] Figure 15 shows the results of synthesizing 100 genes in Example 9. Figure 15A shows the expected synthesis process of the 100 genes; Figure 15B shows the electrophoresis results of the intermediates; Figure 15C shows the number of genes synthesized with error-free sequences and the number of genes synthesized with mutated sequences among the 100 genes in the sequencing results, wherein the mutated sequences should be sequences with high homology to the standard sequences but containing a few mutations; Figure 15D shows the frequency of different genes appearing in the sequences with all correct sequencing results.
[0074] Figure 16 shows the agarose gel electrophoresis results of the intermediates when synthesizing the 7 genes that were not successfully synthesized out of 100 genes in Example 9 individually.
[0075] Figure 17A shows a schematic diagram of the two-step method for synthesizing plasmids; Figure 17B shows the electrophoresis diagrams of the assembly products in the two reaction chambers of the two-step method and the intermediates synthesized in the one-step / two-step methods respectively during the synthesis of a 1200-bp small plasmid; Figure 17C shows the electrophoresis diagrams of the assembly products in the two reaction chambers during the two-step method for synthesizing a small plasmid with GFP; Figure 17D shows the electrophoresis diagrams of the assembly products in the two reaction chambers during the two-step method for synthesizing the pUC19 plasmid.
[0076] Figure 18A shows a schematic diagram of plasmid synthesis using the MOSAIC method; Figure 18B shows the results of verifying the 1200-bp small plasmid through enzyme digestion; Figure 18C shows the sampling and sequencing results after the synthesis and assembly of the 1200-bp small plasmid, the small plasmid with inserted gfp, and the pUC19 plasmid; Figure 18D shows the electrophoresis results and a schematic diagram of the enzyme digestion sites after enzyme digestion of randomly selected yeast colonies.
[0077] Figure 19 shows the correlation between gene length, GC content on short single-stranded nucleic acids, and the error-free rate of the MOSAIC synthesis method. Figure 19A shows the correlation between gene length and the error-free rate of the MOSAIC synthesis method; Figures 19B-19E show the correlation between GC content on short single-stranded nucleic acids and the error-free rate of the MOSAIC synthesis method.
[0078] Figure 20 shows the chain displacement assembly method for assembling the intermediate and the carrier in this application.
[0079] Figure 21 shows the positions of the five primer pairs designed for the synthesized 3.3kb plasmid in Example 10 on the plasmid and the corresponding detection results.
[0080] Figure 22 shows an exemplary schematic diagram of the 16S rRNA synthesized in Example 13.
[0081] Figure 23 shows the detection pattern of the long double-stranded nucleic acid containing the specific sequence synthesized in Example 13. Figure 23A shows the electrophoresis pattern of the synthesized intermediate containing 100% AT. Figure 23B shows the electrophoresis pattern of the long double-stranded nucleic acid fragment containing the specific sequence synthesized using the PCA method; the left image shows the electrophoresis pattern of synthesized 100% AT and 100% GC, the middle image shows the electrophoresis pattern of synthesized DNA encoding rRNA, and the right image shows the electrophoresis pattern of encoding repetitive sequences.
[0082] Figure 24 shows the length distribution of human genes and the assembly diagram and electrophoresis diagram of intermediates containing special sequences. Figure 24A shows the proportion of human genes with lengths below 1200 bp and below 3000 bp, respectively. Figure 24B shows the assembly diagram and corresponding electrophoresis diagram of intermediates containing self-paired sequences, high GC sequences, and repetitive sequences.
[0083] Figure 25 shows the electrophoresis diagram of intermediates synthesized from short single-stranded nucleic acids of different lengths in Example 14.
[0084] Figure 26 shows the effects of the length of the synthesized sequence, the number of cleavages in the intermediate, and the length distribution of the human gene in Example 15. Figure 26A shows the electrophoresis diagram of the intermediate when synthesizing long double-stranded nucleic acids of 300-3000 bp using the MOSAIC method. Figure 26B shows the electrophoresis diagram of the intermediate when synthesizing long double-stranded nucleic acids of 9900 bp using the MOSAIC method. Figure 26C shows the ability of intermediates containing different numbers of cleavages to form long double-stranded nucleic acids through host repair function after treatment with or without ligase. Figure 26D shows the length distribution of the human gene.
[0085] Figure 27 shows the detection results of the 9.9kb plasmid synthesized in Example 15. Figure 27A shows the distribution of the seven pairs of detection primers used in Example 15 on the plasmid, the electrophoresis diagrams of nucleic acid fragments amplified by different primers, and the assembly efficiency of the clones formed by the 9.9kb plasmid. Figure 27B shows the number and distribution of mutation types contained in three randomly selected successfully synthesized 9.9kb plasmids after sequencing.
[0086] Figure 28 shows a schematic diagram of multi-gene synthesis and the detection results. Figure 28A shows a schematic diagram of multi-gene synthesis. Figures 28B-28D show the results of parallel synthesis of 100 genes. Figure 28B shows the electrophoresis diagram of the synthesized intermediates. Figure 28C shows the sequencing results of synthesized genes with and without barcodes. Figure 28D shows the logarithmic plot of the number of reads for each gene in the barcode-based case. Figures 28E-28G show the results of parallel synthesis of 1793 genes. Figure 28E shows the electrophoresis diagram of the synthesized intermediates. Figure 28F shows the percentage of synthesized genes without mutations, with fewer than 10 mutations, and without sequences in the total synthesized genes after sequencing. Figure 28G shows the number of error-free reads between different fragments.
[0087] Figure 29 shows a simplified flowchart of NGS sequencing of genes containing barcodes in a 96-well plate. Detailed Implementation
[0088] The following specific embodiments illustrate the implementation of the invention. Those skilled in the art can easily understand other advantages and effects of the invention from the content disclosed in this specification.
[0089] Terminology Definition
[0090] Unless otherwise defined in this application, the scientific and technical terms used herein shall have the meanings commonly understood by one of ordinary skill in the art. Furthermore, unless the context requires otherwise, singular terms shall include plural terms, and plural terms shall include singular terms. Generally, the cell and tissue culture, molecular biology, immunology, microbiology, genetics, and proteins and nucleic acids described herein are concepts recognized by those skilled in the art.
[0091] In this application, the term "nucleic acid" can be understood both macroscopically and microscopically. Macroscopically, nucleic acid can be understood as nucleic acid molecules and mixtures containing nucleic acid molecules. The mixture may contain solvents (e.g., enzyme-free water), other necessary components (e.g., additives to stabilize the nucleic acid structure, adjuvants to aid subsequent reactions), and other possible impurities. Microscopically, nucleic acid is primarily understood as the molecular structure of nucleic acids. Nucleic acids are mainly composed of nucleotide chains, nucleotide rings, or other possible structures formed by the dehydration condensation of nucleotides to form phosphodiester bonds, where the phosphodiester bonds exist between the phosphate group and the pentose group of the nucleotide. In this application, the term "pentose phosphate chain" refers to a portion of the nucleic acid containing the phosphodiester bonds. The pentose phosphate chain contains the phosphate group, pentose group, and necessary linking structures (such as covalent bonds) of the nucleic acid, but does not contain the bases of the nucleic acid. Since the bases cannot form covalent bonds with each other, but only form hydrogen bonds with their opposite bases when the nucleic acid is double-stranded, the region where the base is located can be referred to as the base side.
[0092] In this application, the term "nucleotide sequence" refers to the sequence of nucleotides that make up a nucleic acid molecule (including DNA and RNA molecules). This sequence is typically represented by the base sequence on one side of the nucleic acid. The starting point of the nucleotide sequence can vary depending on the context. For multiple nucleic acids that are 100% homologous to the same double-stranded nucleic acid, their orientation can be determined from the context.
[0093] In this application, the term "length" generally refers to the length of nucleic acids, including short-chain nucleic acids, long-chain nucleic acids, and single-chain and double-chain nucleic acids. The length can be expressed as the number of bases in a single-chain nucleic acid; or as the total number of bases or base pairs between the two nucleotides at both ends of a double-chain nucleic acid. The two nucleotides at both ends can be on the same single strand or on different single strands. The double-chain nucleic acid may or may not contain dangling sequences. A dangling sequence refers to a segment on the double strand that lacks complementary pairing; it is called a dangling sequence because it usually extends beyond the double strand and is located at the end of the nucleic acid. In this application, the unit of length can be bp, nt, or mer / mers, where bp represents the number of base pairs in the complementary region of the double-chain nucleic acid; nt and mer / mers represent the number of bases in the single-chain nucleic acid. A base pair refers to a pairing formed by a one-to-one correspondence between bases in regions of high homology between the two single strands of a double-chain nucleic acid. Within the base pairs, the pairing between bases can be completely complementary, or a base on one chain can pair with a position on another chain where there is no base, or both of the above can coexist.
[0094] In this application, the terms "short-chain nucleic acid" and "long-chain nucleic acid" are generally distinguished by the length of the nucleic acid.
[0095] In this application, the term "short-chain nucleic acid" generally refers to a nucleic acid molecule that can be synthesized using only existing and potentially future de novo synthesis methods in the art. The de novo synthesis method refers to a nucleic acid synthesis method starting from a single nucleotide, currently mainly including phosphoramide chemical synthesis, photochemical synthesis, electrochemical synthesis, inkjet printing, and other de novo synthesis methods. The maximum length of the short-chain nucleic acid should be the limit of nucleic acids with high fidelity that can be synthesized by existing and potentially future de novo synthesis methods in the art. In some embodiments, the maximum length may be 300 bp. The minimum length of the short-chain nucleic acid should be the minimum length at which it can stably exist; when the short-chain nucleic acid is shorter than the minimum length, it will be subject to significant degradation or other potential risks. In some embodiments, the minimum length may be 10 bp.
[0096] In this application, the term "long nucleic acid" generally refers to a nucleic acid molecule that cannot be synthesized solely by de novo synthesis or whose synthesis by de novo synthesis is costly. The long nucleic acid is typically assembled from the short nucleic acid molecules. The minimum length of the long nucleic acid is generally greater than, but may be less than, the maximum length of nucleic acid molecules that can be synthesized by the de novo synthesis method. In some embodiments, the minimum length may be 300 bp, 100 bp, or 200 bp. In some embodiments, the minimum length should be greater than 60 bp.
[0097] In this application, the terms "double-stranded nucleic acid" and "single-stranded nucleic acid" are generally distinguished based on the structure of the nucleic acid.
[0098] In this application, the term "single-stranded nucleic acid" should be interpreted broadly, generally referring to a nucleic acid molecule from start to finish without any other nucleic acid strands. It is worth noting that single-stranded nucleic acids can also form reverse complementary structures, such as stem-loop structures, in regions of high homology, but this should not be used to classify them as double-stranded nucleic acids. The single-stranded nucleic acid can be a linear nucleic acid, a circular nucleic acid, a single-stranded nucleic acid containing a stem-loop structure, or any other possible structure.
[0099] In this application, the term "double-stranded nucleic acid" generally refers to a nucleic acid composed of two or more single-stranded nucleic acids. The double-stranded nucleic acid typically includes regions with high homology capable of forming reverse complementary regions. The term "double-stranded nucleic acid" should be interpreted broadly; for example, it is not limited to linear double-stranded nucleic acids composed of two reverse complementary strands, but may also include circular double-stranded nucleic acids composed of two reverse complementary strands, and may also include more complex nucleic acid structures containing double-stranded nucleic acids. Examples of complex nucleic acid structures may be, for example, a trifurcate Y-shaped double-stranded structure formed by the pairwise complementarity of three DNA single strands, or a cross-shaped double-stranded structure formed by the pairwise complementarity of four DNA single strands. Within the reverse complementary regions of the double-stranded nucleic acid, all bases may or may not have corresponding complementary bases; that is, homology may be 100% or less than 100%. The homology can refer to the ratio of complementary base pairs within the entire region to the total theoretical number of bases on a single strand within that region (including non-existent bases that pair with bases on another strand). The theoretical number of bases includes the actual number of bases on the single strand and the number of bases skipped due to satisfying a sufficient number of anticomplementary pairs (i.e., the aforementioned "non-existent bases"). Homology can be calculated manually or with computer assistance. An example of computer-aided calculation is the BLAST system. The double-stranded nucleic acid may contain only the region forming anticomplementary pairs, or it may contain structures other than the anticomplementary regions, such as dangling sequences. A dangling sequence refers to a nucleotide chain extending from the anticomplementary region.
[0100] In this application, the term "short-chain single-stranded nucleic acid" refers to a nucleic acid or nucleic acid molecule that meets both the characteristics of a short-chain nucleic acid and a single-stranded nucleic acid. The terms "short-chain single-stranded nucleic acid," "oligonucleotide," and "oligonucleic acid" are used interchangeably. The term "long-chain double-stranded nucleic acid" refers to a nucleic acid or nucleic acid molecule that meets both the characteristics of a long-chain nucleic acid and a double-stranded nucleic acid. It should be noted that the number of long-chain double-stranded nucleic acids in this application refers to the number of long-chain double-stranded nucleic acids with different sequences; that is, long-chain double-stranded nucleic acids with the same sequence should not be considered as different long-chain double-stranded nucleic acids. It is worth noting that if the sequence difference is due to mutations during synthesis, whether it should be considered as different long-chain double-stranded nucleic acids should be determined based on the specific context.
[0101] In this application, the term "crack" generally refers to a structure present on the pentose phosphate chain of a nucleic acid that causes a discontinuity in the nucleic acid chain. The crack may be formed due to the absence of a phosphodiester bond.
[0102] In this application, the term "intermediate" generally refers to a defective long double-stranded nucleic acid that occurs during the synthesis of the long double-stranded nucleic acid from the short single-stranded nucleic acid. Specifically, compared to the long double-stranded nucleic acid, the intermediate has a cleft in the pentose phosphate chain of the nucleic acid, preferably at the interface of the short single-stranded nucleic acid. The short single-stranded nucleic acid can be formed into the intermediate through an annealing process well known in the art, or other possible methods. In the context of this application, the term "intermediate" should be interpreted broadly, meaning that any long double-stranded nucleic acid containing a cleft can be an intermediate. Specific examples include intermediates formed by annealing the short single-stranded nucleic acid, intermediates formed by annealing the short single-stranded nucleic acid and then treating them with a ligase to ligate part of the cleft, intermediates not connected to a vector, intermediates connected to a vector, intermediates containing a pendant sequence, and intermediates not containing a pendant sequence.
[0103] In this application, the term "misaligned" generally refers to the presence of other sequences (such as dangling sequences) between two short single-stranded nucleic acids that have anticomplementary regions, in addition to the anticomplementary regions. These other sequences may or may not be anticomplementary to the other short single-stranded nucleic acid. Furthermore, on each pair of anticomplementary short single-stranded nucleic acids, one short single-stranded nucleic acid may contain the other sequence, or both short single-stranded nucleic acids may contain the other sequence. In the intermediate, after the misaligned sequences form an anticomplementary double strand, they can be referred to as overlapping sequences.
[0104] In this application, the terms "linear nucleic acid" and "circular nucleic acid" are generally distinguished by the presence or absence of linked ends. "Linear nucleic acid" typically refers to nucleic acid without linked ends. "Circular nucleic acid" refers to nucleic acid containing linked ends, which can be composed of a start and an end point on the nucleic acid, or any two points on the nucleic acid. The term "linked" should be interpreted in a narrower sense, meaning linked by covalent bonds on the pentose phosphate chain of the nucleic acid.
[0105] In this application, the term "vector" or "expression vector" generally refers to any molecule used to transfer nucleic acid information to a host (e.g., plasmids or viruses, virus-like particles, polycations, peptide vectors, liposomes, and / or hybridization vectors). The term "vector" includes nucleic acid molecules capable of transporting another nucleic acid linked to them. One type of vector is a "plasmid," which refers to a circular double-stranded DNA molecule into which a target DNA fragment can be inserted. Another type of vector is a viral vector, in which a target DNA fragment can be inserted into a viral genome. Some vectors are capable of autonomous replication in the host cells to which they are introduced (e.g., bacterial vectors with bacterial origins of replication and augmented mammalian vectors). Generally, expression vectors useful in recombinant nucleic acid technologies are typically in the form of plasmids. The terms "plasmid" and "vector" are used interchangeably herein because plasmids are the most commonly used form of vector. However, the disclosure of this application may include other forms of expression vectors, such as viral vectors (e.g., replication-defective retroviruses, adenoviruses, and adeno-associated viruses), which have equivalent functions.
[0106] In this application, the term "and / or" should be understood as any one of the multiple elements used for connection or any combination of elements.
[0107] In this application, the term "comprising" or "including" generally means including the explicitly specified features, but does not exclude other elements.
[0108] In this application, the term “selected from” generally refers to the selection of objects and all combinations thereof. For example, “selected from A, B and C” means all combinations of A, B and C, such as A, B, C, A+B, A+C, B+C, or A+B+C.
[0109] In this application, the term "about" generally refers to a variation within a range of 0.5% to 10% above or below a specified value, such as a variation within a range of 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10% above or below a specified value.
[0110] Invention Details
[0111] In this application, since the short single-stranded nucleic acid synthesis intermediates are assembled through complementary pairing between short single-stranded nucleic acids, this method is named Molecular Self-Assembly Induced Cloning (MOSAIC).
[0112] intermediate
[0113] The most important feature of the synthesis method involved in this application is the formation of the intermediate.
[0114] The intermediate refers to a defective intermediate state formed during the synthesis of a long double-stranded nucleic acid from the short single-stranded nucleic acid, and is therefore called an intermediate. Except for the lack of phosphodiester bonds (creating a gap) between adjacent nucleotides, the intermediate can be considered to have the same characteristics as the long double-stranded nucleic acid. It should be noted that the lack of phosphodiester bonds between adjacent nucleotides does not mean that every pair of adjacent nucleotides is missing, but only a portion of the adjacent nucleotides, and a very small portion at that. Preferably, the number of adjacent nucleotides lacking phosphodiester bonds is 10% or less of the total nucleotides.
[0115] The intermediate is typically formed by annealing the short single-stranded nucleic acids that make up the long double-stranded nucleic acid in the same mixing system. In biology, especially in nucleic acid synthesis, annealing generally refers to a process used to pair two complementary sequences, including heating and cooling.
[0116] During annealing, the purpose of heating is to break all hydrogen bonds between bases, causing all short-chain single-stranded nucleic acids in the system to unfold; this process is called denaturation. Heating temperatures can reach approximately 80°C and above, approximately 85°C and above, approximately 90°C and above, approximately 91°C and above, approximately 92°C and above, approximately 93°C and above, approximately 94°C and above, approximately 95°C and above, approximately 96°C and above, approximately 97°C and above, approximately 98°C and above, approximately 99°C and above, and approximately 100°C and above.
[0117] After denaturation breaks the hydrogen bonds between bases, the system needs to be cooled. The purpose of this process is to allow different short-chain single-stranded nucleic acids to form inter-base hydrogen bonds, thus enabling the dispersed short-chain single-stranded nucleic acids to form a cohesive whole through these bonds. During cooling, due to complementary overlapping sequences between the short-chain single-stranded nucleic acids, the overlapping sequences between the denatured and unfolded short-chain single-stranded nucleic acids will form inter-base hydrogen bonds with their corresponding overlapping sequences. Simultaneously, in the inventors' design, most short-chain single-stranded nucleic acids have overlapping sequences with at least two other short-chain single-stranded nucleic acids. That is, if only two short-chain single-stranded nucleic acids are complementary, a dangling sequence will be formed. This dangling sequence usually overlaps with another short-chain single-stranded nucleic acid, thus forming an inter-base hydrogen bond with that other short-chain single-stranded nucleic acid. This process gradually increases the number of complementary base pairs. This process repeats until the resulting product has no dangling sequences or the length of the dangling sequences is insufficient to support a stable whole with another short-chain single-stranded nucleic acid. This process may or may not involve the breaking of other chemical bonds, but it typically does not involve the breaking of phosphodiester bonds. The resulting short-chain single-stranded nucleic acid, formed by hydrogen bonds, is the intermediate described in this application. This process can be slow cooling or rapid cooling. More specifically, it can be a cooling rate of approximately 0.01°C per second, approximately 0.1°C per second, approximately 1°C per second, approximately 2°C per second, approximately 3°C per second, approximately 4°C per second, or approximately 5°C per second. The final cooling temperature can be a relatively high temperature, such as approximately 40-60°C (e.g., approximately 40°C, approximately 41°C, approximately 42°C, approximately 43°C, approximately 44°C, approximately 45°C, approximately 46°C, approximately 47°C, approximately 48°C, approximately 49°C, approximately 50°C, approximately 51°C, approximately 52°C, approximately 53°C, approximately 54°C, approximately 55°C, approximately 56°C, approximately 57°C, approximately 58°C, approximately 59°C, or approximately 60°C); or it can be a relatively low temperature. For example, approximately 15-40°C (e.g., approximately 15°C, approximately 16°C, approximately 17°C, approximately 18°C, approximately 19°C, approximately 20°C, approximately 21°C, approximately 22°C, approximately 23°C, approximately 24°C, approximately 25°C, approximately 26°C, approximately 27°C, approximately 28°C, approximately 29°C, approximately 30°C, approximately 31°C, approximately 32°C, approximately 33°C, approximately 34°C, approximately 35°C, approximately 36°C, approximately 37°C, approximately 38°C, approximately 39°C, approximately 40°C).
[0118] In some embodiments, the annealing process may further include a pre-cooling process between heating and cooling, wherein the final temperature of the pre-cooling may be lower than the heating temperature and higher than the final cooling temperature.
[0119] During the formation of the intermediate, since it only involves the formation of hydrogen bonds, and hydrogen bonds are non-covalent bonds, they are less stable than covalent bonds and are easily broken by external forces or dynamically changed. Furthermore, the fewer hydrogen bonds a component has, the less stable it is. Therefore, the intermediate connected by hydrogen bonds can also be an unstable structure, especially those components where the homology between overlapping sequences is less than 100%. In overlapping sequences with less than 100% homology, not every pair of corresponding bases has a hydrogen bond. Therefore, hydrogen bonds are relatively fewer compared to overlapping sequences with 100% homology, and the structure is relatively unstable. These overlapping sequences are more prone to hydrogen bond breakage, leading to unpairing of the overlapping sequences. In this case, the two nucleic acids formed by the breakage have dangling sequences that can be complementary to other nucleic acids again. These dangling sequences can potentially complement nucleic acids containing 100% homology to form intermediates. Alternatively, they may continue to complement nucleic acids containing less than 100% homology to form intermediates. In this case, the intermediate is more likely to repeat the previous hydrogen bond breakage process until all overlapping sequences in the intermediate have 100% homology or the break in the intermediate is repaired. Through this process of hydrogen bond breakage and reconnection, the intermediate can spontaneously form a high-fidelity nucleotide sequence.
[0120] The intermediate can be considered as a long double-stranded nucleic acid lacking a partial phosphodiester bond, which is typically located at the interface between the short single-stranded nucleic acid and another adjacent short single-stranded nucleic acid on the same single strand. Therefore, the intermediate should be considered to have approximately the same length and molecular weight as the long double-stranded nucleic acid. Thus, methods characterizing nucleic acid length or molecular weight can be used to detect the formation of the intermediate. The method for characterizing nucleic acid length can be selected from one or more of the following: gel electrophoresis, real-time quantitative PCR, ultracentrifugation, capillary electrophoresis, nanoparticle tracking analysis, nucleic acid sequencing, spectroscopic methods, mass spectrometry, atomic force microscopy, optical microscopy, flow cytometry, nuclear magnetic resonance, and electron microscopy, preferably gel electrophoresis. The gel electrophoresis can be agarose gel electrophoresis or polyacrylamide gel electrophoresis.
[0121] In some embodiments, the intermediate for the same long double-stranded nucleic acid can be one, two, or more. In a preferred embodiment, the presence of two or more intermediates occurs when the long double-stranded nucleic acid is a circular nucleic acid or a long linear nucleic acid.
[0122] The intermediate is usually present in the form of a mixture after its formation. The mixture may contain unreacted short single-stranded nucleic acids, or it may contain any intermediate state from which the intermediate is formed. Therefore, the mixture can be separated and purified to obtain a purer intermediate, or it can be directly proceeded to the downstream steps without separation and purification.
[0123] The separation method can be based on the following properties of nucleic acids: molecular weight, charge, density, specific binding to a specific ligand, and isoelectric point, preferably molecular weight. The molecular weight-based separation method can be selected from one or more of the following: gel electrophoresis and mass spectrometry, preferably gel electrophoresis. The gel electrophoresis can be agarose gel electrophoresis or polyacrylamide gel electrophoresis, preferably agarose gel electrophoresis.
[0124] The purification method is typically performed after the separation and generally includes: (1) separating bands containing the same molecular weight as the intermediate using an instrument (e.g., an instrument emitting ultraviolet light); (2) removing the electrophoretic medium from the system; and (3) removing other impurities. In (1), the separation method can be physical separation, such as cutting with a knife, or chemical separation, or biological separation; in (2), the method for removing the electrophoretic medium can be heating to dissolve, enzymatic digestion, or transferring the target band to another material before purification; in (3), the removal method can be using a filter column or adsorption column, or adding a washing solution. The purification method can be performed using a commercial kit or by performing the purification in-house.
[0125] The evaluation criteria for the intermediates are also related to the synthetic methods involved in this application. The evaluation criteria for the intermediates can be yield, which is the ratio of the mass of the intermediate obtained after isolation and / or purification to the total mass of the short single-stranded nucleic acid from which the synthesis began. The evaluation criteria for the intermediates can also be accuracy, which refers to the percentage of the final synthesized product that completely conforms to the expected sequence. This completely conforming product can be a product that is completely identical to the long double-stranded nucleic acid sequence.
[0126] Short single-stranded nucleic acids
[0127] Short single-stranded nucleic acids are the smallest units that make up intermediates or long double-stranded nucleic acids in this application. The short single-stranded nucleic acid that makes up the long double-stranded nucleic acid should be completely identical in sequence to a certain segment of the long double-stranded nucleic acid to be synthesized.
[0128] The short single-stranded nucleic acid can be designed by dividing the sequence of the long double-stranded nucleic acid to be synthesized into an appropriate number of parts, and the division method should meet the following conditions.
[0129] (1) Except for the short single-stranded nucleic acids at one end of a single strand and the other end of the opposing single strand, which may have overlapping sequences with only one other short single-stranded nucleic acid, the remaining short single-stranded nucleic acids should have overlapping sequences with at least two other short single-stranded nucleic acids besides themselves. The number of overlapping sequences or the number of hydrogen bonds formed should maintain a relatively stable complementary pairing state. Preferably, the overlapping sequence is greater than or equal to 10 base pairs or the number of hydrogen bonds formed is greater than or equal to 20. More preferably, the overlapping sequence is greater than or equal to 16 base pairs or the number of hydrogen bonds formed is greater than or equal to 32.
[0130] (2) All short single-stranded nucleic acids may be of the same length or of different lengths. The length of all short single-stranded nucleic acids should be greater than or equal to the overlapping sequence of one short single-stranded nucleic acid or greater than or equal to the sum of the overlapping sequences of one short single-stranded nucleic acid and the other short single-stranded nucleic acids, preferably equal to the overlapping sequence of one short single-stranded nucleic acid or the sum of the overlapping sequences of one short single-stranded nucleic acid and the other short single-stranded nucleic acids.
[0131] (3) The length of the short single-stranded nucleic acid can be between 10 and 300 nucleotides, preferably between 20 and 300 nucleotides, more preferably between 30 and 300 nucleotides, even more preferably between 40 and 300 nucleotides, and even more preferably between 54 and 300 nucleotides, and even more preferably between 60 and 300 nucleotides. The short single-stranded nucleic acid at one end of one single strand and the other end of the opposing single strand should be greater than or equal to 10 nucleotides, and the remaining short single-stranded nucleic acids should be greater than or equal to 20 nucleotides.
[0132] The long double-stranded nucleic acid composed of the short single-stranded nucleic acid can be a nucleic acid with a classic linear double-stranded structure or a non-classical structure, such as a triangular Y-shaped or cross-shaped structure. In some embodiments, the long double-stranded nucleic acid is a long double-stranded nucleic acid with a classic linear structure, and the short single-stranded nucleic acid may have overlapping sequences with one or two other short single-stranded nucleic acids. In some embodiments, the long double-stranded nucleic acid is a triangular Y-shaped or cross-shaped structure, and the short single-stranded nucleic acid may have overlapping sequences with one, two, three, or even four or more other short single-stranded nucleic acids.
[0133] The short single-stranded nucleic acid may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The short single-stranded nucleic acid may or may not contain modifications, and examples of modifications may be selected from one or more of the following: methylation, phosphorylation, acetylation, adenylation, uridineation, glycosylation, sulfur modification, pseudouracilization, and ribourethridineization.
[0134] The short single-stranded nucleic acid may contain classical nucleotides that make up an organism. For example, the classical nucleotides may be selected from one or more of the following bases: adenine, guanine, cytosine, thymine, and uracil. The short single-stranded nucleic acid may contain non-classical nucleotides. For example, the non-classical nucleotides may be selected from one or more of the following bases: methylcytosine (e.g., 5-methylcytosine), methyladenine (e.g., N6-methyladenine), pseudouracil, ribouracil, xanthine, and hypoxanthine.
[0135] The short single-stranded nucleic acid can be a conventional short single-stranded nucleic acid or a short single-stranded nucleic acid with a special structure. For example, the short single-stranded nucleic acid can contain sequences that are complementary to other sequences on the same short single-stranded nucleic acid, called self-pairing sequences, and the short single-stranded nucleic acid containing self-pairing sequences can form a stem-loop structure. The content of guanine and cytosine on the short single-stranded nucleic acid can be arbitrary, for example, 0-about 10%, about 10%-about 20%, about 20%-about 30%, about 30%-about 40%, about 40%-about 50%, about 50%-about 60%, about 60%-about 70%, about 70%-about 80%, about 80%-about 90%, about 90%-about 100%.
[0136] The short single-stranded nucleic acid can be synthesized de novo using a single nucleotide as a starting material after designing the target sequence. The synthesis method can be selected from one or more of the following: solid-phase phosphoramidite method, photochemical deprotection method, electrochemical deprotection method, inkjet printing method, integrated circuit control method, high-throughput parallel synthesis based on sorting principle, enzymatic synthesis method, mixed enzyme-mediated enzymatic reaction, TdT-dNTP cross-linked enzymatic reaction, and dibasic monomer synthesis method. The short single-stranded nucleic acid can also be obtained by finding an existing nucleic acid fragment containing the target sequence and obtaining the fragment through a truncation method, which can be selected from one or more of the following: restriction endonuclease digestion method, PCR method, gene gun method, and magnetic bead capture method. The short single-stranded nucleic acid can also be synthesized using a single nucleic acid and nucleotide that are 100% homologous to the target sequence but shorter than the target sequence as starting material, and the synthesis method can refer to the above-described de novo synthesis methods. The short single-stranded nucleic acid can also be synthesized using two or more nucleic acids that are 100% homologous to the target sequence but shorter than the target sequence as starting material.
[0137] After synthesis, the short single-stranded nucleic acid can be purified before downstream steps, or it can be directly proceeded to downstream steps without purification. The method for purifying the short single-stranded nucleic acid can be selected from one or more of the following: high-performance affinity adsorption (HAP), polyacrylamide gel electrophoresis (PAGE), enhanced polyacrylamide gel electrophoresis (ULTRAPAGE), high-performance liquid chromatography (HPLC), high-performance liquid chromatography combined with capillary electrophoresis (HPLC-CE), reversed-phase purification filter cartridge (RPC), specific adsorption of whole sequence (ePAGE), oligonucleotide purification column (OPC), reversed-phase chromatography (RP-HPLC), ion exchange chromatography (IE-HPLC), with polyacrylamide gel electrophoresis (PAGE) being the preferred method.
[0138] Long double-stranded nucleic acid
[0139] In this application, the intermediate is repaired to form the long double-stranded nucleic acid, which can be accomplished by forming phosphodiester bonds between the nucleotides on both sides of the cleavage on the intermediate. The phosphodiester bonds can be formed by enzymatic reactions (e.g., polymerases, ligases, etc.), or not by enzymatic reactions, or by in vivo nucleic acid repair, or by other possible methods, preferably by in vivo nucleic acid repair and / or enzymatic reactions.
[0140] The in vivo nucleic acid repair refers to introducing the intermediate into a host with phosphodiester bond synthesis function, and using the host's built-in repair system to repair the cracks on the intermediate. The host may contain cells in the division phase. The host may be a multicellular organism or an in vitro cell line derived from a multicellular organism, or a single-celled organism, preferably a single-celled organism. The single-celled organism may be a single-celled animal, a single-celled plant, or a single-celled microorganism, preferably a single-celled microorganism. The single-celled microorganism may be selected from one or more of the following: *Escherichia coli*, *Radiata-resistant Cocci*, *Bacillus subtilis*, thermophilic archaea, *Saccharomyces cerevisiae*, and *Bacillus buddingus*, preferably *Escherichia coli* and *Bacillus buddingus*.
[0141] In the nucleic acid repair method, the intermediate can be directly transferred into the host, or the intermediate can be first linked to a vector before being transferred into the host. The vector can be a biological or non-biological organism that helps the intermediate repair phosphodiester bonds within the host or helps the intermediate enter the host; preferably, it is an expression vector. The expression vector can be selected from one or more of the following: plasmids, bacteriophages, artificial chromosomes, transposons, introns, and viruses; plasmids are preferred.
[0142] The ligation method can be through ligation of homologous fragments on the vector and the intermediate, or through ligation of complementary pendant sequences on the vector and the intermediate. The ligation method using homologous fragments can be performed using homologous recombinases or without homologous recombinases. The ligation method using pendant sequences can be performed using ligases or without ligases. The method of introducing the homologous fragment into the intermediate can be by adding the target homologous fragment to the long double-stranded nucleic acid used as a design template during the design of the short single-stranded nucleic acid. The homologous fragment can be added to both ends of the long double-stranded nucleic acid or to one end. The homologous fragment can also be added after the intermediate is formed, for example, through a nucleotide-based synthetic method. The method for introducing the pendant sequence into the intermediate can be as follows: when designing the short single-stranded nucleic acid, the target pendant sequence is added to the long double-stranded nucleic acid, which serves as the design template; or when designing the short single-stranded nucleic acid, the target pendant sequence and its complementary pairing sequence are added to the long double-stranded nucleic acid, which serves as the design template, and the complementary pairing sequence of the pendant sequence is subjected to enzymatic digestion or digestion (e.g., enzymatic hydrolysis, acidic hydrolysis, alkaline hydrolysis, pyrolysis, chemical hydrolysis, microbial fermentation, photolysis, oxidizing agents, reducing agents, electrochemical methods, etc.) after the intermediate is formed.
[0143] In the host, after the intermediate is repaired to form the long double-stranded nucleic acid, the long double-stranded nucleic acid can be amplified within the host or not. The amplification of the long double-stranded nucleic acid can be performed within the host or after isolating the long double-stranded nucleic acid from the host. The isolation method can be selected from one or more of the following: chemical extraction, enzymatic digestion, ultrasonic disruption, freeze-thaw reaction, sodium dodecyl sulfate (SDS) method, hexadecyltrimethylammonium chloride (CTAB) method, commercial kits, magnetic bead method, gel electrophoresis purification, and density gradient centrifugation. The amplification method can be selected from one or more of the following: polymerase chain reaction (PCR), cycle-mediated isothermal amplification (LAMP), single-primer isothermal amplification (SPIA), sequence-dependent amplification (NASBA), multiple substitution amplification (MDA), CRISPR-Cas9 system, and in vitro transcription (IVT).
[0144] The enzymatic reaction refers to the formation of phosphodiester bonds through the ligation activity of a polymerase or the ligase itself. The polymerase may be selected from one or more of the following: DNA polymerase, RNA polymerase, reverse transcriptase, terminal deoxynucleotidyl transferase, polyadenylate polymerase, uracil-DNA glycosidase, and Qβ-RNA-dependent RNA polymerase. The ligase may be selected from one or more of the following: T4 DNA ligase, E. coli DNA ligase, T7 DNA ligase, RNase H ligase, ligase I, ligase II, ligase III, and ligase IV, preferably T4 DNA ligase.
[0145] In the aforementioned nucleic acid repair method, in vivo nucleic acid repair and enzymatic reactions can be combined. Preferably, the enzyme used in the enzymatic reaction can be a ligase. In specific embodiments, the enzymatic reaction can occur before the intermediate or the vector containing the intermediate is transferred into the host, or after the intermediate or the vector containing the intermediate is transferred into the host; it can occur before the intermediate is ligated to the vector, or it can occur after the intermediate is ligated to the vector. In a preferred embodiment, the combined reaction can occur when the number of nicks is greater than or equal to 41, 51, 61, 71, 81, 91, or 101.
[0146] The long double-stranded nucleic acid can be a long double-stranded nucleic acid without a pendant sequence, or it can be a long double-stranded nucleic acid containing one or more pendant sequences. Therefore, the intermediate may or may not contain pendant sequences. The long double-stranded nucleic acid can be a linear nucleic acid or a circular nucleic acid. Therefore, the intermediate can also be a linear nucleic acid or a circular nucleic acid.
[0147] The long double-stranded nucleic acid can be a conventional long double-stranded nucleic acid, or it can be a long double-stranded nucleic acid containing multiple fragments with the same sequence or the short single-stranded nucleic acid.
[0148] The long double-stranded nucleic acid can be of any length, preferably greater than or equal to 500 bases or base pairs, 800 bases or base pairs, 1000 bases or base pairs, 1200 bases or base pairs, 1500 bases or base pairs, 1800 bases or base pairs, 2000 bases or base pairs, 2500 bases or base pairs, 3000 bases or base pairs, 3300 bases or base pairs, 4000 bases or base pairs, 6000 bases or base pairs, 9000 bases or base pairs, or 9900 bases or base pairs.
[0149] The long double-stranded nucleic acid may contain one sequence or two or more sequences. In some embodiments, the long double-stranded nucleic acid contains two or more sequences, which may be on the same long double-stranded nucleic acid or on different long double-stranded nucleic acids; the sequences may be contained in the same expression frame to form a fusion gene or fusion protein, or they may be contained in different expression frames. The sequences may be protein-coding sequences (including DNA and RNA), non-coding RNA (e.g., transfer RNA, ribosomal RNA, etc.) or sequences expressing said non-coding RNA, or regulatory sequences (e.g., promoters, enhancers, silencers, insulators, and response elements, etc.). In some embodiments, the sequences are protein-coding sequences, and the long double-stranded nucleic acid may also contain nonsense sequences, which are sequences that do not encode proteins. Examples of nonsense sequences may include sequences from one or more of the following: introns, intergenic regions, non-coding RNA or sequences expressing said non-coding RNA, regulatory sequences, repetitive sequences, transposons, and pseudogenes.
[0150] In this application, the long double-stranded nucleic acid can be evaluated by fidelity. The fidelity can refer to the proportion of the actual sequence of the synthesized long double-stranded nucleic acid that has a higher homology to a standard sequence than a specific value when the sequence of the synthesized long double-stranded nucleic acid is sampled and tested. This specific value can be 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5%. The actual sequence of the long double-stranded nucleic acid can be obtained by sequencing. The standard sequence can be a sequence of the long double-stranded nucleic acid recognized by those skilled in the art, preferably a standard sequence from an authoritative website (e.g., NCBI) gene database.
[0151] In this application, the long double-stranded nucleic acid can be evaluated by an error-free rate. The error-free rate can refer to the proportion of times the actual sequence of the synthesized long double-stranded nucleic acid shows 100% homology with a standard sequence when the sequence of the synthesized long double-stranded nucleic acid is sampled and tested. The actual sequence of the long double-stranded nucleic acid can be obtained by sequencing. As noted by those skilled in the art, "error-free" and "all correct" are used interchangeably.
[0152] In this application, when the synthesized long double-stranded nucleic acid is a linear nucleic acid, the intermediate usually needs to be assembled with a vector and transferred into the host to complete the synthesis or even amplification of the long double-stranded nucleic acid.
[0153] In some embodiments, the assembly can be performed using the Gibson assembly method. In some embodiments, the assembly can be performed using the chain substitution method.
[0154] In some embodiments, the assembly can be performed by direct, nonspecific ligation using a ligase.
[0155] Parallel synthesis of multiple genes
[0156] In some embodiments, the long double-stranded nucleic acid may contain two or more sequences located on different long double-stranded nucleic acids, and the short single-stranded nucleic acids that make up the different long double-stranded nucleic acids are annealed and assembled in the same mixture system to form the intermediate that is identical to the different long double-stranded nucleic acid sequences.
[0157] In some embodiments, the short single-stranded nucleic acid can be a conventional short single-stranded nucleic acid or a short single-stranded nucleic acid with a special structure. For example, the short single-stranded nucleic acid may contain sequences that pair with other sequences on the same short single-stranded nucleic acid, referred to as self-pairing sequences, and the short single-stranded nucleic acid containing self-pairing sequences can form a stem-loop structure. The content of guanine and cytosine on the short single-stranded nucleic acid can be arbitrary, for example, 0-about 10%, about 10%-about 20%, about 20%-about 30%, about 30%-about 40%, about 40%-about 50%, about 50%-about 60%, about 60%-about 70%, about 70%-about 80%, about 80%-about 90%, about 90%-about 100%.
[0158] In some implementations, the different long double-stranded nucleic acids may be of different lengths or of the same length.
[0159] In some embodiments, the different long double-stranded nucleic acids may each contain sequences encoding different homologous proteins. Homologous proteins are proteins that have sequence homology, structural homology, functional homology, and / or evolutionary homology. Even if the homologous proteins may originate from the same ancestral protein, they should not be confused with different mutants of the same protein.
[0160] In a specific implementation, when the long double-stranded nucleic acid is not the same as the sequence length of the protein encoding it, the long double-stranded nucleic acid may also contain the aforementioned homologous fragment or the overhanging sequence used for linking the intermediate to the vector, and may also contain the aforementioned nonsense sequence.
[0161] In this application, the performance metrics for parallel multi-gene synthesis can be evaluated using coverage. Coverage refers to the ratio of the actual number of genes appearing in the sequencing results to the theoretically expected number of genes. The gene sequences appearing in the sequencing results should have at least 60% homology with the standard sequence of that gene, preferably at least 70%, more preferably at least 80%, even more preferably at least 90%, and still more preferably at least 95%. For example, for the short-chain single-stranded nucleic acid mixture, theoretically 100 genes should be synthesized; if 93 genes are detected in the final product, the coverage is 93%.
[0162] Without being limited by any theory, the embodiments described below are merely for illustrating the various technical solutions of the present invention and are not intended to limit the scope of the present invention.
[0163] Example
[0164] The water used to dissolve the nucleic acids in this application was DEPC water purchased from Invitrogen, the primers were purchased from Sangon Biotech, and all enzymes (unless otherwise noted) were purchased from New England Laboratories. The process for synthesizing long double-stranded nucleic acids from short single-stranded nucleic acids is shown in Figure 1. This synthetic method is named MOSAIC.
[0165] As shown in Examples 1-4, the simplified flow of the MOSAIC method described in this application is as follows:
[0166] (1) Synthesize and purify short-chain single-stranded nucleic acids;
[0167] (2) The purified short-chain single-stranded nucleic acid was annealed and assembled into an intermediate, which was then separated and purified.
[0168] (3) After the separated intermediate is connected to the carrier, it is transformed into the host and the host's built-in repair system is used to repair the cracks on the intermediate.
[0169] Example 1: Synthesis and Purification of Short Single-Stranded Nucleic Acids
[0170] Except for the short single-stranded nucleic acids used for parallel synthesis of multiple genes, which were purchased from Integrated DNA Technologies, all short single-stranded nucleic acids used in this application were artificially synthesized nucleic acid molecules. Specifically, the artificial synthesis method was column synthesis based on the phosphoramide method. All short single-stranded nucleic acids required purification after synthesis.
[0171] To purify short single-stranded nucleic acids using PAGE purification, samples to be used in the same mixture system are purified in batches in one reaction. The specific steps are as follows.
[0172] Prepare 50 ml of 8% denaturing polyacrylamide gel [8% acrylamide-bisacrylamide (19:1), 8M urea, 1xTBE (89mM Tris-borate, 89mM boric acid, 2mM EDTA, pH 8), 0.4% APS, 0.04% TEMED]. Prepare short-chain single-stranded nucleic acid loading samples by adding an equal volume of 2xTBE, 8M urea loading buffer to the mixed sample. Vortex the sample and incubate at 95°C for 5 min to denature the DNA, then maintain at 50°C until ready for loading. After pre-running the gel at 300V (i.e., preheating) for 15 min, perform gel electrophoresis at 300V for 1 hour in a 50°C water circulation system. Cut the target band from the gel using a UV lamp. Place the strip into a 1.5 ml centrifuge tube, crush the gel, and add 800 μL of 1x TE (10 mM Tris-Cl, 1 mM EDTA, pH 8). Freeze at -20°C until coagulated, then shake overnight at 1,900 rpm in a shaker at 4°C. Centrifuge at 12,000 rpm for 2 min and collect the supernatant. Resuspend the precipitate in 700 μL of 1x TE, centrifuge at 12,000 rpm for 2 min, and collect the supernatant. Repeat this step once. Collect all the supernatant into a 10 ml centrifuge tube, add 3 volumes of n-butanol, vortex, centrifuge at 7,800 rpm for 2 min, collect the aqueous phase, and repeat this step until the aqueous phase is less than 500 μL. Transfer the aqueous phase to a new 1.5 ml tube, add 1 / 10 volume of 3M NaOAc and 2 volumes of 100% ethanol. Incubate at -20°C for 30 min, then centrifuge at 12,000 rpm for 20 min at 4°C. Discard the supernatant and air dry for 10 min to obtain purified DNA sample. Resuspend the DNA in 400 μl of DEPC water and transfer to a 10K filter column. Centrifuge at 12,000 rpm for 15 min, then add another 400 μl of DEPC water and centrifuge twice at 12,000 rpm for 15 min. Determine the concentration of the remaining solution using a nanodrop microvolume spectrophotometer.
[0173] As shown in Figure 2, the short single-stranded nucleic acids purified by PAGE showed better assembly results than the unpurified ones. The assembly method is shown in Example 2. The intermediates of the purified short single-stranded nucleic acid assembly formed clearly visible bands on the gel, while the unpurified ones did not have any bands.
[0174] Example 2: Short-chain single-stranded nucleic acid synthesis intermediates
[0175] In this application, short-chain single-chain nucleic acids are mainly synthesized as intermediates through an annealing process.
[0176] In a 1x TE buffer containing 10 mM MgCl2, the short single-stranded nucleic acids required for the synthesis of the long double-stranded nucleic acid are mixed and adjusted to a total concentration of approximately 2 μM (calculated using mass concentration and the average molecular weight of the short single-stranded nucleic acids) or 20 μL of solution containing 2–10 μg of short single-stranded nucleic acids are used. The short single-stranded nucleic acid mixture is then annealed in a gradient thermal cycler to form the intermediate.
[0177] When the number of long double-stranded nucleic acids is 1 or 2, the annealing procedure is as follows: react at 95°C for 5 min, react at 70°C for 30 min, and then cool to 16°C.
[0178] When the number of long double-stranded nucleic acids is greater than or equal to 10, the annealing procedure is as follows: react at 95°C for 5 min, react at 80°C for 1 h, react at 75°C for 1 h, react at 70°C for 2 h, react at 65°C for 2 h, react at 60°C for 1 h, react at 55°C for 1 h, and then cool to 16°C.
[0179] After further optimization by the inventors, the following annealing procedure can be applied to the synthesis of all long double-stranded nucleic acids: react at 95°C for 5 min, react at 80°C for 20 min, and then decrease the temperature by 1°C every 15 min until it is cooled to 16°C.
[0180] Example 3: Isolation and purification of intermediates from annealing products and characterization of intermediates.
[0181] A 1% agarose gel was prepared using 0.5×TBE buffer supplemented with 10 mM MgCl2 as the solvent and pre-stained with SYBR Safe (Thermo Scientific). Annealed samples were loaded and electrophoresed at 90 V in an ice-water bath. The gel was imaged using Typhoon. Target gel bands were then cut, carefully crushed in a Freeze'N Squeeze column (Bio-Rad), and directly centrifuged at 1,000 rpm for 2 minutes at room temperature. Concentrations were measured using a nanodrop.
[0182] Example 4: Repairing cracks on the intermediate body
[0183] Vector preparation: The vector was pUC19 plasmid, purchased from / from Tiangen. The plasmid vector was generated using PCR with 2x Taq Master Mix (Sangon Biotech) according to the manufacturer's instructions. The PCR product was digested with the restriction endonuclease DpnI, and after electrophoresis (100V, 30 min), it was extracted from an agarose gel and recovered using the GeneJET PCR Purification Kit (Thermos Scientific) according to the manufacturer's instructions. If the vector is used for the strand substitution method shown in Figure 3B, it needs to be digested with a nick enzyme (37℃ for 1 h; the enzyme type is determined by the recognition sequence).
[0184] Insert preparation: To reduce the truncation rate of sequencing results, the obtained product can be pre-ligated. The reaction system includes 17 μl of intermediate, 2 μl of 10x T4 ligase buffer, and 1 μl of T4 polynucleotide kinase (NEB). After reacting at 37°C for 2 h, 1 μl of T4 ligase is added, and the reaction is carried out at 25°C for 2 h, followed by a final reaction at 52°C for 1 h.
[0185] Recombination: Gibson assembly was performed using the HiFi DNAAssembly Kit according to the instructions. For strand displacement recombination, the vector and the treated intermediate were mixed in a 0.2 ml tube and reacted at 50°C for 20 min. The ratio of vector to intermediate was 1:1.
[0186] Transformation: Gently mix 10 μl of reaction solution with 100 μl of competent Escherichia coli DH5α (TIANGEN). Incubate the mixture on ice for 20 minutes, heat shock at 42°C for 1 minute, place on ice for 3 minutes, then add 800 μl of LB bacterial liquid medium and incubate at 37°C for 40 minutes. Treat the incubated solution as follows: (1) Spread the solution on LB agar plates containing ampicillin or bleomycin, and send them for sequencing after the clones grow to detect gene synthesis; (2) Add the solution directly to 5 mL of LB liquid medium and incubate for 8 hours, and then perform plasmid extraction.
[0187] The schematic diagram of the Gibson assembly method is shown in Figure 3A and Figure 4. Gibson assembly mainly involves connecting homologous fragments between the vector and the intermediate.
[0188] The chain substitution recombination assembly method is shown in Figure 3B and Figure 20, and is briefly described as follows: (1) Double-strand cleavage is performed on the two nickase cleavage sites at the ends of the vector; (2) Single-stranded nucleic acid longer than the overlapping sequence is used to extend the cleavage to form the overhanging sequences at both ends of the vector; (3) The overhanging sequences of the intermediate are replaced with the shorter strand on the vector to assemble a circular double-stranded nucleic acid; (4) Subsequent tear repair work is performed.
[0189] Transformation method of yeast: The product of vector and fragment ligation and competent yeast were co-incubated at 30℃ for one hour. After incubation, the yeast was centrifuged at 2000rpm for 2min, resuspended in 100μl of pure water, spread on selective plates, and cultured at 30℃ for two days. The surviving clones were observed, and the surviving clones were selected for amplification culture. The plasmid was extracted and transformed into E. coli. The E. coli was then sent for sequencing to detect whether gene synthesis was carried out.
[0190] Example 5: Synthesis of a single conventional long-chain double-stranded nucleic acid
[0191] As shown in Figure 5A, the MOSAIC method described in this application mainly includes two stages: (1) Designing the target gene sequence into short single-stranded nucleic acids with overlapping sequences and synthesizing these short single-stranded nucleic acids, and then assembling them into intermediates through annealing hybridization. In one case, the intermediates, except for containing a cleft and possibly containing homologous ends that are assembled with the vector, should have other characteristics (such as the base sequence) consistent with the target gene. (2) After assembling the intermediates with the vector, transforming them into a host (such as Escherichia coli or yeast) to repair the cleft and amplify the target gene, and then randomly selecting clones for DNA sequencing to verify the synthesis yield of the target gene.
[0192] This application first tests the feasibility of MOSAIC by synthesizing the green fluorescent protein (GFP) encoding gene sequence as shown in SEQ ID NO:1. As shown in Figures 4 and 5B, the inventors first divided the 810 bp DNA sequence containing the GFP coding sequence and two homologous ends, generating 26 60 nt short single-stranded nucleic acids containing overlapping sequences and two 30 nt short single-stranded nucleic acids (26 x 60-mers and 2 x 30-mers), wherein the overlap between each group of nucleic acids is 30. Simultaneously, the GFP was filled with nonsense sequences to achieve the designed 810 bp fragment; this nonsense sequence filling was also applied to the synthesis of other sequences.
[0193] The short single-stranded nucleic acids used to synthesize the full-length GFP sequence are shown in Table 1. F0-F13 are short single-stranded nucleic acids obtained by partitioning the coding strand of the GFP coding gene from the 5' end to the 3' end according to the first 30nt segment, with the remainder being 60nt. R0-R13 are short single-stranded nucleic acids obtained by partitioning the template strand of the GFP coding gene from the 5' end to the 3' end according to the first 30nt segment, with the remainder being 60nt. It should be noted that the coding strand and the template strand are inversely complementary; therefore, F0-F13 and R0-R13 are also inversely complementary. For example, F0 is anticomplementary to the 3' 30nt portion of R13, and F1 is anticomplementary to the 3' 30nt portion of R12 and the 5' 30nt portion of R13, as shown below. The left half of the first row represents F0, the right half of the first row and the third row represent F1, the second row represents R13, and the fourth row represents R12. R12 and R13 are arranged from 3' to 5' from left to right. Sequences within the same frame are complementary. Other genes and sequences in this application are also divided into short single-stranded nucleic acids using the same logic.
[0194] Table 1. Short single-stranded nucleic acid sequence listing for synthesized GFP - 30nt / 60nt
[0195] Next, 28 short single-stranded nucleic acids of gfp were synthesized and purified in batches using the method in Example 1. The purified short single-stranded nucleic acids were annealed using the method in Example 2 to assemble intermediates, which were then characterized using the method in Example 3. The characterization results are shown in Figure 5C. After inserting the vector and transforming *E. coli*, 16 single transformants were randomly selected for sequencing. Five of the 16 sequenced clones were confirmed to have gene sequences completely identical to the target gene. Furthermore, as shown in Figure 6, green fluorescent colonies were observed on a plate under the induction of isopropyl β-d-thiogactopyranoside, further confirming the successful synthesis of gfp.
[0196] Example 6: Synthesis of Long Double-Stranded Nucleic Acids from Short Single-Stranded Nucleic Acids with Overlapping Sequences of Different Lengths
[0197] To investigate whether the length of overlapping sequences affects the assembly of intermediates, the inventors designed six groups of short single-stranded nucleic acids with different overlapping lengths, including those containing multiple different lengths. These were then annealed as described in Example 2 to assemble intermediates. Characterization was performed using agarose gel electrophoresis (3% magnesium chloride) at 120V on ice for 1 hour. The characterization results and the overlapping sequence lengths of the short single-stranded nucleic acids composing the different intermediates are shown in Figure 7. The right side of Figure 7 shows the long double-stranded nucleic acid divided into multiple short single-stranded nucleic acids, with the length of each overlapping region listed below its location in each design. For example, in Group 1, the first overlapping region is 64 nt long, the second is 16 nt long, and so on. Correspondingly, the first short single-stranded nucleic acid separated from the upper single strand is 80 nt long, with overlapping sequences of 64 nt and 16 nt with the short single-stranded nucleic acids on the left and right sides. Since the purpose of this experiment was to investigate whether the length of overlapping sequences affects assembly, it was not strictly stipulated that the assembled intermediates should not contain overhanging sequences. As shown in Figure 7, the inventors set the minimum overlapping sequence to be 16nt and the maximum to be 40nt.
[0198] The gel electrophoresis results in Figure 7 show that all short single-stranded nucleic acids were successfully assembled into intermediates, so the reduction of overlapping sequences does not affect the assembly of intermediates.
[0199] Example 7: Synthesis of Long Double-Stranded Nucleic Acids from Short Single-Stranded Nucleic Acids with Special Sequences
[0200] To further evaluate the capabilities of MOSAIC, the inventors performed assembly tests on some short single-stranded nucleic acids containing specific sequences. As shown in Figures 5B, 8A, and 8B, the inventors synthesized three structures: (1) a 630 bp fragment (20 × 60-mers and 2 × 30-mers) with 74.3% self-paired hairpin structure (Figures 5B and 8A); (2) a 1650 bp fragment (54 × 60-mers and 2 × 45-mers) with a GC content of approximately 77% (Figure 5B); and (3) a 330 bp human microsatellite region (10 × 60-mers and 2 × 30-mers) composed of six tandem repeat sequences of the same 45-bp sequence (Figures 5B and 8B).
[0201] As shown in Figures 5C and 8C, using MOSAIC, the inventors successfully assembled all three special fragments. After more accurate calculation of the sequencing results, the error-free rates were 62.5% (5 / 8), 66.7% (4 / 6), and 17.6% (3 / 17), respectively.
[0202] Meanwhile, the inventors used the PCA (polymerase cycle assembly) method to assemble the above three short single-stranded nucleic acids containing special sequences and characterized the products, the results of which are shown in Figure 9. As shown in Figure 9, the gfp sequence and self-paired sequence in Example 5 could be successfully assembled into the final product; high GC did not assemble or generated fragmented products; and repetitive sequences generated a series of products of indefinite length. A comparison of Figure 9 with Figures 5C and 8C shows that the long double-stranded nucleic acid synthesis method protected in this application has advantages in synthesizing special sequences.
[0203] Example 8: Testing the mismatch probability when synthesizing multiple sequences
[0204] To further explore whether MOSAIC can be adapted to a wider range of applications, the inventors hope to apply the MOSAIC system to a multi-gene synthesis platform, that is, to synthesize a large number of gene sequences from different sources in a single reaction. Before performing multi-gene synthesis, the inventors want to explore whether there are any interferences between different gene synthesizers that would lead to a decrease in yield or error-free rate.
[0205] The purpose of this embodiment is to test whether cross-pairing will occur between short single-stranded nucleic acids originating from different target long double-stranded nucleic acids when synthesizing multiple long double-stranded nucleic acids in the same reaction due to interference, resulting in chimeric sequences in the product. This embodiment tests the cross-pairing from two perspectives.
[0206] Testing long double-stranded nucleic acids of different lengths
[0207] The inventors investigated the cross-pairing of two genes (gfp and bacterial rhodopsin) of different lengths synthesized in parallel using MOSAIC. Figure 10A shows the short single-stranded nucleic acids after dividing the 810-bp gfp fragment (26×60-mers and 2×30-mers) and the 1,160-bp bacterial rhodopsin fragment (28×80-mers and 2×40-mers). These short single-stranded nucleic acids were mixed together in the same reaction system for annealing assembly as described in Example 2. The assembled products were then subjected to electrophoresis to observe whether the target intermediate was generated and to purify it. As shown in Figure 10B, two bands of the same length as the target sequence appeared on the agarose gel, indicating that there was no significant cross-pairing between the two sequences. The products were then purified separately, and the purified assembly products were mixed and ligated with a vector for downstream transformation and sequencing. As shown in Figure 10C, some sequences in the sequencing results showed one or more point mutations, but no large-scale sequence changes, thus confirming that the two genes were synthesized in parallel and that the probability of cross-pairing was low.
[0208] Testing the coding sequence of homologous genes
[0209] Considering that short single-stranded nucleic acid sequences derived from highly homologous fragments are more likely to have high similarity and are more likely to cross-pair, the inventors also evaluated the applicability of the MOSAIC method for assembling highly homologous sequence fragments.
[0210] As shown in Figure 11A, the two nr2f1 and nr2f2 gene fragments, each 930 bp in length and with a sequence similarity of up to 85%, were divided into 32 short single-stranded nucleic acids (30 × 60-mers and 2 × 30-mers), and 16-nt dangling sequences were added to each end.
[0211] Precise sequence alignment was performed on corresponding homologous fragments in the two fragments to determine the degree of homology between the corresponding short single-stranded nucleic acids in the two fragments. Figure 11B shows the number of identical bases on corresponding homologous fragments in the nr2f1 and nr2f2 fragments. As shown in Figure 11B, the nr2f1 and nr2f2 fragments have high homology, and the homologous sites are relatively dispersed, with almost all short single-stranded nucleic acids showing more than 70% homology.
[0212] The synthesized and purified short single-stranded nucleic acids were annealed and assembled in the same reaction, and the assembled products were subjected to electrophoresis to observe whether the target intermediate was generated and to purify it. As shown in Figure 11C, a band of the same length as the target sequence appeared on the agarose gel. The purified assembly product was then ligated to a vector for downstream transformation and sequencing. As shown in Figure 11D, some sequences in the sequencing results had one or more point mutations, but no cross-pairing sequence formation was observed, thus confirming that the two genes were synthesized in parallel and that the probability of cross-pairing was low.
[0213] Taking homologous gene coding sequences as an example, this study investigates the advantages of the MOSAIC method compared to traditional synthesis methods.
[0214] Since each short single-stranded nucleic acid in a homologous sequence has high homology, the probability of mismatch is theoretically much higher than that of other sequences. Therefore, the inventors decided to use homologous genes as an example to illustrate the advantages of the synthesis method protected in this application in reducing sequence mismatch compared to existing methods.
[0215] The inventors used the PCA method to assemble and characterize the short single-stranded nucleic acids derived from the homologous genes nr2f1 and nr2f2 in the same reaction, and then sequenced the products. The results are shown in Figure 12. As shown in Figure 12A, the PCA method successfully synthesized sequences of the same length as the final product. As shown in Figure 12B, the inventors compared each sequenced sequence with the standard sequences of nr2f1 and nr2f2, and displayed the comparison results of a portion of the sequences horizontally in Figure 12B, where mismatches are shown as white vertical lines. Among the obtained sequences, there were no sequences that completely matched the standard sequences of nr2f1 and nr2f2, i.e., the error rate was 0. As can be seen from Figure 12B, mismatches that appeared in the nr2f1 comparison results did not appear in the nr2f2 sequence, and vice versa. Therefore, it can be determined that the homologous genes synthesized by the PCA method had a large number of cross-pairings.
[0216] Therefore, it is evident that using MOSAIC for gene sequence synthesis has the advantage of a low mismatch rate.
[0217] Example 9: Parallel Synthesis Method of Multiple Genes
[0218] Example 8 confirmed that the mismatch rate was low when short single-stranded nucleic acids from different genes were synthesized into long double-stranded nucleic acids in the same reaction, especially the mismatch rate for homologous fragments. This shows that the mispairing rate of gene fragments from different sources synthesized in one reaction using the MOSAIC method is within an extremely low range, proving the feasibility of a method for parallel synthesis of multiple genes in one reaction.
[0219] Therefore, the inventors set out to test whether the multi-gene parallel synthesis method could complete the synthesis task. They evaluated the synthesis efficiency for 10 genes and 100 genes respectively.
[0220] Results of parallel synthesis of 10 genes
[0221] Figure 13 shows the length distribution of the 10 genes. These 10 genes were padded with nonsense sequences to adjust their length to 1170-bp or 1050-bp, while ensuring they contained the same homologous ends, for assembly with a common vector.
[0222] Using a 1170-bp length strategy, 10 genes were divided into equimolar mixtures of 140 short single-stranded nucleic acids (12 × 180-mers and 2 × 90-mers per gene) and assembled and validated in the same reaction system. The electrophoresis results of the intermediates are shown in Figure 13A, indicating that the intermediates showed bands of the designed size on agarose gels. The intermediates were then assembled into vectors, and 200 clones were sequenced. This confirmed that all 10 genes successfully synthesized sequences with high homology to the standard sequences, with an overall error-free rate of 44%.
[0223] Using a 1050-bp length strategy, 10 genes were divided into an equimolar mixture of 342 short single-stranded nucleic acids (34 × 60-mers per gene and two 46-mer ends shared by all genes) to synthesize 10 1050 bp gene fragments. After observing the target bands on a gel, the intermediates were purified and assembled. The resulting 140 clones were sequenced, yielding all 10 error-free genes, with an overall error-free rate of 36%.
[0224] Results of parallel synthesis of 100 genes
[0225] Next, 100 1,650-bp fragments were annealed and assembled in a single reaction. The same oligonucleotide design strategy described above was used. As shown in Figure 14, the 100 genes had different length distributions and were filled into a uniform 1,650-bp sequence using nonsense sequences (54 × 60-mers per gene and 2 46-mers shared by all genes). Among all the designed short single-stranded nucleic acids, 12 sequences were identical, resulting in 5,390 different short single-stranded nucleic acids (total nucleotide count of 323,492). As shown in Figure 15A, the inventors anticipated that all short single-stranded nucleic acids would self-classify into their corresponding intermediates.
[0226] The short-chain single-stranded nucleic acid mixture was annealed and assembled into intermediates using MOSAIC in the same reaction system, and electrophoresis was performed to detect whether the intermediates were synthesized. As shown in Figure 15B, the presence of bands of the target size on the agarose gel indicates that the intermediates have been successfully synthesized.
[0227] The intermediate was purified and assembled into a vector for gap repair and amplification. 2,050 grown clones were sequenced. Sequencing results confirmed that 93 genes had been successfully assembled (the number of point mutations in all synthesized genes was less than 5 compared to the standard sequence). As shown in Figure 15C, the inventors tested 2,050 clones, of which 718 clones contained error-free sequences, resulting in an overall error-free rate of 35%. Furthermore, among the 93 successfully assembled genes, 76 yielded completely error-free constructs.
[0228] In the remaining seven genes, we assembled each gene separately and performed electrophoresis on the intermediates. As shown in Figure 16, nkx-2 and nifc failed to form the target band, and the en1 band was shorter than the target band. The five fragments that formed bands (excluding nkx-2 and nifc) were then subjected to the subsequent MOSAIC step. However, only bocr-1 successfully produced viable colonies. Sequencing of 30 clones from the bcor-1 group yielded one error-free bcor-1 and 29 mutated bcor-1s. This indicates that only one of the seven genes was successfully synthesized, and the error-free rate of the successfully synthesized bcor-1 gene after assembly was extremely low (1 / 30). This suggests that the synthesis of certain fragments is inherently challenging, rather than a defect in the parallel synthesis of multiple genes.
[0229] In summary, MOSAIC can achieve highly parallel gene synthesis in the same reaction with coverage of over 90% and accuracy of over 30%.
[0230] Example 10 Synthesis of Circular Nucleic Acids
[0231] Because current synthetic methods commonly used by those skilled in the art present difficulties in synthesizing plasmids, such as requiring multiple rounds of enzymatic reactions to synthesize the desired construct, the inventors, after studying linear nucleic acids, began to study the MOSAIC synthesis of circular nucleic acids (mainly plasmids).
[0232] The first plasmid synthesized was a small circular plasmid of 1200-bp, which the inventors divided into 120-mer segments with overlapping sequences of 20 bases each.
[0233] The inventors designed two strategies for using MOSAIC for plasmid assembly and synthesis.
[0234] The first method is a one-step method, in which the short single-stranded nucleic acids constituting the plasmid are assembled into a circular intermediate in the same reaction system, then transferred into the host to repair the crack and amplify. The second method is a two-step or stepwise method, as shown in Figures 17A and 18A. This involves dividing the short single-stranded nucleic acids constituting the plasmid into two reaction systems, assembling them separately into two intermediates with pendant sequences, and then mixing the two intermediates to form a circular intermediate. As shown in Figure 17B, both methods can assemble intermediates of the correct length, but the two-step method produces a more concentrated product and a higher yield. After purification, the circular intermediate is transferred into *E. coli*, and randomly selected colonies are verified using the expected restriction enzyme digestion pattern shown in Figure 18B. As shown in Figures 18B and 18C, the electrophoresis results after restriction enzyme digestion are as expected; 80% (16 / 20) of the clones contain error-free plasmids. In addition, the remaining four clones with errors contain the full-length sequence of the plasmid, but compared to the standard sequence, they contain 1-3 mutations, indicating a high accuracy rate.
[0235] The aforementioned small plasmid with the GFP insert was further synthesized. To increase assembly yield, the entire 2,040 bp circular nucleic acid (34 × 120 meters, with a 60 bp overlap sequence) was synthesized into an intermediate using a two-step method, as shown in Figure 17C. Both reaction systems synthesized intermediates of the correct length. The mixed intermediates were transformed into *E. coli*, and sequencing was performed on the colonies after growth, as shown in Figure 18C. The synthesis had a high error-free rate. In addition, the remaining selected colonies contained the full-length plasmid sequence, but compared to the standard sequence, the sequence contained 1-3 mutations, with a high accuracy rate.
[0236] Using a similar method, the inventors also synthesized a 2,700 bp pUC19 plasmid from 30 × 180-mers (with an overlap sequence of 90 bp), and the assembly results are shown in Figures 17D and 18C.
[0237] However, difficulties were encountered in synthesizing a 3300bp plasmid containing 110 × 60-mers (110 cleavages). The inventors hypothesized that there was an upper limit to the cleavage recovery and cloning capabilities of *E. coli*. Therefore, the inventors chose to test this challenging synthetic task in budding yeast. Two 1650bp reaction systems were assembled to form two linear intermediates, which were purified separately, mixed, and co-transfected into budding yeast.
[0238] The inventors randomly selected two yeast colonies for plasmid isolation, sequencing, and enzyme digestion verification. As shown in Figure 18D, both plasmids were full-length products, and the length of the enzyme digestion products met expectations.
[0239] In addition to enzyme digestion verification, the inventors selected five pairs of primers for different parts of the plasmid, randomly selected clones, and performed PCR analysis on the selected clones using the five pairs of primers in different reactions. After PCR, clones with the specified PCR amplification fragment were considered successfully assembled plasmids. The inventors performed three rounds of PCR analysis, independently selecting eight clones for analysis in each round. As shown in Figure 21, all selected clones exhibited the correct PCR fragment, indicating that all clones were correctly assembled.
[0240] Furthermore, 50% (1 / 2) of the plasmids were error-free, with an error rate of 0.015% (1 / 6,600) per base. The inventors noted that the error-free rate of the synthesized plasmids was higher than that of the linear fragments in this application. Regarding this phenomenon, the inventors speculated that plasmids with mutations in essential elements such as replication regions or resistance genes might not produce growable clones, meaning that such mutations might cause the plasmids to lose their proliferative capacity.
[0241] In summary, the inventors' results of assembling plasmids in Escherichia coli and yeast demonstrate that MOSAIC is suitable for generating full-length plasmids in prokaryotes or eukaryotes.
[0242] Example 11: Effect of the number of cracks in the intermediate on nucleic acid synthesis
[0243] In Example 5, when synthesizing GFP, the inventors noticed that many fragments were truncated (62.5%, 10 / 16) when using the Gibson assembly method, while no fragment truncation occurred when using the chain substitution method. Therefore, it is hypothesized that the number of cracks on the intermediate may be positively correlated with the truncation rate.
[0244] To verify this hypothesis, the inventors implemented two strategies to reduce the number of cleavages on the intermediate (in the initial design, the intermediate contained 26 cleavages): (1) designing longer short single-stranded nucleic acids, using 180-nt short single-stranded nucleic acids with 90-nt overlap (9×180mers and 2×90mers, 12 cleavages); (2) pre-ligating the intermediate by contacting it with a ligase (such as T4 ligase) after its formation and before it is assembled into the vector, and ligating some of the cleavages first.
[0245] An example of the aforementioned pre-connection method and the method without pre-connection is shown below:
[0246] The short single-stranded nucleic acid was divided into two subpools of 1650bp.
[0247] The unligated group is the aforementioned (1): 3 μg of short single-stranded nucleic acid mixtures in each subpool were placed in 20 μl of 1×TE buffer containing 10 mM MgCl2 and annealed in a thermal cycler (95℃ for 10 min, 70℃ for 30 min, 65℃ for 30 min, 60℃ for 30 min, 10℃ ∞).
[0248] The ligation group, as described in (2) above, consisted of a mixture of short-chain single-stranded nucleic acids totaling approximately 3.3 μg, annealed in a thermal cycler using T4 polynucleotide kinase (NEB, M0201) (37°C for 10 h, 95°C for 10 min, 70°C for 2 h, 16°C at ∞). Then, 18 μl of the assembled fragment was treated with NEB T4 DNA ligase (NEB, M0202) in a total volume of 20 μl overnight at 4°C. The assembly product or the ligation product of the two subpools was mixed at a 1:1 molar ratio and directly transformed into the BY4742 strain according to the protocol in Example 4. The transformation mixture was inoculated onto SC-his plates and cultured at 30°C for two days, followed by PCR analysis.
[0249] As shown in Table 2, both methods effectively reduced the truncation rate to 12% (2 / 17) or 0% (0 / 18), respectively. Since altering the length of short single-stranded nucleic acids is relatively more complex, pre-ligation of intermediates is a better method for reducing the truncation rate. Therefore, this method was widely applied in this application.
[0250] Table 2. Effects of different treatment strategies on gene truncation rate
[0251] Example 12 Other possible influencing factors
[0252] In the parallel synthesis of multiple genes in Example 9, the inventors explored whether the length of the gene before filling with nonsense sequences and the GC content in the short single-stranded nucleic acid affected the error-free rate of synthesis.
[0253] Different short-chain single-stranded nucleic acids were designed to meet exploratory purposes. As shown in Figure 19A, the correlation between gene length and error-free rate was studied when 100 genes were synthesized in parallel. It can be seen that there is no direct correlation between the two, indicating that the MOSAIC method is limited by gene length. As shown in Figure 19B, the correlation between GC content and error-free rate is not high. Dividing the horizontal axis of Figure 19B into two parts to form Figures 19C and 19D, and statistically analyzing them to obtain Figure 19E, as shown in Figures 19C-19E, even after segmenting the GC content, it does not show a correlation with the error-free rate.
[0254] Therefore, the length of the gene before filling with nonsense sequences and the GC content in short single-stranded nucleic acids have no significant impact on the error-free synthesis rate.
[0255] Example 13 Synthesis of long double-stranded nucleic acids containing specific sequences
[0256] Regarding the synthesis effect of short single-stranded nucleic acids containing special sequences in Example 7, the inventors supplemented the synthesis experiments with 16S rRNA (a very complex nucleic acid containing many self-pairing structures, an exemplary structure of which is shown in Figure 22), AT sequences, and short single-stranded nucleic acids of different lengths designed for repetitive sequences.
[0257] As shown in Figures 23 and 24B, the inventors synthesized the following four structures: (1) a 1590bp fragment containing the 16S rRNA encoding gene (Figure 24B, top); (2) a 1050bp fragment with 100% GC content (Figure 23B, left and Figure 24B, middle); (3) a 1050bp fragment with 100% AT content (Figure 23A and Figure 23B, left); and (4) a 330bp human microsatellite region containing six tandem repeats of the same 45bp sequence (the reaction contains 10 short single-stranded nucleic acids of 57-66-mers, 1 of 41-mers, and 1 of 44-mers, Figure 24B, bottom).
[0258] As shown in Figures 23 and 24B, using MOSAIC, the inventors successfully assembled all four special fragments. After more accurate calculation of the sequencing results, the error-free synthesis rates were 50% (5 / 10, determined by nanopore sequencing), 30% (3 / 10, determined by nanopore sequencing), 28% (7 / 25), and 16% (4 / 25), respectively.
[0259] Meanwhile, the inventors used the PCA (polymerase cycle assembly) method to assemble the above four short single-stranded nucleic acids containing specific sequences and characterized the products, the results of which are shown in Figure 23B. As shown in Figure 23B, the PCA method cannot effectively synthesize the desired nucleic acid fragments. Therefore, compared with the traditional PCA method, the MOSAIC method has advantages in synthesizing the above-mentioned specific sequences.
[0260] Example 14 Synthesis of Long Double-Stranded Nucleic Acids from Short Single-Stranded Nucleic Acids of Different Lengths
[0261] To clarify the differences between short single-stranded nucleic acids of different lengths in the synthesis of long double-stranded nucleic acids and to explore the most suitable length of short single-stranded nucleic acids for the synthesis of long double-stranded nucleic acids, the inventors designed short single-stranded nucleic acids of different lengths for the full-length GFP sequence synthesized in Example 5. That is, the full-length GFP sequence was synthesized in different reactions, and the length of the short single-stranded nucleic acids used in different reactions was different. However, the length of the short single-stranded nucleic acids used in the same reaction was the same except for the head and tail.
[0262] The inventors designed short single-stranded nucleic acids of lengths of 20nt, 40nt, 60nt, 80nt, 100nt, and 120nt (with overlapping sequences of lengths of 10nt, 20nt, 30nt, 40nt, 50nt, and 60nt) for the same full-length GFP sequence, and prepared long double-stranded nucleic acids according to the methods in Examples 1-4. The electrophoresis results of the intermediates synthesized from the short single-stranded nucleic acids of different lengths are shown in Figure 25.
[0263] As shown in Figure 25, except for 20 nt, all short single-stranded nucleic acids of various lengths can be used to synthesize intermediates of the correct length. Given that 60 nt of short single-stranded nucleic acids offers the best cost-effectiveness in commercial synthesis, the inventors have determined it to be the preferred length for the short single-stranded nucleic acid building blocks in the MOSAIC synthesis method.
[0264] Example 15: Maximum gene length synthesized using the MOSAIC method
[0265] In this embodiment, the inventors explored the maximum length of sequences that can be synthesized by the nucleic acid synthesis method protected in this application.
[0266] The inventors used short single-stranded nucleic acids (60 nt in length) to synthesize long double-stranded nucleic acids ranging from 300 bp to 9900 bp. As shown in Figure 26A, the yield of the corresponding intermediate gradually decreased as the length of the long double-stranded nucleic acid increased. As shown in Figure 26B, when the long double-stranded nucleic acid was 9900 bp, the corresponding band was weak, but the intermediate could still be successfully formed.
[0267] Regarding crack repair, the inventors also adjusted the repair strategy to improve the success rate of long-chain double-stranded nucleic acid synthesis.
[0268] In linear nucleic acids, without the use of ligases according to Examples 1-4, the inventors can stably synthesize long double-stranded nucleic acids with fewer than 40 cleavages and a length not exceeding 1200 bp. When a ligase (preferably T4 ligase) is used for the intermediate, the inventors are able to successfully synthesize long double-stranded nucleic acids with a sequenced length of 3000 bp.
[0269] After research and analysis, the inventors discovered that the main obstacle to the synthesis of long double-stranded nucleic acids (LLCs) lies in the number of cleavages exceeding the repair capacity of the host. Therefore, the inventors investigated the effects of the number of cleavages on the intermediate and the use of ligases on the synthesis of LCCs, the results of which are shown in Figure 26C. As shown in Figure 26C, an increase in the number of cleavages reduces the success rate of LCC synthesis. When the number of cleavages reaches 70 or 100, LCC synthesis becomes impossible. However, the use of ligases can improve the success rate of LCC synthesis when the number of cleavages increases.
[0270] As shown in Figures 24A and 26D, more than 88% of human gene coding sequences are less than 3000 bp, which is within the synthesis range of MOSAIC.
[0271] In summary, the MOSAIC method can synthesize long double-stranded nucleic acids with a maximum length of 3000 bp, indicating that most human gene coding sequences can be synthesized using the MOSAIC method.
[0272] In the circular nucleic acid, the inventors successfully synthesized a 9.9 kb plasmid. Specifically, the inventors divided the required short single-stranded nucleic acid into three portions, forming three 3300 bp reaction systems, and assembled them into three intermediates. These intermediates were then purified separately, mixed in a 1:1:1 molar ratio, and transformed into the BY4742 yeast strain according to the method in Example 4. The transformation mixture was inoculated onto SC-his plates and cultured at 30°C for two days, followed by PCR analysis.
[0273] For the synthesized 9.9kb plasmid, the inventors designed seven pairs of primers to synthesize different fragments. The positions of the seven primer pairs are shown in Figure 27A. Six clones were selected for PCR analysis, and the results are shown in Figure 27A. As shown in Figure 27A, in the six selected clones, all seven primer pairs synthesized fragments of the corresponding lengths.
[0274] Finally, the inventors randomly selected three successfully synthesized 9.9kb plasmids and used Sanger sequencing to obtain their full sequences. As shown in Figure 27B, all three plasmids contain mutations that do not affect their normal function, including base deletions, substitutions, and insertions. The number of different mutation types in each plasmid varies.
[0275] Example 16: Supplement to the Multi-Gene Parallel Synthesis Method
[0276] To improve the efficiency of multi-gene synthesis, in a method for parallel synthesis of 100 genes, the inventors introduced a barcode at the end of each gene and then performed NGS sequencing on the synthesized 100 genes using the barcode.
[0277] The synthesis results of 100 genes containing barcodes are shown in Figures 28C-28D and Figure 29.
[0278] As shown in Figure 28C, the inventors had already sequenced 76 error-free genes before introducing barcodes. To obtain the remaining 24 genes with greater efficiency, the inventors introduced specific barcodes at the ends of each gene. As shown in Figure 29, colonies on a single plate are picked into individual wells of a 96-well plate, meaning each well will ultimately contain only one type of colony. The inventors amplified the barcode region sequence in situ within the wells, introducing well location information during the amplification process using specific primers. Finally, the amplified DNA products were mixed and subjected to NGS sequencing, allowing for rapid identification of which gene is present in each well, thus achieving efficient selection.
[0279] As shown in Figure 28D, the barcode sequences of all 100 genes were present in these bacteria, proving that all 100 genes were successfully assembled. As shown in Figure 28C, the inventors ultimately obtained 94 error-free genes using this method.
[0280] Results of parallel synthesis of 1793 genes
[0281] The inventors further expanded the number of genes that MOSAIC can synthesize, exploring the synthesis of 1793 genes in a single reaction.
[0282] For the 1793 500bp fragments to be synthesized, each fragment containing a gene sequence of 200bp-450bp, the inventors designed approximately 43,000 40nt short single-stranded nucleic acids and synthesized long double-stranded nucleic acids according to Examples 1-4, followed by sequencing verification. The results are shown in Figures 28E-28G. As shown in Figures 28E-28G, in one reaction synthesizing 1793 genes, the synthesized intermediates were of correct length, 1522 synthesized sequences were free of mutations (85%), and 1560 synthesized sequences contained fewer than 10 mutations (87%), demonstrating reliable coverage and accuracy.
[0283] Example 17: Exemplary PCA Synthesis Procedure
[0284] Mix 1 μL of 1.2 μM short-chain single-stranded nucleic acid mixture, 10 μL of 2×PrimeSTAR polymerase mixture (Takara Bio), and 9 μL of DEPC water in a PCR reaction tube and perform the reaction using the following procedure: (1) React at 98℃ for 2 min; (2) React at 98℃ for 10 s, react at 65-72℃ (determine the specific reaction temperature according to the sequence) for 20 s, and react at 72℃ for 1 min; (3) Repeat step (2) 54 times; (4) After reacting at 72℃ for 10 min, store the reaction tube at 4℃.
[0285] The foregoing detailed description is provided by way of explanation and example and is not intended to limit the scope of the appended claims. Various variations of the embodiments listed herein will be apparent to those skilled in the art and are reserved within the scope of the appended claims and their equivalents.
Claims
A method for synthesizing long double-stranded nucleic acids from short single-stranded nucleic acids, the method comprising the following steps: a) The short single-stranded nucleic acids are complementary to form an intermediate, the intermediate being the long double-stranded nucleic acid with a cleavage, the cleavage referring to the absence of a phosphodiester bond between two adjacent nucleic acids in the long double-stranded nucleic acid; b) Repair the cracks on the intermediate to obtain the long double-stranded nucleic acid. According to the method of claim 1, the short single-stranded nucleic acid comprises 10-300 nucleotides. According to the method of claim 1 or 2, the length of the long double-stranded nucleic acid is more than twice that of the short single-stranded nucleic acid, preferably more than five times, and more preferably more than ten times. According to the method of any one of claims 1-3, the short single-stranded nucleic acid is designed and synthesized by the following method: (1) Design a nucleotide sequence containing the long double-stranded nucleic acid; (2) Divide each single strand of the double helix of the nucleotide sequence into several short single-stranded nucleic acids; and (3) Synthesize the short single-stranded nucleic acid. According to the method of claim 4, in step (2), the split points in the two chains are staggered by at least 10 nucleotides. According to any one of claims 1-5, in step b), the repair method is to form a phosphate diester bond at the crack. According to the method of claim 6, the intermediate is a linear nucleic acid, and the repair method includes: i. Attach the intermediate to the carrier; ii. Transfer the vector into the host body. The intermediate is a circular nucleic acid according to any one of claims 1-6. According to the method of claim 8, the circular nucleic acid is a plasmid, and the repair method includes transferring the plasmid into the host. According to any one of claims 1-9, the number of the cracks is 2-50. According to any one of claims 1-9, the number of the cracks is 51 or more. The method according to any one of claims 7-11 further includes, prior to step i, treating the intermediate with a ligase. Use of the method according to any one of claims 1-12 in the synthesis of long double-stranded nucleic acids. According to the use described in claim 13, the number of long double-stranded nucleic acids is 1. According to the use described in claim 13, the number of long double-stranded nucleic acids is greater than or equal to 2. According to any one of claims 13-15, the long double-stranded nucleic acid is a linear nucleic acid. According to any one of claims 13-15, the long double-stranded nucleic acid is a circular nucleic acid.