High sequence fidelity nucleic acid synthesis and assembly
Thermostable mismatch-recognition proteins and high-fidelity DNA polymerases improve the sequence fidelity of nucleic acid assembly by correcting errors in nucleic acid molecules, achieving high-throughput synthesis of accurate larger molecules from shorter oligonucleotides.
Patent Information
- Application Number
- JP2022578942
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-06
- Filing Date
- 2021-03-05
- Publication Date
- 2026-02-05
- Estimated Expiration
- 2041-03-05
AI Technical Summary
Existing nucleic acid synthesis and assembly methods struggle to achieve high-throughput production of nucleic acid molecules with high sequence fidelity, particularly in the assembly of larger molecules from shorter oligonucleotides.
The use of thermostable mismatch-recognition proteins, such as mismatch-binding proteins and mismatch endonucleases, in combination with high-fidelity DNA polymerases, to detect and correct errors in nucleic acid molecules during assembly and amplification processes, including primary and secondary assembly PCR steps.
This approach results in the generation of nucleic acid molecules with fewer than two errors per 1,000 base pairs, enhancing the sequence fidelity and accuracy of nucleic acid assembly.
Smart Images

Figure 0007811559000064 
Figure 0007811559000065 
Figure 0007811559000066
Abstract
Description
[Technical Field]
[0001] The present disclosure generally relates to compositions and methods for the synthesis of nucleic acid molecules with low error rates. For example, in many cases, compositions and methods are provided for the high-throughput synthesis and assembly of nucleic acid molecules with high sequence fidelity. In many cases, a thermostable mismatch-recognition protein (e.g., a thermostable mismatch-binding protein, a thermostable mismatch endonuclease) is present in the composition and used in the provided method. [Background technology]
[0002] Over the years, efforts have been made to develop high-throughput synthesis platforms that make gene synthesis more cost-effective and produce nucleic acid molecules with high sequence fidelity.
[0003] Biological materials that can be used in processes to generate nucleic acid molecules produced with high sequence fidelity have co-evolved with the organisms that produce them, including DNA polymerases with proofreading capabilities and materials involved in various pathways for the correction of nucleic acid sequence errors (e.g., mismatch endonucleases, mismatch-binding proteins, etc.).
[0004] With the advancement of genetic engineering, the production of larger nucleic acid molecules has become necessary.In many cases, nucleic acid assembly methods begin with the synthesis of relatively short nucleic acid molecules (e.g., chemically synthesized oligonucleotides), followed by the generation of double-stranded fragments or subassemblies (e.g., by annealing and extending multiple overlapping oligonucleotides), and often proceed to the construction of larger assemblies such as genes, operons, and even functional biological pathways (e.g., by ligation, enzymatic extension, recombination, or a combination thereof).The present disclosure generally relates to compositions and methods for the assembly of nucleic acid molecules with high sequence fidelity. Summary of the Invention
[0005] The present disclosure relates, in part, to compositions and methods for the assembly (e.g., by assembly PCR) and amplification of nucleic acid molecules with high nucleotide sequence fidelity. The compositions and methods described herein may contain or employ proteins (e.g., DNA polymerases, mismatch endonucleases, mismatch-binding proteins, etc.) that can detect and / or remove error-containing nucleic acid molecules.
[0006] In some embodiments, provided herein are methods for generating an error-corrected population of nucleic acid molecules. Such methods may include (a) assembling oligonucleotides having regions of terminal sequence complementarity (single-stranded regions that, upon hybridization, form double-stranded regions of about 10 to about 30, about 12 to about 30, about 15 to about 30, about 20 to about 30, about 15 to about 40, about 6 to about 20, or about 8 to about 25 base pairs in length) by primary assembly PCR to form a population of assembled nucleic acid molecules, and (b) amplifying the population of assembled nucleic acid molecules formed in step (a) by primary amplification to form a population of amplified assembled nucleic acid molecules. In some cases, the population of amplified assembled nucleic acid molecules may contain fewer than two errors per 1,000 base pairs (e.g., about 2 to about 0.01, about 2 to about 0.05, about 2 to about 0.08, about 2 to about 0.1, about 2 to about 0.5, about 2 to about 0.75, about 1 to about 0.01, about 1 to about 0.05, about 1 to about 0.1, about 2 to about 0.001, about 1 to about 0.001, about 0.5 to about 0.001, about 0.1 to about 0.001, etc.). In some cases, steps (a) and / or (b) above may be performed in the presence of one or more (e.g., 1 to 10, 1 to 8, 1 to 5, 1 to 3, 1 to 2, etc.) thermostable mismatch-recognition proteins. In some embodiments, at least one of the one or more thermostable mismatch recognition proteins is a thermostable mismatch binding protein, such as, for example, a thermostable mismatch binding protein selected from mismatch binding proteins having an amino acid sequence set forth in Table 13 or Table 15. In some embodiments, at least one of the one or more thermostable mismatch recognition proteins is a thermostable mismatch endonuclease, such as a mismatch endonuclease selected from mismatch endonucleases having an amino acid sequence set forth in Table 12 or Table 15 (e.g., TkoEndoMS, PfuEndoMS, etc.).
[0007] In some cases, a high-fidelity DNA polymerase can be used in the methods described herein. Even more particularly, the high-fidelity DNA polymerase can be used in steps (a) and / or (b) described in the above-described methods for generating an error-corrected population of nucleic acid molecules. Furthermore, the high-fidelity DNA polymerase can be a component of an error-reducing polymerase reagent. The error-reducing polymerase reagent can include one or more (e.g., 1-10, 1-8, 1-5, 1-3, 1-2, etc.) amine compounds, such as one or more amine compounds selected from the group consisting of (a) dimethylamine hydrochloride, (b) diisopropylamine hydrochloride, (c) ethyl(methyl)amine hydrochloride, and (d) trimethylamine hydrochloride.
[0008] In certain variations of the methods described herein, and in the above-described methods for generating an error-corrected population of nucleic acid molecules, at least one of the one or more thermostable mismatch-recognition proteins may be present in step (a). Additionally, in some cases, at least one of the one or more thermostable mismatch-recognition proteins may be present in step (b). Furthermore, one or more (e.g., 1-10, 1-8, 1-5, 1-3, 1-2, etc.) error correction steps may be performed after the primary amplification. A primary post-amplification of the amplified population of assembled nucleic acid molecules may be performed after step (b). In some cases, the amplified population of assembled nucleic acid molecules may be contacted with one or more mismatch-recognition proteins prior to the primary post-amplification. Additionally, at least one of the one or more mismatch recognition proteins can be a mismatch endonuclease, such as one or more (e.g., 1-10, 1-8, 1-5, 1-3, 1-2, etc.) non-thermostable mismatch endonucleases (e.g., T7 endonuclease I, CEL II nuclease, CEL I nuclease, and / or T4 endonuclease VII).
[0009] The methods described herein also relate to generating a population of amplified, assembled nucleic acid molecules that comprise subfragments of a larger nucleic acid molecule. Furthermore, in some cases, such a population of amplified, assembled nucleic acid molecules can be combined with one or more (e.g., 1-10, 1-8, 1-5, 1-3, 1-2, etc.) additional nucleic acid molecules that are also subfragments of the larger nucleic acid molecule to form a nucleic acid molecule pool. In some cases, the nucleic acid molecules of such a nucleic acid molecule pool can be assembled by secondary assembly PCR to form a larger nucleic acid molecule. In some cases, the subfragments can be contacted with one or more mismatch-recognition proteins before or during assembly by secondary assembly PCR. Furthermore, the larger nucleic acid molecule can be heat-denatured, then renatured, and subsequently contacted with one or more (e.g., 1-10, 1-8, 1-5, 1-3, 1-2, etc.) mismatch-recognition proteins. Additionally, at least one of the one or more mismatch recognition proteins (e.g., 1-10, 1-8, 1-5, 1-3, 1-2, etc.) can be a mismatch-binding protein, such as a mismatch-binding protein bound to a solid support. Thus, the methods described herein include methods for separating error-containing nucleic acid molecules from error-free nucleic acid molecules. In some cases, the population of amplified, assembled nucleic acid molecules can be sequenced. Such sequencing can be performed to determine whether errors are present, and if so, the number and type of errors.
[0010] Provided herein are compositions, such as compositions that can be used in the methods described herein. In some cases, the compositions described herein can include one or more (e.g., 1-10, 1-8, 1-5, 1-3, 1-2, etc.) thermostable mismatch-recognition proteins, one or more (e.g., 1-10, 1-8, 1-5, 1-3, 1-2, etc.) DNA polymerases, and one or more (e.g., 1-10, 1-8, 1-5, 1-3, 1-2, etc.) amine compounds. Furthermore, at least one of the one or more amine compounds can be selected from the group consisting of (a) dimethylamine hydrochloride, (b) diisopropylamine hydrochloride, (c) ethyl(methyl)amine hydrochloride, and / or (d) trimethylamine hydrochloride.
[0011] The compositions described herein can further comprise two or more nucleic acid molecules (e.g., the two or more nucleic acid molecules are subfragments of a larger nucleic acid molecule). Additionally, the two or more nucleic acid molecules can be single-stranded. Such single-stranded nucleic acid molecules can vary widely in length, but are often less than 100 nucleotides in length (e.g., between about 35 and about 90, about 35 and about 80, about 35 and about 70, about 35 and about 65, about 40 and about 90, about 30 and about 60, about 30 and about 65, etc.).
[0012] The compositions described herein may further comprise two or more nucleic acid molecules, wherein at least one of the two or more nucleic acid molecules is single-stranded and at least one of the two or more nucleic acid molecules is double-stranded.
[0013] In some compositions described herein, at least one of the thermostable mismatch recognition proteins can be a thermostable mismatch endonuclease, such as a thermostable mismatch endonuclease having an amino acid sequence set forth in Table 12 or Table 15 (e.g., TkoEndoMS, PfuEndoMS, etc.), and variants thereof having at least 80% sequence identity thereto (e.g., at least about 80% to about 99%, about 80% to about 95%, about 80% to about 90%, about 85% to about 95%, about 90% to about 99%, about 92% to about 99%, about 95% to about 99%, about 97% to about 99%, etc.).
[0014] In some specific cases, the compositions and methods provided herein may contain or use a mismatch-specific endonuclease that shares at least 30%, 40%, 50%, or 60% (e.g., about 30% to about 70%, about 30% to about 60%, about 30% to about 50%, about 30% to about 45%, about 30% to about 40%, etc.) amino acid sequence identity with TkoEndoMS (SEQ ID NO: 3). Examples of such mismatch-specific endonucleases are PisEndoMS (SEQ ID NO: 11) or SacEndoMS (SEQ ID NO: 12).
[0015] In some compositions described herein, at least one of the thermostable mismatch recognition proteins can be a thermostable mismatch recognition protein, such as a thermostable mismatch recognition protein having an amino acid sequence set forth in Table 13 or Table 15, and variants thereof having at least 80% sequence identity thereto (e.g., at least about 80% to about 99%, about 80% to about 95%, about 80% to about 90%, about 85% to about 95%, about 90% to about 99%, about 92% to about 99%, about 95% to about 99%, about 97% to about 99%, etc.).
[0016] Also described herein are methods for generating nucleic acid molecules having a predetermined sequence. In some cases, such methods may include: (a) providing a plurality of single-stranded oligonucleotides having complementary overlapping regions, each of the single-stranded oligonucleotides comprising a sequence region of a target nucleic acid molecule, the plurality of single-stranded oligonucleotides including (i) a plurality of internal oligonucleotides having a sequence region overlapping with two other oligonucleotides in the plurality, and (ii) two terminal oligonucleotides designed to be located at the 5' and 3' ends of the full-length nucleic acid molecule and having a sequence region overlapping with one of the internal oligonucleotides in the plurality; (b) assembling the plurality of oligonucleotides by primary assembly PCR to obtain an assembled double-stranded nucleic acid assembly product; and (c) combining at least a portion of the assembly product obtained in step (b) with a pair of primers. In some cases, the pair of primers may be designed to bind to the 5' and 3' ends of the assembly product and perform a PCR amplification reaction to produce an amplified assembly product. Furthermore, in some cases, step (b) and / or step (c) may be performed in the presence of one or more thermostable mismatch-recognition proteins.
[0017] Also described herein is a method for generating a nucleic acid molecule having a predetermined sequence, further comprising (d) performing one or more error correction steps. In some cases, such error correction steps may include (i) denaturing and reannealing the amplified assembly product of step (c) to generate one or more mismatch-containing double-stranded nucleic acids, (ii) treating the mismatch-containing double-stranded nucleic acids with one or more mismatch-recognition proteins, and (iii) optionally performing an amplification reaction. In some cases, the mismatch-recognition protein used in step (d) is a mismatch endonuclease (e.g., T7 endonuclease I) or a mismatch-binding protein (e.g., MutS). Furthermore, the thermostable mismatch endonuclease used may be derived from a hyperthermophilic archaeon, and optionally, the hyperthermophilic archaeon is Pyrococcus furiosus or Pyrococcus abyssi. Additionally, the thermostable mismatch-recognition protein may be selected from the group of proteins having the amino acid sequences set forth in Tables 12, 13, or 15, and variants thereof having at least 80% sequence identity thereto (e.g., at least about 80% to about 99%, about 80% to about 95%, about 80% to about 90%, about 85% to about 95%, about 90% to about 99%, about 92% to about 99%, about 95% to about 99%, about 97% to about 99%, etc.).
[0018] In some cases, one or more of the thermostable mismatch recognition proteins used can be produced and / or obtained by in vitro transcription / translation, while in other cases, one or more of the thermostable mismatch recognition proteins used can be produced and / or obtained by cellular expression.
[0019] When polymerase is present in composition and used in the methods described herein, these polymerases can be high fidelity DNA polymerases.Therefore, provided herein is a method, such as the method of producing the nucleic acid molecule with predetermined sequence described above, wherein one or more of step (b), (c) and (d) (iii) are carried out in the presence of high fidelity DNA polymerase, and optionally, polymerase can be selected from the group consisting of PHUSION™ DNA polymerase (PHUSION™), PLATINUM™ SUPERFI™ II DNA polymerase (SUPERFI™ II), Q5 DNA polymerase and PRIMESTAR GXL DNA polymerase. Additionally, one or more of steps (b), (c), and (d)(iii) can be performed in the presence of a high-fidelity DNA polymerase, optionally wherein the polymerase has an amino acid sequence selected from the group consisting of (1) DNA polymerase 1, (2) DNA polymerase 2, (3) DNA polymerase 3, (4) DNA polymerase 4, (5) DNA polymerase 5, (6) DNA polymerase 6, and (7) DNA polymerase 7, as set forth in Table 14.
[0020] For example, in some variations of the above-described methods for generating a nucleic acid molecule having a predetermined sequence, two or more amplified assembly products may be pooled before performing one or more error correction steps. Additional variations may further include treating the amplified assembly products with an exonuclease before the one or more error correction steps, where the exonuclease is optionally exonuclease I.
[0021] A better understanding of the features and advantages of the subject matter described herein can be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the subject matter described herein are utilized, and the accompanying drawings. [Brief explanation of the drawings]
[0022] [Figure 1A-1B]A comparison of two nucleic acid assembly workflows is shown. Figure 1A is a schematic diagram of a standard workflow for assembling nucleic acid molecules from single-stranded overlapping oligonucleotides, including steps of oligonucleotide synthesis, oligonucleotide assembly PCR, and assembly PCR of the reaction mixture to generate subfragments (collectively referred to as primary assembly PCR); amplification of the assembly products (primary amplification); purification of the amplified products; nuclease treatment, for example, to generate complementary overhangs (e.g., by type IIs endonuclease-mediated cleavage); and vector insertion and conversion. Figure 1B is a schematic diagram of one variation of the sequence extension and ligation reaction according to the method described herein. Because the assembly PCR (primary assembly PCR), amplification (primary amplification), and vector insertion steps can be performed in a single closed vessel (e.g., a single sealed tube), such reactions are often performed as "one-pot" reactions. In the workflow of Figure 1B, the vector ends serve as amplification primers. [Figure 2] Schematic diagram of a PCR-based process for assembling and amplifying nucleic acid molecules. (a) Overlapping forward and reverse oligonucleotides are extended in the first PCR cycle. (b) The extended assembly products anneal to each other and are further extended in the second cycle. (c) Further extension occurs in subsequent PCR cycles, resulting in the accumulation of assembly products. The assembly process in this diagram is referred to herein as "primary assembly PCR" (labeled "A"). The two terminal oligonucleotides (1) and (2) can also be universal primers. Furthermore, the terminal oligonucleotides can be added to the primary assembly PCR product, or the primary assembly PCR product can be added to a separate tube and then mixed with the terminal primers. Furthermore, vector ends can be used instead of the terminal oligonucleotides (see Figure 1B). The final amplification step in this diagram, which uses the two terminal oligonucleotides, is referred to herein as "primary amplification" (labeled "B"). [Figure 3] FIG. 1 is a schematic diagram of an exemplary workflow for the synthesis of error-corrected nucleic acid molecules. [Figure 4] FIG. 1 shows a workflow schematic in which oligonucleotides are amplified, then error corrected and assembled into longer nucleic acid molecules. [Figure 5A-5B] 5A-5C show workflow schematics involving dual error correction and amplification-based assembly of PCR-generated nucleic acid molecules (e.g., previously assembled nucleic acid molecules). In one variation (FIG. 5A), error correction is performed using one or more endonucleases at two locations in the workflow. For reference purposes, nine line number labels are included in FIG. 5A. In another variation (FIG. 5B), error correction at two different locations in the workflow is performed using one or more endonucleases in the first round and a mismatch-binding protein in the second round. Similar to FIG. 5A, nine line number labels are included in FIG. 5B for reference purposes. [Figure 6] 1 shows a schematic representation of a workflow using beads conjugated with mismatch-binding proteins to separate nucleic acid molecules containing mismatches from nucleic acid molecules that do not contain mismatches. NMM refers to non-mismatched nucleic acid molecules, and MM refers to mismatched nucleic acid molecules. [Figure 7]The figure shows error rate data (total errors) generated using various experimentally determined conditions. For this figure, the term "assembly" refers to the primary assembly PCR (see, e.g., the top of Figure 2, labeled A). The term "amplification" refers to the primer-based primary amplification of the assembled nucleic acid molecule (see, e.g., the bottom of Figure 2, labeled B). The term "error correction" refers to whether a post-primary amplification T7 endonuclease I (T7NI)-mediated error correction step, in this case a secondary amplification, was performed. The designations in the "assembly" and "amplification" columns indicate whether the wild-type mismatched endonuclease from Thermococcus kodakarensis (Ishino et al., Nucl. Acids Res. 44:2977-2989 (2016)) (referred to herein as "TkoEndoMS") was included during the assembly PCR and / or amplification. The column labeled "sequenced fragments" refers to the number of sets of fragments with different sequences tested. The "error rate" shown is the average of the data. The term "benchmark" refers to the error rate determined in a separate experiment using the same oligonucleotides but without error correction. The table also shows the numerical average of all eight benchmark values. Note: Runs 1 through 8 were each performed with a set of oligonucleotides with different nucleotide sequences to allow for single-run, next-generation sequencing. [Figure 8] 8 is a graphical representation showing the total error data points used to generate the data of FIG. 7. The numbers and letter descriptions on the bottom axis of FIG. 8 correlate with the two columns on the left of FIG. 7. Each data point represents the number of errors per base pair for each of the analyzed nucleic acid molecule populations. The boxes above each vertical line represent the area of the vertical line where half of the data points fall. The horizontal line within the box represents the median. This diagram shows the total number of errors present in individual nucleic acid molecules. Thus, each data point represents the average number of errors for nucleic acid molecules designed to have the same nucleotide sequence. Analysis shows that the further away from the bottom axis you are, the fewer errors there are. [Figure 9]A graphical representation similar to that of Figure 8, but showing the number of deletions instead of total errors. [Figure 10] A graphical representation similar to Figure 8, but showing the number of insertions instead of total errors. [Figure 11] A graphical representation similar to Figure 8, but showing the number of substitutions instead of total errors. [Figures 12A-12D] The specific types of errors present in the two samples are shown. In one sample (Figures 12A and 12B), the nucleic acid molecules were assembled and amplified without error correction. In the other sample (Figures 12C and 12D), the nucleic acid molecules were assembled and amplified with TkoEndoMS error correction. T7NI error correction was not performed in either sample. The mismatch types listed in Figures 12B and 12D are as follows: TS1 = GT, CA; TS2 = AC, GT; TV1 = CT, GA; TV2 = AA, TT; TV3 = GG, CC; and TV4 = TC, AG. "TS" refers to transition, and "TV" refers to transversion. The overall error rates were as follows: Figure 12A—1 in 349 bases (standard deviation (SD): 1 in 99 bases) and Figure 12C—1 in 488 bases (SD: 1 in 210 bases). The overall substitution rates were as follows: Figure 12B - 1 in 647.8 bases, Figure 12D - 1 in 242.5 bases. [Figure 13] Data generated when a sample set of nucleic acid molecules was assembled and amplified with no error correction and with PHUSION™ DNA polymerase (Thermo Fisher Scientific, catalog number F530S) (A-C) and when TkoEndoMS was used during assembly PCR and amplification and PLATINUM™ SUPERFI™ II DNA polymerase (Thermo Fisher Scientific, catalog number 12361010) (D-F) was used for both assembly PCR and amplification are shown. T7NI error correction was not performed on either sample. [Figures 14A-14D]The specific types of errors present in the two samples are shown. In one sample (Figures 14A and 14B), nucleic acid molecules were assembled (primary assembly PCR) and amplified (primary amplification without error correction and using PHUSION™ DNA polymerase). In the other sample (Figures 14C and 14D), nucleic acid molecules were assembled and amplified with TkoEndo MS error correction and PLATINUM™ SUPERFI™ II DNA polymerase. T7NI error correction was not performed in either sample. The overall error rates were as follows: Figure 14A - 1 in 251 bases (standard deviation (SD): 1 in 25 bases) and Figure 14C - 1 in 670 bases (SD: 1 in 112 bases). The overall substitution rates were as follows: Figure 14B - 1 in 462.4 bases, Figure 14D - 1 in 565.2 bases. [Figure 15] The amino acid sequence of TkoEndoMS (SEQ ID NO: 1) with an N-terminal signal peptide and a C-terminal histidine purification tag, and the nucleotide sequence of the codon-optimized nucleic acid molecule encoding this protein (SEQ ID NO: 2) are shown. [Figure 16] 1 shows an amino acid sequence alignment of Thermococcus kodakarensis EndoMS (herein referred to as "TkoEndoMS") (SEQ ID NO: 3) and Pyrococcus furiosus EndoMS (herein referred to as "PfuEndoMS") (SEQ ID NO: 4). The amino acid sequences of these two proteins share 69% sequence identity. [Figure 17A]Data from 30 nucleic acid molecules assembled using PHUSION™ ("before") or PLATINUM™ SUPERFI™ II ("after") DNA polymerase is shown. The figure shows the relative change in error rate of individual fragments before vs. after. The actual error rate and standard deviation for individual fragments was 1 in 339±52 base pairs (bps) for before and 1 in 447±89 bps for after, representing an average improvement in error rate of 32.3±20.1%. PLATINUM™ SUPERFI™ II DNA polymerase is shown to result in a lower error rate compared to PHUSION™ DNA polymerase. [Figure 17B] The same data as in Figure 17A is shown separated by error type (deletion, insertion, and substitution). PLATINUM™ SUPERFI™ II Polymerase is shown to have a similar positive effect on all error types. The overall deletion rate change is 40.4 ± 55.1% (1 / 1157 ± 840 bps to 1 / 1429 ± 547 bps). The overall insertion rate change is 41.9 ± 90.6% (1 / 2875 ± 1201 bps to 1 / 3803 ± 2841 bps). The overall substitution rate change is 32.7 ± 21.2% (1 / 666 ± 115 bps to 1 / 873 ± 152 bps). [Figure 17C] Data from 25 nucleic acid molecules assembled using PHUSION™ ("before") or PLATINUM™ SUPERFI™ II ("after") DNA polymerase and TkoEndoMS ("after") are shown. These 25 fragments differed from the 30 fragments used to generate the data shown in Figures 17A and 17B. The figure shows the relative change in error rate for individual fragments before versus after. The actual error rate and standard deviation for individual fragments was 1 in 332 ± 68 bp for before and 1 in 534 ± 161 bp for after, representing an average improvement in error rate of 60.3 ± 32.9%. The addition of TkoEndoMS is shown to further improve the error rate. [Figure 17D]The same data as in Figure 17C are shown separated by error type (deletion, insertion, and substitution). The addition of TkoEndoMS demonstrates a positive effect on insertions and substitutions. The overall deletion mutation rate is 44.4 ± 51.3% (1 / 1019 ± 261 bps to 1 / 1397 ± 392 bps). The overall insertion mutation rate is 78.3 ± 109.7% (1 / 2690 ± 1191 bps to 1 / 4075 ± 1517 bps). The overall substitution mutation rate is 77.6 ± 36.5% (1 / 681 ± 150 bps to 1 / 1217 ± 380 bps). DETAILED DESCRIPTION OF THE INVENTION
[0023] definition The term "nucleic acid molecule," as used herein, refers to a covalently linked sequence of nucleotides or bases (e.g., ribonucleotides in RNA and deoxyribonucleotides in DNA, but also DNA / RNA hybrids where the DNAs are on separate strands or the same strand), in which the 3' position of the pentose of one nucleotide is joined to the 5' position of the pentose of the next nucleotide by a phosphodiester linkage. Nucleic acid molecules can be single-stranded, double-stranded, or partially double-stranded. Nucleic acid molecules may appear in linear or circular form, in supercoiled or relaxed form, with blunt or sticky ends, and may contain "nicks." Nucleic acid molecules may consist of fully complementary single strands or partially complementary single strands that form at least one base mismatch. Nucleic acid molecules may further comprise two self-complementary sequences, which may form a double-stranded stem region, optionally separated at one end by a loop sequence. The two regions of a nucleic acid molecule comprising the double-stranded stem region are substantially complementary to each other, resulting in self-hybridization. However, the stem may contain one or more mismatches, insertions, or deletions.
[0024] Nucleic acid molecules can include chemically, enzymatically, or metabolically modified forms of nucleotides or combinations thereof. Chemically synthesized nucleic acid molecules can refer to nucleic acids typically 200 nucleotides or less in length (e.g., 5-200, 10-150, 15-100, or 20-50 nucleotides in length), while enzymatically synthesized nucleic acid molecules can include smaller and larger nucleic acid molecules described elsewhere herein. Enzymatic synthesis of nucleic acid molecules can involve a stepwise process using enzymes such as polymerases, ligases, exonucleases, endonucleases, recombinases, etc., or combinations thereof. Accordingly, compositions and associated methods related to the enzymatic assembly of chemically synthesized nucleic acid molecules are provided in part herein.
[0025] Because the phosphodiester linkage occurs between the 5' and 3' carbons of the pentose ring of the substituted mononucleotide, the nucleic acid molecule has a "5' end" and a "3' end." The end of a nucleic acid molecule where the new linkage is at the 5' carbon is its 5'-terminal nucleotide. The end of a nucleic acid molecule where the new linkage is at the 3' carbon is its 3'-terminal nucleotide. As used herein, a terminal nucleotide or base is the nucleotide at the end of the 3' or 5' end. A nucleic acid molecule region can be said to have a 5' end and a 3' end, even if it is internal to a larger nucleic acid molecule (e.g., a sequence region within a nucleic acid molecule). Nucleic acid molecules also refer to short nucleic acid molecules, often referred to as primers or probes, for example. The terms "5'-" and "3'-" also refer to the strands of a nucleic acid molecule. Thus, a linear single-stranded nucleic acid molecule has a 5' end and a 3' end. However, a linear double-stranded nucleic acid molecule has a 5' end and a 3' end of each strand. Thus, a nucleic acid molecule encoding a protein can be referred to, for example, as the 3' end of the sense strand.
[0026] The term "oligonucleotide" as used herein refers to DNA and RNA, as well as any other type of nucleic acid molecule, typically DNA but also N-glycosides of purine or pyrimidine bases. Thus, oligonucleotides are a subset of nucleic acid molecules and can be single-stranded or double-stranded. Oligonucleotides (including primers, as described below) may be referred to as "forward" or "reverse" to indicate their orientation relative to a given nucleic acid sequence. For example, a forward oligonucleotide may represent a portion of the sequence of the first strand (e.g., the "sense" strand) of a nucleic acid molecule, while a reverse oligonucleotide may represent a portion of the sequence of the second strand (e.g., the "antisense" strand) of the nucleic acid molecule, or vice versa. In many cases, a set of oligonucleotides used to assemble a longer nucleic acid molecule includes both forward and reverse oligonucleotides that can hybridize to each other through complementary regions. Oligonucleotides are typically less than 200 nucleotides in length, more typically less than 100 nucleotides in length. Thus, "primers" generally fall within the category of oligonucleotides. Oligonucleotides can be prepared by any suitable method, including direct chemical synthesis by methods such as the phosphotriester method of Narang et al., Meth. Enzymol. 68:90-99 (1979), the phosphodiester method of Brown et al., Meth. Enzymol. 68:109-151 (1979), the diethyl phosphoramidite method of Beaucage et al., Tetrahedron Letters 22:1859-1862 (1981), and the solid support method of U.S. Patent No. 4,458,066. A review of the synthesis methods for conjugates of oligonucleotides and modified nucleotides is provided in Goodchild, Bioconjugate Chemistry 1:165-187 (1990). Where necessary, the term oligonucleotide can refer to primer or probe, and these terms can be used interchangeably herein.
[0027] The term "primer," as used herein, refers to a short nucleic acid molecule capable of acting as an initiation point for nucleic acid synthesis under suitable conditions. Such conditions include those in which synthesis of a primer extension product complementary to a nucleic acid strand is induced in an appropriate buffer and at a suitable temperature in the presence of different nucleoside triphosphates (e.g., A, C, G, T, and / or U) and an agent for extension (e.g., DNA polymerase or reverse transcriptase). Primers generally consist of single-stranded DNA, but can also be provided as double-stranded molecules for certain applications (e.g., blunt-end ligation). Optionally, primers can be naturally occurring or synthesized using recombinant chemical synthesis. The appropriate length of a primer depends on the intended use of the primer but typically ranges from about 6 to about 200 nucleotides, with intermediate ranges such as about 10 to about 50 nucleotides, about 15 to about 35 nucleotides, about 18 to about 75 nucleotides, and about 25 to about 150 nucleotides. The design of primers suitable for amplifying a given target sequence is well known in the art and described in the literature (see, for example, OLIGOPERFECT™ Designer, Thermo Fisher Scientific).Primer can incorporate additional features that allow the primer to be detected or immobilized, but do not change the basic property of primer, which is to act as the starting point for DNA synthesis.Therefore, primer can comprise detectable moiety or label.For example, label can comprise fluorescent, luminescent or radioactive moiety.
[0028] A set of primers used in the same amplification reaction can have substantially the same melting temperature, where the melting temperatures are within about 10-5°C of each other, or within about 5-2°C of each other, or within about 2-0.5°C of each other, or less than about 0.5°C of each other.
[0029] The terms "complementary" or "complementarity," as used herein, refer to the natural binding of nucleic acid molecules (such as primers, oligonucleotides, or polynucleotides) under permissive salt and temperature conditions through base pairing. For example, the sequence "AGT" binds to the complementary sequence "TCA." Complementarity between two single-stranded molecules can be "partial," in which only a portion of the nucleic acid binds, or "complete," in which complete complementarity exists between the single-stranded molecules. The degree of complementarity between nucleic acid strands significantly affects the efficiency and strength of hybridization between nucleic acid strands. This is particularly important in amplification reactions, which rely on binding between nucleic acid strands. Complementary regions between nucleic acid molecules, such as oligonucleotides, are sometimes referred to as "overlap" or "overlapping" regions, as defined below.
[0030] The term "hybridization," as used herein, refers to any process by which a strand of nucleic acid binds with a complementary strand through base pairing. Hybridization and the strength of hybridization (e.g., the strength of the association between nucleic acids) depend on the degree of complementarity between the nucleic acids, the stringency of the associated conditions, the T of the hybrid formed, and the degree of complementarity between the nucleic acids. m and the G:C ratio within the nucleic acid.
[0031] The term "homologous," as used herein, refers to the degree of complementarity. Nucleic acid sequences can be partially or fully homologous (identical). A partially complementary sequence, which at least partially inhibits a fully complementary sequence from hybridizing to a target nucleic acid, is referred to using the functional term "substantially homologous."
[0032] The term "overlap" or "overlapping," as used herein, refers to sequence homology or sequence identity shared by portions of two or more oligonucleotides.
[0033] The term "gene" or "gene sequence," as used herein, generally refers to a nucleic acid sequence that encodes a distinct cellular product. In many cases, a gene or gene sequence comprises a DNA sequence containing an open reading frame (ORF), which can be transcribed into mRNA, which can be translated into a polypeptide chain, or into rRNA or tRNA, or which can serve as a recognition site for enzymes and proteins involved in DNA replication, transcription, and regulation. These genes include, but are not limited to, structural genes, immune genes, regulatory genes, and secretory (transport) genes. However, as used herein, "gene" refers not only to a nucleotide sequence that encodes a specific protein, but also to any adjacent 5' and 3' non-coding nucleotide sequences involved in regulating the expression of the protein encoded by the gene of interest. These non-coding sequences include terminator sequences, promoter sequences, upstream activator sequences, regulatory protein binding sequences, and the like. In many cases, genes are assembled from shorter oligonucleotides or nucleic acid fragments.
[0034] The terms "fragment," "subfragment," "segment," or "building block," or similar terms, when used herein in reference to nucleic acid molecules or sequences, refer to either a product or intermediate product resulting from one or more process steps (e.g., synthesis, assembly PCR, amplification, etc.), or to a portion, part, or template of a longer or modified nucleic acid product to be obtained by one or more process steps (e.g., assembly PCR, amplification, ligation, cloning). In some cases, a nucleic acid fragment or subfragment can represent both an assembly product (e.g., assembled from multiple oligonucleotides) and a starting compound for higher-order assembly (e.g., a gene assembled from multiple fragments or a fragment assembled from multiple subfragments, etc.).
[0035] As used herein, "amine" or "amine compound," as used herein, includes a chemical entity of Formula I immediately below, or a salt thereof: [ka]
[0036] wherein R1 is H, R2 is selected from alkyl, alkenyl, alkynyl, or (CH2)n-R5, where n=1-3, R5 is aryl, amino, thiol, mercaptan, phosphate, hydroxy, alkoxy, and R3 and R4 may be the same or different and are independently selected from H or alkyl, with the proviso that when R2 is (CH2)n-R5, at least one of R3 and / or R4 is alkyl. Thus, amines include diethylamine hydrochloride, diisopropylamine hydrochloride, ethyl(methyl)amine hydrochloride, trimethylamine hydrochloride, and dimethylamine hydrochloride.
[0037] The term "vector," as used herein, refers to any nucleic acid molecule capable of transferring genetic material to a host organism. Vectors can be linear or circular in topology and include, but are not limited to, plasmids, viruses, and bacteriophages. Vectors can contain amplification genes, enhancers, or selection markers and may or may not integrate into the genome of the host organism.
[0038] The term "plasmid," as used herein, refers to a vector that can be genetically modified to insert one or more nucleic acid molecules (e.g., assembly products). Plasmids typically contain one or more regions that allow them to replicate in at least one cell type.
[0039] The term "amplification," as used herein, refers to the production of additional copies of a nucleic acid molecule. Amplification is often carried out using polymerase chain reaction (PCR) techniques well known in the art (see, e.g., Dieffenbach, C. Wand and G. S. Deksler (1995) PCR Primer, a Laboratory Manual, Cold Spring Harbor Press, Plainview, NY), but may also be carried out by other means, including isothermal amplification methods such as transcription-mediated amplification, strand displacement amplification, rolling circle amplification, loop-mediated isothermal amplification, helicase-dependent amplification, single-primer isothermal amplification, or recombinase polymerase amplification (see, e.g., Fakruddin et al., "Nucleic acid amplification: Alternative methods of polymerase chain reaction", J. Pharm Bioallied Sci, 2013, v. 5(4), 245-252, or Gill and Ghaemi, "Nucleic acid isothermal amplification technologies: a review", Nucleosides Nucleotides Nucleic Acids. 2008 27(3), 224-43). To reconstitute each strand of the denatured double-stranded nucleic acid molecule, an amplification reaction can be carried out using the terminal primers.
[0040] The term "assembly chain reaction," also referred to herein as "assembly PCR," as used herein, refers to the assembly of larger nucleic acid molecules from smaller nucleic acid molecules by polymerase-mediated extension of overlapping, partially complementary nucleic acid molecules. The overlapping, partially complementary nucleic acid molecules can be single-stranded or double-stranded. Furthermore, the double-stranded nucleic acid molecules are typically denatured prior to or as part of their use in the assembly chain reaction. An example of an assembly chain reaction is depicted at the top of Figure 2, in which overlapping, partially complementary nucleic acid molecules are used to generate larger nucleic acid molecules at each polymerase-mediated extension step.
[0041] The term "primary post-amplification error correction," as used herein, refers to an amplification-based error correction step that occurs after the completion of the workflow shown in FIG. 2. In the workflow of FIG. 2, oligonucleotides are first assembled (primary assembly PCR) and then amplified using terminal primers (primary amplification). When this occurs, additional rounds of error correction (e.g., error correction involving PCR-based fragment assembly and amplification) may occur. For example, in the workflow of FIG. 5A, if the three subfragments / PCR products in step 1 were generated using the workflow of FIG. 2, then all error correction steps in FIG. 5A are primary post-amplification error correction.
[0042] Error correction often involves the use of a mismatch endonuclease. An exemplary error correction process is depicted in Figure 4. In this diagram, double-stranded nucleic acid molecules assembled from amplified oligonucleotides are denatured and then reannealed (lines 4 and 5). The reannealed nucleic acid molecules, some of which may contain one or more mismatches, are then contacted with, for example, a mismatch endonuclease (line 6) to cleave the nucleic acid molecules at or near the sites of the mismatches. The cleaved nucleic acid molecules in the reaction mixture on line 6 are then reassembled and amplified by overlap-extension PCR to obtain error-free nucleic acid molecules (output of the process on line 7) that are intended to be the same length as the "uncorrected" starting nucleic acid molecule (line 3).
[0043] The term "non-amplification error correction," as used herein, refers to an error correction process that does not involve nucleic acid amplification. An example of such a method is a method in which nucleic acid strands are hybridized to each other, followed by the use of mismatch-binding proteins to remove double-stranded nucleic acid molecules containing mismatches (see, e.g., Figure 3).
[0044] The term "adjacent," as used herein, refers to a position in a nucleic acid molecule immediately 5' or 3' to a reference region.
[0045] The term "sequence fidelity" as used herein refers to the level of sequence identity of a nucleic acid molecule compared with a reference sequence.Complete identity is 100% identical across the entire length of the nucleic acid molecule that is scored for sequence identity.Sequence fidelity can be measured in many ways, for example, by comparing the actual nucleotide sequence of a nucleic acid molecule with the desired nucleotide sequence (for example, the nucleotide sequence that is intended to be used to generate a nucleic acid molecule).Another way to measure sequence fidelity is by comparing the sequences of two nucleic acid molecules in a reaction mixture.In many cases, the difference between each base is the same on average.
[0046] The error rate of a DNA polymerase can be measured by quantifying the total number of errors or different types of errors. For the high-fidelity DNA polymerases described herein, the error rate "benchmark" is set based on the substitution rate. In particular, the high-fidelity DNA polymerases have an error rate of 1.0 × 10 per base. -5 This indicates a lower substitution error rate.Examples of high-fidelity polymerases include PHUSION™ DNA polymerase, PLATINUM™ SUPERFI™ II DNA polymerase, Q5™ DNA polymerase, and PRIMESTAR™ GXL DNA polymerase (Takara).Methods for determining error rate are known in the art, and are described, for example, in Potapov et al., "Examining Sources of Error in PCR by Single Molecule Sequencing", PLOS ONE, DOI:10.1371 / journal.pone.0169774 January 6, 2017.
[0047] The term "transition," when used in reference to a nucleotide sequence of a nucleic acid molecule, refers to a transition from a purine nucleotide to another purine nucleotide.
number
number
[0048] The term "transversion", when used in reference to a nucleotide sequence of a nucleic acid molecule, refers to a point mutation involving a substitution of a (bicyclic) purine for a (monocyclic) pyrimidine or a (monocyclic) pyrimidine for a (bicyclic) purine.
[0049] The term "indel," as used herein, refers to an insertion or deletion of one or more bases in a nucleic acid molecule.
[0050] The term "mismatch," as used herein, refers to two bases in different strands of a double-stranded nucleic acid molecule that do not form Watson-Crick base pairs, but have sequence complementarity with surrounding bases in different nucleic acid strands, forming Watson-Crick base pairing. The length of the complementary region can vary, but is often at least 20 base pairs. For each strand of a nucleic acid molecule containing only the four standard DNA bases, there are four correct (Watson-Crick base pairing) complementary matches (i.e., A / T, T / A / G / C, and C / G) and 12 "mismatches" (i.e., A / A, A / C, A / G, T / T, T / C, T / G, G / G, G / A, G / T, C / C, C / T, and C / A). In terms of base pairing, in the absence of strand reference, there are two correct complementary matches (i.e., A / T and G / C) and eight "mismatches" (i.e., A / A, A / C, A / G, T / T, T / C, T / G, G / G, and C / C). In terms of substitution, these mismatches can be represented as: (1) A to G and T to C, (2) G to A and C to T, (3) A to C and T to G, (4) A to T and T to A, (5) G to C and C to G, and (6) G to T and C to A.
[0051] The term "thermostable," as used herein with respect to a protein, refers to a protein that retains at least 85% of its biological activity after being heated to 95°C for 5 minutes. A thermostable protein may or may not have biological activity at 95°C. Thus, depending on the protein, an assay of retained biological activity may be performed after incubation at 95°C for 5 minutes or at another (e.g., lower) temperature, using as a "benchmark" the same protein that has not been heated to 95°C for 5 minutes.
[0052] The term "mismatch recognition protein," as used herein, refers to a protein that has specific biological activity toward mismatched bases in double-stranded DNA. These activities can include nuclease activity and / or binding activity. Such proteins include resolvases, MutS and MutS homologs, MutM and MutM homologs, MutY and MutY homologs, and members of the RecB nuclease family of proteins. Mismatch binding proteins and mismatch endonucleases are both mismatch recognition proteins. Mismatch recognition proteins can be thermostable or non-thermostable. Some exemplary mismatch recognition proteins are listed in Table 15 and other tables provided herein.
[0053] The term "mismatch endonuclease" or "MME" (also referred to as "mismatch repair endonuclease"), as used herein, refers to a nuclease that has the activity of cleaving a double-stranded nucleic acid molecule (one or both strands) at or near (e.g., within about 1 to about 5 base pairs) a mismatch site. Mismatch endonuclease activity includes the ability to cleave the phosphodiester bond at or near the nucleotide that forms the mismatched base pair, as well as the activity of cleaving the phosphodiester bond adjacent to the nucleotide located 1 to 5, and often 1 to 3, base pairs away from the mismatched base pair. Examples of proteins with mismatch endonuclease activity are listed below in Tables 13 and 15. Specific examples of mismatched endonucleases include CEL I (Till et al., Nucl. Acid Res. 32:2632-2641 (2004)) and CEL II (U.S. Patent No. 7,129,075), bacteriophage resolvases such as T7NI and T4 endonuclease VII (Mashal, et al., Nature Genetics 9:177-183 (1995)), E. coli endonuclease V (Yao and Kow, J. Biol. Chem. 272:30774-30779 (1997)), TkoEndoMS (Ishino et al., Nucl. Acids Res. 44:2977-2986 (2016)), and Pyrococcus furiosus EndoMS (referred to herein as "PfuEndoMS"). The mismatched endonuclease can be thermostable (TsMME) or non-thermostable.
[0054] The term "EndoMS," as used herein, refers to a mismatch-specific endonuclease that shares at least 50% amino acid sequence identity with one or more of the EndoMS proteins listed in Table 15 and has mismatch-specific endonuclease activity. "Nucs" has been used in the art as an alternative term for EndoMS. Thus, the terms "EndoMS" and "Nucs" can be used interchangeably.
[0055] The term "mismatch-binding protein" (also referred to as "mismatch repair-binding protein"), as used herein, refers to a protein that has specific binding activity to mismatched bases in double-stranded DNA. Examples of such proteins are listed in Tables 12 and 15 below. Many of these proteins are MutS homologs. Mismatch-binding proteins can be thermostable or non-thermostable.
[0056] The term "error correction," as used herein, refers to a process designed to reduce the total number of nucleotide sequence defects in a population of nucleic acid molecules. These defects can be mismatches, insertions, deletions, and / or substitutions. A defect can occur when nucleic acid molecules generated (e.g., by chemical or enzymatic synthesis) are each intended to contain a particular base at a certain position, but a different base is present at that position in one or more of the nucleic acid molecules.
[0057] An example of error correction is as follows: Assume there is a population of double-stranded nucleic acid molecules with a desired length of 100 base pairs. Assume also that the two strands of the double-stranded nucleic acid molecules are each synthesized separately and hybridized with each other to form the double-stranded nucleic acid molecules of the population. Assume further that nucleic acid synthesis produces an average of one error per 200 nucleotides. In this case, there is one "error" per 100 base pairs. Therefore, on average, each double-stranded nucleic acid molecule of the population contains one error. Of course, some double-stranded nucleic acid molecules in the population are error-free, while other double-stranded nucleic acid molecules have more than one error. If the error correction process removes half of the nucleic acid molecules from the population and none of the error-free nucleic acid molecules are removed, the error rate of the remaining double-stranded nucleic acid molecules in the population will be less than one per 200 base pairs. This is because, as suggested above, some of the removed nucleic acid molecules have more than one error, and no "correct" nucleic acid molecules are removed.
[0058] As used herein, the phrases "error correction round" and "round of error correction" refer to a series of steps that result in the cleavage or removal of nucleic acid molecules with errors from a population of nucleic acid molecules. Using FIG. 4 for illustrative purposes, lines 4-7 describe one round of error correction. While the process described in FIG. 4 includes a series of amplification reactions (e.g., PCR cycles), a round of error correction does not necessarily require this. For example, a modification of the process described in FIG. 4 is where a mismatch-binding protein can be used to separate nucleic acid molecules with mismatches from nucleic acid molecules without mismatches (see line 5).
[0059] As used herein, an "error-reducing polymerase reagent" is a composition comprising a polymerase (e.g., a DNA polymerase) and an additional component that reduces the number of errors in an amplified nucleic acid molecule (e.g., by about 5% to about 30%, about 5% to about 30%, about 5% to about 30%, about 10% to about 40%, about 10% to about 70%, etc.), where the additional component is not a mismatch-recognition protein. One category of such compounds are amines, such as the amines described herein.
[0060] The term "transformation," as used herein, describes the process by which exogenous nucleic acid molecules enter and change recipient cells. It can occur under natural or artificial conditions using a variety of methods well known in the art. Transformation can rely on any known method for inserting foreign nucleic acid sequences into prokaryotic or eukaryotic host cells. The method is selected based on the host cell to be transformed and can include, but is not limited to, viral infection, electroporation, lipofection, and particle bombardment. Such "transformed" cells include stably transformed cells in which the inserted nucleic acid can replicate as an autonomously replicating plasmid or as part of the host chromosome. They also include cells that transiently express the inserted DNA or RNA for a limited period of time.
[0061] The term "solid support," as used herein, refers to a porous or non-porous material on which polymers, such as oligonucleotides or nucleic acid molecules, can be synthesized and / or immobilized. As used herein, "porous" means that the material contains pores that can be of non-uniform or uniform diameter (e.g., in the nm range). Porous materials include paper, synthetic filters, and the like. In such porous materials, reactions can occur within the pores. Solid supports can have any one of many shapes, such as pins, strips, plates, disks, rods, fibers, bends, cylindrical structures, flat, concave or convex surfaces, or capillaries or columns. Solid supports can be particles, including beads, microparticles, nanoparticles, and the like. Solid supports can be non-bead-type particles (e.g., filaments) of similar size. Supports can vary in width and size. For example, the size of beads (e.g., magnetic beads) that can be used in practicing embodiments of the methods described herein can vary widely, but can be in the range of 0.01 μm to 100 μm, 0.005 μm to 100 μm, 0.005 μm to 10 μm, 0.01 μm to 100 μm, 0.01 μm to 1,000 μm, 1.0 μm to 2.0 μm, 1.0 μm to 100 μm, 15 μm to 200 μm, 20 μm to 300 μm, 30 μm to 400 μm, 40 μm to 500 μm, 50 μm to 600 μm, 60 μm to 700 μm, 70 μm to 800 μm, 80 μm to 900 μm, 90 μm to 1000 μm, 100 μm to 1500 μm, 100 μm to 20 ... The beads may have a diameter of 2.0 μm to 100 μm, 3.0 μm to 100 μm, 0.5 μm to 50 μm, 0.5 μm to 20 μm, 1.0 μm to 10 μm, 1.0 μm to 20 μm, 1.0 μm to 30 μm, 10 μm to 40 μm, 10 μm to 60 μm, 10 μm to 80 μm, or 0.5 μm to 10 μm.
[0062] The support can be hydrophobic or can bind to molecules via hydrophobic interactions. The support can be hydrophilic or hydrophilic, and includes inorganic powders such as silica, magnesium sulfate, and alumina, natural polymeric materials, particularly cellulose and cellulose-derived materials, such as fibers containing paper, such as filter paper and chromatography paper. The support can be immobilized at addressable positions on a carrier, such as a multi-well plate or a microchip. The support can be discrete or particulate (e.g., resin material or beads in a well), or can be reversibly immobilized or linked to the carrier (e.g., by cleavable chemical bonds or magnetic forces). In some embodiments, the solid support can be fragmentable. The solid support can be a synthetic or modified natural polymer, such as nitrocellulose, carbon, cellulose acetate, polyvinyl chloride, polyacrylamide, cross-linked dextran, agarose, polyacrylate, polyethylene, polypropylene, poly(4-methylbutene), polystyrene, polymethacrylate, poly(ethylene terephthalate), nylon, poly(vinyl butyrate), polyvinylidene difluoride (PVDF) membrane, glass, controlled pore glass, magnetically controlled pore glass, magnetic or non-magnetic beads, ceramic, or metal, used alone or in combination with other materials. In some embodiments, the support can be in a chip, array, microarray, or microwell plate format. In many cases, the support used in the methods or compositions described herein is one on which individual nucleic acid molecules are synthesized separately or in distinct regions to generate features (i.e., locations containing individual nucleic acid molecules) on the support. In some embodiments, the size of the defined features is selected to allow for the formation of microvolume droplets or reaction volumes on the features, with each droplet or reaction volume being kept separate from each other. As described herein, features are typically, but not necessarily, separated by an intermediate feature space to prevent droplet or reaction volume or fusion between two adjacent features. The intermediate feature typically does not carry nucleic acid molecules on its surface and corresponds to inert space.In some embodiments, features and intermediate features may differ in their hydrophilic or hydrophobic properties. In some embodiments, features and intermediate features may include modifiers. In some cases described herein, the features are wells, microwells, or notches. Nucleic acid molecules may be covalently or non-covalently bound to the surface, or deposited, synthesized, or assembled on the surface.
[0063] The singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.
[0064] overview The compositions and methods described herein relate, in part, to the preparation of nucleic acid molecules with high sequence fidelity. Numerous embodiments and variations can be used, but in many cases, nucleic acid molecules are synthesized (e.g., chemically, enzymatically, etc.). These synthesized nucleic acid molecules can then optionally be assembled, for example, by assembly PCR (e.g., primary assembly PCR) to form one or more larger nucleic acid molecules. Figures 1A and 1B are schematic diagrams showing exemplary assembly PCR steps that can be used in the methods described herein.
[0065] Generally, the occurrence of sequence errors in synthesized oligonucleotides is relatively low, and has a semi-random distribution.In many cases, when the nucleic acid molecule with erroneous base (for example, deletion, insertion, substitution) hybridizes with the nucleic acid molecule with correct base, a region that does not show standard Watson-Crick base pairing is formed.These "non-standard" regions can be used to recognize the nucleic acid molecule that contains errors.Furthermore, when these "non-standard" regions are detected in a group of nucleic acid molecules, the nucleic acid molecules that contain these regions can be removed from the group, or can be modified in a way that prevents their amplification or reduces their ability to be amplified.
[0066] Many methods can be used to reduce the percentage of nucleic acid molecules in a population that contain errors (e.g., deletions, insertions, substitutions). These methods include: 1. Cleavage of nucleic acid molecules containing errors, 2. Separating error-containing nucleic acid molecules from error-free nucleic acid molecules; 3. Suppressing / inhibiting the amplification of error-containing nucleic acid molecules compared to error-free nucleic acid molecules.
[0067] Additionally, two or more of the above methods may be used to reduce the number of errors present in a nucleic acid molecule.
[0068] Much of the disclosure herein relates to compositions and methods for the synthesis, assembly (e.g., assembly PCR), and amplification of nucleic acid molecules. Provided herein are compositions and methods for generating nucleic acid molecules with high sequence fidelity.
[0069] For some applications, the use of nucleic acid molecules with low error rates is important. To illustrate, consider the situation in which 100 nucleic acid molecules are assembled. Each molecule is 100 base pairs long, with one error every 200 base pairs. The net result is an average of 50 sequence errors in each 10,000 base pairs of assembled nucleic acid molecules. For example, if the intention is to express one or more proteins from the assembled nucleic acid molecules, the number of amino acid sequence errors may therefore be considered too high. Furthermore, nucleotide sequence errors in many protein-coding regions can result in "frameshift" mutations, which generally result in undesirable proteins. Non-frameshift coding regions can also result in the formation of proteins with point mutations. All of this reduces the purity of the desired protein expression product, and many of the "contaminating" proteins produced will be carried over into the final expression product mixture, even with affinity purification.
[0070] High sequence fidelity can be achieved by several means, including sequencing nucleic acid fragments or partially assembled nucleic acid molecules before assembly, sequencing fully assembled nucleic acid molecules to identify nucleic acid molecules with the correct sequence, and / or error correction.
[0071] Errors can be introduced into nucleic acid molecules in many ways, including chemical synthesis errors, amplification / polymerase-mediated errors (especially when proofreading polymerases are used), and assembly PCR-mediated errors (usually occurring at nucleic acid fragment junctions).
[0072] Sequence errors in nucleic acid molecules can be referred to in many ways. For example, there is the error rate associated with synthetic nucleic acid molecules, the error rate associated with nucleic acid molecules after error correction and / or selection, and the error rate associated with final product nucleic acid molecules (for example, (1) the error rate of synthetic nucleic acid molecules that have been selected for the correct sequence, or (2) the error rate of assembled chemically synthesized nucleic acid molecules). These errors can be caused by the chemical synthesis process, assembly process, and / or amplification process. Errors can be removed or prevented by methods such as the selection of nucleic acid molecules with the correct sequence, error correction, and / or improved chemical synthesis methods.
[0073] In some cases, the methods described herein can combine error removal and prevention methods to produce nucleic acid molecules with a relatively low number of errors. Thus, the assembled nucleic acid molecules produced by the methods described herein can have an error rate of about 1 in 1,500 to about 1 in 30,000, about 1 in 2,000 to about 1 in 30,000, about 1 in 4,000 to about 1 in 30,000, about 1 in 8,000 to about 1 in 30,000, about 1 in 10,000 to about 1 in 30,000, about 1 in 15,000 to about 1 in 30,000, about 1 in 10,000 to about 1 in 20,000, etc.
[0074] Two methods for reducing the number of errors in assembled nucleic acid molecules are by (1) selecting nucleic acid molecules (e.g., oligonucleotides, subfragments, etc.) for assembly that have the correct sequence, and (2) correcting errors in nucleic acid molecules, partially assembled subassemblies, or fully assembled nucleic acid molecules.
[0075] Errors can be incorporated into nucleic acid molecules regardless of how they are generated. Even when nucleic acid molecules known to have the correct sequence are used in assembly PCR, errors can still be introduced into the final assembly product. Therefore, reducing errors is often desirable.
[0076] In many cases, regardless of the method for generating larger nucleic acid molecules from chemically synthesized oligonucleotides, errors from the chemical synthesis process will exist. While sequencing of individual nucleic acid molecules can be performed to identify and select error-free nucleic acid molecules, alternative approaches can include one or more error correction or removal steps. Therefore, in many cases, error correction is desirable. Error correction can be achieved in various ways. Typically, such error removal steps are performed after the first round of assembly PCR. Thus, in some embodiments, the methods described herein can include (in this order or in a different order): (i) fragment amplification and / or assembly PCR (e.g., by the methods described herein), (ii) error correction, (iii) final assembly (e.g., by the in vitro or in vivo methods described herein, e.g., using a protocol such as that described in Figure 1A or 1B).
[0077] Errors can be removed or otherwise avoided from nucleic acid molecules at one or more locations in the workflow used to generate these molecules. Using the illustrative workflow depicted in FIG. 1A, oligonucleotide synthesis can be performed under conditions that introduce few sequence errors. Nucleic acid assembly PCR (e.g., oligonucleotide assembly) can be performed in conjunction with mismatch recognition-based error correction. The assembled nucleic acid molecule can be amplified in conjunction with mismatch recognition-based error correction. The assembled nucleic acid molecule can also be subjected to mismatch recognition-based error correction in the absence of assembly PCR or amplification. This is often accomplished by heat denaturation of the nucleic acid molecule of interest, followed by renaturation of the nucleic acid molecule, and then contacting it with one or more mismatch recognition proteins.
[0078] Furthermore, the introduction of errors into nucleic acid molecules can be avoided or reduced by many methods.Some of these methods include using nucleic acid starting materials that contain few errors.As described in Example 2 and Table 10 and 11, using nucleic acid starting materials that contain few errors results in fewer errors existing in the assembled error-corrected molecules.It is believed that this is because error correction methods cannot always correct 100% of existing errors.Therefore, generally, the fewer errors that exist for correction, the fewer errors there will be after error correction.
[0079] In many cases, the nucleic acid molecule starting material will be from about 1 in 250 to about 1 in 2,000 (e.g., from about 1 in 250 to about 1 in 1,900, from about 1 in 250 to about 1 in 1,500, from about 1 in 250 to about 1 in 1,200, from about 1 in 250 to about 1 in 1,000, from about 1 in 250 to about 1 in 800, have an initial average number of sequence errors that is about 1 in 400 to about 1 in 1,900, about 1 in 400 to about 1 in 1,500, about 1 in 400 to about 1 in 1,100, about 1 in 650 to about 1 in 2,000, about 1 in 650 to about 1 in 1,700, about 1 in 650 to about 1 in 1,500, etc.
[0080] As also described in Example 2, error correction efficiency varies to some extent with the thermal cycling conditions used. Thus, one factor that can be altered to obtain product nucleic acid molecules with low error numbers is the thermal cycling conditions.
[0081] Another way to avoid introducing errors into nucleic acid molecules is, for example, by using synthetic methods to generate nucleic acid subunits with fewer errors. Another method is to use high-fidelity polymerases and high-fidelity amplification methods for low-error replicative assembly and amplification of nucleic acid molecules.
[0082] Using the workflow in Figure 2 for illustrative purposes, synthetically produced oligonucleotides are assembled by DNA polymerase through a series of heating and cooling steps, resulting in large nucleic acid molecules with each assembly PCR cycle. Hybridization of complementary regions of single-stranded nucleic acid molecules occurs during each assembly PCR cycle. During these hybridization reactions, regions that do not exhibit standard Watson-Crick base pairing may form, and when this occurs, the resulting double-stranded nucleic acid molecules are "marked" as containing errors. The methods for generating "error-corrected" populations of nucleic acid molecules described herein use DNA polymerase and mismatch-recognition proteins to eliminate or reduce the prevalence of error-containing nucleic acid molecules from a mixed population ("error correction").
[0083] Again using the workflow in Figure 2 for illustration, error correction can be performed at any one or more steps and elsewhere in the larger workflow (e.g., after the primary amplification shown), and can include multiple error-correction reagents and mechanisms, as well as other error-reduction methods. Furthermore, Figure 2 shows a series of assembly PCR and amplification reactions. Error correction can occur in none, some, or all of these steps. For example, Figure 2 shows four overlap extension cycles of an assembly PCR reaction (based on the number of downward arrows (a)-(c) shown). For example, if a thermostable mismatch-recognition protein is used, it can be added before the first assembly PCR cycle or during the assembly PCR reaction (i.e., after one or more of the extension cycles are complete). Examples of error-correction reagents that can be used include mismatch endonucleases and mismatch-binding proteins.
[0084] The reagents that can be used to perform error correction include mismatch endonuclease, mismatch binding protein, and high-fidelity polymerase, and reagents that contain high-fidelity polymerase.In addition, the proteins used in the methods described herein can be thermostable or non-thermostable.One example of a reagent that contains high-fidelity polymerase is PLATINUM™ SUPERFI™ II DNA polymerase (Thermo Fisher Scientific, Cat. No. 12361010).
[0085] One common workflow for error correction of nucleic acid molecules is that single-stranded nucleic acid molecules with sequence complementarity regions are hybridized with each other, or double-stranded nucleic acid molecules are denatured and then hybridized with each other.In this case, when two nucleic acid strands whose nucleotide sequences differ by one or more nucleotides are hybridized with each other, the resulting double-stranded nucleic acid molecule generally forms a region that does not show Watson-Crick base pairing.In some cases, the error correction process can be based on the recognition of the region that does not show Watson-Crick base pairing.Therefore, in many cases, the error correction process involves hybridization of single-stranded nucleic acid molecules to form double-stranded nucleic acid molecules.Although error correction can be performed in the absence of DNA polymerase, assembly PCR and amplification processes that may include error correction are shown in Figure 1A, Figure 1B and Figure 2.
[0086] The methods described herein include various combinations of error mitigation, error correction associated with assembly PCR and / or amplification steps. Furthermore, error correction processes can be integrated into such steps or occur before or after such steps.
[0087] The methods described herein can include any number and combination of steps of the workflows described herein. Using the workflows of Figures 1A, 2, and 5A and 5B to illustrate exemplary embodiments of the methods described herein, oligonucleotides with overlapping sequence-complementary ends can be generated (Figure 1A). These oligonucleotides can then be assembled through a series of assembly PCR cycles, referred to as primary assembly PCR (Figures 1A and 2). The assembly products are then amplified using terminal primers, referred to as primary amplification (Figures 1A and 2). For example, as described in Figure 2, assembly products generated in separate assembly PCR reactions with complementary end sequences can be further assembled as described at the top of Figures 5A and 5B, referred to as secondary assembly PCR. In these examples, subfragment PCR products A, B, and C are combined in a vessel to perform one-cup mismatch cleavage-based error correction, followed by PCR steps to fuse and extend the error-corrected fragments (referred to as the third PCR in line 3, respectively), resulting in a longer nucleic acid assembly product containing fragments A, B, and C. Error correction can occur before and / or after each assembly and / or amplification step.
[0088] Using the data in Figure 7 for illustration, primary assembly PCR was performed in the presence or absence of TkoEndoMS. In each case, primary amplification was followed, also in the presence or absence of TkoEndoMS. Error correction using T7NI was then performed, which included secondary amplification.
[0089] Figure 1B shows the workflow in which only the primary assembly PCR and primary amplification occur.
[0090] In summary, in some embodiments, provided herein are methods that include a combination of assembly PCR and / or amplification steps, and error correction can occur during or between any of such steps. In many cases, one or more thermostable mismatch-recognition proteins can be present during the assembly PCR and / or amplification steps.
[0091] The term "primary assembly PCR" refers to an assembly PCR reaction in which single-stranded nucleic acid molecules are assembled to form double-stranded nucleic acid molecules that are longer in length than the individual single-stranded nucleic acid molecules. Although the workflow in Figure 1B shows an assembly reaction in which single-stranded nucleic acid molecules are assembled with double-stranded nucleic acid molecules (i.e., vectors), this is considered to include primary assembly PCR because vector inserts are formed from single-stranded nucleic acid molecules. Thus, in such cases, the vector inserts are assembled via primary assembly PCR.
[0092] The term "secondary assembly PCR" refers to an assembly PCR reaction in which initial double-stranded nucleic acid molecules are assembled to form product double-stranded nucleic acid molecules that are longer in length than the initial double-stranded nucleic acid molecules.
[0093] The term "primary amplification" refers to the first set of amplification reactions performed on the products of an assembly PCR reaction in which single-stranded nucleic acid molecules are assembled to form double-stranded nucleic acid molecules. Subsequent amplification cycles are referred to as "secondary," "tertiary," "quaternary," etc. As an example, step 3 in Figure 5A is a secondary amplification. Amplification cycles after primary amplification may or may not result in amplification products of different lengths than the starting nucleic acid molecules. The workflow distinguishes between amplification cycles. For example, Figure 7 shows data obtained from a primary amplification occurring with or without TkoEndoMS. Additionally, Figure 7 shows data including error correction using T7NI followed by secondary amplification.
[0094] Generation of nucleic acid molecules One of the first steps in producing a desired nucleic acid molecule or protein is the design of the nucleic acid molecule after the molecule has been identified. Many factors influence the design of the nucleic acid sequence to be synthesized and the oligonucleotides used to generate the nucleic acid molecule. These factors include one or more of the following: (1) the AT / GC content of all or part of the nucleic acid molecule (e.g., the coding region); (2) the presence or absence of restriction endonuclease cleavage sites (including the addition and / or removal of restriction sites); (3) the preferred codon usage for the particular protein production or host expression system used; (4) the junctions of the assembled oligonucleotides; (5) the number and length of the oligonucleotides used to produce the desired nucleic acid molecule; (6) minimization of undesired regions (e.g., "hairpin" sequences, regions of sequence homology with cellular nucleic acids, repetitive sequences, inhibitory cis-acting elements, restriction enzyme cleavage sites, internal splicing sites, etc.); and (7) flanking segments of the coding region that can be used to link 5' and 3' components (e.g., restriction endonuclease sites, primer binding sites, sequencing adapters or barcodes, recombination sites, etc.).
[0095] In many cases, parameters are input into a computer, and the software generates an in silico nucleotide sequence that balances the input parameters.The software may assign "weights" to the input parameters, for example, in that what is considered to be a nucleic acid molecule that closely matches some of the input criteria may be difficult or impossible to assemble.An exemplary nucleic acid design method is described in U.S. Patent No. 8,224,578.As further explained below, sequence design can also take into account the requirement for multiplexing oligonucleotides belonging to different subfragments of the product nucleic acid molecule.
[0096] Furthermore, the design factor of nucleic acid molecule can be considered throughout the length of nucleic acid molecule or in specific regions of the molecule.For example, the GC content can be restricted throughout the length of nucleic acid molecule to prevent synthesis "failure" due to specific positions within the molecule.Therefore, the synthesis potential of nucleic acid molecule is a feature of the entire nucleic acid molecule, in that regional "assembly failure" will result in the designed nucleic acid molecule not being assembled.From a regional perspective, the codon for optimal translation can be selected, which may conflict with, for example, the local restriction of GC content.
[0097] Successful assembly often involves multiple parameters and regional characteristics of the nucleic acid molecule of interest. Global and regional GC content are just some examples of parameters. For example, the total GC content of a nucleic acid molecule may be 50%, but the GC content in a specific region of the same nucleic acid molecule may be 75%. Thus, in many cases, the GC content is "balanced" throughout the nucleic acid molecule, and regionally, the total GC content may vary by less than 15%, 10%, 8%, 7%, or 5%.
[0098] Therefore, the goal is to reach the best possible compromise between satisfying various requirements. If the product nucleic acid molecule encodes a protein, the large number of amino acids in the protein, based on the degeneracy of the genetic code, would in principle result in a combinatorial explosion of the number of possible DNA sequences that can express the protein of interest. For this reason, various computer-assisted methods have been proposed to identify optimal codon sequences.
[0099] The oligonucleotides or nucleic acid subfragments used in PCR assembly of desired nucleic acid molecules can be obtained from many sources, for example, they can be derived from cloning, polymerase chain reaction, chemically synthesized, or purchased. In many cases, chemically synthesized nucleic acids tend to be less than 100 nucleotides in length. PCR and cloning can be used to generate much longer nucleic acids. Furthermore, the percentage of error bases present in a nucleic acid (e.g., a nucleic acid fragment) is somewhat related to the method by which it is produced. Typically, chemically synthesized nucleic acids have the highest error rate.
[0100] Many methods for chemically synthesizing oligonucleotides are known. In many cases, oligonucleotide synthesis is carried out by stepwise adding nucleotides to the 5' end of a growing chain until an oligonucleotide of desired length and sequence is obtained. Furthermore, each nucleotide addition can be referred to as a synthesis cycle, which often consists of four chemical reactions: (1) deblocking / deprotecting, (2) coupling, (3) capping, and (4) oxidation.
[0101] EGA and PGA deprotection reagents and methods for generating such acids, as well as their use in oligonucleotide synthesis, are described, for example, in Maurer et al., "Electrochemically Generated Acid and Its Containment to 100 Micron Reaction Areas for the Production of DNA Microarrays", PLoS, Issue 1, e34 (2006), or PCT Publication Nos. 2013 / 049227 and 2016 / 094512. Thus, in some cases, EGA is generated as part of the deprotection process. Furthermore, in certain cases, all or part of the oligonucleotide synthesis reaction can be carried out in aqueous solution. In other examples, organic solvents are used.
[0102] In many cases, a typical nucleic acid assembly PCR protocol may involve a combination of methods described herein, such as, for example, single-stranded overhang exonuclease-mediated generation followed by PCR-based assembly ("standard workflow"). In some embodiments, such a standard workflow may include at least the following steps: (i) synthesizing together single-stranded oligonucleotides comprising the sequence of the desired assembly product, each oligonucleotide having a sequence region that is complementary to a sequence region of another oligonucleotide; (ii) hybridizing the oligonucleotides through their complementary sequence regions and extending the oligonucleotides in an overlap-extension PCR reaction (primary assembly PCR) to assemble one or more double-stranded nucleic acid molecules; and (iii) amplifying the assembled nucleic acid molecules in the presence of terminal primers. (primary amplification), (iv) purifying the amplified nucleic acid molecules, (v) generating single-stranded overhangs at the ends of one or more amplified nucleic acid molecules and, optionally, generating single-stranded overhangs at the ends of a linearized target vector for subsequent cloning (e.g., by treating the fragments with restriction endonucleases and / or exonucleases), (vi) inserting one or more nucleic acid molecules into the target vector via the complementary single-stranded overhangs, optionally followed by a ligation step, and (vii) transforming a host cell (e.g., E. coli) with the resulting vector construct. In some embodiments, the assembled nucleic acid molecules can be ligated "in vivo" by endogenous enzymatic activity of the transformed cell. For example, gapped or nicked assembly products can be directly transformed into E. coli, and the gaps or nicks can be repaired by E. coli endogenous repair mechanisms.
[0103] Two methods for assembling nucleic acid molecules are shown in Figures 1A and 1B. Both of these methods involve using PCR to start with oligonucleotides or subfragments containing overlapping sequences at their ends, which are generally "stitched" together via their complementary sequence regions. In some embodiments, the overlap is about 10 base pairs; in other embodiments, the overlap can be 15, 25, 30, 50, 60, 70, 80, or 100 base pairs, etc. (e.g., about 10 to about 120, about 15 to about 120, about 20 to about 120, about 25 to about 120, about 30 to about 120, about 40 to about 120, about 10 to about 40, about 15 to about 50, about 40 to about 80, about 60 to about 90, about 20 to about 50, about 15 to about 35, etc.). To avoid assembly errors, individual overlaps are typically not duplicated or closely matched between subfragments. Because hybridization does not require 100% sequence identity between the nucleic acid molecules or regions involved, each end must be sufficiently different to prevent misassembly. Furthermore, ends intended to undergo homologous recombination with each other must share at least 90%, 93%, 95%, or 98% sequence identity.
[0104] Additionally, multiple cycles of polymerase chain reaction can be used to generate successively larger nucleic acid molecules. In many cases, the stitched oligonucleotides are chemically synthesized and are less than 100 nucleotides in length (e.g., about 40-100, about 50-100, about 60-100, about 40-90, about 40-80, about 40-75, about 50-85, etc. nucleotides). If insertion into a cloning vector is desired, primers containing restriction sites can be used. If desired, the assembled nucleic acid molecule can be directly inserted into a vector and host cell. If the desired construct is fairly small (e.g., less than 5 kilobases), PCR-based insertion into a target vector may be appropriate.
[0105] The standard workflow is represented in Figure 1A by the basic steps of oligonucleotide synthesis, primary assembly PCR to assemble the oligonucleotides, primary amplification to amplify the assembled product, followed by purification of the amplified product, treatment with nucleases to generate single-stranded overlaps between the purified insert and the target vector, insertion of the insert into the target vector, followed by a transformation step.
[0106] Another assembly PCR method involves a combined sequence extension and ligation reaction (Figure 1B), combining steps (ii), (iii), and (vi) of the standard workflow described above into a single (“one-pot”) reaction, while other steps (such as steps (iv) and (v)) may be omitted. In particular, such methods involve the direct assembly of single-stranded overlapping oligonucleotides (primary assembly PCR) onto a linearized target vector via overlap-extension PCR and amplification (primary amplification) of the resulting subfragment-vector fusion construct in a single step. According to some embodiments, a separate PCR reaction is not required to generate double-stranded subfragments prior to vector insertion. Instead, single-stranded oligonucleotides that together represent at least a portion of the polynucleotide to be assembled can be used directly in the overlap-extension reaction. After an initial denaturation step to separate the strands of a given linearized vector, the single-stranded oligonucleotides are annealed via their complementary ends. Two of the oligonucleotides are designed to retain sequence homology with the vector backbone, allowing hybridization with one end of the denatured vector strand. The 3' end of the annealed oligonucleotide and / or the 3' end of the vector strand serves as a primer for the synthesis of a complementary nucleic acid strand. When the 5' end of the hybridized oligonucleotide is encountered, the polymerase terminates its extension, resulting in the production of a nicked circularized double-stranded nucleic acid molecule. The fusion and amplified assembly product can be directly transformed into a host cell without further purification. In some embodiments, no ligation step is performed before transformation. The final ligation of the nicked fusion construct is achieved endogenously within the host cell.
[0107] In the assembly chain reaction, overlapping oligonucleotides are assembled into linear double-stranded DNA fragments by successive cycles of oligonucleotide denaturation, annealing, and mutual extension (primary assembly PCR) (see Figure 2). In a subsequent amplification reaction, the nucleic acid molecules formed by assembly PCR are amplified by PCR using terminal primers to generate and / or amplify assembled nucleic acid molecules (primary amplification), which can be used "as is" or in downstream processes (e.g., insertion into a vector, see Figure 1A).
[0108] In some embodiments described herein, one or more thermostable mismatch recognition proteins are present in the assembly PCR and / or amplification reaction (see, e.g., FIG. 2). The inclusion of a thermostable mismatch recognition protein requires the addition of the mismatch recognition protein after the denaturation step, allowing multiple rounds of error correction and / or error suppression to be performed. Thus, the mismatch recognition protein can be used to reduce the number and / or percentage of nucleic acid molecules in a population that includes correct nucleic acid molecules and nucleic acid molecules that contain errors.
[0109] A schematic diagram of one process for correcting errors in nucleic acid molecules during amplification (primers not shown) is shown in Figure 3. This schematic shows single-stranded nucleic acid molecules at the top left, some of which contain point mutations (denoted by ovals and circles). During hybridization, single-stranded nucleic acid molecules with point mutations are likely to hybridize with nucleic acid molecules that do not contain the same point mutation. The net result is a "mismatch." The population of double-stranded nucleic acid molecules is then contacted with a mismatch endonuclease, which cleaves nucleic acid molecules containing recognized mismatches, rendering the cleaved nucleic acid molecules unsuitable for exponential amplification. Of course, other methods can also be used to inhibit the exponential amplification of nucleic acid molecules containing mismatches. For example, mismatch-binding proteins can be used to remove nucleic acid molecules containing mismatches or to inhibit the amplification of such nucleic acid molecules. Additionally, error-reducing polymerase reagents can be used during amplification.
[0110] More specifically, Figure 3 shows the workflow of an exemplary process for synthesizing error-minimized nucleic acid molecules. In the first step, nucleic acid molecules shorter than the assembled nucleic acid molecule are obtained. Each of the smaller nucleic acid molecules is intended to have a desired nucleotide sequence that comprises a portion of the assembled nucleic acid molecule. In the second to final step of the process described in Figure 3, the annealed nucleic acid molecules are reacted with one or more exonucleases as part of an error correction process. Some variations of this process are as follows: First, two or more rounds of error correction (e.g., two, three, four, five, six, etc.) can be performed. Second, more than one endonuclease can be used in one or more rounds of error correction. For example, T7NI and Cel II can be used in each round of error correction. Third, different endonucleases can be used in different rounds of error correction. For example, T7NI and Cel II can be used in the first round of error correction, and TkoEndoMS can be used alone in the second round of error correction.
[0111] In many cases, a ligase may be present in the reaction mixture during error correction. Some endonucleases used in the error correction process are believed to have nickase activity. The inclusion of one or more ligases is believed to increase the yield of error-corrected nucleic acid molecules after amplification, as the enzymes cause sealed nicks. Exemplary ligases that can be used are T4 DNA ligase, Taq ligase, and PBCV-1 DNA ligase. The ligase used in carrying out the methods described herein can be thermolabile or thermostable (e.g., Taq ligase). When a thermolabile ligase is used, it typically needs to be re-added to the reaction mixture after each error correction round. A thermostable ligase typically does not need to be re-added during each round, as long as the temperature is maintained below its denaturing point.
[0112] In many cases, the error correction of nucleic acid molecules can be mediated by one or more different mismatch recognition proteins.Examples of such protein categories are mismatch binding proteins and mismatch endonucleases.In addition, mismatch binding proteins and mismatch endonucleases can be thermostable or non-thermostable, which often depends on the conditions under which the protein is used and the biological activity of the specific protein (for example, the type of error that is recognized).
[0113] One exemplary method of error correction that can be used in the methods described herein is described in Figures 4 and 5A. Figure 4 is a flow chart of an exemplary process for synthesizing error-minimized nucleic acid molecules. In the first step (line 1), nucleic acid molecules (e.g., oligonucleotides) of a shorter length than the assembled nucleic acid molecule are obtained. Each oligonucleotide is intended to have a desired nucleotide sequence that includes a portion of the nucleotide sequence of the assembled nucleic acid molecule. Each oligonucleotide may also have a nucleotide sequence that includes one or more of the following: (1) an adapter primer for PCR amplification of the nucleic acid molecule, a recognition site for a restriction enzyme, (2) a tethering sequence for attachment to a microchip or solid support, or (3) any other nucleotide sequence determined by experimental or other purposes. Oligonucleotides can be obtained in one or more ways, for example, via synthesis, purchase, etc., as described elsewhere herein.
[0114] In the optional second step (Figure 4, line 2), the oligonucleotides are amplified to obtain more of each oligonucleotide. However, in many cases, sufficient numbers of oligonucleotides are produced, making amplification unnecessary. If used, amplification can be achieved by any method known in the art, such as PCR, rolling circle amplification (RCA), loop-mediated isothermal amplification (LAMP), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), ligase chain reaction (LCR), self-sustained sequence replication (3SR), or solid-phase PCR reactions (SP-PCR) such as bridge PCR (e.g., for a review of various amplification techniques, see Fakruddin et al., J. Pharm. Bioallied. Sci. 5(4):245-252 (2013)). The introduction of additional errors into the nucleotide sequence of any of the nucleic acid molecules may occur during amplification. In some cases, it may be preferable to avoid post-synthesis amplification. If the nucleic acid molecules are produced in sufficient yield in step 1, the optional amplification step may be omitted. This can be achieved, for example, by using optimized bead formats designed to enable synthesis of nucleic acid molecules with sufficient yield and quality, as described, for example, in PCT Publication No. 2016 / 094512.
[0115] In the third step (line 3 in Figure 4), the amplified nucleic acid molecules are optionally assembled into a first set of nucleic acid molecules intended to have a desired length (primary assembly PCR). Of course, in some cases, the nucleic acid molecules in line 3 may be subfragments of a larger nucleic acid molecule.
[0116] In the fourth step (line 4 in Figure 4), the first set of assembled nucleic acid molecules is denatured. Denaturation converts the double-stranded molecules into single-stranded molecules. Denaturation can be achieved by any means. In some embodiments, denaturation is achieved by heating the molecules.
[0117] In the fifth step (line 5 in Figure 4), the denatured molecules are annealed. Annealing converts the single-stranded molecules into a second set of double-stranded nucleic acid molecules. Annealing can be achieved by any means. In some embodiments, annealing is achieved by cooling the molecules. Some of the annealed molecules may contain one or more mismatches that represent sites of sequence errors.
[0118] In the sixth step (line 6 in Figure 4), the second set of molecules is reacted with one or more mismatch-cleaving endonucleases to obtain a third set of nucleic acid molecules intended to have a length shorter than the length of the complete desired gene sequence. Exemplary mismatch-binding and / or cleaving enzymes are described elsewhere herein, but include T7NI, endonuclease VII (encoded by T4 gene 49), RES I endonuclease, CEL I endonuclease, EndoMS (e.g., PfuEndoMS, TkoEndoMS, etc.), and SP endonucleases or endonucleases containing enzyme complexes. These endonucleases generally function by cleaving one or more of the second set of molecules into shorter molecules (single- or double-stranded cleavage). Cleavage at the site of any nucleotide sequence errors is particularly desirable in that assembly of fragments of one or more molecules cleaved at the error site offers the possibility of removing the cleavage error in the final step of the process.
[0119] In the seventh step (line 7 in Figure 4), the third set of molecules is assembled into a fourth set of molecules of a length intended to represent the full length of the desired nucleotide sequence. Typically, in the seventh step, which is based on overlap-extension PCR, the 3'->5' exonuclease activity of a DNA polymerase removes the 3' overhangs generated by endonuclease cleavage in step 6 at mismatched sites, thereby eliminating errors. Therefore, the inherent exonuclease activity of a DNA polymerase can be used to remove errors in the assembly that were not removed in step 6 (e.g., by using a combination of nucleases with mismatch cleavage and exonuclease activity). This principle is outlined, for example, in Saaem et al. ("Error correction of microchip synthesized genes using surveyor nuclease," Nucl. Acids Res., 40:e23 (2012)). Such a final assembly step can be performed in the presence of terminal primers, thereby including functions required for downstream processes such as cloning or protein expression. Each PCR reaction can be set up to first assemble error-corrected fragments by overlap extension to full length in approximately 15 cycles of denaturation, annealing, and extension in the absence of terminal primers, followed by an additional 20 cycles in the presence of terminal primers.
[0120] The process described above and in Figure 4 is also described in U.S. Patent No. 7,704,690. Additionally, the above-described process may be encoded on a computer-readable medium as processor-executable instructions.
[0121] One representative workflow that can be used in the method described in Figure 5A. In this workflow, three nucleic acid subfragments (line 1) are pooled and subjected to error correction using the enzyme T7 endonuclease I ("T7NI") (line 2). The resulting products are then assembled by PCR (line 3) and then subjected to a second round of error correction (line 4). After another round of PCR (line 5), the resulting nucleic acid molecules are transformed into E. coli (step 6), and then screened for full-length sequences (line 7), followed by DNA preparation (line 8). These nucleic acid molecules can then be screened for remaining errors, for example, by sequencing (line 9). In a first variation of the workflow in Figure 5A, the pooled subfragments can be treated with an exonuclease (e.g., exonuclease I) before being subjected to the error correction process. The exonuclease treatment removes single-stranded primer molecules remaining in the PCR reaction product, which may interfere with subsequent PCR reactions and generate nonspecific amplification products. In a second variation of the workflow, the first error correction step may use two or more endonucleases, such as, for example, T7NI in combination with RES I. Optionally, the workflow may include a third error correction or error removal step to eliminate any remaining mismatches after fragment fusion PCR. Such a third step may be performed with a mismatch-binding protein, such as MutS. Those skilled in the art will understand that various orders and combinations of the first, second, and / or third, and possibly further, error correction and / or removal rounds may be applied to further reduce the error rate of the assembled nucleic acid molecule.
[0122] Another process for achieving error correction of chemically synthesized nucleic acid molecules that can be used in the methods described herein is by a commercial process called ERRASE™ (Novici Biotech).
[0123] A variation of the workflow in Figure 5A is outlined in Figure 5B. In this embodiment, the three subfragments (Figure 5B, line 1) are pooled and treated with an exonuclease (e.g., exonuclease I, line 2a in the right-hand workflow) to undergo a double error correction process (Figure 5B, lines 2b and 4). The exonuclease removes single-stranded primer molecules remaining in the PCR reaction product, which could interfere with the subsequent PCR reaction (line 3) and generate nonspecific amplification products. In another variation of the workflow, the first error correction step can use more than one endonuclease, such as T7NI in combination with RES I (Figure 5B, line 2b). Optionally, the workflow can include a third error correction step to eliminate mismatches remaining after segment assembly PCR (line 3, in this example, a secondary assembly PCR). Such a third error correction step can be performed with a mismatch-binding protein, such as MutS (line 4). Of course, various permutations and combinations of the first, second and / or third, and possibly further, rounds of error correction may be applied to further reduce the error rate of the assembled nucleic acid molecule.
[0124] Using the workflow shown in Figure 5A for illustrative purposes, error-containing nucleic acid molecules can be removed in one or more steps. For example, "mismatched" nucleic acid molecules can be removed between steps 1 and 2 of Figure 5A and / or before step 1. This results in treatment of a "preselected" population of nucleic acid molecules with a mismatched endonuclease. Furthermore, two error correction steps such as these can be used in combination. As an example, nucleic acid molecules can be denatured, then reannealed, followed by removal of mismatched nucleic acid molecules via binding with immobilized MutS, and then nucleic acid molecules not separated by MutS binding can be contacted with a mismatched endonuclease without the intervening denaturation and reannealing steps. While not wishing to be bound by theory, it is believed that amplification of nucleic acid molecules introduces errors into the amplified molecules. One means of avoiding the introduction of amplification-mediated errors and / or removing such errors is by selecting nucleic acid molecules that have the correct sequence after most or all of the amplification steps have been performed. Again using the workflow described in Figure 5B for illustration, after step 5, nucleic acid molecules with mismatches can be separated from nucleic acid molecules without mismatches by an additional separation step using a mismatch-binding protein (not shown in Figure 5B).
[0125] Variations of this process are as follows: First, two or more rounds of error correction (e.g., two, three, four, five, six, etc.) can be performed, with each round using a thermostable mismatch-recognition protein. Second, more than one endonuclease can be used in one or more rounds of error correction. For example, T7NI and Cel II can be used in each round of error correction. Third, different endonucleases can be used in different error correction rounds, or can be combined with an error filtering step using a mismatch-binding protein. For example, a pool of reannealed oligonucleotides can be subjected to an error filtering step using a mismatch-binding protein (such as MutS) to remove a first plurality of oligonucleotides with errors from the pool (see Figure 5B), and then the remaining (unbound) oligonucleotides can be subjected to an error correction step using an endonuclease such as T7NI to correct the remaining errors.
[0126] In some cases, for example, T7NI and Cel II can be used in the first round of error correction, and Cel II can be used alone in the second round of error correction. Of course, other mismatched endonucleases can also be used. In another exemplary embodiment, the molecule is cleaved with only one endonuclease (which can be a single-stranded nuclease such as mung bean endonuclease or a resolvase such as T7NI, or another endonuclease with similar function). In yet another embodiment, the same endonuclease (e.g., T7NI) can be used in two subsequent error correction rounds (line 4 in Figure 5A). In yet another embodiment, an enzyme with mismatched cleavage activity can be combined with an enzyme with exonuclease activity to enable removal of errors contained in single-stranded overhangs after mismatched cleavage. In certain aspects, a mismatched endonuclease with inherent exonuclease activity can be used to achieve cleavage and subsequent error removal in a single step. Enzymes with both endonuclease and exonuclease activity include, for example, mung bean nuclease, Cel I, or SP1 endonuclease. In other embodiments, removal of errors can be achieved by a separation step that includes further exonuclease treatment, for example, as described in PCT Publication No. 2005 / 095605(A1).
[0127] In many cases, one or more ligases may be present in the reaction during error correction. Some endonucleases used in the error correction process are believed to have nickase activity. The inclusion of one or more ligases is believed to increase the yield of error-corrected nucleic acid molecules after amplification, as the enzymes cause sealed nicks. Exemplary ligases that can be used are T4 DNA ligase, Taq ligase, and PBCV-1 DNA ligase. The ligase used in carrying out the methods described herein can be thermolabile or thermostable (e.g., Taq ligase). When a thermolabile ligase is used, it typically needs to be added to the reaction mixture in each error correction round. A thermostable ligase typically does not need to be added again during each round, as long as the temperature is kept below its denaturing point.
[0128] If the second set of molecules represents a subfragment of a larger nucleic acid molecule, two or more subfragments (e.g., two, three, or more subfragments) that together represent the larger nucleic acid molecule can be combined and reacted with one or more mismatch-cutting endonucleases in a single reaction mixture. For example, if the open reading frame to be assembled is longer than 1 kb, it can be divided into two or more subfragments assembled separately in parallel reactions in step 3, as shown in Figure 5A, and the resulting two or more subfragments can be combined and error-corrected in a single reaction. The amount of subfragments combined in a single error-correction round may depend on the length of the individual subfragments. For example, up to three subfragments approximately 1 kb in length can be efficiently combined in a single reaction mixture. Of course, more than three subfragments (e.g., four, five, six, seven, eight, nine, etc.) can also be combined. As long as at least one correctly assembled, amplifiable and / or replicable nucleic acid molecule is obtained, the assembly efficiency may decrease. Thus, multiple subfragments (eg, subfragments of about 1 kb in length) can be assembled as long as a correctly assembled product nucleic acid molecule results from the assembly process.
[0129] Nucleic acid molecules with mismatches can be separated from nucleic acid molecules without mismatches by binding with a mismatch-binding agent in a number of ways. For example, a mixture of nucleic acid molecules, some of which have mismatches, can be (1) passed through a column containing bound mismatch-binding protein, or (2) contacted with a surface (e.g., beads (such as magnetic beads), a plate surface, etc.) to which a mismatch-binding protein has been bound.
[0130] Exemplary formats and related methods include those that use beads or other supports to which mismatch binding proteins are bound. For example, a solution of nucleic acid molecules can be contacted with beads to which mismatch binding proteins are bound. The nucleic acid molecules bound to the mismatch binding proteins then become attached to the surface and cannot be easily removed or transferred from the solution.
[0131] In a particular format described in Figure 6, beads with mismatch-binding proteins attached can be placed in a container (e.g., a well of a multiwell plate) in which the nucleic acid molecules are in solution under conditions that allow binding of the mismatch-bearing nucleic acid molecules to the mismatch-binding proteins (e.g., 5 mM MgCl, 100 mM KCl, 20 mM Tris-HCl (pH 7.6), 1 mM DTT, 25°C for 10 minutes). The fluid can then be transferred to another container (e.g., a well of a multiwell plate) without transferring the beads and / or the mismatched nucleic acid molecules. One particular type of bead that can be used is magnetic mismatch-binding beads (M2B2), MAGDETECT™ (United States Biological, Salem, MA, catalog number M9557-01A). Additionally, mismatch-binding proteins used in a workflow similar or identical to that described in Figure 6 can be thermostable or non-thermostable.
[0132] As an example, a protein that has been shown to bind to double-stranded nucleic acid molecules containing mismatches is E. coli MutS (Wagner et al., Nucleic Acids Res., 23:3944-3948 (1995)). Wan et al., Nucleic Acids Res., 42:e102 (2014) demonstrated that chemically synthesized nucleic acid molecules containing errors can be retained on a MutS-immobilized cellulose column, while error-free nucleic acid molecules are not.
[0133] Thus, the subject matter described herein includes methods and related compositions in which nucleic acid molecules are denatured, then reannealed, and then the reannealed nucleic acid molecules containing mismatches are separated. In some embodiments, the mismatch-binding protein used is MutS (e.g., E. coli MutS). Of course, other mismatch-binding proteins, such as those listed in Tables 12 and 15, can also be used.
[0134] Additionally, mixtures of mismatch-binding proteins can be used in the practice of the methods described herein. Different mismatch-binding proteins have been found to have different activities with respect to the type of mismatch they bind. For example, Thermus aquaticus MutS has been shown to effectively remove insertion / deletion errors but is less effective at removing substitution errors than E. coli MutS. Furthermore, the combination of two MutS homologs has been shown to further improve error correction efficiency for the removal of both substitution and insertion / deletion errors and also reduce the effects of biased binding. Thus, the subject matter described herein includes mixtures of two or more (e.g., about 2 to about 10, about 3 to about 10, about 4 to about 10, about 2 to about 5, about 3 to about 5, about 4 to about 6, about 3 to about 7, etc.) mismatch-binding proteins.
[0135] The subject matter described herein further includes the use of multiple rounds of error correction using mismatch binding proteins (e.g., about 2 to about 10, about 3 to about 10, about 4 to about 10, about 2 to about 5, about 3 to about 5, about 4 to about 6, about 3 to about 7, etc.). One or more of these rounds of error correction may employ the use of two or more mismatch binding proteins. Alternatively, a single mismatch binding protein may be used in the first round of error correction, and the same or a different mismatch binding protein may be used in the second round of error correction.
[0136] Once oligonucleotide synthesis is complete, the resulting oligonucleotides are typically subjected to a series of post-processing steps, including one or more of the following: (a) cleavage of the oligonucleotides or elution from the support on which they were synthesized; (b) concentration measurement; (c) concentration adjustment or dilution of the oligonucleotide solution, often referred to as "normalization," to obtain equally concentrated dilutions of each oligonucleotide species; and / or (d) pooling or mixing aliquots of two or more normalized oligonucleotide samples to obtain an equimolar mixture of all oligonucleotides required to assemble one or more specific nucleic acid molecules; the foregoing steps may be combined in different orders.
[0137] Yet another process for reducing errors during nucleic acid synthesis that can be used in embodiments of the subject matter described herein is called circular assembly amplification and is described in PCT Publication No. 2008 / 112683(A2).
[0138] Synthetically produced nucleic acid molecules typically have an error rate of approximately 1 base per 300-500 bases. Conditions can be adjusted to reduce synthesis errors substantially below 1 base per 300-500 bases. Furthermore, in many cases, over 80% of errors are single-base frameshift deletions and insertions. Furthermore, when high-fidelity PCR amplification is used, polymerase action results in less than 2% errors. Therefore, the error correction process using the PCR-based assembly step described above can be combined with one or more error correction methods that do not involve polymerase activity. In many cases, mismatch endonuclease (MME) correction is performed using a fixed protein:DNA ratio. Non-PCR-based error correction can be achieved, for example, by separating nucleic acid molecules with mismatches from those without by binding mismatch binders in a number of ways. For example, a mixture of nucleic acid molecules, some of which have mismatches, can be (1) passed through a column containing bound mismatch-binding proteins, or (2) contacted with a surface (e.g., beads (such as magnetic beads), a plate surface, etc.) to which mismatch-binding proteins are bound.
[0139] Exemplary formats and related methods include those that use a surface or support (e.g., beads) to which a mismatch-binding protein is bound. For example, a solution of nucleic acid molecules can be contacted with beads to which a mismatch-binding protein is bound. One mismatch-binding protein that can be used in various embodiments of the methods described herein is MutS from Thermus aquaticus, the gene sequence of which is described in Biswas and Hsieh, J. Biol. Chem. 271:5040-5048 (1996) and available at GenBank, accession number U33117. Furthermore, mismatch-cutting endonucleases such as EndoMS (e.g., PfuEndoMS, TkoEndoMS, etc.), such as T7NI or CeI from celery, can be genetically engineered to inactivate the cleavage function for use in mismatch-based error filtering processes. Nucleic acid molecules bound to the mismatch-binding protein can be actively removed from the pool of nucleic acid molecules (e.g., via magnetic force using magnetic beads coated with the mismatch-binding protein), or unbound nucleic acids can be removed or displaced from the sample (by pipetting, acoustic liquid processing, etc.) while they remain in the sample and can be immobilized or linked to a surface. Such methods are described, for example, in PCT Publication No. 2016 / 094512.
[0140] As mentioned above, mismatch recognition protein can be used in combination with the hybridization of nucleic acid molecules.The mismatch recognition protein contained in the composition and used in the methods described herein can be thermostable or non-thermostable.In addition, the methods described herein include the method in which more than one mismatch recognition protein is used at more than one position in nucleic acid-related workflow (for example, assembly PCR, amplification, error correction alone, or one or more combinations of these processes).
[0141] Thermostable mismatch recognition proteins (e.g., one or more thermostable mismatch endonucleases) can eliminate sequence errors during processes such as assembly PCR, amplification, and error correction without the need to re-add the mismatch recognition protein after each heat denaturation step. Thus, the compositions and methods described herein allow for multiple rounds of error correction in which the mismatch recognition protein is not added after each nucleic acid denaturation step. Of course, non-thermostable mismatch recognition proteins can also be used in such workflows, but the mismatch recognition activity of such proteins is generally eliminated or substantially reduced by each heat denaturation cycle. In many cases, it is necessary or desirable to add additional non-thermostable mismatch recognition proteins with each heat denaturation cycle.
[0142] The type of mismatch recognition protein used in the workflow can vary. In some cases, error correction can be performed at one or more locations in the workflow. In some cases, a thermostable mismatch recognition protein is often used in combination with a non-thermostable mismatch recognition protein.
[0143] One method for removing nucleic acid molecules with errors is by separating these nucleic acid molecules from the nucleic acid molecules that do not contain errors.Therefore, the present specification provides a workflow that uses the agent that binds to the nucleic acid molecules that contain errors and separates them from the nucleic acid molecules that do not contain errors, and the composition that is used in this workflow.The example of this agent is mismatch binding protein.
[0144] The mismatch-binding protein can be bound to a support and, for example, contacted with a sample containing nucleic acid molecules with and without mismatches under conditions whereby the nucleic acid molecules with mismatches are bound to the support. The support to which the nucleic acid molecules with mismatches are bound can then be removed from contact with the nucleic acid molecules without mismatches, thereby separating the nucleic acid molecules with mismatches from the nucleic acid molecules without mismatches.
[0145] Another method for increasing the percentage of correct nucleic acid molecules in a composition is by suppressing the amplification of nucleic acid molecules containing errors (e.g., deletions, insertions, mismatches, etc.). In some cases, one or more proteins (e.g., one or more mismatch-binding proteins) can be used that reduce the number of errors in a population of nucleic acid molecules by inhibiting assembly PCR and / or amplification of nucleic acid molecules containing one or more errors. In some cases, polymerase reagents can be used that reduce the number of errors in a population of nucleic acid molecules by disfavoring assembly PCR and / or amplification of nucleic acid molecules containing one or more errors.
[0146] Some examples of workflows that can be implemented are listed in Table 1. [Table 1]
[0147] For example, as illustrated by the workflow variations set forth in Table 1, provided herein are compositions and methods for generating a population of nucleic acid molecules. In some such methods, the workflows include two or more different types of processes (e.g., nucleic acid assembly, nucleic acid amplification, nucleic acid denaturation / renaturation, etc.) in which single-stranded nucleic acid molecules hybridize to each other to form double-stranded nucleic acid molecules. Error correction or error reduction can occur during all or part of such workflows. In some cases, error correction can occur between the steps referenced in Table 1. For example, when one or more non-thermostable mismatched endonucleases (e.g., T7NI) are used after primary amplification, they are typically contacted with the amplification products before secondary amplification. This is because thermal cycling usually denatures non-thermostable mismatched endonucleases. Mismatch-binding proteins can also be used during the amplification step, where the mismatch-binding proteins are used to separate mismatched nucleic acid molecules from non-mismatched nucleic acid molecules.
[0148] In some cases, the collective effect of the processes described herein results in less than 1 error per 500 base pairs (e.g., about 1 per 500 base pairs to about 1 per 2,000 base pairs, about 1 per 600 base pairs to about 1 per 2,000 base pairs, about 1 per 700 base pairs to about 1 per 2,000 base pairs, about 1 per 800 base pairs to about 1 per 2,000 base pairs, about 1 per 900 base pairs to about 2 A population of nucleic acid molecules containing about 1 per 1,000 base pairs, about 1 per 1,000 base pairs to about 1 per 2,000 base pairs, about 1 per 700 base pairs to about 1 per 1,500 base pairs, about 1 per 700 base pairs to about 1 per 1,200 base pairs, about 1 per 700 base pairs to about 1 per 1,000 base pairs, about 1 per 800 base pairs to about 1 per 1,200 base pairs, etc. may be obtained.
[0149] The addition of one or more mismatch-binding proteins (e.g., thermostable mismatch-binding proteins) to the assembly PCR mixture can be used to functionally remove oligonucleotides containing sequence errors by blocking extension by the polymerase when the mismatch-binding proteins bind to mismatches formed during annealing (see Fukui et al., "Simultaneous Use of MutS and RecA for Suppression of Nonspecific Amplification during PCR," J. Nucleic Acids, Volume 2013, Article ID 823730).
[0150] Mismatch-binding proteins and mismatch endonucleases often exhibit specificity for a particular type of mismatch. Therefore, in some cases, more than one mismatch-recognition protein may be used in the workflow described herein. Furthermore, when more than one mismatch-recognition protein is present, the proteins often have different error-recognition activities. For example, the mismatch endonucleases TkoEndoMS and T7NI differ in that T7NI appears to have higher activity with respect to deletions and insertions than TkoEndoMS (see Figures 9-11). Furthermore, when more than one mismatch-recognition protein is used, these proteins may have different activities with respect to different types of mismatches.
[0151] Figure 7 shows data in which oligonucleotides were assembled by primary assembly PCR. The assembled nucleic acid molecules were then subjected to primary amplification in the presence of TkoEndoMS and secondary amplification with or without T7NI after incubation of the primary amplification product. The resulting nucleic acid molecules were then sequenced to determine the error rate.
[0152] Sample No. 1 (No Std-EC) was a control run in which 66 fragments were assembled without error correction. As can be seen from this figure, the median error rate for Sample No. 1 is 1 in 308. This increases to 1 in 456 when primary post-amplification T7NI-mediated error correction is used (Sample No. 2). Sample Nos. 1 and 2 represent the error correction baseline for the no error correction condition and the error correction condition using T7NI primary post-amplification of the assembled fragments.
[0153] The data for samples 3 and 4 in Figure 7 were generated under conditions in which the thermostable mismatch endonuclease (TkoEndoMS) was present only in the amplification process and not in the assembly PCR process. Additionally, for sample 4, primary post-amplification T7NI-mediated error correction was used, while for sample 3, primary post-amplification T7NI-mediated error correction was not used. As can be seen from Figure 7, the error rate for sample 3 is 1 in 353. This increases to 1 in 716 when primary post-amplification T7NI-mediated error correction is used (sample 4).
[0154] The data for samples 5 and 6 in Figure 7 were generated under conditions in which a thermostable mismatch endonuclease (TkoEndoMS) was present in the assembly PCR process but not in the amplification process. Additionally, for sample 6, primary post-amplification T7NI-mediated error correction was used, while for sample 5, primary post-amplification T7NI-mediated error correction was not used. As can be seen from Figure 7, the median error rate for sample 5 is 1 in 398. This increases to 1 in 830 when primary post-amplification T7NI-mediated error correction is used (sample 6).
[0155] The data for samples 7 and 8 in Figure 4 were generated under conditions in which a thermostable mismatch endonuclease (TkoEndoMS) was present in both the assembly PCR and amplification processes. Additionally, for sample 8, primary post-amplification T7NI-mediated error correction was used, while for sample 7, primary post-amplification T7NI-mediated error correction was not used. As can be seen in Figure 7, the median error rate for sample 7 is 1 in 488. This increases to 1 in 803 when primary post-amplification T7NI-mediated error correction is used (sample 8).
[0156] The data presented in Figure 7 show that assembled amplified nucleic acid molecules prepared using a thermostable mismatch endonuclease and subjected to T7NI-mediated error correction have the lowest overall error rate.
[0157] Table 1 below shows data derived from Figure 7. From Table 2, it can be seen that the lowest levels of total errors present in nucleic acid molecules prepared using the TkoEndoMS method described in Example 1 below were found in sample numbers 4, 6, and 8. These samples share in common that TkoEndoMS was present during (1) the assembly PCR process, (2) the amplification process, or (3) both the assembly PCR and amplification processes. Additionally, all three of these samples were also subjected to primary post-amplification T7NI-mediated error correction. [Table 2]
[0158] The data in Figure 7 and Table 2 suggest that (1) the presence of a mismatched endonuclease in the assembly PCR process alone results in a lower error rate than the presence of a mismatched endonuclease in the amplification process alone, and (2) the inclusion of a primary post-amplification mismatched endonuclease-mediated error correction step provides enhanced error correction when used in combination with the use of thermostable mismatched endonuclease activity in the assembly PCR process and / or the amplification process.
[0159] As used herein, the error rate of the assembled and amplified nucleic acid molecule is defined as about 1 in 500 base pairs to about 1 in 5,000 base pairs (e.g., about 1 in 550 base pairs to about 1 in 1,500 base pairs, about 1 in 600 base pairs to about 1 in 1,500 base pairs, about 1 in 650 base pairs to about 1 in 1,500 base pairs, about 1 in 700 base pairs to about 1 in 1,500 base pairs, about 1 in 800 base pairs to about 1,500 base pairs). Approximately 1 in 0 base pairs, approximately 1 in 500 to 1,400 base pairs, approximately 1 in 500 to 1,350 base pairs, approximately 1 in 500 to 1,300 base pairs, approximately 1 in 500 to 1,250 base pairs, approximately 1 in 500 to 1,200 base pairs, approximately 1 in 500 to 1,150 base pairs, approximately 1 in 500 to 1,000 base pairs about 1 in 600 base pairs, about 1 in 600 to about 1 in 1,000 base pairs, about 1 in 650 to about 1 in 1,000 base pairs, about 1 in 600 to about 1 in 900 base pairs, about 1 in 650 to about 1 in 900 base pairs, about 1 in 700 to about 1 in 850 base pairs, about 1 in 550 to about 1 in 2,000 base pairs, about 1 in 550 to about 1 in 2,500 base pairs, Compositions and methods are provided in which the nucleic acid sequence is between about 1 in 550 base pairs and about 1 in 3,500 base pairs, between about 1 in 550 base pairs and about 1 in 4,500 base pairs, between about 1 in 900 base pairs and about 1 in 3,500 base pairs, between about 1 in 1,500 base pairs and about 1 in 5,000 base pairs, between about 1 in 2,000 base pairs and about 1 in 5,000 base pairs, between about 1 in 2,500 base pairs and about 1 in 5,000 base pairs, etc. Such nucleic acid molecules can be produced by primary assembly PCR and primary assembly, optionally followed by secondary amplification.
[0160] As used herein, a fold reduction ("X") in the error rate of an assembled and amplified nucleic acid molecule is greater than 1.75 (e.g., about 1.75 to about 8, about 1.75 to about 7, about 1.75 to about 8, about 1.75 to about 5, about 1.75 to about 4, about 1.75 to about 5, about 1.75 to about 6, about 1.75 to about 7, about 1.75 to about 8, about 1.75 to about 9, about 1.75 to about 10, about 1.75 to about 11, about 1.75 to about 12, about 1.75 to about 13, about 1.75 to about 14, about 1.75 to about 15, about 1.75 to about 16, about 1.75 to about 17, about 1.75 to about 18, about 1.75 to about 19, about 1.75 to about 20, about 1.75 to about 21, about 1.75 to about 22, about 1.75 to about 23, about 1.75 to about 24, about 1.75 to about 25, about 1.75 to about 26, about 1.75 to about 27, about 1.75 to about 28, about 1.75 to about 29, about 1.75 to about 30, about 1.75 to about 31, about 1.75 to about 32, about 1.75 to about 33, about 1.75 to about 34, about 1.75 to about 35, about 1.75 to about 36, about 1.75 to about 37, about 1.75 to about 38, about 1.75 to about 39, about Compositions and methods are provided for fold reduction in error rate (e.g., 1.75 to about 3, about 2.0 to about 8, about 2.1 to about 8, about 2.2 to about 8, about 2.3 to about 8, about 2.5 to about 8, about 2.75 to about 8, about 2.0 to about 7, about 2.0 to about 6, about 2.0 to about 5, about 2.0 to about 4.5, about 2.2 to about 8, about 2.2 to about 7, about 2.2 to about 6, about 2.2 to about 5, about 2.2 to about 3, about 2.2 to about 2.8, about 2.1 to about 2.8, etc.) (see data in Figure 7 and Table 2). A formula that can be used to calculate the fold reduction in error rate is as follows:
number
[0161] 9, 10, and 11 show detailed data relating to error rates associated with deletions, insertions, and substitutions using the experimental data used to generate FIGS.
[0162] Samples no. 8, 6, 4, and 2 (T7NI-treated) all show similarly low levels of deletions and insertions in Figures 9 and 10. These data indicate that deletions and insertions not removed by TkoEndoMS during assembly PCR and amplification are removed by primary post-amplification T7NI-mediated error correction.
[0163] FIG. 10 shows that TkoEndoMS eliminates substitution errors when included in the assembly PCR process, the amplification process, or both the assembly PCR and amplification processes.
[0164] Many different types of substitutions can be found in double-stranded nucleic acid molecules. Furthermore, mismatch recognition proteins often vary in the specificity of the type of substitutions they exhibit activity. This specificity can vary depending on specific conditions, such as the presence or absence of divalent metal ions and the surrounding nucleic acid region. Some of these EndoMS variations are described in Ishino et al., Nucl. Acids Res. 44:2977-2989 (2016). Additional EndoMS proteins are listed in Table 15. Also, modified forms of wild-type thermostable mismatch endonuclease from Pyrococcus furiosus have been generated (see U.S. Pat. No. 10,196,618 and U.S. Patent Publication No. 2017 / 253909). Furthermore, modified forms of wild-type mismatch recognition proteins (e.g., mismatch endonucleases) can be generated to alter their mismatch recognition activity. Such modified forms of wild-type mismatch recognition proteins can be included in and / or used in the methods described herein.
[0165] Figures 12A-12D show some error correction properties of TkoEndoMS under the conditions used in Example 1. Figures 12A and 12C compare the deletion, insertion, and substitution levels found in assembled and amplified nucleic acid molecules generated without error correction (Figure 12A) and when TkoEndoMS was included in both the assembly PCR and amplification processes (Figure 12C). As can be seen, the number of deletions and insertions is similar under both sets of conditions. Although there is considerable variability in the data, these data indicate that substitution rates are lower when TkoEndoMS is present.
[0166] Figures 12B and 12D show some of the error-correcting activity of TkoEndoMS for specific substitutions. TkoEndoMS appears to be effective at correcting most transitions and transversions, but appears to have low activity associated with TV1 (CT and GA) and TV4 (CT and GA) mismatches (Figure 12D). Furthermore, T7NI also appears to have low activity associated with TV1 (CT and GA) and TV4 (CT and GA) mismatches (Figure 12B).
[0167] SURVEYOR® nuclease cleaves all types of mismatches, although some are believed to be preferred over others. In particular, CT, AC, and CC are equally preferred over TT, followed by AA and GG, and finally AG and GT, which have the lowest preference.
[0168] Many mismatch-recognition proteins (e.g., the mismatch-recognition proteins listed in Table 15) are known to have recognition activity for different types of mismatches. The error-correction specificities of some mismatch-recognition proteins are shown in Table 3. [Table 3]
[0169] The methods described herein include the combined use of more than one mismatch-recognition protein. Using the workflow shown in Figure 1A for illustrative purposes, PfuEndoMS and TkoEndoMS can be used together in the oligonucleotide assembly process. This results in the presence of two distinct mismatch endonucleases with overlapping but distinct error-recognition activities. Furthermore, one or both of TaqMutS and TthMutS can be used in combination with each other, or with, for example, PfuEndoMS and TkoEndoMS, to remove double-stranded nucleic acid molecules containing errors recognized by them.
[0170] Provided herein are methods for the correction of errors in nucleic acid molecules that involve the sequence or simultaneous use of mismatch-recognizing proteins that differ in the type of error they recognize.
[0171] Suitable error correction methods and reagents for use in the methods provided herein are described in U.S. Pat. Nos. 7,838,210 and 7,833,759, U.S. Patent Publication No. 2008 / 0145913(A1) (mismatched endonucleases), PCT Publication No. 2011 / 102802(A1), and Ma et al., Trends in Biotechnology, 30(3):147-154 (2012). Additionally, one of skill in the art will recognize that other methods of error correction and / or error filtering (i.e., specifically removing molecules containing errors), such as those described in U.S. Patent Publication Nos. 2006 / 0127920(AA), 2007 / 0231805(AA), 2010 / 0216648(A1), or 2011 / 0124049(A1), can be implemented in certain embodiments of the subject matter described herein.
[0172] Provided herein are compositions and methods that contain and use many different error-correcting agents. Such error-correcting agents have activity related to correcting one or more of the following types of errors, also referred to as mismatches: deletion, insertion, and substitution. Furthermore, with respect to substitutions, activity is generally directed to different types of substitutions.
[0173] Many different polymerases and different types of polymerases can be included and used in the compositions and methods described herein. The type of polymerase used in one or more steps of the assembly PCR and amplification workflow is believed to affect the number of errors present in the assembled nucleic acid molecule.
[0174] Figures 13 and 14A-14D show data generated using different types of polymerases. Figure 13 shows data generated without error correction in combination with PHUSION™ DNA polymerase, while assembly PCR and amplification error correction were performed using TkoEndoMS in combination with PLATINUM™ SUPERFI™ II DNA polymerase reagent.
[0175] A representative workflow of the methods provided herein is depicted in Figure 5A. In this workflow, three nucleic acid segments (referred to as "subfragments") are pooled and subjected to error correction using the enzyme T7 endonuclease I ("T7NI") (Figure 5A, line 2). The three nucleic acid segments are then assembled by PCR (secondary assembly PCR) (Figure 5A, line 3) and then subjected to a second round of error correction (Figure 5A, line 4). After another round of PCR (tertiary assembly PCR) (line 5), the resulting nucleic acid molecules are screened against full-length versions (Figure 5A, line 7). These nucleic acid molecules can then be screened for remaining errors, for example, by nucleotide sequencing.
[0176] After synthesis, oligonucleotides can be assembled into larger nucleic acid molecules in stages (primary assembly PCR), and optionally amplified. The method used to assemble nucleic acid molecules can vary (see, for example, Figures 1A and 1B). Furthermore, regardless of the method used, error correction can be integrated into the appropriate assembly process. In many cases, error correction can be performed using mismatch recognition proteins (e.g., mismatch binding proteins and thermostable mismatch recognition proteins such as mismatch endonucleases).
[0177] In some embodiments, the length of the assembled nucleic acid molecule is from about 20 base pairs to about 10,000 base pairs, from about 100 base pairs to about 5,000 base pairs, from about 150 base pairs to about 5,000 base pairs, from about 200 base pairs to about 5,000 base pairs, from about 250 base pairs to about 5,000 base pairs, from about 300 base pairs to about 5,000 base pairs, from about 350 base pairs to about 5,000 base pairs, from about 400 base pairs to about 5,000 base pairs, from about 500 base pairs to about 5,000 base pairs, from about 700 base pairs to about 5,000 base pairs, from about 800 base pairs to about 5,000 base pairs, from about 1,000 base pairs to about 5,000 base pairs, from about 100 base pairs to about 4,000 base pairs, from about 150 base pairs to about 5,000 base pairs, The length may vary from about 1 base pair to about 4,000 base pairs, from about 200 base pairs to about 4,000 base pairs, from about 300 base pairs to about 4,000 base pairs, from about 500 base pairs to about 4,000 base pairs, from about 50 base pairs to about 3,000 base pairs, from about 100 base pairs to about 3,000 base pairs, from about 200 base pairs to about 3,000 base pairs, from about 250 base pairs to about 3,000 base pairs, from about 300 base pairs to about 3,000 base pairs, from about 400 base pairs to about 3,000 base pairs, from about 600 base pairs to about 3,000 base pairs, from about 800 base pairs to about 3,000 base pairs, from about 100 base pairs to about 2,000 base pairs, from about 200 base pairs to about 2,000 base pairs, from about 300 base pairs to about 1,500 base pairs, etc.
[0178] Any number of methods can be used for amplifying and assembling nucleic acids. One exemplary method is described in Yang et al., Nucleic Acids Research 21:1889-1893 (1993) and U.S. Patent No. 5,580,759. In the process described in Yang et al., a linear vector is mixed with double-stranded nucleic acid molecules that share sequence homology at their ends. An enzyme with exonuclease activity (i.e., T4 DNA polymerase, T5 exonuclease, T7 exonuclease, etc.) is added, which generates single-stranded overhangs at all ends present in the mixture. The nucleic acid molecules with single-stranded overhangs are then annealed and incubated with DNA polymerase and deoxynucleotide triphosphates under conditions that allow the filling of single-stranded gaps. Nicks in the resulting nucleic acid molecules can be repaired by introducing the molecules into cells or by adding ligase. Of course, the vector may be omitted depending on the application and workflow. Additionally, the resulting nucleic acid molecule or a subportion thereof can be amplified by polymerase chain reaction.
[0179] Other methods of nucleic acid assembly include those described in U.S. Patent Publication Nos. 2010 / 0062495(A1), 2007 / 0292954(A1), 2003 / 0152984(AA), and 2006 / 0115850(AA), U.S. Patent Nos. 6,083,726, 6,110,668, 5,624,827, 6,521,427, 5,869,644, and 6,495,318, and WO2020 / 001783(A1).
[0180] A method for isothermal assembly of nucleic acid molecules is described in U.S. Patent Publication No. 2012 / 0053087. In one aspect of this method, nucleic acid molecules for assembly are contacted with a thermolabile protein having exonuclease activity (e.g., T5 polymerase), and optionally with a thermostable polymerase and / or a thermostable ligase under conditions in which the exonuclease activity decreases over time (e.g., 50°C). The exonuclease "bites back" one strand of the nucleic acid molecule, and if there is sequence complementarity, the nucleic acid molecules anneal to each other. In one embodiment, a thermostable polymerase can be used to fill gaps, and a thermostable ligase can be provided to seal nicks. In another embodiment, the annealed nucleic acid product can be used directly to transform a host cell, and gaps and nicks are repaired "in vivo" by endogenous enzymatic activity of the transformed cell.
[0181] Single-stranded binding proteins such as T4 gene 32 protein and RecA, as well as other nucleic acid binding or recombination proteins known in the art, can be included, for example, to facilitate annealing of nucleic acid molecules.
[0182] In some cases, standard ligase-based ligation of partially and completely assembled nucleic acid molecules can be used. For example, assembled nucleic acid molecules can be generated to have restriction enzyme sites near their ends. These nucleic acid molecules can then be treated with one of the more suitable restriction enzymes to generate, for example, one or two "sticky ends." These sticky-end molecules can then be introduced into vectors by standard restriction enzyme-ligase methods. If an inactive nucleic acid molecule has only one sticky end, ligase can be used to blunt-end ligate the "non-sticky" end.
[0183] Multiplex assembly of nucleic acid molecules The complexity of an oligonucleotide population is determined, in part, by the number of different oligonucleotides present. In some cases, the number of oligonucleotides present that are designed to have different nucleotide sequences can be from about 2,000 to about 20,000 (e.g., from about 2,000 to about 20,000, from about 2,000 to about 20,000, from about 2,000 to about 20,000, from about 2,000 to about 20,000, from about 2,000 to about 20,000, from about 2,000 to about 20,000, from about 2,000 to about 20,000, 00, approximately 2,000 to approximately 20,000, approximately 2,000 to approximately 20,000, approximately 2,000 to approximately 20,000, approximately 2,000 to approximately 20,000, approximately 2,000 to approximately 20,000, approximately 2,000 to approximately 20,000, approximately 2,000 to approximately 20,000, approximately 2,000 to approximately 20,000, approximately 2,000 to approximately 20,000, etc.).
[0184] Additionally, the oligonucleotides in a reaction mixture may represent subfragments of more than one larger nucleic acid molecule. For example, if it is desired to assemble three assembled nucleic acid molecules in one reaction mixture, and 10 oligonucleotides are required to assemble each of the assembled nucleic acid molecules, the reaction mixture would initially contain at least 30 oligonucleotides.
[0185] Provided herein are compositions and methods useful for assembling more than one assembled error-corrected nucleic acid. In some cases, the number of assembled error-corrected nucleic acid molecules produced by these methods is about 2 to about 100 (e.g., about 2 to about 90, about 2 to about 80, about 2 to about 70, about 2 to about 50, about 5 to about 90, about 5 to about 60, about 8 to about 90, about 8 to about 50, about 8 to about 35, about 10 to about 90, about 2 to about 60, about 15 to about 90, about 15 to about 55, etc.).
[0186] Polymerases and Polymerase Reagents There are many different types of DNA polymerases. For example, many prokaryotic cells contain DNA polymerases type I, type II, and type III. DNA polymerases may or may not have proofreading activity. Proofreading DNA polymerases typically also have 3' to 5' exonuclease activity. Furthermore, DNA polymerases may be thermostable or non-thermostable.
[0187] Although any type of DNA polymerase can be included and used in the compositions and methods described herein, in many cases, proofreading polymerases are used herein. In some cases, the DNA polymerase is formulated for "hot start," in which the DNA polymerase is bound to an antibody that releases the DNA polymerase when heated.
[0188] DNA polymerases that can be contained in and used in the compositions and methods described herein. Exemplary DNA polymerases and DNA polymerase reagents include Phi29 DNA polymerase or a derivative thereof, Bsm, Bst, T4, T7, DNA Pol I, or Klenow Fragment, or mutants, variants, and derivatives thereof. Further exemplary DNA polymerases and DNA polymerase reagents include Taq, Tbr, Tfl, Tth, Tli, Tfi, Tne, Tma, Pfu, Pwo, and Kod DNA polymerases, as well as VENT® DNA polymerase (New England Biolabs), DEEP VENT® DNA polymerase (New England Biolabs), PHUSION™ DNA polymerase, PHUSION™ U DNA polymerase, SUPERFI™ II DNA polymerase, SUPERFI™ U DNA polymerase, or mutants, variants, and derivatives thereof, and / or GoTaq G2 Hot Start Polymerase (Promega), ONETAQ® Hot Start DNA Polymerase (New England Biolabs), TAKARA TAQ™ DNA Polymerase Hot Start (Takara), KAPA2G Robust Hot Start DNA Polymerase (KAPA), FASTSTART™ Taq DNA Polymerase (Sigma-Aldrich), Hot Start Taq DNA polymerase (New England Biolabs), Q5® DNA polymerase (New England Biolabs), KAPA HiFi DNA polymerase (Roche), PRIMESTAR® Max DNA polymerase (Takara), and PRIMESTAR® GXL DNA polymerase (Takara).
[0189] In some cases, the DNA polymerase may comprise a chimeric DNA polymerase. Furthermore, the chimeric DNA polymerase may comprise a sequence-nonspecific double-stranded DNA (dsDNA) binding domain. In some cases, the dsDNA binding domain may comprise Sso7d from Sulfolobus solfataricus; Sac7d, Sac7a, Sac7b; and Sac7e from S. acidocaldarius, Ssh7a and Ssh7b from Sulfolobus shibatae; Pae3192; Pae0384; Ape3192; HMf family archaeal histone domain; or archaeal proliferating cell nuclear antigen (PCNA) homolog. In addition, the DNA polymerase present in the compositions described herein and used in the methods described herein may also comprise exonuclease activity and / or an exonuclease domain.
[0190] Additionally, DNA polymerases that can be contained in and used in the compositions and methods described herein include all or part of the DNA polymerases listed in Table 14, as well as modified forms of such polymerases (e.g., DNA polymerases that are at least 90%, at least 95%, or at least 97.5% identical to a DNA polymerase listed in Table 14).
[0191] PHUSION™ U DNA Polymerase (Thermo Fisher Scientific, Catalog No. F555S) is an engineered high-fidelity enzyme developed using fusion technology. Due to a mutation in the dUTP-binding pocket of PHUSION™ U, PHUSION™ U overcomes the limitations of proofreading enzymes in that it can incorporate dUTP and read uracil present in DNA templates. In addition to this property, PHUSION™ U can amplify long amplicons up to 20 kb.
[0192] The DNA polymerases that can be present in the compositions described herein and can be used in the methods described herein include those modified to reduce the effects of inhibitors and / or those formulated with one or more compounds that reduce the effects of inhibitors.For example, PLATINUM™ II Taq Hot Start DNA Polymerase (Thermo Fisher Scientific, Catalog No. 14966001) is a "hot start" polymerase formulation in which the DNA polymerase is modified to reduce the effects of interfering compounds (e.g., humic acid, xylan, hemin, etc.).Furthermore, it is formulated to allow primer annealing at 60°C.
[0193] DNA polymerase reagents can be formulated to reduce the effects of interfering compounds. One category of compounds that can be used in such formulations are "amines." Amines have been found to improve (1) the yield of nucleic acid synthesis products and / or (2) resistance to inhibitors of nucleic acid synthesis. Amines include compounds that can be contained and used in the compositions and methods described herein, including compounds containing one or more amines of formula I or salts thereof: [ka]
[0194] wherein R1 is H, R2 is selected from alkyl, alkenyl, alkynyl, or (CH2)n-R5, where n=1 to 3, R5 is aryl, amino, thiol, mercaptan, phosphate, hydroxy, alkoxy, and R3 and R4 may be the same or different and are independently selected from H or alkyl, with the proviso that when R2 is (CH2)n-R5, at least one of R3 and / or R4 is alkyl.
[0195] Specific amine-containing compounds that can be included and used in the compositions and methods described herein include dimethylamine hydrochloride, diethylamine hydrochloride, diisopropylamine hydrochloride, ethyl(methyl)amine hydrochloride, and / or trimethylamine hydrochloride.
[0196] When one or more amine compounds are present in the formulation, the concentration of this or these compounds generally ranges from 5 mM to 500 mM (e.g., about 5 mM to about 500 mM, about 10 mM to about 500 mM, about 20 mM to about 500 mM, about 30 mM to about 500 mM, about 40 mM to about 500 mM, about 5 mM to about 300 mM, about 5 mM to about 250 mM, about 5 mM to about 200 mM, about 5 mM to about 100 mM, about 10 mM to about 250 mM, about 20 mM to about 200 mM, about 25 mM to about 180 mM, about 50 mM to about 110 mM, etc.).
[0197] One particular example of a DNA polymerase reagent that can be used in the methods described herein is PLATINUM™ SUPERFI™ II DNA polymerase (Thermo Fisher Scientific, catalog number 12361010).
[0198] vector Vectors that can be used in the methods described herein can be any vector suitable for cloning and transforming host cells. In many cases, high-copy-number vectors can be used to obtain high yields of the desired polynucleotide. Common high-copy-number vectors include pUC (about 500 to about 700 copies), pBLUESCRIPT®, or PGEM® (about 300 to about 500 copies, respectively), or their derivatives. In some cases, for example, when high expression of a given insert may be toxic to the transformed cells, low-copy-number vectors may be used. Such low-copy-number vectors, having copy numbers of about 5 to about 30, include, for example, pBR322, various pET vectors, pGEX, pColE1, pR6K, pACYC, or pSC101.
[0199] An exemplary list of vectors that may be used in any of the assembly or cloning methods disclosed herein includes: BACULODIRECT™ Linear, DNA Cloning Fragment DNA, BACULODIRECT™ N-terminal Linear DNA, BACULODIRECT™ C-terminal Baculovirus Linear DNA, BACULODIRECT™ N-terminal Baculovirus Linear DNA, CHAMPION™ pET100 / D-TOPO®, CHAMPION™ pET 101 / D-TOPO (registered trademark), CHAMPION (trademark) pET104-DEST, CHAMPION (trademark) pcDN3.1A / 5-His-TOPO, pcDNA3.1(-), pcDNA3.1(+), pcDNA3.1(+) / myc-HisA, pcDNA3.1(+) / myc-His series, pcDNA3.1 / His series, pcDNA3. 1 / Hygro(-), pcDNA3.1 / Hygro(+), pcDNA3.1 / NT-GFP-TOPO, pcDNA3.1 / nV5-DEST, pcDNA3.1A / 5-His series, pcDNA3.1 / Zeo(+), pcDNA3.1 / Zeo(+), pcDNA3.1DA / 5-His-TOPO, pcDNA3.2 / V5-DEST, pcDNA3.2-DEST, pcDNA4 / Hisシリーズ, pcDNA4 / HisMax-TOPO, pcDNA4 / HisMax-TOPO, pcDNA4 / myc-Hisシリーズ, pcDNA4 / TO, pcDNA4 / T O. pcDNA4 / TO / myc-Hisシリーズ, pcDNA4 / V5-Hisシリーズ, pcDNA5 / FRT, pcDNA5 / FRT / TO / CAT, pcDNA5 / FRT / TO-TOPO, pcDNA-D EST47, pcDNA-DEST53, PDEST (trademark)10, PDEST (trademark)14, PDEST (trademark)15, pDEST (trademark)17, pDEST (trademark)20, pDEST (trademark)22, PDEST (trademark)24, pDEST (trademark)26, pDES (trademark)27, pDEST (trademark)32, pDEST (trademark)8, pDEST (trademark)38, pDEST (trademark)39, pDisplay, pDONR (trademark)P2R P3, PDONR (trademark) P2R-P3, pDONR (trademark) P4-P1R, pDONR (trademark) P4-P1R, pDONR (trademark) / Zeo, pDONR (trademark) 201, pDONR (trademark) 207, pDONR (trademark) 221, pEF / myc / cyto, pEF / myc / mito, pEF / myc / nuc, pEFi / His series, pEF4 / V5-His series, pEF5 / FRT V5 D-TOPO, pEF5 / FRT / V5-DEST (trademark), pEF6 / His series, pEF6 / myc-His series, pEF6A / 5-His-TOPO, pEF-DEST51, pENTR-TEV / D-TOPO, pENTR (trademark) / D-TOPO, pENTR (trademark) / D-TOPO, pHybLex / Zeo, pHyBLex / Zeo-MS2, pIB / His series, pIBA / 5-His Topo, pYES2.1A / 5-His-TOPO, pYES2 / CT, pYES2 / NT, pYES2 / NTシリーズ, pYES3 / CT, pYES6 / C T, pYES-DEST (trademark) 52, pYESTrp, pYESTrp2, pZeoSV2, pZeoSV2(+), pZErO-1, and およびpZErO-2. .
[0200] In some embodiments, the vector may have a size restriction to allow PCR-mediated extension of the full-length fusion construct. Under certain conditions, full-length extension and / or amplification of the fusion construct may not be necessary. In such circumstances, the size of the target vector may not be a restriction. Thus, in some embodiments, the target vector may have a size of about 0.5 to about 5 kb, or about 1 kb to about 3 kb, while in other embodiments, the target vector may have a size of about 2 kb to about 10 kb, or about 5 kb to about 20 kb.
[0201] The assembled nucleic acid molecule may also contain functional elements that confer desirable properties. These elements may be provided by either multiple oligonucleotides or the targeting vector. Examples of such elements include origins of replication, long terminal repeats, resistance markers (such as antibiotic resistance genes), selectable markers and antidote coding sequences (e.g., a ccdA coding sequence to counter the toxic effects of ccdB), promoters, enhancers, polyadenylation signal coding sequences, 5' and 3' UTRs, and other components suitable for the particular use of the nucleic acid molecule (e.g., enhancing mRNA or protein production efficiency). In embodiments in which nucleic acid molecules are assembled to form an operon, the assembled nucleic acid product often contains promoter and terminator sequences. Additionally, the assembled nucleic acid molecule may contain multiple cloning sites, such as type II or type IIs cleavage sites and / or GATEWAY® recombination sites, as well as other sites for interconnecting nucleic acid molecules.
[0202] A vector can be linearized by any means, including PCR amplification of a closed circular template vector molecule. Alternatively, a vector can be linearized by restriction enzyme cleavage with one or more enzymes that produce either blunt or sticky ends. Such enzymes include type II restriction endonucleases, which cleave nucleic acids at fixed positions relative to their recognition sequences. Restriction enzymes that can be selected to produce either "blunt" or "sticky" ends when cleaving double-stranded nucleic acids are known to those skilled in the art and can be selected by those skilled in the art depending on the vector sequence and assembly requirements. In some cases, a vector can be linearized using a restriction endonuclease that generates blunt ends.
[0203] After cleavage, the vector can be used directly in, for example, an assembly PCR reaction (e.g., sequence extension and ligation reaction), or it can be purified using gel extraction or amplified in a PCR reaction before use in the assembly PCR reaction. Purification of the linearized vector produced by PCR amplification is often not necessary, and the PCR product can be used directly in the assembly PCR reaction. Alternatively, a circular vector containing a type IIS restriction enzyme cleavage site can be used and subjected to a one-step cleavage and ligation process to seamlessly clone one or more assembled nucleic acid molecules into a vector, commonly known as the Golden Gate cloning system, described below.
[0204] After assembly PCR, the reaction mixture containing the assembled circular construct or an aliquot thereof can be directly used to transform a suitable competent host cell, such as a common E. coli strain, according to standard protocols. Those skilled in the art will be able to select a suitable host cell depending on the size and nucleotide composition of the construct, the copy number of the plasmid, the selection criteria, etc. Useful strains are commercially available from the American Type Culture Collection and Yale's E. coli Genetic Stock Center, as well as suppliers such as Agilent, Promega, Merck, Thermo Fisher Scientific, and New England Biolabs, respectively.
[0205] In many cases, the nucleic acid molecules prepared by the methods provided herein are replicable. Furthermore, many of these replicable nucleic acid molecules are circular (e.g., plasmids). Regardless of whether they are circular or not, replicable nucleic acid molecules are generally formed from the assembly of two or more (e.g., 3, 4, 5, 8, 10, 12, etc.) nucleic acid fragments. In some cases, the methods provided herein use selection based on the reconstitution of one or more (e.g., 2, 3, 4, etc.) selectable markers or one or more (e.g., 2, 3, 4, etc.) replication origins resulting from the ligation of different nucleic acid fragments. If circularity is required for replication, further selection may result from the formation of circular nucleic acid molecules.
[0206] In another embodiment, the single-stranded oligonucleotides used in the sequence extension and ligation reaction (Figure 1B) can be replaced by one or more double-stranded nucleic acid fragments with complementary ends to enable overlap-extension PCR on the linearized target vector (and between fragments if two or more fragments are assembled simultaneously into the target vector). The complementary ends (i.e., overlap) can have a size of about 15 bp to about 50 bp, about 20 bp to about 40 bp, e.g., 40 bp. The required size of the overlap may depend on the size of the fragments to be fused and their melting temperatures. Double-stranded fragments are first assembled from the single-stranded oligonucleotides and amplified in the presence of terminal primers, as described above in steps (ii) and (iii), respectively, of a workflow such as that depicted in Figure 1A. The amplified fragments can then be subjected to one or more error correction and / or error removal rounds (e.g., by mismatch endonuclease treatment as described above) and then used in combination with the insertion and extension reactions, as described above for the sequence extension and ligation reaction. In some embodiments, the overlap of interconnected adjacent fragments and / or the overlap of terminal fragments onto the linearized vector can be about 15 to about 40 or about 18 to about 30 nucleotides in length. In embodiments where hybridization over a longer region is required to ensure successful assembly, the overlap can be about 30 to about 60 nucleotides in length, or even greater than 60 nucleotides in length.
[0207] Assembled constructs obtained by an assembly workflow can be further combined with other assembly workflow products or nucleic acid molecules obtained from other sources to assemble larger nucleic acid molecules (e.g., genes). Larger constructs can be assembled by any means known to those skilled in the art. For example, if larger constructs (e.g., 5-100 kilobases) are desired, multiple fragments (e.g., 2, 3, 5, 8, 10, etc.) can be assembled using type IIs restriction site-mediated assembly methods. One suitable cloning system is called Golden Gate and is described in various forms in U.S. Patent Publication No. 2010 / 0291633(A1) and PCT Publication No. 2010 / 040531.
[0208] At many points during the workflows provided herein, it may be desirable to separate nucleic acid molecules or assembly products from reaction mixture components (e.g., dNTPs, primers, truncated oligonucleotides, tRNA molecules, buffers, salts, proteins, etc.). This can be done in a number of ways, such as by enzymatically removing unwanted nucleic acid by-products using exonucleases, restriction enzymes, or UNG glycosylase, as described above. In some cases, nucleic acid molecules may be precipitated or attached to a solid support (e.g., magnetic beads). Once separated from reaction components to facilitate processes (e.g., pooling or multiplexing of selected oligonucleotides, nucleic acid synthesis, error correction, etc.), the nucleic acid molecules can then be used in additional reactions (e.g., assembly PCR reactions, amplification, cloning, etc.).
[0209] Larger nucleic acid molecules can also be assembled in vivo. In in vivo assembly methods, a mixture of all subfragments to be assembled is often used to transfect host cells using standard transfection techniques. The ratio of the number of subfragment molecules in the mixture to the number of cells in the transfected culture must be high enough to ensure that at least some cells take up more subfragment molecules than there are different subfragments in the mixture. Thus, in most cases, the higher the transfection efficiency, the greater the number of cells that contain all of the nucleic acid subfragments necessary to form the final desired assembly product. Technical parameters along these lines are described in U.S. Patent Publication No. 2009 / 0275086(A1).
[0210] Large nucleic acid molecules are relatively fragile and therefore easily sheared. One way to stabilize such molecules is by maintaining them intracellularly. Thus, in some embodiments, the subject matter described herein involves the assembly and / or maintenance of large nucleic acid molecules in host cells. Large nucleic acid molecules are typically 20 kb or larger (e.g., greater than 25 kb, greater than 35 kb, greater than 50 kb, greater than 70 kb, greater than 85 kb, greater than 100 kb, greater than 200 kb, greater than 500 kb, greater than 700 kb, greater than 900 kb, etc.).
[0211] Methods for producing and further analyzing large nucleic acid molecules are known in the art. For example, Karas et al., "Assembly of eukaryotic algal chromosomes in yeast," Journal of Biological Engineering 7:30 (2013) demonstrates the assembly of algal chromosomes in yeast and pulsed-field gel analysis of such large nucleic acid molecules.
[0212] As alluded to above, one group of organisms known to carry out homologous recombination quite efficiently are yeast, and therefore, host cells used in practicing the methods described herein can be yeast cells (e.g., Saccharomyces cerevisiae, Schizosaccharomyces pombe, Pichia pastoris, etc.).
[0213] Yeast hosts are particularly suitable for manipulating donor genome material due to their unique genetic manipulation toolset. The natural capabilities of yeast cells and decades of research have created a rich set of tools for manipulating yeast DNA. These advantages are well known in the art. For example, yeast, with its rich genetic system, can assemble and reassemble nucleotide sequences through homologous recombination, an ability not shared by many readily available organisms. Yeast cells can be used to clone larger fragments of DNA, such as whole cells, organelles, and viral genomes that cannot be cloned in other organisms. Thus, in some embodiments, the enormous ability of yeast genetics to generate large nucleic acid molecules (e.g., synthetic genomics) can be utilized by using yeast as a host cell for assembly and maintenance. [Example]
[0214] Example 1 A codon-optimized coding sequence for TkoEndoMS containing an amino-terminal signal peptide (METDTLLLWV LLLWVPGSTG SKDKVTVIT (SEQ ID NO: 5)) and a carboxy-terminal six-histidine purification tag (FIG. 15) was designed using the following parameters: Codon usage was adjusted for the codon bias of Homo sapiens genes. Additionally, regions with very high (>80%) or very low (<30%) GC content were avoided whenever possible.
[0215] During the optimization process, the following cis-acting sequence motifs were avoided, where applicable: (1) internal TATA boxes, Chi sites, and ribosome entry sites, (2) AT-rich or GC-rich sequence stretches, (3) RNA instability motifs, (4) repeat sequences and RNA secondary structures, and (5) (cryptic) splicing donor and acceptor sites in higher eukaryotes. The result is the nucleotide sequence shown in Figure 15, which encodes a protein having the amino acid sequence shown in Figure 15.
[0216] The nucleotide sequence shown in Figure 15 was transfected and expressed in EXPI™ 293 cells. EXPI™ 293 cells were cultured for 6 days after transfection, and the expressed protein was then harvested. The secreted TkoEndoMS protein was purified using the His-tag on a HisTrap column using a linear gradient of 20 to 500 mM imidazole in Tris-HCl, 500 mM NaCl. The purified TkoEndoMS protein was dialyzed against 50 mM Tris-HCl pH 8.0, 0.5 mM DTT, 0.1 mM EDTA, 0.5 M NaCl for 16 hours. Purity was assessed by Coomassie blue staining, and the resulting TkoEndoMS was determined to be 95% pure. TkoEndoMS was stored at a final concentration of 130 ng / μl in 50 mM Tris-HCl pH 8.0, 0.5 mM DTT, 0.1 mM EDTA, 0.5 M NaCl, 50% glycerol.
[0217] Benchmark Oligonucleotide Assembly Protocol Assembly PCR [Table 4] A master mix of all reaction components was created, except for the oligonucleotide mix for assembly. 730 nl of the master mix was transferred to the wells of a 384-well plate using an ECHO® 555 liquid handler (Labcyte Inc.). 500 nl of the oligonucleotide mix was then added using the ECHO® 555. Thermal cycling was then performed using the cycler protocol shown below. [Table 5]
[0218] amplification [Table 6] A master mix of all components except the assembly PCR product was prepared. 8.8 μl of the master mix was then transferred using a multi-step pipettor to the wells of a 384-well plate containing the assembly PCR product. Thermal cycling was then performed using the cycler protocol shown below. [Table 7]
[0219] EndoMS Oligonucleotide Assembly Protocol Using PHUSION™ DNA Polymerase A. Assembly PCR This is the same as the benchmark protocol, except that the reaction mixture contains 0.020 μl of TkoEndoMS (130 ng / μl), and therefore 0.420 μl of HO.
[0220] B. Amplification This is the same as the benchmark protocol, except that the reaction mixture contains 0.140 μl of TkoEndoMS (130 ng / μl), and therefore 6.386 μl of HO.
[0221] Oligonucleotide Assembly Protocol using SUPERFI™ II DNA Polymerase (EndoMS Optional) A. Assembly PCR [Table 8] [Table 9] A master mix of all reaction components was created, except for the oligonucleotide mix for assembly. 730 nl of the master mix was transferred to a well of a 384-well plate using an ECHO® 555 liquid handler. 500 nl of the oligonucleotide mix was then added using the ECHO® 555. Thermal cycling was then performed using the cycler protocol shown below. [Table 10]
[0222] B. Amplification [Table 11] [Table 12] A master mix of all components except the assembly PCR product was prepared. 8.8 μl of the master mix was then transferred using a multi-step pipettor to the wells of a 384-well plate containing the assembly PCR product. Thermal cycling was then performed using the cycler protocol shown below. [Table 13]
[0223] Error correction protocol using T7 endonuclease I (T7NI) A. Error Correction I (Denaturation and Reannealing) [Table 14] [Table 15] Error Correction II (Disconnect Mismatch) [Table 16]
[0224] B. Error Correction III (Amplification) [Table 17] [Table 18]
[0225] Example 2 Thermostable mismatch endonuclease (TsMME) After demonstrating in Example 1 that the use of TkoEndoMS during assembly and / or amplification results in the production of nucleic acid molecules with reduced error rates, conditions were tested for further reduction in error rate. These conditions included the use of different thermostable mismatch endonucleases (abbreviated herein as "TsMME"), such as homologs of TkoEndoMS, different DNA polymerases, and different cycler protocols.
[0226] material and method: The "TsMME" used in the experiments described in this example, listed in Table 4, with the amino acid sequences of these enzymes shown in Table 15, was produced in Expi293 for thermostable error correction (abbreviated herein as "TsEC"). These enzymes, produced by Thermo Fisher Scientific GeneArt GmbH (Regensburg, DE), were greater than 95% pure and stored in the following buffer: 50 mM Tris-HCl pH 8.0, 0.5 mM DTT, 0.1 mM EDTA, 0.5 M NaCl, 50% glycerol, respectively.
[0227] In the experiments set out in this example, error correction using T7 endonuclease I was not performed. [Table 19]
[0228] Benchmark Oligonucleotide Assembly Protocol The benchmark data described in this example was generated using PHUSION™ DNA polymerase and either no error correction or error correction mediated by the specified thermostable enzyme.Unless otherwise specified, "benchmark" data was generated using PHUSION™ DNA polymerase without error correction.Before error correction was performed, benchmarking was performed because oligonucleotides with different sequences contained different numbers of errors.To correct for this variable, benchmark data was generated using the same oligonucleotides as those used to generate comparative data, unless otherwise specified herein.
[0229] Assembly PCR [Table 20] A master mix containing all components except the oligonucleotide mixture was created. 730 nL of the master mix was transferred to individual wells of a 384-well plate using a Labcyte ECHO® 555 acoustic liquid handler. 500 nL of the oligonucleotide mixture was then added to the same wells using a Labcyte ECHO® 555 acoustic liquid handler. [Table 21]
[0230] amplification [Table 22] A master mix containing all components except the assembly reaction product was prepared, and 8.8 μl of this master mix was then transferred with a multi-step pipettor to individual wells of a 384-well plate containing the assembly reaction products. [Table 23]
[0231] TsEC Oligonucleotide Assembly Protocol Using PHUSION™ DNA Polymerase assembly The method used was the same as the benchmark protocol described earlier in this example, except that the reaction mixture contained 0.020 μl of TkoEndoMS (130 ng / μl) and 0.420 μl of H2O.
[0232] amplification The method used was the same as the benchmark protocol above, except that the reaction mixture contained 0.140 μl of TkoEndoMS (130 ng / μl) and 6.386 μl of H2O.
[0233] Oligonucleotide Assembly Protocol using PLATINUM™ SUPERFI™ II DNA Polymerase (TsMME Optional) [Table 24] [Table 25] [Table 26] A master mix containing all components except the oligonucleotide mixture was created. 730 nL of the master mix was transferred to individual wells of a 384-well plate using a Labcyte ECHO® 555 acoustic liquid handler. 500 nL of the oligonucleotide mixture was then added to the same wells using a Labcyte ECHO® 555 acoustic liquid handler. [Table 27] [Table 28] [Table 29] [Table 30] [Table 31]
[0234] amplification [Table 32] [Table 33] A master mix containing all components except the assembly reaction product was prepared, and 8.8 μl of this master mix was then transferred with a multi-step pipettor to the wells of a 384-well plate containing the assembly reaction product. [Table 34]
[0235] result: Assembly of 20 individual fragments using the "Benchmark Oligonucleotide Assembly Protocol" and PHUSION™ DNA polymerase (PHUSION™) was used to establish a "Benchmark" / reference number of errors. The same 20 individual fragments were also assembled using the "Oligonucleotide Assembly Protocol," PLATINUM™ SUPERFI™ II DNA polymerase ("SUPERFI™ II"), with error correction using PhoNucS or SacEndoMS and Cycler Protocol C. The resulting data are shown in Tables 5 and 6 below. [Table 35]
[0236] The data in Table 5 show that treatment with SUPERFI™ II and PhoNucS improved the overall error rate, on average, compared to treatment with SUPERFI™ II and SacEndoMS. While SacEndoMS primarily corrected substitutions and had a smaller effect on deletions and insertions, PhoNucS was found to have significant error-correcting activity for deletions and insertions, in addition to greater activity for substitutions. The data also show that sequence errors in some nucleic acid fragments are easier to correct than others. For example, treatment with SUPERFI™ II and PhoNucS improved the overall error rate by 100% for two fragments and by 275% for three fragments, while treatment with SUPERFI™ II and SacEndoMS improved the overall error rate by 25% for one fragment and by 100% for four fragments. This variability is believed to be due, in part, to differences in the sequences of the nucleic acid fragments. Differences in nucleotide sequence can result in variations in the prevalence of different error types in nucleic acid fragments, and as discussed elsewhere herein, error-correcting enzymes differ in their ability to recognize and interact with (e.g., bind and / or cleave) different error types. [Table 36]
[0237] The data in Table 6 show that nucleic acid molecules assembled and amplified using PLATINUM™ SUPERFI™ II DNA polymerase with error correction mediated by the PhoNucS and SacEndoMS enzymes are almost completely devoid of four of the six substitution types, while the benchmark sample contains significant amounts of all six substitution types. Upon hybridization with wild-type molecules, the substitutions removed by the enzymes form mismatches for which their homolog, TkoEndoMS, has significant cleavage activity (Ishino et al., Nucl. Acids Res. 44:2977-2989 (2016)).
[0238] The data in Table 6 also suggest that the PhoNucS and SacEndoMS enzymes do not exhibit high levels of cleavage activity for (1) A>C and T>G and (2) G>T and C>A transversions. When hybridized with wild-type molecules, these transversions form mismatches for which their homolog, TkoEndoMS, has low cleavage activity (Ishino et al., Nucl. Acids Res. 44:2977-2989 (2016)). [Table 37]
[0239] Table 7 shows a comparison of error rate data for assembly and amplification of nucleic acid fragments with SUPERFI™ II versus PHUSION™ DNA polymerase. Two different thermal cycler protocols were used (Protocols A and C). As can be seen from the data, in the two runs described in Table 7, assembly and amplification of nucleic acid fragments with SUPERFI™ II was found to result in a lower error rate compared to PHUSION™ DNA polymerase. The data also show that the error rate improvement seen in Table 5 appears to be due in small part to the use of SUPERFI™ II. This suggests that the error rate improvement seen in Table 5 is due in large part to the use of TsMME. [Table 38]
[0240] As can be seen in Table 8, use of the "Benchmark Oligonucleotide Assembly Protocol" with PHUSION™ DNA polymerase and TkoEndoMS for error correction resulted in a substantial reduction in the number of sequence errors in the product nucleic acid molecules generated. [Table 39]
[0241] The data in Table 9 show that, compared to benchmark samples containing significant amounts of all six substitution types, nucleic acid molecules assembled and amplified using PHUSION™ DNA polymerase with error correction mediated by the TkoEndoMS enzyme exhibited significantly reduced rates of four of the six substitution types. When hybridized with wild-type molecules, the substitutions removed by TkoEndoMS form mismatches for which the enzyme has significant cleavage activity (Ishino et al., Nucl. Acids Res. 44:2977-2989 (2016)).
[0242] The data in Table 9 also suggest that the TkoEndoMS enzyme does not exhibit high levels of cleavage activity for (1) A>C and T>G and (2) G>T and C>A transversions. When hybridized to wild-type molecules, these transversions form mismatches at which TkoEndoMS has low cleavage activity (Ishino et al., Nucl. Acids Res. 44:2977-2989 (2016)). [Table 40]
[0243] Table 10 shows many effects, one of which is that the use of different thermostable error-correcting enzymes results in different error rates in the assembled and amplified product nucleic acid molecules. The number of errors present in the assembled and amplified nucleic acid molecules also varies to some extent depending on the cycler protocol used. Therefore, two factors that can be altered to obtain assembled and amplified nucleic acid molecules with low error rates are (1) the error-correcting enzyme(s) used, and (2) the method of assembling and amplifying the subcomponents of the nucleic acid molecule (e.g., thermal cycler protocol, buffers used / present, buffer components, etc.). [Table 41]
[0244] The data in Table 11 also demonstrate that efficient error rate reduction is achieved regardless of the initial error rate. For assembly and amplification using SUPERFI™ II polymerase and PhoNucS, a 2.1- to 2.6-fold error reduction was achieved when the benchmark error rate was 1 in 222 to 1 in 303 (Table 10), and a 1.9-fold error reduction was achieved when the benchmark error rate was 1 in 1092. For assembly and amplification using SUPERFI™ II polymerase and TkoEndoMS, a 1.5- to 1.8-fold error reduction was achieved when the benchmark error rate was 1 in 205 to 1 in 283 (Table 10), and a 2.1-fold error reduction was achieved when the benchmark error rate was 1 in 1092.
[0245] While specific aspects of the subject matter described herein have been shown and described herein, it will be apparent to those skilled in the art that such aspects are provided by way of example only. Those skilled in the art will recognize numerous variations, changes, and substitutions without departing from the subject matter described herein. It should be understood that various alternatives to the aspects of the subject matter described herein may be used in practicing the subject matter described herein. The following claims define the scope of the subject matter described herein, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby. [Table 42] [Table 43-1] [Table 43-2] [Table 44-1] [Table 44-2] [Table 44-3] [Table 45-1] [Table 45-2] [Table 45-3] [Table 45-4] [Table 45-5] [Table 45-6] [Table 45-7] [Table 45-8] [Table 45-9] [Table 45-10] [Table 45-11]
[0246] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. This includes the following patent documents: U.S. Patent Publication Nos. 2003 / 0152984, 2006 / 0115850, 2006 / 0127920, 2007 / 0231805, 2007 / 0292954, 2009 / 0275086, 2010 / 0062495, 2010 / 0216648, 2010 / 0291633, 2011 / 0124049, 2012 / 0053087, and 2017 / 253909. U.S. Patent Nos. 5,580,759, 5,624,827, 5,869,644, 6,110,668, 6,495,318, 6,521,427, 7,704,690, 7,833,759, 7,838,210, 8,224,578, 10,626,383, and 10,196,618. PCT Publication Nos. 2005 / 095605, 2010 / 040531, 2011 / 102802, 2013 / 049227, 2016 / 094512, and 2020 / 001783.
[0247] Exemplary subject matter of the present invention is represented by the following clauses:
[0248] Clause 1. A method for generating an error-corrected population of nucleic acid molecules, the method comprising: (a) assembling oligonucleotides having regions of terminal sequence complementarity by primary assembly PCR to form a population of assembled nucleic acid molecules; (b) amplifying the population of assembled nucleic acid molecules formed in step (a) by a primary amplification to form a population of amplified assembled nucleic acid molecules; A method wherein steps (a) and / or (b) are carried out in the presence of one or more thermostable mismatch-recognition proteins.
[0249] Clause 2. The method of clause 1, wherein at least one of the one or more thermostable mismatch recognition proteins is a thermostable mismatch binding protein.
[0250] Clause 3. The method of Clause 2, wherein the thermostable mismatch binding protein is selected from a mismatch binding protein having an amino acid sequence set forth in Table 13 or Table 15.
[0251] Clause 4. The method of clause 1, wherein at least one of the one or more thermostable mismatch recognition proteins is a thermostable mismatch endonuclease.
[0252] Clause 5. The method of clause 1 or 4, wherein the thermostable mismatch endonuclease is selected from an endonuclease having an amino acid sequence set forth in Table 12 or Table 15.
[0253] Clause 6. The method of clause 4 or 5, wherein the thermostable mismatch endonuclease is TkoEndoMS.
[0254] Clause 7. The method of any one of clauses 1-6, wherein a high-fidelity DNA polymerase is used in steps (a) and / or (b).
[0255] Clause 8. The method of clause 7, wherein the high-fidelity DNA polymerase is a component of an error-reducing polymerase reagent.
[0256] Clause 9. The method of clause 7 or 8, wherein the high-fidelity DNA polymerase is a polymerase having an amino acid sequence selected from the group consisting of (1) DNA polymerase 1, (2) DNA polymerase 2, (3) DNA polymerase 3, (4) DNA polymerase 4, (5) DNA polymerase 5, (6) DNA polymerase 6, and (7) DNA polymerase 7, as set forth in Table 14.
[0257] Clause 10. The method of clause 8 or 9, wherein the error-reducing polymerase reagent comprises one or more amine compounds.
[0258] Clause 11. One or more amine compounds (a) Dimethylamine hydrochloride (b) diisopropylamine hydrochloride, (c) ethyl(methyl)amine hydrochloride, and (d) trimethylamine hydrochloride.
[0259] Clause 12. The method of any one of clauses 1 to 11, wherein at least one of the one or more thermostable mismatch-recognition proteins is present in step (a).
[0260] Clause 13. The method of any one of clauses 1 to 12, wherein at least one of the one or more thermostable mismatch-recognition proteins is present in step (b).
[0261] Clause 14. The method of any one of clauses 1 to 13, wherein one or more error correction steps are performed after primary amplification.
[0262] Clause 15. The method of any one of clauses 1-14, wherein a primary post-amplification of the population of amplified assembled nucleic acid molecules is performed after step (b).
[0263] Clause 16. The method of any one of clauses 1-15, wherein the population of amplified assembled nucleic acid molecules is contacted with one or more mismatch-recognizing proteins prior to primary post-amplification.
[0264] Clause 17. The method of Clause 16, wherein at least one of the one or more mismatch recognition proteins is a mismatch endonuclease.
[0265] Clause 18. The method of Clause 17, wherein the mismatched endonuclease is a non-thermostable mismatched endonuclease.
[0266] Clause 19. The non-thermostable mismatch endonuclease (a) T7 endonuclease I, (b) CEL II nuclease; (c) CEL I nuclease, and (d) T4 endonuclease VII.
[0267] Clause 20. The method of any one of clauses 1-19, wherein the population of amplified assembled nucleic acid molecules comprises a subfragment of a larger nucleic acid molecule and is combined with another nucleic acid molecule that is also a subfragment of the larger nucleic acid molecule to form a pool of nucleic acid molecules.
[0268] Clause 21. The method of Clause 20, wherein the nucleic acid molecules of the pool of nucleic acid molecules are assembled by secondary assembly PCR to form larger nucleic acid molecules.
[0269] Clause 22. The method of clause 21, wherein the subfragments are contacted with one or more mismatch-recognition proteins prior to or during assembly by secondary assembly PCR.
[0270] Clause 23. The method of any one of clauses 20 to 22, wherein the larger nucleic acid molecule is heat denatured, then renatured, and subsequently contacted with one or more mismatch-recognition proteins.
[0271] Clause 24. The method of clause 23, wherein at least one of the one or more mismatch recognition proteins is a mismatch binding protein.
[0272] Clause 25. The method of Clause 24, wherein the mismatch binding protein is bound to a solid support.
[0273] Clause 26. The method of any one of clauses 1-25, wherein the population of amplified assembled nucleic acid molecules is sequenced.
[0274] Clause 27. The method of any one of clauses 1-26, wherein the population of amplified assembled nucleic acid molecules contains fewer than two errors per 1,000 base pairs.
[0275] Clause 28. A composition comprising a thermostable mismatch recognition protein, a DNA polymerase, and one or more amine compounds.
[0276] Clause 29. The composition of clause 28, wherein the DNA polymerase is a high-fidelity DNA polymerase.
[0277] Clause 30. The composition of clause 29, wherein the high-fidelity DNA polymerase is a component of an error-reducing polymerase reagent.
[0278] Clause 31. The composition of clause 29 or 30, wherein the high-fidelity DNA polymerase comprises an amino acid sequence set forth in Table 14.
[0279] Article 32. One or more amine compounds (a) dimethylamine hydrochloride, (b) diisopropylamine hydrochloride, (c) ethyl(methyl)amine hydrochloride, and (d) trimethylamine hydrochloride.
[0280] Clause 33. The composition of any one of clauses 28 to 32, further comprising two or more nucleic acid molecules.
[0281] Clause 34. The composition of clause 33, wherein the two or more nucleic acid molecules are subfragments of a larger nucleic acid molecule.
[0282] Clause 35. The composition of clause 33 or 34, wherein the two or more nucleic acid molecules are single-stranded.
[0283] Clause 36. The composition of clause 35, wherein two or more single-stranded nucleic acid molecules are less than 100 nucleotides in length.
[0284] Clause 37. The composition of Clause 35, wherein the two or more single-stranded nucleic acid molecules are from about 35 to about 90 nucleotides in length.
[0285] Clause 38. The composition of Clause 35, wherein the two or more single-stranded nucleic acid molecules are from about 30 to about 65 nucleotides in length.
[0286] Clause 39. The composition of any one of clauses 28 to 38, wherein the thermostable mismatch recognition protein is a mismatch endonuclease.
[0287] Clause 40. The composition of Clause 39, wherein the thermostable mismatch endonuclease is selected from an endonuclease having an amino acid sequence set forth in Table 12 or Table 15.
[0288] Clause 41. The composition of clause 40, wherein the thermostable mismatch endonuclease is TkoEndoMS.
[0289] Clause 42. The composition of any one of clauses 28 to 38, wherein the thermostable mismatch recognition protein is a mismatch binding protein.
[0290] Clause 43. The composition of clause 42, wherein the thermostable mismatch binding protein is selected from mismatch binding proteins having an amino acid sequence set forth in Table 13 or Table 15.
[0291] Clause 44. The composition of clause 33 or 34, wherein at least one of the two or more nucleic acid molecules is single-stranded and at least one of the two or more nucleic acid molecules is double-stranded.
[0292] Clause 45. A method for generating a nucleic acid molecule having a predetermined sequence, the method comprising: (a) providing a plurality of single-stranded oligonucleotides having complementary overlapping regions, each of the single-stranded oligonucleotides comprising a sequence region of a target nucleic acid molecule, the plurality of single-stranded oligonucleotides comprising: (i) a plurality of internal oligonucleotides, the plurality having a sequence region that overlaps with two other oligonucleotides; and (ii) providing two terminal oligonucleotides designed to be located at the 5' and 3' ends of the full-length nucleic acid molecule, the terminal oligonucleotides having sequence regions that overlap in plurality with one of the internal oligonucleotides; (b) assembling a plurality of oligonucleotides by primary assembly PCR to obtain an assembled double-stranded nucleic acid assembly product; (c) combining at least a portion of the assembly product obtained in step (b) with a pair of primers, the primers designed to bind to the 5' and 3' ends of the assembly product, and performing a PCR amplification reaction to produce an amplified assembly product; A method wherein step (b) and / or step (c) is carried out in the presence of one or more thermostable mismatch-recognition proteins.
[0293] Clause 46.(d) further comprising performing one or more error correction steps, the error correction steps comprising: (iii) denaturing and reannealing the amplified assembly product of step (c) to generate one or more mismatch containing double-stranded nucleic acids; and (iv) treating the mismatch containing double-stranded nucleic acid with one or more mismatch-recognition proteins; and (v) optionally performing an amplification reaction.
[0294] Clause 47. The method of Clause 46, wherein the mismatch recognition protein used in step (d) is a mismatch endonuclease or a mismatch binding protein.
[0295] Clause 48. The method of Clause 47, wherein the mismatched endonuclease is T7 endonuclease I.
[0296] Clause 49. The method of Clause 47, wherein the mismatch binding protein is MutS.
[0297] Clause 50. The method of clause 45 or 46, wherein the thermostable mismatch recognition protein is a thermostable mismatch endonuclease.
[0298] Clause 51. The method of Clause 50, wherein the thermostable mismatch endonuclease is derived from a hyperthermophilic archaeon, and optionally the hyperthermophilic archaeon is Pyrococcus furiosus or Pyrococcus abyssi.
[0299] Clause 52. The method of Clause 45 or 46, wherein the thermostable mismatch recognition protein is selected from the group of proteins having an amino acid sequence set forth in Table 12, 13, or 15, and variants thereof having at least 95% sequence identity thereto.
[0300] Clause 53. The method of any one of clauses 49 to 52, wherein the thermostable mismatch recognition protein is obtained by in vitro transcription / translation.
[0301] Clause 54. The method of any one of clauses 45-53, wherein one or more of steps (b), (c) and (d)(iii) are carried out in the presence of a high-fidelity DNA polymerase, optionally wherein the polymerase is selected from the group consisting of PHUSION™ DNA polymerase, PLATINUM™ SUPERFI™ II DNA polymerase, Q5 DNA polymerase, and PRIMESTAR GXL DNA polymerase.
[0302] Clause 55. The method of any one of clauses 45-53, wherein one or more of steps (b), (c), and (d)(iii) are carried out in the presence of a high-fidelity DNA polymerase, and optionally, the polymerase is a polymerase having an amino acid sequence selected from the group consisting of (1) DNA polymerase 1, (2) DNA polymerase 2, (3) DNA polymerase 3, (4) DNA polymerase 4, (5) DNA polymerase 5, (6) DNA polymerase 6, and (7) DNA polymerase 7, as set forth in Table 14.
[0303] Clause 56. The method of any one of clauses 45 to 53, wherein two or more amplified assembly products are pooled before performing one or more error correction steps.
[0304] Clause 57. The method of any one of clauses 46 to 53, further comprising treating the amplified assembly product with an exonuclease prior to the one or more error correction steps, optionally wherein the exonuclease is exonuclease I. Another aspect of the present invention may be as follows. [1] A method for generating an error-corrected population of nucleic acid molecules, the method comprising: (a) assembling oligonucleotides having regions of terminal sequence complementarity by primary assembly PCR to form a population of assembled nucleic acid molecules; (b) amplifying the population of assembled nucleic acid molecules formed in step (a) by a primary amplification to form a population of amplified assembled nucleic acid molecules; A method wherein steps (a) and / or (b) are carried out in the presence of one or more thermostable mismatch-recognition proteins. [2] The method according to [1], wherein at least one of the one or more thermostable mismatch recognition proteins is a thermostable mismatch binding protein. [3] The method according to [2], wherein the thermostable mismatch binding protein is selected from mismatch binding proteins having an amino acid sequence listed in Table 13 or Table 15. [4] The method according to [1], wherein at least one of the one or more thermostable mismatch recognition proteins is a thermostable mismatch endonuclease. [5] The method according to [1] or [4], wherein the thermostable mismatch endonuclease is selected from endonucleases having an amino acid sequence listed in Table 12 or Table 15. [6] The method according to [4] or [5], wherein the thermostable mismatch endonuclease is TkoEndoMS. [7] The method according to any one of [1] to [6] above, wherein a high-fidelity DNA polymerase is used in steps (a) and / or (b). [8] The method described in [7], wherein the high-fidelity DNA polymerase is a component of an error-reducing polymerase reagent. [9] The method according to [7] or [8], wherein the high-fidelity DNA polymerase is a polymerase having an amino acid sequence selected from the group consisting of (1) DNA polymerase 1, (2) DNA polymerase 2, (3) DNA polymerase 3, (4) DNA polymerase 4, (5) DNA polymerase 5, (6) DNA polymerase 6, and (7) DNA polymerase 7 listed in Table 14.
[10] The method of [8] or [9], wherein the error-reducing polymerase reagent comprises one or more amine compounds.
[11] The one or more amine compounds are (a) Dimethylamine hydrochloride (b) diisopropylamine hydrochloride, (c) ethyl(methyl)amine hydrochloride, and (d) trimethylamine hydrochloride.
[12] The method according to any one of [1] to
[11] , wherein at least one of the one or more thermostable mismatch-recognition proteins is present in step (a).
[13] The method according to any one of [1] to
[12] , wherein at least one of the one or more thermostable mismatch-recognition proteins is present in step (b).
[14] The method according to any one of [1] to
[13] above, wherein one or more error correction steps are performed after the primary amplification.
[15] The method according to any one of [1] to
[14] above, wherein a post-primary amplification of the amplified population of assembled nucleic acid molecules is carried out after step (b).
[16] The method according to any one of [1] to
[15] , wherein the amplified population of assembled nucleic acid molecules is contacted with one or more mismatch-recognition proteins prior to the post-primary amplification.
[17] The method according to
[16] , wherein at least one of the one or more mismatch recognition proteins is a mismatch endonuclease.
[18] The method according to
[17] , wherein the mismatched endonuclease is a non-thermostable mismatched endonuclease.
[19] The non-thermostable mismatch endonuclease is (a) T7 endonuclease I, (b) CEL II nuclease; (c) CEL I nuclease, and (d) T4 endonuclease VII.
[20] The method of any one of [1] to
[19] , wherein the population of amplified assembled nucleic acid molecules comprises a subfragment of a larger nucleic acid molecule and is combined with another nucleic acid molecule that is also a subfragment of the larger nucleic acid molecule to form a pool of nucleic acid molecules.
[21] The method of
[20] , wherein the nucleic acid molecules of the nucleic acid molecule pool are assembled by secondary assembly PCR to form the larger nucleic acid molecule.
[22] The method of
[21] , wherein the subfragments are contacted with the one or more mismatch-recognition proteins before or during assembly by secondary assembly PCR.
[23] The method according to any one of
[20] to
[22] , wherein the larger nucleic acid molecule is heat-denatured, then renatured, and then contacted with the one or more mismatch-recognition proteins.
[24] The method according to
[23] , wherein at least one of the one or more mismatch recognition proteins is a mismatch binding protein.
[25] The method according to
[24] , wherein the mismatch-binding protein is bound to a solid support.
[26] The method according to any one of [1] to
[25] above, wherein the population of amplified assembled nucleic acid molecules is sequenced.
[27] The method according to any one of [1] to
[26] above, wherein the population of amplified assembled nucleic acid molecules contains less than two errors per 1,000 base pairs.
[28] A composition comprising a thermostable mismatch recognition protein, a DNA polymerase, and one or more amine compounds.
[29] The composition described in
[28] , wherein the DNA polymerase is a high-fidelity DNA polymerase.
[30] The composition described in
[29] , wherein the high-fidelity DNA polymerase is a component of an error-reducing polymerase reagent.
[31] The composition described in
[29] or
[30] , wherein the high-fidelity DNA polymerase comprises an amino acid sequence listed in Table 14.
[32] The one or more amine compounds are (a) dimethylamine hydrochloride, (b) diisopropylamine hydrochloride, (c) ethyl(methyl)amine hydrochloride, and (d) The composition according to
[28] , wherein the compound is selected from the group consisting of trimethylamine hydrochloride.
[33] The composition according to any one of
[28] to
[32] above, further comprising two or more nucleic acid molecules.
[34] The composition described in
[33] , wherein the two or more nucleic acid molecules are subfragments of a larger nucleic acid molecule.
[35] The composition described in
[33] or
[34] , wherein the two or more nucleic acid molecules are single-stranded.
[36] The composition described in
[35] , wherein the two or more single-stranded nucleic acid molecules are less than 100 nucleotides in length.
[37] The composition according to
[35] above, wherein the two or more single-stranded nucleic acid molecules are about 35 to about 90 nucleotides in length.
[38] The composition according to
[35] above, wherein the two or more single-stranded nucleic acid molecules are about 30 to about 65 nucleotides in length.
[39] The composition described in any one of
[28] to
[35] , wherein the thermostable mismatch-recognition protein is a mismatch endonuclease.
[40] The composition described in
[39] , wherein the thermostable mismatch endonuclease is selected from endonucleases having an amino acid sequence listed in Table 12 or Table 15.
[41] The composition described in
[40] , wherein the thermostable mismatch endonuclease is TkoEndoMS.
[42] The composition described in any one of
[28] to
[38] , wherein the thermostable mismatch-recognition protein is a mismatch-binding protein.
[43] The composition described in
[42] , wherein the thermostable mismatch binding protein is selected from mismatch binding proteins having an amino acid sequence listed in Table 13 or Table 15.
[44] The composition described in
[33] or
[34] , wherein at least one of the two or more nucleic acid molecules is single-stranded and at least one of the two or more nucleic acid molecules is double-stranded.
[45] A method for generating a nucleic acid molecule having a predetermined sequence, the method comprising: (a) providing a plurality of single-stranded oligonucleotides having complementary overlapping regions, each of the single-stranded oligonucleotides comprising a sequence region of a target nucleic acid molecule, the plurality of single-stranded oligonucleotides comprising: (i) a plurality of internal oligonucleotides, the plurality having a sequence region that overlaps with two other oligonucleotides in the plurality; and (ii) comprising two terminal oligonucleotides designed to be located at the 5' and 3' ends of the full-length nucleic acid molecule, the terminal oligonucleotides having sequence regions that overlap with one of the internal oligonucleotides in the plurality; (b) assembling the plurality of oligonucleotides by primary assembly PCR to obtain an assembled double-stranded nucleic acid assembly product; (c) combining at least a portion of the assembly product obtained in step (b) with a pair of primers, the primers designed to bind to the 5' and 3' ends of the assembly product, and performing a PCR amplification reaction to produce an amplified assembly product; A method wherein step (b) and / or step (c) is carried out in the presence of one or more thermostable mismatch-recognition proteins.
[46] (d) further comprising performing one or more error correction steps, the error correction steps comprising: (iii) denaturing and reannealing the amplified assembly product of step (c) to generate one or more mismatch containing double-stranded nucleic acids; and (iv) treating the mismatch containing double-stranded nucleic acid with one or more mismatch-recognition proteins; and (v) optionally, performing an amplification reaction.
[47] The method according to
[46] , wherein the mismatch recognition protein used in step (d) is a mismatch endonuclease or a mismatch binding protein.
[48] The method described in
[47] , wherein the mismatched endonuclease is T7 endonuclease I.
[49] The method described in
[47] , wherein the mismatch binding protein is MutS.
[50] The method according to
[50] , wherein the thermostable mismatch recognition protein is a thermostable mismatch endonuclease.
[51] The method described in
[50] , wherein the thermostable mismatch endonuclease is derived from a hyperthermophilic archaeon, and optionally the hyperthermophilic archaeon is Pyrococcus furiosus or Pyrococcus abyssi.
[52] The method of
[45] or
[46] , wherein the thermostable mismatch-recognition protein is selected from the group consisting of proteins having an amino acid sequence set forth in Table 12, 13, or 15, and variants thereof having at least 95% sequence identity thereto.
[53] The method according to any one of
[49] to
[52] , wherein the thermostable mismatch recognition protein is obtained by in vitro transcription / translation.
[54] The method of any one of
[45] to
[53] , wherein one or more of steps (b), (c), and (d)(iii) are carried out in the presence of a high-fidelity DNA polymerase, optionally wherein the polymerase is selected from the group consisting of PHUSION™ DNA polymerase, PLATINUM™ SUPERFI™ II DNA polymerase, Q5 DNA polymerase, and PRIMESTAR GXL DNA polymerase.
[55] The method of any one of
[45] to
[53] , wherein one or more of steps (b), (c), and (d)(iii) are performed in the presence of a high-fidelity DNA polymerase, and optionally the polymerase is a polymerase having an amino acid sequence selected from the group consisting of (1) DNA polymerase 1, (2) DNA polymerase 2, (3) DNA polymerase 3, (4) DNA polymerase 4, (5) DNA polymerase 5, (6) DNA polymerase 6, and (7) DNA polymerase 7 listed in Table 14.
[56] The method according to any one of
[45] to
[53] , wherein two or more amplified assembly products are pooled before performing the one or more error correction steps.
[57] The method of any one of
[46] to
[53] , further comprising treating the amplified assembly product with an exonuclease prior to the one or more error correction steps, optionally wherein the exonuclease is exonuclease I.
Claims
1. 1. A method for generating an error-corrected population of nucleic acid molecules, said method comprising: (a) assembling single-stranded oligonucleotides by primary assembly PCR to form a population of assembled nucleic acid molecules; wherein each of the oligonucleotides comprises a fragment of the assembled nucleic acid molecule and a terminal region comprising a complementary sequence; the oligonucleotide hybridizes to another oligonucleotide via the complementary sequence; assembly of all said oligonucleotides results in a sequence corresponding to said assembled nucleic acid molecule; and, (b) amplifying the population of assembled nucleic acid molecules formed in step (a) by a primary amplification to form a population of amplified assembled nucleic acid molecules; A method wherein step (a) is carried out in the presence of one or more thermostable mismatch endonucleases, wherein said thermostable mismatch endonucleases retain at least 85% of their biological activity after being heated at 95°C for 5 minutes.
2. 2. The method of claim 1, wherein the thermostable mismatch endonuclease is selected from endonucleases having an amino acid sequence set forth in Table 12 or Table 15.
3. The method of claim 1 or 2, wherein the thermostable mismatched endonuclease is TkoEndoMS.
4. 4. The method of claim 1, wherein a high-fidelity DNA polymerase is used in steps (a) and / or (b), and wherein the high-fidelity DNA polymerase exhibits a substitution error rate of less than 1.0 x 10 substitutions per base.
5. The method of claim 4, wherein the high-fidelity DNA polymerase is a polymerase having an amino acid sequence selected from the group consisting of (1) DNA polymerase 1, (2) DNA polymerase 2, (3) DNA polymerase 3, (4) DNA polymerase 4, (5) DNA polymerase 5, (6) DNA polymerase 6, and (7) DNA polymerase 7 listed in Table 14.
6. 6. The method of any one of claims 1 to 5, wherein at least one of the one or more thermostable mismatch endonucleases is present in step (b).
7. The method according to any one of claims 1 to 6, wherein one or more error correction steps are performed after the primary amplification.
8. The method of any one of claims 1 to 7, wherein a secondary amplification of the amplified population of assembled nucleic acid molecules is carried out after step (b).
9. 9. The method of claim 8, wherein the amplified population of assembled nucleic acid molecules is contacted with one or more mismatched endonucleases prior to the secondary amplification.
10. 10. The method of claim 9, wherein the mismatched endonuclease is a non-thermostable mismatched endonuclease.
11. the non-thermostable mismatch endonuclease (a) T7 endonuclease I, (b) CEL II nuclease; (c) CEL I nuclease, and (d) T4 endonuclease VII.
12. The method of any one of claims 1 to 11, further comprising combining the amplified population of assembled nucleic acid molecules with one or more additional nucleic acid molecules to form a pool of nucleic acid molecules.
13. The method of claim 12, further comprising assembling the amplified population of assembled nucleic acid molecules and the one or more additional nucleic acid molecules by secondary assembly PCR to form a larger nucleic acid molecule.
14. The method described in claim 13, wherein the amplified population of assembled nucleic acid molecules and the one or more additional nucleic acid molecules are contacted with the one or more mismatch endonucleases prior to or during assembly by secondary assembly PCR.
15. 15. The method of any one of claims 13 to 14, wherein the larger nucleic acid molecule is heat denatured, then renatured, and subsequently contacted with the one or more mismatch-binding proteins.
16. 16. The method of claim 15, wherein the mismatch binding protein is bound to a solid support.
17. The method of any one of claims 1 to 16, wherein the amplified population of assembled nucleic acid molecules is sequenced.
18. 18. The method of any one of claims 1 to 17, wherein the population of amplified assembled nucleic acid molecules contains fewer than 2 errors per 1,000 base pairs.
19. (c) performing one or more error correction steps, the error correction steps comprising: (i) denaturing and reannealing the amplified assembly products of step (b) to generate one or more double-stranded nucleic acids containing mismatches; and (ii) treating the double-stranded nucleic acid containing the mismatch with one or more mismatch endonucleases or mismatch-binding proteins; and 10. The method of claim 1, comprising: (iii) performing an amplification reaction.
20. 20. The method of claim 19, wherein the mismatched endonuclease is T7 endonuclease I.
21. 20. The method of claim 19, wherein the mismatch binding protein is MutS.
22. The method of claim 1 , wherein the mismatched endonuclease is derived from a hyperthermophilic archaeon.
23. The method described in claim 22, wherein the hyperthermophilic archaea is Pyrococcus furiosus or Pyrococcus abyssi.
24. The method of any one of claims 1 to 23, wherein the mismatched endonuclease is obtained by in vitro transcription / translation.
25. 25. The method of any one of claims 19 to 24, wherein one or more of steps (a), (b) and (c)(iii) are carried out in the presence of a high-fidelity DNA polymerase, wherein the high-fidelity DNA polymerase exhibits a substitution error rate of less than 1.0 x 10 substitutions per base.
26. The method described in claim 25, wherein the high-fidelity polymerase is a polymerase having an amino acid sequence selected from the group consisting of (1) DNA polymerase 1, (2) DNA polymerase 2, (3) DNA polymerase 3, (4) DNA polymerase 4, (5) DNA polymerase 5, (6) DNA polymerase 6, and (7) DNA polymerase 7 listed in Table 14.
27. 27. The method of any one of claims 19 to 26, wherein two or more amplified assembly products are pooled before performing the one or more error correction steps.
28. 28. The method of any one of claims 19 to 27, further comprising treating the amplified assembly product with an exonuclease prior to the one or more error correction steps.
29. The method described in claim 28, wherein the exonuclease is exonuclease I.
Citation Information
Patent Citations
Quantitative amplification method using labeled probes and 3'→5' exonuclease activity
JP2007531527A
Methods of Reducing Errors in Nucleic Acid Populations
JP2008517586A
Gene synthesis using pooled DNA
JP2008534016A
Compositions and methods for high-fidelity assembly of nucleic acids
JP2014526899A
Materials and methods for the synthesis of nucleic acid molecules with minimal errors
JP2015509005A