Method for introducing mutations

By using low biased DNA polymerase and nucleotide analogs, combining sample tag design and primer binding sites, the problem of uneven mutations of target nucleic acid molecules in the prior art is solved, and the efficiency and accuracy of sequencing and protein engineering are improved.

CN112088219BActive Publication Date: 2025-07-18ILLUMINA SINGAPORE PTE LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980027144.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-02-20
Filing Date
2019-02-19
Publication Date
2025-07-18
Estimated Expiration
2039-11-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively and uniformly introduce mutations into target nucleic acid molecules, especially in long nucleic acid sequences, resulting in difficulty in identifying sample sources and sequence assembly during sequencing and protein engineering.

Method used

Amplification was performed using low biased DNA polymerase and binding to nucleotide analogs such as dPTP, sample tag groups were designed to ensure random uniformity and discrimination of mutations while prioritizing long target nucleic acid molecules to achieve mutations by introducing specific primer binding sites and adaptors.

Benefits of technology

It realizes uniformly and randomly introduced mutations into target nucleic acid molecules, improves the effectiveness of sequencing methods and sample recognition capabilities, simplifies the assembly process of long nucleic acid sequences, and enhances the accuracy of protein engineering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0002733781270000201
    Figure BDA0002733781270000201
  • Figure BDA0002733781270000211
    Figure BDA0002733781270000211
  • Figure BDA0002733781270000221
    Figure BDA0002733781270000221
Patent Text Reader

Abstract

The present invention relates to a method for introducing mutations into at least one target nucleic acid molecule, the method comprising: (a) providing a sample comprising at least one of at least one target nucleic acid molecule; and (b) amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase. The present invention relates to: the use of a low-bias DNA polymerase in a method for introducing mutations into one or more nucleic acid molecules, a sample tag set, a method for designing a sample tag set, a computer-readable medium, and a method for preferentially amplifying a target nucleic acid molecule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to: methods for introducing mutations into one or more nucleic acid molecules, the use of low-bias DNA polymerases in methods for introducing mutations into one or more nucleic acid molecules, sample tag sets, methods for designing sample tag sets, computer-readable media, and methods for preferentially amplifying target nucleic acid molecules. Background Art

[0002] DNA polymerases can be used to introduce mutations into nucleic acid sequences. This can be useful in a number of applications. For example, mutagenesis techniques can be useful in applications including Sequencing assisted by mutagenesis (SAM) techniques and for introducing mutations into protein sequences to discover mutations that affect protein activity.

[0003] DNA polymerases with low fidelity can be used to introduce mutations. Low-fidelity DNA polymerases make errors during replication, resulting in the introduction of mutations. However, many low-fidelity DNA polymerases introduce mutations at a rate of less than 2% per mutagenesis reaction (one round of replication), and for some applications, a higher mutagenesis rate is useful. Additionally, low-fidelity DNA polymerases can introduce mutations in a biased manner. Such DNA polymerases can be referred to as High bias DNA polymerases.

[0004] In the presence of nucleotide analogs (such as dPTP), DNA polymerases can be used to introduce mutations by replicating a sequence. The DNA polymerase can incorporate the nucleotide analog in place of the natural nucleotide. Then, in subsequent replication cycles, the nucleotide analog can pair with a natural nucleotide that was not present in the original sequence, thereby introducing a mutation. Introducing mutations by replicating a sequence in the presence of a nucleotide analog can be used to achieve a higher mutation rate.

[0005] Common DNA polymerases (such as Taq polymerase) can be used to incorporate nucleotide analogs in place of natural nucleotides. However, these polymerases are high-bias polymerases. High-bias DNA polymerases can exhibit two possible biases: mutation bias and template amplification bias.

[0006] Some highly biased polymerases have a high mutation bias because they do not mutate all four natural nucleotides (adenine, cytosine, guanine, and thymine) randomly and equally. For example, a highly biased DNA polymerase may mutate some nucleotides at a higher frequency than other nucleotides. The adenine / thymine pair is linked by two hydrogen bonds, while the guanine / cytosine pair is linked by three hydrogen bonds. Therefore, highly biased DNA polymerases are more likely to introduce mutations into the adenine / thymine pair compared to the guanine / cytosine pair.

[0007] Highly biased polymerases with a high mutation bias may not incorporate nucleotide analogs randomly. For example, a highly biased polymerase may tend to substitute certain bases with nucleotide analogs. DPTP can interconvert between two different tautomeric forms (the imino form and the amino form). The imino tautomer can form Watson-Crick base pairs with adenine, while the amino form can form Watson-Crick base pairs with guanine (Kong Thoo Lin P, Brown D M (1989). “Synthesis and duplex stability of oligonucleotides containing cytosine-thymine analogues”. Nucleic Acids Research. 17:10373–10383; Stone M J et al. (1991). “Molecular basis for methoxyamine-initiated mutagenesis: 1 H nuclear magnetic resonance studies of base-modified oligodeoxynucleotides.” Journal of Molecular Biology. 222:711–723; Nedderman A N R et al. (1993). “Molecular basis for methoxyamine initiated mutagenesis: 1"Hnuclear magnetic resonance studies of oligonucleotide duplexes containing base-modified cytosine residues”. Journal of Molecular Biology. 230:1068–1076; Moore M H et al. (1995). “Direct observation of two base-pairing modes of a cytosine-thymine analogue with guanine in a DNA Z-form duplex. Significance for base analogue mutagenesis”. Journal of Molecular Biology. 251:665–673). This effectively means that replication in the presence of dPTP can be used to introduce substitutions in the nucleotide sequence in place of adenine, cytosine, guanine, or thymine. However, in aqueous solution, the ratio of the imino form to the amino form of dPTP has been shown to be approximately 10:1 (Harris V H et al. (2003). “The effect of tautomeric constant on the specificity of nucleotide incorporation during DNA replication: support for the rare tautomer hypothesis of substitution mutagenesis”. Journal of Molecular Biology. 326:1389-1401).Thus, when introducing mutations using dPTP with a polymerase (such as Taq polymerase), it introduces substitutions of adenine and thymine more frequently than substitutions of guanine and cytosine (Zaccolo M et al. (1996). “An approach to random mutagenesis of DNA using mixtures of triphosphate derivatives of nucleoside analogues”. Journal of Molecular Biology. 255:589–603; Harris V H et al. (2003). “The effect of tautomeric constant on the specificity of nucleotide incorporation during DNA replication: support for the rare tautomer hypothesis of substitution mutagenesis”. Journal of Molecular Biology. 326:1389-1401).

[0008] Second, a high-bias polymerase can exhibit template amplification bias, i.e., a high-bias polymerase can replicate some template nucleic acid molecules with a higher success rate than other nucleic acid molecules in each PCR cycle. In multiple cycles of PCR, this bias creates extremely high copy number differences between templates. Regions of the template nucleic acid molecule can form secondary structures or can contain a higher proportion of some nucleotides (such as guanine or cytosine nucleotides) than other nucleotides. Compared to template nucleic acid molecules rich in adenine and thymine, a high-bias polymerase can more effectively amplify, for example, template nucleic acid molecules rich in guanine and cytosine, or can more effectively amplify template nucleic acid molecules that do not form secondary structures.

[0009] Many applications of mutagenesis would be more effective if mutagenesis could be performed with low bias (mutation bias and template amplification).

[0010] The accurate assembly of genomic sequences has proven difficult because many second-generation sequencing platforms are only capable of sequencing short nucleic acid fragments and require amplification of the target nucleic acid sequence during the sequencing process to provide sufficient nucleic acid molecules for the sequencing step. If a user wishes to sequence a larger nucleic acid sequence, this can be achieved by sequencing regions of the target nucleic acid molecule. The user must then computationally assemble the sequence of the complete nucleic acid sequence from the regional sequences.

[0011] Assembling nucleic acid sequences using regional sequences can be difficult. In particular, in cases where long regional sequences are very similar to each other, it may be difficult to determine whether the sequences of two regions are replicate sequences of the same original template nucleic acid molecule or correspond to sequences from two different original template nucleic acid molecules. Similarly, it may be difficult to determine whether the sequences of two regions are replicate sequences corresponding to the same part of the template nucleic acid molecule or are actually two different replicates within the template nucleic acid molecule. These difficulties can be circumvented by introducing mutations into the target nucleic acid molecule prior to amplification. The user can then identify that fragments with the same mutation pattern are likely to originate from the same part of the same original template nucleic acid molecule. This type of sequencing method is sometimes referred to as Sequencing aided by mutagenesis (SAM). Summary of the Invention

[0012] The above sequencing method is more effective when the mutations introduced into the target nucleic acid molecule are uniformly random. If the mutations are uniformly random, then for example any given part of the template nucleic acid molecule will have a higher likelihood of having a unique mutation pattern. Therefore, there is a need to identify DNA polymerases that can introduce mutations randomly and uniformly (with low mutation bias).

[0013] In addition, sequencing methods using DNA polymerases with high template amplification bias may be limited. DNA polymerases with high template amplification bias will replicate and / or mutate some target nucleic acid molecules better than others, so sequencing methods using such high-bias DNA polymerases may not sequence some target nucleic acid molecules well.

[0014] The inventors have identified polymerases that are low-bias polymerases (with low template amplification bias and low mutation bias) and are thus particularly useful in methods for introducing mutations into at least one target nucleic acid molecule.

[0015] The user may wish to use the method of the present invention on more than one sample simultaneously. In such cases, it would be advantageous for the user to be able to identify which target nucleic acid molecule came from which original sample. This identification can be achieved by labeling the target nucleic acid molecules with sample tags. However, the sample tags themselves can mutate during the method, so the inventors have determined how to design sample tags that can be distinguished from each other even if they mutate.

[0016] The user may also wish to ensure that the method of the present invention is preferentially used to mutate and amplify long target nucleic acid molecules compared to short nucleic acid molecules. The inventors have found that this can be achieved by introducing special primer binding sites at each end of the target nucleic acid molecule.

[0017] Accordingly, in a first aspect of the present invention, there is provided a method for introducing a mutation into at least one target nucleic acid molecule, the method comprising:

[0018] a. providing at least one sample comprising at least one target nucleic acid molecule; and

[0019] b. amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase.

[0020] In a second aspect of the present invention, there is provided the use of a low-bias DNA polymerase in a method for introducing a mutation into at least one target nucleic acid molecule.

[0021] In a third aspect of the present invention, there is provided a method for determining the sequence of at least one target nucleic acid molecule, the method comprising the method for introducing a mutation of the present invention.

[0022] In a fourth aspect of the present invention, there is provided a method for protein engineering, the method comprising the method for introducing a mutation of the present invention.

[0023] In a fifth aspect of the present invention, there is provided a set of sample tags, wherein each sample tag differs from substantially all other sample tags in the set by at least one low-probability mutation difference or at least three high-probability mutation differences.

[0024] In a sixth aspect of the present invention, there is provided a method for designing a set of sample tags suitable for use in a method for introducing a mutation into at least one target nucleic acid molecule, the method comprising:

[0025] a. analyzing the method for introducing a mutation into at least one target nucleic acid molecule and determining the average number of low-probability mutations that occur during the method for introducing a mutation into at least one target nucleic acid molecule; and

[0026] b. determining the sequences of the set of sample tags, wherein each sample tag differs from substantially all sample tags in the set by more low-probability differences than the average number of low-probability mutations that occur during the method for introducing a mutation into at least one target nucleic acid molecule.

[0027] In a seventh aspect of the present invention, there is provided a method for introducing a mutation into at least one target nucleic acid molecule, comprising:

[0028] a. providing at least one sample comprising at least one target nucleic acid molecule; and

[0029] b. introducing a mutation into at least one target nucleic acid molecule by amplifying the at least one target nucleic acid molecule using a DNA polymerase to provide a mutated at least one target nucleic acid molecule,

[0030] wherein step b. is carried out using unequal concentrations of dNTPs.

[0031] In an eighth aspect of the present invention, there is provided a sample tag set, which can be obtained by a method for designing the sample tag set of the present invention.

[0032] In a ninth aspect of the present invention, there is provided a computer-readable medium configured to execute a method for designing the sample tag set of the present invention.

[0033] In a tenth aspect of the present invention, there is provided a method for preferentially amplifying a target nucleic acid molecule having a length greater than 1 kbp, the method comprising:

[0034] a. providing at least one sample comprising the target nucleic acid molecule;

[0035] b. introducing a first adaptor at the 3'-end of the target nucleic acid molecule and a second adaptor at the 5'-end of the target nucleic acid molecule; and

[0036] c. amplifying the target nucleic acid molecule using a primer complementary to a part of the first adaptor,

[0037] wherein the first adaptor and the second adaptor can anneal to each other. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 Shows the mutation levels achieved using three different polymerases in the presence or absence of dPTP. Panel A shows the data obtained using Taq (Jena Biosciences), panel B shows the data obtained using LongAmp (New England Biolabs), and panel C shows the data obtained using Primestar GXL (Takara). The dark grey bars show the results obtained in the absence of dPTP, and the light grey bars show the results obtained in the presence of 0.5 mM dPTP;

[0039] Figure 2 Describes the mutation rates obtained by dPTP mutagenesis of templates with various G+C contents using a Thermococcus polymerase (Primestar GXL; Takara). For a low GC template (33% GC) from Staphylococcus aureus, the median mutation observed was ~7%, while the median for the other templates was approximately 8%);

[0040] Figure 3 is a sequence listing;

[0041] Figure 4 Depicts the self-annealing of a nucleic acid molecule when using a first primer binding site and a second primer binding site that anneal to each other;

[0042] Figure 6 Depicts the sizes of target nucleic acid molecules amplified using adaptors annealed to each other (right line) or using standard adaptors (left line);

[0043] Figure 7 Provides a graphical representation of mutagenesis using the nucleotide analogue dPTP (referred to as "P" in Figure 7 ). DETAILED DESCRIPTION

[0044] General Definitions

[0045] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0046] Generally, the term "comprising" is intended to mean including but not limited to. For example, the phrase "a method for introducing a mutation into at least one target nucleic acid molecule comprising" certain steps should be interpreted to mean that the method includes the recited steps, but other steps may be performed.

[0047] In some embodiments of the invention, the word "comprising" is replaced with "consisting of". The term "consisting of" is intended to be limiting. For example, the phrase "a method for introducing a mutation into at least one target nucleic acid molecule consisting of" certain steps should be understood to mean that the method includes the recited steps and no other steps are performed.

[0048] For the purposes of the present invention, to determine the percentage identity of two sequences (e.g., two polynucleotide sequences), the sequences are aligned for optimal alignment purposes (e.g., gaps may be introduced in the first sequence for optimal alignment with the second sequence). Then the nucleotides or amino acid residues at each position are compared. When the position in the first sequence is occupied by the same residue as at the corresponding position in the second sequence, the residues are identical at that position. The percentage identity between two sequences is a function of the number of positions shared by the sequences (i.e., identity % = number of identical positions / total number of positions × 100). Generally, sequence alignment is performed over the length of the reference sequence. For example, to evaluate whether a test sequence is at least 95% identical to SEQ ID NO.2 (reference sequence), one of ordinary skill in the art would align over the length of SEQ ID NO.2 and identify how many positions in the test sequence are identical to SEQ ID NO.2. If at least 80% of the positions are identical, the test sequence is at least 80% identical to SEQ ID NO.2. If the sequence is shorter than SEQ ID NO.2, gaps should be considered non-identical positions.

[0049] Those skilled in the art are aware of different computer programs that can be used to determine the homology or identity between two sequences. For example, mathematical algorithms can be used to perform sequence comparison and determination of the percentage of identity between two sequences. In one embodiment, the percentage of identity between two amino acid or nucleic acid sequences is determined using the Blosum 62 matrix or the PAM250 matrix, and gap weights of 16, 14, 12, 10, 8, 6, or 4, length weights of 1, 2, 3, 4, 5, or 6, using the Needleman and Wunsch (1970) algorithm, which has been incorporated into the GAP program of the Accelrys GCG software package (available at http: / / www.accelrys.com / products / gcg / ).

[0050] Method for introducing mutations into at least one target nucleic acid molecule

[0051] In a first aspect, the present invention provides a method for introducing mutations into at least one target nucleic acid molecule. In another aspect, the present invention provides the use of a low-bias DNA polymerase in a method for introducing mutations into at least one target nucleic acid molecule.

[0052] The mutation can be a substitution mutation, an insertion mutation, or a deletion mutation. For the purposes of the present invention, the term "substitution mutation" shall be construed to mean that a nucleotide is replaced by a different nucleotide. For example, the conversion of the sequence ATCC to the sequence AGCC is a substitution mutation. For the purposes of the present invention, the term "insertion mutation" shall be construed to mean that at least one nucleotide is added to the sequence. For example, the conversion of the sequence ATCC to the sequence ATTCC is an example of an insertion mutation (an additional T nucleotide is inserted). For the purposes of the present invention, the term "deletion mutation" shall be construed to mean that at least one nucleotide is removed from the sequence. For example, the conversion of the sequence ATTCC to ATCC is an example of a deletion mutation (the T nucleotide is removed). Preferably, the mutation is a substitution mutation.

[0053] For the purposes of the present invention, a "nucleic acid molecule" refers to a polymeric form of nucleotides of any length. The nucleotides can be deoxyribonucleotides, ribonucleotides, or analogs thereof. Preferably, the target nucleic acid molecule consists of deoxyribonucleotides or ribonucleotides. Even more preferably, the target nucleic acid molecule consists of deoxyribonucleotides, i.e., the target nucleic acid molecule is a DNA molecule.

[0054] At least one "target nucleic acid molecule" can be any nucleic acid molecule into which a user of the method desires to introduce a mutation. The target nucleic acid molecule can form part of a larger nucleic acid molecule such as a chromosome. The target nucleic acid molecule can contain a gene, multiple genes, or a fragment of a gene. The size of the target nucleic acid molecule can be greater than 1 kbp, greater than 1.5 kbp, greater than 2 kb, greater than 4 kbp, greater than 5 kbp, greater than 7 kbp, greater than 8 kbp, from 1 kbp to 50 kbp, or from 1 kbp to 20 kbp.

[0055] The term "at least one target nucleic acid molecule (molecule)" is considered interchangeable with the term "at least one target nucleic acid molecule (molecules)".

[0056] "At least one target nucleic acid molecule" can be single-stranded or can be part of a double-stranded complex. For example, if at least one target nucleic acid molecule consists of deoxyribonucleotides, it can form part of a double-stranded DNA complex. In such a case, one strand (e.g., the coding strand) will be considered the at least one target nucleic acid molecule, and the other strand is a nucleic acid molecule complementary to the at least one target nucleic acid molecule.

[0057] A method for introducing a mutation into at least one target nucleic acid molecule can include:

[0058] a. Providing at least one sample that contains at least one target nucleic acid molecule; and

[0059] b. Amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase.

[0060] Providing at least one sample that contains at least one target nucleic acid molecule

[0061] A method for introducing a mutation into at least one target nucleic acid molecule can include the step of providing at least one sample that contains at least one target nucleic acid molecule.

[0062] The at least one sample can include any sample that contains at least one target nucleic acid molecule. The at least one sample can be obtained from any source. For example, the at least one sample can include a nucleic acid sample derived from a human, e.g., a sample extracted from a skin swab of a human patient. Alternatively, the at least one sample can be derived from other sources, e.g., a sample from a water source. Such a sample may contain billions of template nucleic acid molecules. Each of these billions of target nucleic acid molecules can be mutated simultaneously using the method of the present invention, and thus there is no upper limit to the number of target nucleic acid molecules that can be used in the method of the present invention.

[0063] In one embodiment, step a. includes providing more than one sample. For example, step a. may include providing 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 15, 20, 25, 50, 75, or 100 samples. Alternatively, step a. includes providing less than 2000, less than 1000, less than 750, or less than 500 samples. In another embodiment, step a. includes providing 2 to 100, 2 to 75, 2 to 50, 2 to 25, 5 to 15, or 7 to 15 samples.

[0064] Amplification of at least one target nucleic acid molecule using a low bias DNA polymerase

[0065] The methods of the invention may comprise amplifying at least one target nucleic acid molecule using a low bias DNA polymerase.

[0066] Amplification of at least one target nucleic acid molecule refers to duplicating the at least one target nucleic acid molecule to provide at least one nucleic acid molecule complementary to at least one target nucleic acid molecule and / or a replica of at least one target nucleic acid molecule. Amplification of at least one target nucleic acid molecule using a low-bias DNA polymerase increases the number of replicas of at least one target nucleic acid molecule and introduces mutations into at least one target nucleic acid molecule. Due to the introduction of mutations, the replicas are not necessarily the same as the original at least one target nucleic acid molecule. The original at least one target nucleic acid molecule and the replicas of at least one target nucleic acid molecule can be collectively referred to as "at least one mutated target nucleic acid molecule".

[0067] For example, amplifying at least one target nucleic acid molecule using a low bias DNA polymerase can include incubating a sample containing at least one target nucleic acid molecule with a low bias DNA polymerase and appropriate primers under conditions suitable for the low bias DNA polymerase to catalyze the generation of copies of the at least one target nucleic acid molecule.

[0068] Suitable primers include short nucleic acid molecules that are complementary to such regions: regions flanking the at least one target nucleic acid molecule, or regions flanking the nucleic acid molecule () that is complementary to the at least one target nucleic acid molecule. For example, if the target nucleic acid molecule is part of a chromosome, the primer can be complementary to the following: a region in the chromosome that is close to the 3' end of the target nucleic acid molecule to the 3' end, and a nucleic acid molecule that is complementary to the region close to the 5' end of the target nucleic acid molecule to the 5' end, or the primer will be complementary to the following: a region in the chromosome that is close to the 3' end of the nucleic acid molecule that is complementary to the target nucleic acid molecule to the 3' end, and a nucleic acid molecule that is complementary to the region close to the 5' end of the nucleic acid molecule that is complementary to the target nucleic acid molecule to the 5' end. Alternatively, the user can introduce a primer binding site (short nucleic acid sequence) into a region flanking at least one target nucleic acid molecule. This is described in detail in the section entitled "Barcodes, Samples, and Adapters".

[0069] Suitable conditions include temperatures at which a low-bias DNA polymerase can catalyze the production of replicas of at least one target nucleic acid molecule. For example, temperatures of 40 °C to 90 °C, 50 °C to 80 °C, 60 °C to 70 °C, or approximately 68 °C can be used.

[0070] The step of amplifying at least one target nucleic acid molecule can include multiple rounds of replication. For example, the step of amplifying at least one target nucleic acid molecule preferably includes:

[0071] i) performing rounds of replication on at least one target nucleic acid molecule to provide at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule; and

[0072] ii) performing rounds of replication on at least one target nucleic acid molecule to provide replicas of the at least one target nucleic acid molecule.

[0073] Optionally, the step of amplifying at least one target nucleic acid molecule includes replicating the at least one target nucleic acid molecule at least 2 rounds, at least 4 rounds, at least 6 rounds, at least 8 rounds, or at least 10 rounds. Some of these rounds of replication performed on the at least one target nucleic acid molecule can be carried out in the presence of nucleotide analogs. Optionally, the step of amplifying at least one target nucleic acid molecule includes replicating at least 1 round, at least 2 rounds, at least 3 rounds, at least 4 rounds, at least 5 rounds, or at least 6 rounds at a temperature of 60 °C to 80 °C.

[0074] Optionally, the step of amplifying at least one target nucleic acid molecule is carried out using polymerase chain reaction (PCR). PCR is a process that involves performing multiple rounds of the following steps to replicate nucleic acid molecules:

[0075] a) denaturation;

[0076] b) annealing;

[0077] c) extension; and

[0078] d) elongation.

[0079] Mix the nucleic acid molecule (e.g., at least one target nucleic acid molecule) with a suitable primer and polymerase (e.g., the low-bias DNA polymerase of the present invention). In the denaturation step, heat the nucleic acid molecule to a temperature above 90 °C to denature the double-stranded nucleic acid molecule (separate into two strands). In the annealing step, cool the nucleic acid molecule to a temperature below 75 °C, such as 55 °C to 70 °C, about 55 °C, or about 68 °C, to anneal the primer to the nucleic acid molecule. In the extension step, heat the nucleic acid molecule to a temperature above 60 °C to enable the DNA polymerase to catalyze primer extension and add nucleotides complementary to the template strand. In the elongation step, heat the nucleic acid molecule to a temperature at which the DNA polymerase has high activity (e.g., a temperature of 60 °C to 70 °C) to catalyze the addition of other complementary nucleic acids to complete the new nucleic acid strand.

[0080] Optionally, the method of the present invention includes multiple rounds of PCR using a low-bias DNA polymerase.

[0081] Low-bias DNA polymerase

[0082] The method of the present invention may include the step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase.

[0083] According to the present invention, a "low-bias DNA polymerase" is a DNA polymerase that: (a) exhibits a low mutation bias, and / or (b) exhibits a low template amplification bias.

[0084] Low mutation bias

[0085] A low-bias DNA polymerase that exhibits a low mutation bias is a DNA polymerase that can mutate adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine at a similar rate. In one embodiment, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at a similar rate.

[0086] Optionally, the low-bias DNA polymerase is capable of mutating adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or approximately 1:1, respectively. Preferably, the low-bias DNA polymerase is capable of mutating guanine or adenine at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or approximately 1:1, respectively. Preferably, the low-bias DNA polymerase is capable of mutating thymine and cytosine at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or approximately 1:1, respectively.

[0087] In this embodiment, in the step of amplifying at least one target nucleic acid molecule using the low-bias DNA polymerase, the DNA polymerase mutates adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine nucleotides in the at least one target nucleic acid molecule at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or approximately 1:1, respectively. Preferably, the low-bias DNA polymerase mutates guanine and adenine nucleotides in the at least one target nucleic acid molecule at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or approximately 1:1, respectively. Preferably, the low-bias DNA polymerase mutates thymine and cytosine nucleotides in the at least one target nucleic acid molecule at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or approximately 1:1, respectively.

[0088] Optionally, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at a ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or approximately 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3.

[0089] In this embodiment, in the step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase, the DNA polymerase mutates the adenine, thymine, guanine, and cytosine nucleotides in the at least one target nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or approximately 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase mutates the adenine, thymine, guanine, and cytosine nucleotides in the at least one target nucleic acid molecule at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3.

[0090] Adenine, thymine, cytosine, and / or guanine can be replaced by another nucleotide. For example, if the low-bias DNA polymerase can mutate adenine, then amplifying at least one target nucleic acid molecule in the presence of the low-bias DNA polymerase can replace at least one adenine nucleotide in the nucleic acid molecule with thymine, guanine, or cytosine. Similarly, if the low-bias DNA polymerase can mutate thymine, then amplifying at least one target nucleic acid molecule in the presence of the low-bias DNA polymerase can replace at least one thymine nucleotide with adenine, guanine, or cytosine. If the low-bias DNA polymerase can mutate guanine, then amplifying at least one target nucleotide in the presence of the low-bias DNA polymerase can replace at least one guanine nucleotide with thymine, adenine, or cytosine. If the low-bias DNA polymerase can mutate cytosine, then amplifying at least one target nucleotide in the presence of the low-bias DNA polymerase can replace at least one cytosine nucleotide with thymine, guanine, or adenine.

[0091] A low-bias DNA polymerase may not be able to directly displace nucleotides, but it may still be able to mutate nucleotides by replacing the corresponding nucleotides on the complementary strand. For example, if the target nucleic acid molecule contains thymine, there will be an adenine nucleotide at the corresponding position in at least one nucleic acid molecule complementary to at least one target nucleic acid molecule. The low-bias DNA polymerase may be able to replace the adenine nucleotide in at least one nucleic acid molecule complementary to at least one target nucleic acid molecule with guanine. Thus, when at least one nucleic acid molecule complementary to at least one target nucleic acid molecule is replicated, this will result in cytosine being present at the position in the corresponding replicated at least one target nucleic acid molecule where there was originally thymine (thymine-to-cytosine substitution).

[0092] In one embodiment, the low-bias DNA polymerase mutates 1% to 15%, 2% to 10%, or about 8% of the nucleotides in at least one target nucleic acid. In such an embodiment, the step of amplifying at least one target nucleic acid molecule using the low-bias DNA polymerase is carried out such that 1% to 15%, 2% to 10%, or about 8% of the nucleotides in at least one target nucleic acid are mutated. For example, if a user wishes to mutate about 8% of the nucleotides in a target nucleic acid molecule and the low-bias DNA polymerase mutates about 1% of the nucleotides per round of replication, the step of amplifying at least one target nucleic acid molecule using the low-bias DNA polymerase may comprise 8 rounds of replication.

[0093] In one embodiment, the low-bias DNA polymerase is capable of mutating 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5% of the nucleotides in at least one target nucleic acid molecule per round of replication. In one embodiment, the low-bias DNA polymerase mutates 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5% of the nucleotides in at least one target nucleic acid molecule per round of replication. The actual amount of mutation that occurs per round may vary, but may average 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5%.

[0094] Whether a DNA polymerase is able to mutate nucleotides, and if so, at what rate

[0095] Whether a low-bias DNA polymerase can mutate a certain proportion of nucleotides in at least one target nucleic acid molecule per round of replication can be determined by amplifying a nucleic acid molecule of a known sequence for a certain number of replication rounds in the presence of the low-bias DNA polymerase. The resulting amplified nucleic acid molecules can then be sequenced, and the percentage of mutated nucleotides per round of replication can be calculated. For example, a nucleic acid molecule of a known sequence can be amplified using 10 rounds of PCR in the presence of a low-bias DNA polymerase. The resulting nucleic acid molecules can then be sequenced. If the resulting nucleic acid molecules contain 10% nucleotides that are different from the corresponding nucleotides in the original known sequence, the user will understand that the low-bias DNA polymerase can mutate 1% of the nucleotides in at least one target nucleic acid molecule on average per round of replication. Similarly, to see whether a low-bias DNA polymerase mutates a certain proportion of nucleotides in at least one target nucleic acid molecule in a given method, the user can perform the method on a nucleic acid molecule of a known sequence and use sequencing to determine the percentage of mutated nucleotides after the method is completed.

[0096] If, when a low-bias DNA polymerase is used to amplify a nucleic acid molecule, it provides nucleic acid molecules in which a nucleotide (such as adenine) is substituted or deleted in some cases, then the low-bias DNA polymerase can mutate that nucleotide. Preferably, the term "mutate" refers to the introduction of a substitution mutation, and in some embodiments, the term "mutate" can be replaced with "introduce a substitution in...".

[0097] If, when performing the step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase, the low-bias DNA polymerase mutates a nucleotide such as adenine in at least one target nucleic acid molecule in the method of the present invention, then this step results in at least one target nucleic acid molecule that is mutated with the nucleotide being mutated in some cases. For example, if, when performing the step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase, the low-bias DNA polymerase mutates adenine in at least one target nucleic acid molecule, then this step results in at least one target nucleic acid molecule that is mutated with at least one adenine being substituted or deleted.

[0098] To determine whether a DNA polymerase can introduce certain mutations, a technician only needs to test the DNA polymerase using a nucleic acid molecule of known sequence. A suitable nucleic acid molecule of known sequence is a fragment from a bacterial genome of known sequence, such as that of Escherichia coli MG1655. The technician can use PCR to amplify the nucleic acid molecule of known sequence in the presence of a low-bias DNA polymerase. The technician can then sequence the amplified nucleic acid molecule and determine whether its sequence is the same as the original known sequence. If it is not the same, the technician can determine the nature of the mutation. For example, if the technician wishes to use a nucleotide analogue to determine whether a DNA polymerase can mutate adenine, the technician can use PCR to amplify the nucleic acid molecule of known sequence in the presence of the nucleotide analogue and sequence the resulting amplified nucleic acid molecule. If the amplified DNA has a mutation at the position corresponding to the adenine nucleotide in the known sequence, the technician will know that the DNA polymerase can use the nucleotide analogue to mutate adenine.

[0099] The rate ratio can be calculated in a similar manner. For example, if the technician wishes to determine the rate ratio of guanine and cytosine nucleotide mutations, the technician can use PCR to amplify a nucleic acid molecule with a known sequence in the presence of a low-bias DNA polymerase. The technician can then sequence the resulting amplified nucleic acid molecule and identify how many guanine nucleotides have been substituted or deleted, and how many cytosine nucleotides have been substituted or deleted. The rate ratio is the ratio of the number of guanine nucleotides that have been substituted or deleted to the number of cytosine nucleotides that have been substituted or deleted. For example, if 16 guanine nucleotides have been substituted or deleted and 8 cytosine nucleotides have been substituted or deleted, the guanine and cytosine nucleotides have mutated at a rate ratio of 16:8 or 2:1, respectively.

[0100] Using nucleotide analogues

[0101] A low-bias DNA polymerase may not be able to directly replace nucleotides with other nucleotides (at least not at high frequency), but a low-bias DNA polymerase may still be able to mutate a nucleic acid molecule using a nucleotide analogue. A low-bias DNA polymerase may be able to replace nucleotides with other natural nucleotides (i.e., cytosine, guanine, adenine, or thymine) or nucleotide analogues.

[0102] For example, a low-bias DNA polymerase can be a high-fidelity DNA polymerase. Generally, high-fidelity DNA polymerases tend to introduce few mutations because they are highly accurate. However, the inventors have found that some high-fidelity DNA polymerases may still be able to mutate a target nucleic acid molecule because they may be able to introduce a nucleotide analogue into the target nucleic acid molecule.

[0103] In one embodiment, in the absence of nucleotide analogs, a high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, from 0% to 0.0015%, or from 0% to 0.001% mutations per round of replication.

[0104] In one embodiment, a low-bias DNA polymerase is capable of incorporating a nucleotide analog into at least one target nucleic acid molecule. In one embodiment, a low-bias DNA polymerase incorporates a nucleotide analog into at least one target nucleic acid molecule. In one embodiment, a low-bias DNA polymerase can use a nucleotide analog to mutate adenine, thymine, guanine, and cytosine. In one embodiment, a low-bias DNA polymerase uses a nucleotide analog to mutate adenine, thymine, guanine, and / or cytosine in at least one target nucleic acid molecule. In one embodiment, a DNA polymerase replaces guanine, cytosine, adenine, and / or thymine with a nucleotide analog. In one embodiment, a DNA polymerase can replace guanine, cytosine, adenine, and / or thymine with a nucleotide analog.

[0105] Incorporating a nucleotide analog into at least one target nucleic acid molecule can be used to mutate nucleotides because the nucleotide analog can be incorporated in place of an existing nucleotide and the nucleotide analog can pair with a nucleotide in the opposite strand. For example, dPTP can be incorporated into a nucleic acid molecule in place of a pyrimidine nucleotide (which can replace thymine or cytosine); see Figure 7 . Once in a nucleic acid strand, when in its imino tautomeric form, dPTP can pair with adenine. Thus, when a complementary strand is formed, the complementary strand can have adenine at the position complementary to dPTP. Similarly, once in a nucleic acid strand, when in its amino tautomeric form, dPTP can pair with guanine. Thus, when a complementary strand is formed, the complementary strand can have guanine at the position complementary to dPTP.

[0106] For example, if dPTP is introduced into at least one target nucleic acid molecule of the present invention, then when forming at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule, the at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule will contain adenine or guanine at the position complementary to dPTP in the at least one target nucleic acid molecule (depending on whether dPTP is in its amino form or imino form). When replicating at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule, the resulting replica of the at least one target nucleic acid molecule will contain thymine or cytosine at the position complementary to dPTP in the at least one target nucleic acid molecule. Thus, a mutation to thymine or cytosine can be introduced into the mutated at least one target nucleic acid molecule.

[0107] Alternatively, if dPTP is introduced into at least one nucleic acid molecule complementary to at least one target nucleic acid molecule when forming a replica of the at least one target nucleic acid molecule, the replica of the at least one target nucleic acid molecule will contain adenine or guanine (depending on the tautomeric form of dPTP) at positions complementary to dPTP in the at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule. Thus, a mutation to adenine or guanine can be introduced into the mutated at least one target nucleic acid molecule.

[0108] In one embodiment, a low-bias DNA polymerase can substitute cytosine or thymine with a nucleotide analogue. In another embodiment, the low-bias DNA polymerase incorporates guanine or adenine nucleotides at a rate ratio of 0.5 - 1.5:0.5 - 1.5, 0.6 - 1.4:0.6 - 1.4, 0.7 - 1.3:0.7 - 1.3, 0.8 - 1.2:0.8 - 1.2, or approximately 1:1, respectively, using a nucleotide analogue. The guanine or adenine nucleotides can be incorporated by a low-bias DNA polymerase that pairs them relative to a nucleotide analogue such as dPTP. In another embodiment, the low-bias DNA polymerase incorporates guanine or adenine nucleotides at a rate ratio of 0.7 - 1.3:0.7 - 1.3, respectively.

[0109] One of ordinary skill in the art can use conventional methods to determine whether a low-bias DNA polymerase can incorporate a nucleotide analogue into at least one target nucleic acid molecule or to mutate adenine, thymine, guanine, and / or cytosine in at least one target nucleic acid molecule using a nucleotide analogue.

[0110] For example, to determine whether a low-bias DNA polymerase can incorporate a nucleotide analogue into at least one target nucleic acid molecule, a technician can use the low-bias DNA polymerase to amplify a nucleic acid molecule for two rounds of replication. The first round of replication should be carried out in the presence of the nucleotide analogue, and the second round of replication should be carried out in the absence of the nucleotide analogue. The resulting amplified nucleic acid molecules can be sequenced to see if mutations have been introduced and, if so, how many. The user should repeat the experiment in the absence of the nucleotide analogue and compare the number of mutations introduced with and without the nucleotide analogue. If the number of mutations introduced in the presence of the nucleotide analogue is significantly higher than the number of mutations introduced in the absence of the nucleotide analogue, the user can conclude that the low-bias DNA polymerase can incorporate the nucleotide analogue. Similarly, a technician can use a nucleotide analogue to determine whether a DNA polymerase incorporates a nucleotide analogue or mutates adenine, thymine, guanine, and / or cytosine. The technician only needs to perform the method in the presence of the nucleotide analogue and see if the method results in mutations at positions initially occupied by adenine, thymine, guanine, and / or cytosine.

[0111] If a user wishes to mutate at least one target nucleic acid molecule using a nucleotide analogue, the method can include the step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase, wherein the step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase is carried out in the presence of the nucleotide analogue and the step of amplifying at least one target nucleic acid molecule provides at least one target nucleic acid molecule comprising the nucleotide analogue.

[0112] Suitable nucleotide analogues include dPTP (2'-deoxy-P-nucleoside-5'-triphosphate), 8-oxo-dGTP (7,8-dihydro-8-oxoguanine), 5Br-dUTP (5-bromo-2'-deoxy-uridine-5'-triphosphate), 2OH-dATP (2-hydroxy-2'-deoxyadenosine-5'-triphosphate), dKTP (9-(2-deoxy-β-D-ribofuranosyl)-N6-methoxy-2,6,-diaminopurine-5'-triphosphate), and dITP (2'-deoxyinosine 5'-triphosphate). The nucleotide analogue can be dPTP. The nucleotide analogue can be used to introduce substitution mutations as described in Table 1.

[0113] Table 1

[0114] Nucleotide Substitution 8-oxo-dGTP A:T to C:G and T:A to G:C dPTP A:T to G:C and G:C to A:T 5Br-dUTP A:T to G:C and T:A to C:G 2OH-dATP A:T to C:G, G:C to T:A and A:T to G:C dITP A:T to G:C and G:C to A:T dKTP A:T to G:C and G:C to A:T

[0115] Different nucleotide analogs can be used alone or in combination to introduce different mutations into at least one target nucleic acid molecule. Thus, a low-bias DNA polymerase can use nucleotide analogs to introduce guanine-to-adenine substitution mutations, cytosine-to-thymine substitution mutations, adenine-to-guanine substitution mutations, and thymine-to-cytosine substitution mutations. A low-bias DNA polymerase may optionally be able to use nucleotide analogs to introduce guanine-to-adenine substitution mutations, cytosine-to-thymine substitution mutations, adenine-to-guanine substitution mutations, and thymine-to-cytosine substitution mutations.

[0116] A low-bias DNA polymerase may be able to introduce guanine-to-adenine substitution mutations, cytosine-to-thymine substitution mutations, adenine-to-guanine substitution mutations, and thymine-to-cytosine substitution mutations at rate ratios of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or approximately 1:1:1:1, respectively. Preferably, a low-bias DNA polymerase is able to introduce guanine-to-adenine substitution mutations, cytosine-to-thymine substitution mutations, adenine-to-guanine substitution mutations, and thymine-to-cytosine substitution mutations at rate ratios of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively. Suitable methods for determining whether a low-bias DNA polymerase is able to introduce substitution mutations and at what rate ratios are described under the heading "Whether a DNA polymerase can mutate nucleotides and, if so, at what rate".

[0117] In some methods, a low-bias DNA polymerase introduces guanine-to-adenine substitution mutations, cytosine-to-thymine substitution mutations, adenine-to-guanine substitution mutations, and thymine-to-cytosine substitution mutations at rate ratios of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or approximately 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase introduces guanine-to-adenine substitution mutations, cytosine-to-thymine substitution mutations, adenine-to-guanine substitution mutations, and thymine-to-cytosine substitution mutations at a rate ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively. Suitable methods for determining whether a substitution mutation has been introduced and at what rate ratio are described under the heading "Can a DNA polymerase mutate nucleotides, and if so, at what rate?"

[0118] Typically, when a low-bias DNA polymerase uses a nucleotide analogue to introduce mutations, this requires more than one round of replication. In the first round of replication, the low-bias DNA polymerase incorporates the nucleotide analogue in place of the nucleotide, and in the second round of replication, the nucleotide analogue pairs with the natural nucleotide to introduce a substitution mutation in the complementary strand. The second round of replication can be carried out in the presence of the nucleotide analogue. However, the method can further include the step of amplifying at least one target nucleic acid molecule containing the nucleotide analogue in the absence of the nucleotide analogue. The step of amplifying at least one target nucleic acid molecule containing the nucleotide analogue in the absence of the nucleotide analogue can be carried out using a low-bias DNA polymerase.

[0119] Optionally, the method provides at least one target nucleic acid molecule with a mutation, and the method includes a further step of amplifying at least one target nucleic acid molecule with a mutation using a low-bias DNA polymerase.

[0120] Low-template amplification bias

[0121] A low-bias DNA polymerase can have low template amplification bias. A low-bias DNA polymerase has low template amplification bias if it can amplify different target nucleic acid molecules with a similar degree of success per cycle. A high-bias DNA polymerase may have difficulty amplifying template nucleic acid molecules that contain a high G:C content or have a high degree of secondary structure. In one embodiment, the low-bias DNA polymerase of the present invention has low template amplification bias for template nucleic acid molecules that are less than 25,000 nucleotides in length, less than 10,000 nucleotides in length, 1 to 15,000 or 1 to 10,000 nucleotides in length.

[0122] In one embodiment, to determine whether a DNA polymerase has low template amplification bias, a person skilled in the art can use the DNA polymerase to amplify a series of different sequences and sequence the resulting amplified DNA to see if the different sequences are amplified at different levels. For example, a person skilled in the art can select a series of short (e.g., 50 nucleotide) nucleic acid molecules with different characteristics, including nucleic acid molecules with a high GC content, nucleic acid molecules with a low GC content, nucleic acid molecules with a high degree of secondary structure, and nucleic acid molecules with a low degree of secondary structure. Then, the user can use the DNA polymerase to amplify those sequences and quantify the amplification level of each nucleic acid molecule in the nucleic acid molecules. In one embodiment, the DNA polymerase has low template amplification bias if the levels are within 25%, 20%, 10%, or 5% of each other.

[0123] Alternatively, in one embodiment, a DNA polymerase has low template amplification bias if it can amplify a 7 kbp - 10 kbp fragment with a Kolmogorov-Smirnov D of less than 0.1, less than 0.09, or less than 0.08. The Kolmogorov-Smirnov D of a particular low-bias DNA polymerase capable of amplifying a 7 kbp - 10 kbp fragment can be determined using the assay provided in Example 4.

[0124] A low-bias DNA polymerase can be a high-fidelity DNA polymerase. A high-fidelity DNA polymerase is a DNA polymerase that is not very error-prone, so when a high-fidelity DNA polymerase is used to amplify a target nucleic acid molecule in the absence of nucleotide analogs, a large number of mutations are usually not introduced. High-fidelity DNA polymerases are not usually used in methods for introducing mutations because error-prone DNA polymerases are generally considered to be more efficient. However, this application demonstrates that certain high-fidelity polymerases can introduce mutations using nucleotide analogs and can introduce those mutations with a lower bias compared to error-prone DNA polymerases such as Taq polymerase.

[0125] High-fidelity DNA polymerases have other advantages. When used with nucleotide analogs, high-fidelity DNA polymerases can be used to introduce mutations, but in the absence of nucleotide analogs, high-fidelity DNA polymerases can replicate target nucleic acid molecules with a high degree of precision. This means that a user can use the same DNA polymerase to efficiently mutate at least one target nucleic acid molecule and amplify the at least one mutated target nucleic acid molecule with high precision. If a low-fidelity DNA polymerase is used to mutate a target nucleic acid molecule, it may be necessary to remove the low-fidelity DNA polymerase from the reaction mixture before amplifying the target nucleic acid molecule.

[0126] High-fidelity DNA polymerases can have proofreading activity. Proofreading activity can help the DNA polymerase amplify a target nucleic acid sequence with high precision. For example, a low-bias DNA polymerase can contain a proofreading domain. The proofreading domain can confirm whether the nucleotide that has been added by the polymerase is correct (checking that it pairs correctly with the corresponding nucleic acid of the complementary strand), and if not, excise that nucleotide from the nucleic acid molecule. The inventors have surprisingly found that in some DNA polymerases, the proofreading domain will accept the pairing of natural nucleotides with nucleotide analogs. The structures and sequences of suitable proofreading domains are known to those skilled in the art. DNA polymerases that contain a proofreading domain include members of DNA polymerase families I, II, and III, such as Pfu polymerase (derived from Pyrococcus furiosus), T4 polymerase (derived from bacteriophage T4), and the Thermococcal polymerases described in detail below.

[0127] In one embodiment, in the absence of nucleotide analogs, the high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, from 0% to 0.0015%, or from 0% to 0.001% mutations per round of replication.

[0128] In addition, a low-bias DNA polymerase can contain a processivity-enhancing domain. The processivity-enhancing domain allows the DNA polymerase to amplify the target nucleic acid molecule more quickly. This is advantageous because it allows the methods of the invention to be carried out more quickly.

[0129] Thermococcal polymerase

[0130] In one embodiment, the low-bias DNA polymerase is a fragment or variant of a polypeptide comprising SEQ ID NO.2, SEQ ID NO.4, SEQ ID NO.6, or SEQ ID NO.7. The polypeptides of SEQ ID NO.2, 4, 6, and 7 are Thermococcus polymerases. The polymerases of SEQ ID NO.2, SEQ ID NO.4, SEQ ID NO.6, or SEQ ID NO.7 are low-bias DNA polymerases with high fidelity, and they can mutate target nucleic acid molecules by incorporating nucleotide analogs (such as dPTP). The polymerases of SEQ ID NO.2, SEQ ID NO.4, SEQ ID NO.6, or SEQ ID NO.7 are particularly advantageous because they have a low mutation bias and a low template amplification bias. The polymerases of SEQ ID NO.2, SEQ ID NO.4, SEQ ID NO.6, or SEQ ID NO.7 are also highly processive and are high-fidelity polymerases containing proofreading domains, which means that in the absence of nucleotide analogs, they can rapidly and accurately amplify mutated target nucleic acid molecules.

[0131] The low-bias DNA polymerase may comprise a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of the following sequences:

[0132] a. The sequence of SEQ ID NO.2;

[0133] b. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.2;

[0134] c. The sequence of SEQ ID NO.4;

[0135] d. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.4;

[0136] e. The sequence of SEQ ID NO.6;

[0137] f. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.6;

[0138] g. The sequence of SEQ ID NO.7; or

[0139] h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.7.

[0140] Preferably, the low-bias DNA polymerase comprises a fragment of at least 700 consecutive amino acids of the following sequences:

[0141] a. The sequence of SEQ ID NO.2;

[0142] b. A sequence that is at least 98% or at least 99% identical to SEQ ID NO.2;

[0143] c. The sequence of SEQ ID NO.4;

[0144] d. A sequence that is at least 98% or at least 99% identical to SEQ ID NO.4;

[0145] e. The sequence of SEQ ID NO.6;

[0146] f. A sequence that is at least 98% or at least 99% identical to SEQ ID NO.6;

[0147] g. The sequence of SEQ ID NO.7; or

[0148] h. A sequence that is at least 98% or at least 99% identical to SEQ ID NO.7.

[0149] The low-bias DNA polymerase may comprise:

[0150] a. The sequence of SEQ ID NO.2;

[0151] b. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.2;

[0152] c. The sequence of SEQ ID NO.4;

[0153] d. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.4;

[0154] e. The sequence of SEQ ID NO.6;

[0155] f. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.6;

[0156] g. The sequence of SEQ ID NO.7; or

[0157] h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.7.

[0158] Preferably, the low-bias DNA polymerase comprises:

[0159] a. The sequence of SEQ ID NO.2;

[0160] b. A sequence that is at least 98% or at least 99% identical to SEQ ID NO.2;

[0161] c. The sequence of SEQ ID NO.4;

[0162] d. A sequence that is at least 98% or at least 99% identical to SEQ ID NO.4;

[0163] e. The sequence of SEQ ID NO.6;

[0164] f. A sequence that is at least 98% or at least 99% identical to SEQ ID NO.6;

[0165] g. The sequence of SEQ ID NO.7; or

[0166] h. A sequence that is at least 98% or at least 99% identical to SEQ ID NO.7.

[0167] The low-bias DNA polymerase can be a Thermococcus polymerase or a derivative thereof. The DNA polymerases of SEQ ID NO.2, 4, 6, and 7 are Thermococcus polymerases. Thermococcus polymerases are advantageous because they are generally high-fidelity polymerases that can be used to introduce mutations with low mutation and template amplification bias using nucleotide analogs.

[0168] A Thermococcus polymerase is a polymerase having the polypeptide sequence of a polymerase isolated from a strain of the genus Thermococcus. A derivative of a Thermococcus polymerase can be a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 contiguous amino acids of the Thermococcus polymerase, or at least 95%, at least 98%, at least 99%, or 100% identical to a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 contiguous amino acids of the Thermococcus polymerase. A derivative of a Thermococcus polymerase can be at least 95%, at least 98%, at least 99%, or 100% identical to the Thermococcus polymerase. A derivative of a Thermococcus polymerase can be at least 98% identical to the Thermococcus polymerase.

[0169] In the context of the present invention, Thermococcus polymerases from any strain can be effective. In one embodiment, the Thermococcus polymerase is derived from a Thermococcus strain selected from the group consisting of Thermococcus kodakarensis, Thermococcus celer, Thermococcus siculi, and Thermococcus sp. KS-1. The Thermococcus polymerases from these strains are described in SEQ ID NO.2, SEQ ID NO.4, SEQ ID NO.6, and SEQ ID NO.7.

[0170] Optionally, the low-bias DNA polymerase is a polymerase that has high catalytic activity at a temperature of 50°C to 90°C, 60°C to 80°C, or approximately 68°C.

[0171] Barcodes, sample tags, and adapters

[0172] The method can further include introducing a barcode into the target nucleic acid molecule. For the purposes of the present invention, a barcode is a degenerate or randomly generated nucleotide sequence. The term "barcode" is synonymous with the term "unique molecular identifier" (UMI) or "unique molecular tag" (UMT). The method can include introducing one, two, or more barcodes into the target nucleic acid molecule. In a preferred embodiment, the method includes introducing a plurality of barcodes into the target nucleic acid molecule such that, after the introduction of the barcodes, most of the original target nucleic acid molecules contain unique barcodes compared to other original target nucleic acid molecules.

[0173] If the method for introducing mutations of the present invention is used as part of a method for determining a sequence, introducing a barcode into the target nucleic acid molecule can be useful. The use of barcodes can help the user identify from which original target nucleic acid molecule each sequence in at least one target nucleic acid molecule (or at least one amplified or fragmented target nucleic acid molecule) is derived. If the barcodes used in each original target nucleic acid molecule are different, the user can sequence the barcodes or the target nucleic acid molecules, and the sequences of the target nucleic acid molecules that contain the same barcode are likely to be the sequences of target nucleic acid molecules that originated from the same original target nucleic acid molecule.

[0174] A method for introducing mutations into at least one target nucleic acid molecule can include introducing a sample tag into the target nucleic acid molecule. A sample tag is a short series of nucleic acids of a known (specified) sequence. For example, the method of the present invention can be performed on a plurality of target nucleic acid molecules taken from different samples. Those samples can be combined, but prior to combination, a sample tag can be introduced into the target nucleic acid molecules in the sample (the target nucleic acid molecules are labeled with the sample tag). Target nucleic acid molecules from different samples can be labeled with different sample tags. Optionally, target nucleic acid molecules from the same sample are labeled with the same sample tag or sample tags from the same subgroup of sample tags. For example, if the user decides to use two samples, the target nucleic acid molecules in the first sample can be labeled with a first sample tag having a specified sequence, while the target nucleic acid molecules in the second sample can be labeled with a second sample tag having a second specified sequence. Similarly, if the user decides to use two samples, the target nucleic acid molecules in the first sample can be labeled with sample tags from a first subgroup of sample tags, while the target nucleic acid molecules in the second sample can be labeled with sample tags from a second subgroup of sample tags. The user will understand that any target nucleic acid molecule containing the first sample tag or sample tags from the first subgroup of sample tags is derived from the first sample, and any target nucleic acid molecule containing the second sample tag or sample tags from the second subgroup of sample tags is derived from the second sample. By sequencing the target nucleic acid sequence, it is possible to determine which tag has been used to label the target nucleic acid sequence. Suitable sequencing methods are described in more detail below.

[0175] In one embodiment, the sample tag is introduced (the target nucleic acid molecule is labeled with the sample tag) prior to the step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase. This is advantageous because it means that the samples can be combined at an early stage of the method, thus reducing the processing time, the amount of reagents required, and the likelihood of introducing sample handling errors. However, if the sample tag is introduced prior to the step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase, then it is possible that the sample tag will be mutated by the low-bias DNA polymerase. The inventors have designed groups of sample tags that are designed to be distinguishable from each other even if they have been mutated.

[0176] In one embodiment, a sample tag set is used, and different sample tags from the set are used to label target nucleic acid molecules from different samples. Target nucleic acid molecules from the same sample can be labeled with the same sample tag from the set or a sample tag from the same subgroup of sample tags from the set. For example, if the sample tag set contains sample tags named A, B, C, and D, then all target nucleic acid molecules in the first sample can be labeled with A or A / B, and all target nucleic acid molecules in the second sample can be labeled with C or C / D. Each sample tag in the sample tag set can differ from substantially all other sample tags in the set by at least 1 low-probability mutation difference. Each sample tag in the sample tag set can differ from all other sample tags in the set by at least 1 low-probability mutation difference.

[0177] In one aspect, the present invention provides a sample tag set, wherein each sample tag in the set differs from substantially all other sample tags in the set by at least 1 low-probability mutation difference. Each sample tag can differ from all other sample tags in the set by at least 1 low-probability mutation difference.

[0178] The term "differs from substantially all other sample tags in the set by at least 1 low-probability mutation difference" means that each tag has been designed such that if the sample tag mutates by at least 1 low-probability mutation, the tags will still be different from each other (from substantially all or all other tags). In one embodiment, the term "substantially all other sample tags" refers to at least 90%, at least 95%, or at least 98% of the other sample tags. A low-probability mutation is a mutation that occurs rarely in the method of introducing mutations of the present invention. For example, a low-probability mutation can be a transversion mutation or an indel mutation. When using dPTP as a nucleotide analogue in the method of introducing mutations of the present invention, transversion mutations and indel mutations rarely occur. A transversion mutation is the replacement of a purine nucleotide with a pyrimidine nucleotide (adenine to cytosine, adenine to thymine, guanine to cytosine, or guanine to thymine), or the replacement of a pyrimidine nucleotide with a purine nucleotide (cytosine to adenine, cytosine to guanine, thymine to adenine, thymine to adenine, thymine to guanine). An indel mutation is a deletion mutation or an insertion mutation. Appropriate tags can be designed using statistical methods by calculation. For example, a person skilled in the art will be able to determine which type of mutation is a low-probability mutation in the method of introducing mutations of the present invention. A person skilled in the art can perform the method of introducing mutations of the present invention and determine the type of mutation that has been introduced by sequencing the nucleic acid molecule product. The most frequently occurring mutations are high-probability mutations, and the least frequently occurring mutations are low-probability mutations.

[0179] A user can use the method of designing a sample tag set of the present invention to generate appropriate sample tags.

[0180] Optionally, each sample tag differs from substantially all of the other sample tags in the set by at least 2, at least 3, at least 4, at least 5, 3 to 50, 3 to 25, or 3 to 10 low-probability mutation differences. Optionally, each sample tag differs from all of the other sample tags in the set by at least 2, at least 3, at least 4, at least 5, 3 to 50, 3 to 25, or 3 to 10 low-probability mutation differences.

[0181] Each sample tag can differ from substantially all of the other sample tags in the set by at least 2 high-probability mutation differences. High-probability mutations are mutations that occur frequently in the method for introducing mutations of the present invention. For example, a high-probability mutation difference can be a transition mutation. A transition mutation is the replacement of a purine nucleotide with another purine nucleotide (adenine to guanine, or guanine to adenine), or the replacement of a pyrimidine nucleotide with another pyrimidine nucleotide (cytosine to thymine, or thymine to cytosine).

[0182] Each sample tag can differ from all of the other sample tags in the set by at least 2 high-probability mutation differences, i.e., each sample tag has been designed such that if the sample tag mutates by at least 2 high-probability mutations, the tags will still be different from each other.

[0183] Optionally, each sample tag differs from substantially all of the other sample tags in the set by at least 3, 2 to 50, 3 to 25, or 3 to 10 high-probability mutation differences. Optionally, each sample tag differs from all of the other sample tags in the set by at least 3, 2 to 50, 5 to 25, or 5 to 10 high-probability mutation differences.

[0184] In one embodiment, the length of each sample tag is at least 8 nucleotides, at least 10 nucleotides, at least 12 nucleotides, 8 to 50 nucleotides, 10 to 50 nucleotides, or 10 to 50 nucleotides.

[0185] Suitable sample tags are those of SEQ ID NO: 8 - 136.

[0186] The method can further include introducing an adaptor into each of the target nucleic acid molecules. The adaptor can contain a primer binding site. For the purposes of the present invention, a primer binding site is a known nucleotide sequence that is long enough for a primer to hybridize specifically. Optionally, the length of the primer binding site is at least 8, at least 10, at least 12, 8 to 50, or 10 to 25 nucleotides.

[0187] The method can include introducing a first adaptor at the 3' end of at least one target nucleic acid molecule and introducing a second adaptor at the 5' end of at least one target nucleic acid molecule, wherein the first adaptor and the second adaptor can anneal to each other.

[0188] In one aspect, the present invention provides a method for preferentially amplifying a target nucleic acid molecule having a length greater than 1 kbp, the method comprising:

[0189] a. providing at least one sample comprising a target nucleic acid molecule;

[0190] b. introducing a first adaptor at the 3' end of the target nucleic acid molecule and introducing a second adaptor at the 5' end of the target nucleic acid molecule; and

[0191] c. amplifying the target nucleic acid molecule using a primer complementary to a portion of the first adaptor,

[0192] wherein the first adaptor and the second adaptor can anneal to each other.

[0193] The second adaptor may comprise a portion complementary to the first primer binding site, and the first adaptor may comprise the first primer binding site.

[0194] The inventors have found that by introducing a first adaptor and a second adaptor that can anneal to each other into at least one target nucleic acid molecule, it can ensure that the method of the present invention preferentially amplifies long target nucleic acid molecules and / or mutates long target nucleic acid molecules. If the first adaptor can anneal to the second adaptor, it can do so in the method of the present invention, resulting in self-annealing of at least one target nucleic acid molecule (as Figure 5 shown). The self-annealed target nucleic acid molecule will not be replicated and thus will not be amplified and / or mutated by the method of the present invention. During the method of the present invention, the probability of the first adaptor and the second adaptor annealing to each other is higher for shorter target nucleic acid molecules than for longer target nucleic acid molecules. For these reasons, adding the first adaptor and the second adaptor to at least one target nucleic acid molecule of the present invention can be used to preferentially amplify a larger at least one target nucleic acid molecule.

[0195] The method for preferentially amplifying a nucleic acid molecule can be a method for preferentially amplifying a target nucleic acid molecule longer than 1.5 kbp. The method can further include a step of sequencing the target nucleic acid molecule. Examples of possible sequencing methods include Maxam Gilbert sequencing, Sanger sequencing, nanopore sequencing, or sequencing including bridge PCR. In a typical embodiment, the sequencing step involves bridge PCR. Optionally, the bridge PCR step is performed using an extension time greater than 5 seconds, greater than 10 seconds, greater than 15 seconds, or greater than 20 seconds. An example of using bridge PCR is in an Illumina genomic analysis sequencer.

[0196] A user can determine whether a first adaptor and a second adaptor can anneal to each other. In one embodiment, the user can identify whether the first adaptor and the second adaptor can anneal to each other by providing a nucleic acid molecule comprising the first adaptor and checking whether a primer comprising the second adaptor can initiate replication of the nucleic acid molecule under PCR conditions.

[0197] Optionally, in one embodiment, the first adaptor and the second adaptor can be considered to be able to anneal to each other if they hybridize under the following conditions: Combine two primers at equimolar concentration (e.g., 50 μM), then incubate at a high temperature (e.g., 95 °C) for 5 minutes to ensure that the primers are single-stranded. Then slowly cool the solution to room temperature (25 °C) over a period of about 45 minutes.

[0198] The method can include using primers that are identical to each other or substantially identical to each other to amplify a target nucleic acid molecule. The primers can be complementary to a portion of the first adaptor. Two primers are "substantially identical" to each other if they have the same sequence or sequences that differ by 1, 2, or 3 nucleotides. In a preferred embodiment, the method of the invention includes using primers that have the same sequence or differ by a single nucleotide difference to amplify a target nucleic acid molecule.

[0199] In one embodiment, the first adaptor and the second adaptor comprise sequences that are complementary to each other or substantially complementary to each other. The first adaptor can be substantially complementary to the second adaptor if the first adaptor is complementary to a nucleic acid molecule that is at least 80%, at least 90%, at least 95%, or at least 99% identical to the second adaptor.

[0200] A user can use primers that contain primer binding sites, and these primers can be used to preferentially amplify at least one replica of a target nucleic acid molecule generated in the last round of replication. For example, a first set of primers containing a third primer binding site can be used in one round of replication. In another round of replication, a second set of primers that bind to the third primer binding site can be used. The second set of primers will only replicate at least one replica of a target nucleic acid molecule generated using the first set of primers in the previous round of replication.

[0201] Third and other sets of primers can be used. Preferably, replicating the replicas of the previous round of replication is advantageous because this can ensure that each amplified target nucleic acid molecule contains a high level of mutations (since only at least one target nucleic acid molecule that has been exposed to at least one round of amplification by a low-bias DNA polymerase will be replicated).

[0202] Thus, the method of the invention can include:

[0203] (a) Introduce a first adaptor comprising a first primer binding site included in at least one target nucleic acid molecule or at the 3' end of the target nucleic acid molecule, and a second adaptor comprising a portion complementary to the first primer binding site at the 5' end of at least one target nucleic acid molecule or the target nucleic acid molecule, wherein the first adaptor and the second adaptor can anneal to each other;

[0204] (b) Optionally, use a low-bias DNA polymerase to amplify the target nucleic acid molecule with a first set of primers complementary to the first primer binding site and comprising a second primer binding site; and

[0205] (c) Optionally, use a low-bias DNA polymerase to amplify the target nucleic acid molecule with a second set of primers complementary to the second primer binding site.

[0206] The second set of primers can comprise a third primer binding site, and a third set of primers or another set of primers complementary to the third primer binding site or an additional primer binding site can be used for further amplification steps.

[0207] Any suitable method can be used to introduce barcodes, sample tags, and / or adaptors, including PCR of the target nucleic acid, tagmentation, and physical shearing or restriction digestion, followed by ligation of the adaptor (optionally, sticky-end ligation). For example, a first set of primers capable of hybridizing to at least one target nucleic acid molecule can be used to perform PCR on at least one target template nucleic acid molecule. Primers can be used to introduce barcodes, sample tags, and adaptors into each of at least one target nucleic acid molecule by PCR, the primer comprising a portion (5' end portion) comprising the barcode, sample tag, and / or adaptor, and a portion (3' end portion) having a sequence capable of hybridizing (optionally complementary) to at least one target nucleic acid molecule. Such a primer will hybridize to the target nucleic acid molecule, and then PCR primer extension will provide a nucleic acid molecule comprising the barcode, sample tag, and / or adaptor. Further cycles of PCR with these primers can be used to add barcodes, sample tags, and / or adaptors to the other end of at least one target nucleic acid molecule. The primers can be degenerate, i.e., the 3' end portions of the primers can be similar but not identical to each other.

[0208] Barcodes, sample labels, and / or adaptors can be introduced using tag fragmentation. Barcodes, sample labels, and / or adaptors can be introduced using direct tag fragmentation, or a defined sequence can be introduced by tag fragmentation, followed by two PCR cycles using a primer that includes a portion capable of hybridizing to the defined sequence and includes a portion containing the barcode, sample label, and / or adaptor. Barcodes, sample labels, and / or adaptors can be introduced by performing a restriction digestion on at least one original target nucleic acid molecule, followed by ligating a nucleic acid containing the barcode, sample label, and / or adaptor. The restriction digestion of the at least one original target nucleic acid molecule should be performed such that the digestion produces a nucleic acid molecule containing the region to be sequenced (the at least one target template nucleic acid molecule). Barcodes, sample labels, and / or adaptors can be introduced by shearing at least one target nucleic acid molecule, followed by end repair, A-tailing, and then ligating a nucleic acid containing the barcode, sample label, and / or adaptor.

[0209] Method for determining the sequence of at least one target nucleic acid molecule

[0210] One aspect of the present invention relates to a method for determining the sequence of at least one target nucleic acid molecule, the method comprising the method of the present invention for introducing mutations.

[0211] As described above, the method of the present invention for introducing mutations can be useful as part of a method for determining the sequence of at least one target nucleic acid molecule, since the mutations can enable a technician to assemble the sequence.

[0212] As described in the background section, sequencing methods can be improved by an incorporation step that introduces mutations into at least one target nucleic acid molecule to be sequenced. A user typically amplifies at least one target nucleic acid molecule and / or fragments at least one target nucleic acid molecule prior to sequencing. Then, the user assembles the consensus sequence of at least one target nucleic acid molecule from the regional sequences of the amplified or fragmented at least one target nucleic acid molecule. Introducing mutations into at least one target nucleic acid molecule prior to amplification or fragmentation can help the user identify from which of the original at least one template nucleic acid molecules each regional sequence of the amplified or fragmented at least one target nucleic acid molecule is derived, thereby improving the accuracy of the consensus sequence.

[0213] The more random the introduced mutations are, the easier it is to identify from which of the original at least one target nucleic acid molecules each sequence of the amplified or fragmented at least one target nucleic acid molecule is derived. The method of introducing mutations of the present invention using a low-bias DNA polymerase can be used to introduce mutations in a substantially random manner and is thus ideal for incorporation into a method for determining the sequence of at least one target nucleic acid molecule.

[0214] A method for determining the sequence of at least one target nucleic acid molecule can include the steps:

[0215] a. Perform the method of the present invention for introducing mutations into at least one target nucleic acid molecule to provide at least one mutated target nucleic acid molecule;

[0216] b. Sequence a region of at least one mutated target nucleic acid molecule to provide mutated sequence reads; and

[0217] c. Use the mutated sequence reads to assemble the sequence of at least a portion of at least one target nucleic acid molecule.

[0218] Generally, the sequencing step can be performed using any sequencing method. Examples of possible sequencing methods include Maxam Gilbert sequencing, Sanger sequencing, nanopore sequencing, or sequencing including bridge PCR. In a typical embodiment, the sequencing step involves bridge PCR. Optionally, the bridge PCR step is performed using an extension time greater than 5 seconds, greater than 10 seconds, greater than 15 seconds, or greater than 20 seconds. An example of using bridge PCR is in the Illumina Genome Analyzer sequencer.

[0219] The method can include sequencing a region of at least one mutated target nucleic acid molecule to provide mutated sequence reads. The region can correspond to a fragment that may comprise most of the at least one mutated target nucleic acid molecule. It may not be possible to sequence the entire at least one mutated target nucleic acid molecule for some reasons, but the user can still find the sequence of a portion of the at least one mutated target nucleic acid molecule useful. The region of the at least one mutated target nucleic acid molecule can include the full-length at least one mutated target nucleic acid molecule.

[0220] The method can include assembling the sequence of at least a portion of at least one target nucleic acid molecule from the mutated sequence reads. The sequence can be assembled by aligning the mutated sequence reads and grouping the reads sharing the same mutation pattern together. The sequence will be assembled from the mutated sequence reads in the same group. Software such as Clustal W2, IDBA-UD, or SOAPdenovo can be used for assembly.

[0221] A method for determining the sequence of at least one target nucleic acid molecule can include the steps:

[0222] a. Perform the method of the present invention for introducing mutations into at least one target nucleic acid molecule to provide at least one mutated target nucleic acid molecule;

[0223] b. Fragment at least one mutated target nucleic acid molecule and / or amplify at least one mutated target nucleic acid molecule to provide at least one fragmented and / or amplified mutated target nucleic acid molecule;

[0224] c. Sequencing a region of at least one fragmented and / or amplified mutant target nucleic acid molecule to provide sequence reads of the mutation; and

[0225] d. Using the sequence reads of the mutation to assemble the sequence of at least a portion of at least one target nucleic acid molecule.

[0226] The step of amplifying at least one mutant target nucleic acid molecule can be carried out by any suitable amplification technique, such as PCR. Suitably, PCR is carried out using a low-bias DNA polymerase under the conditions described, for example, under the heading "Amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase".

[0227] Any suitable method can be used to carry out the step of fragmenting at least one mutant target nucleic acid molecule. For example, fragmentation can be carried out using restriction digestion or by PCR using primers complementary to at least one internal region of at least one mutant target nucleic acid molecule. Preferably, fragmentation is carried out using a technique that generates random fragments. The term "random fragment" refers to a fragment generated randomly, such as a fragment generated by tagmentation. Fragments generated using a restriction endonuclease are not "random" because the restriction digestion occurs at specific DNA sequences defined by the restriction endonuclease used. Even more preferably, fragmentation is carried out by tagmentation. If fragmentation is carried out by tagmentation, the tagmentation reaction optionally introduces an adapter region into at least one mutant target nucleic acid molecule. The adapter region is a short DNA sequence that can encode, for example, an adapter to allow sequencing of at least one mutant target nucleic acid molecule using Illumina technology.

[0228] The fragmentation step can include a further step of enriching at least one fragmented mutant target nucleic acid molecule. The step of enriching at least one fragmented mutant target nucleic acid molecule can be carried out by PCR. Suitably, PCR is carried out using a low-bias DNA polymerase under the conditions described, for example, under the heading "Amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase".

[0229] Methods for protein engineering

[0230] The method for introducing mutations of the present invention can be useful as part of a method for protein engineering. For example, protein engineering may involve finding mutations that increase or decrease protein activity or alter its structure. As part of protein engineering, the user may wish to randomly mutate a protein and see how the mutations affect the protein's activity or structure. This method is a method that results in highly random mutagenesis and can therefore be advantageously used as part of a method for protein engineering.

[0231] Accordingly, in one aspect of the present invention, there is provided a method for engineering a protein, the method comprising the method for introducing mutations of the present invention.

[0232] The method may comprise the steps of:

[0233] a. Performing the method for introducing mutations of the present invention to provide at least one mutated target nucleic acid molecule;

[0234] b. Inserting the at least one mutated target nucleic acid molecule into a vector; and

[0235] c. Expressing the protein encoded by the at least one mutated target nucleic acid molecule.

[0236] The method may comprise the steps of:

[0237] a. Performing the method for introducing mutations of the present invention to provide at least one mutated target nucleic acid molecule;

[0238] b. Amplifying the at least one target nucleic acid molecule in the presence of a nucleotide analogue using a low-bias DNA polymerase to provide a target nucleic acid molecule containing the nucleotide analogue;

[0239] c. Amplifying the target nucleic acid molecule containing the nucleotide analogue in the absence of the nucleotide analogue to provide at least one mutated target nucleic acid molecule;

[0240] d. Inserting the at least one mutated target nucleic acid molecule into a vector; and

[0241] e. Expressing the protein encoded by the at least one mutated target nucleic acid molecule.

[0242] Any suitable buffer may be used. Optionally, the vector is a plasmid, virus, cosmid or artificial chromosome. Generally, the vector further comprises control sequences operably linked to the inserted sequence, thereby allowing the expression of the polypeptide. Preferably, the vectors of the present invention further comprise appropriate initiators, promoters, enhancers and other elements that may be necessary and are positioned in the correct orientation to allow the expression of the polypeptide.

[0243] Optionally, the step of expressing the at least one mutated target nucleic acid molecule is achieved by transforming bacterial cells with the vector, transfecting eukaryotic cells or transducing eukaryotic cells. Optionally, the bacterial cells are Escherichia coli (E. coli) cells.

[0244] For example, the step of expressing at least one mutant target nucleic acid molecule can include inserting the at least one mutant target nucleic acid molecule into a plasmid vector and transforming Escherichia coli with the plasmid. The plasmid can contain control elements suitable for expression in Escherichia coli, such as the lac or T7 promoter (Dubendorff JW, Studier FW (1991). "Controlling basal expression in an inducible T7 expression system by blocking the target T7 promoter with lac repressor". Journal of Molecular Biology. 219(1):45–59.). Suitable expression techniques are described in Sambrook, J. et al., (1989) Molecular Cloning: A Laboratory Manual Second Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York.

[0245] Alternatively, the step of expressing at least one mutant target nucleic acid molecule can include expressing a fragment directly generated by the step of amplifying the target nucleic acid molecule using an in vitro method.

[0246] The method can further include the step of testing the activity of the protein encoded by the at least one mutant target nucleic acid molecule or evaluating the structure of the protein encoded by the at least one mutant target nucleic acid molecule.

[0247] The step of testing the activity of the protein encoded by the at least one mutant target nucleic acid molecule or evaluating the structure of the protein encoded by the at least one mutant target nucleic acid molecule can be carried out using many known techniques. For example, a person skilled in the art will know suitable techniques for evaluating protein structure, including Nuclear magnetic resonance (NMR) techniques, microscopy techniques (such as cryo-electron microscopy), small-angle X-ray scattering techniques, or X-ray crystallography.

[0248] Similarly, a person skilled in the art will know techniques available for assessing the activity of a protein. The method used will depend on the protein encoded by the at least one mutant target nucleic acid molecule. For example, if the protein encoded by the at least one mutant target nucleic acid molecule is a clotting factor, the person skilled in the art will, for example, use a chromogenic clotting assay to test the clotting activity of the protein. Alternatively, if the protein encoded by the at least one mutant target nucleic acid molecule is an enzyme, the person skilled in the art can test the activity of the enzyme by measuring the rate at which the enzyme catalyzes its reaction, for example, by measuring the decrease in the concentration of the starting product or the increase in the concentration of the end product of the reaction catalyzed by the enzyme.

[0249] Method for designing a sample tag set

[0250] In one aspect, the present invention further provides a method for designing a sample tag set suitable for a method of introducing a mutation into at least one target nucleic acid molecule, comprising:

[0251] a. Analyzing a method for introducing a mutation into at least one target nucleic acid molecule and determining the average number of low-probability mutations that occur during the method for introducing a mutation into at least one target nucleic acid molecule; and

[0252] b. Determining the sequence of the sample tag set, wherein each sample tag differs by a greater number of low-probability mutation differences from substantially all of the sample tags in the set compared to the average number of low-probability mutations that occur during the method for introducing a mutation into at least one target nucleic acid molecule.

[0253] For example, a user can generate a first putative sample tag by using a computer program to generate a random sequence. The first putative sample tag is added to the sample tag set. Then, the user can generate a second putative sample tag in the same manner and compare the sequence of the second putative sample tag with the first putative sample tag to see if the second sample tag is different from the first sample tag, such that even if a relevant number of low-probability mutations are introduced into the second putative sample tag, it will still be different from the first putative sample tag. If so, the second putative sample tag is added to the sample tag set. If not, the second putative sample tag is discarded. This can be repeated for a third putative sample tag and other putative sample tags.

[0254] As described above, in a method for introducing a mutation into at least one target nucleic acid molecule, it is advantageous to add a sample tag to the at least one target nucleic acid molecule. However, if the sample tag is added before the mutation is introduced, this may mean that the sample tag has mutated and can then no longer be used to distinguish target nucleic acid molecules derived from the same or different samples. This can be avoided by designing the sample tags such that even if the sample tags mutate, they can still be sufficiently different from each other for the user to distinguish the sample tags.

[0255] The method may further comprise:

[0256] a. (i) Analyzing a method for introducing mutations into at least one target nucleic acid molecule and determining the average number of high-probability mutations that occur during the method for introducing mutations into at least one target

[0257] nucleic acid molecule; and

[0258] (ii) Determining the sequences of a set of sample tags, wherein each sample tag differs by a greater number of high-probability mutation differences from substantially all of the sample tags in the set as compared to the average number of high-probability mutations that occur during the method for introducing mutations into at least one target nucleic acid molecule.

[0259] Low-probability mutations can be transversion mutations or indel mutations. High-probability mutation differences can be transition mutations.

[0260] The method can be a computer-implemented method.

[0261] In another aspect of the invention, there is provided a computer-readable medium configured to execute a method for designing a set of sample tags suitable for a method of introducing mutations into at least one target nucleic acid molecule.

[0262] In another aspect of the invention, there is provided a set of sample tags obtainable by the method for designing the set of sample tags of the invention. Optionally, the set of sample tags is obtained by the method for designing sample tags of the invention.

[0263] Using unequal concentrations of dNTPs

[0264] The step of amplifying at least one target nucleic acid using a low-bias DNA polymerase can be carried out using unequal concentrations of dNTPs.

[0265] In one aspect of the invention, there is provided a method for introducing mutations into at least one target nucleic acid molecule, comprising:

[0266] a. Providing at least one sample comprising at least one target nucleic acid molecule; and

[0267] b. Introducing mutations into at least one target nucleic acid molecule by amplifying at least one target nucleic acid molecule using a DNA polymerase to provide at least one mutated target nucleic acid molecule,

[0268] wherein step b. is carried out using unequal concentrations of dNTPs.

[0269] To be able to use a DNA polymerase (such as a low-bias DNA polymerase) to amplify at least one target nucleic acid, the target nucleic acid can be exposed to the DNA polymerase and dNTPs under conditions suitable for DNA replication to occur, such as in a PCR machine. If unequal concentrations of dNTPs are used for the step of amplifying at least one target nucleic acid, the target nucleic acid is exposed to the DNA polymerase (such as a low-bias DNA polymerase) and dNTPs, wherein the concentrations of the dNTPs are different from each other.

[0270] The term dNTP refers to deoxynucleotides. Specifically, in the context of the present application, the term "dNTP" refers to a solution containing dTTP (deoxythymidine triphosphate) or dUTP (deoxyuridine), dGTP (deoxyguanosine triphosphate), dCTP (deoxycytidine triphosphate), and dATP (deoxyadenosine). Optionally, "dNTP" refers to a solution containing dTTP (deoxythymidine triphosphate), dGTP (deoxyguanosine triphosphate), dCTP (deoxycytidine triphosphate), and dATP (deoxyadenosine triphosphate).

[0271] The phrase "unequal concentrations of dNTPs" means that the concentrations of the four dNTPs relative to each other in the solution are different. For example, one dNTP can be present at a higher concentration than (relative to) the other three dNTPs, two dNTPs can be present at a higher concentration than (relative to) the other two dNTPs, or three dNTPs can be present at a higher concentration than (relative to) the other dNTP.

[0272] DGTP can be present at a higher concentration than (relative to) dCTP, dTTP, and dATP, dGTP can be present at a higher concentration than (relative to) dTTP and dATP, dGTP can be present at a higher concentration than (relative to) dATP, dGTP can be present at a higher concentration than (relative to) dTTP, dCTP can be present at a higher concentration than (relative to) dGTP, dTTP, and dATP, dCTP can be present at a higher concentration than (relative to) dTTP and dATP, dCTP can be present at a higher concentration than (relative to) dATP, dCTP can be present at a higher concentration than (relative to) dTTP, dTTP can be present at a higher concentration than (relative to) dGTP, dCTP, and dATP, dTTP can be present at a higher concentration than (relative to) dGTP and dCTP, dTTP can be present at a higher concentration than (relative to) dCTP, dTTP can be present at a higher concentration than (relative to) dGTP, dATP can be present at a higher concentration than (relative to) dGTP, dTTP, and dCTP, dATP can be present at a higher concentration than (relative to) dGTP and dCTP, dATP can be present at a higher concentration than (relative to) dGTP, dATP can be present at a higher concentration than (relative to) dGTP, dCTP and dATP can be present at a higher concentration than (relative to) dGTP and dCTP, or dGTP and dCTP can be present at a higher concentration than (relative to) dATP and dTTP.

[0273] The user can prepare a solution of dNTPs at unequal concentrations in any convenient manner. Solutions of DATP, dTTP, dGTP, and dTTP are readily available on the market, and the user only needs to mix them in appropriate proportions.

[0274] Optionally, the method:

[0275] (i) includes a further step of amplifying at least one target nucleic acid molecule containing a nucleotide analogue in the absence of the nucleotide analogue, and the further step of amplifying at least one target nucleic acid molecule containing a nucleotide analogue in the absence of the nucleotide analogue is carried out using unequal concentrations of dNTP; or

[0276] (ii) provides at least one mutated target nucleic acid molecule, and includes a further step of amplifying the at least one mutated target nucleic acid molecule using a low-bias DNA polymerase, and the further step of amplifying the at least one mutated target nucleic acid molecule using a low-bias DNA polymerase is carried out using unequal concentrations of dNTP.

[0277] Optionally, introducing mutations into at least one target nucleic acid molecule by amplifying at least one target nucleic acid molecule using a DNA polymerase is carried out in the presence of a nucleotide analogue. Optionally, the method for introducing mutations into at least one target nucleic acid molecule includes the step of amplifying the at least one target nucleic acid molecule with mutations in the absence of a nucleotide analogue, and optionally, this step is carried out using unequal concentrations of dNTPs.

[0278] When introducing mutations into at least one target nucleic acid molecule using a nucleotide analogue, this generally involves two amplification steps. In the first amplification step, the nucleotide analogue is incorporated into the target nucleic acid molecule (mutation step). In the second amplification step, the nucleotide analogue pairs with a natural nucleotide, thereby introducing a mutation into one strand of the target nucleic acid molecule (recovery step). When the target nucleic acid molecule is further amplified, this mutation will be passed on to both strands of the target nucleic acid molecule. Optionally, both the first (mutation) amplification step and the second (recovery) amplification step can be carried out using unequal concentrations of dNTPs. Optionally, the unequal concentrations of dNTPs are different in the first (mutation) amplification step and the second (recovery) amplification step. For example, in the first (mutation) amplification step, the unequal concentrations of dNTPs can include dTTP at a concentration lower than the concentrations of the other dNTPs, while in the second (recovery) amplification step, the unequal concentrations of dNTPs can include dATP at a concentration lower than the concentrations of the other dNTPs. The step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase or the step of providing the at least one target nucleic acid molecule with mutations can correspond to one or more "mutation steps". The further step of amplifying at least one target nucleic acid molecule containing a nucleotide analogue in the absence of a nucleotide analogue, or the further step of amplifying the at least one target nucleic acid molecule with mutations, can correspond to one or more "recovery steps".

[0279] Optionally, the nucleotide analogue is dPTP.

[0280] In one embodiment, unequal concentrations of dNTPs are used to alter the characteristics of the introduced mutations. Unequal concentrations of dNTPs are used in a method including introducing mutations into at least one target nucleic acid molecule. Thus, the method results in a target nucleic acid molecule containing mutations (such as the mutated target nucleic acid molecule described herein). The number of mutations, the type of mutations, and the position of each mutation introduced into a given target nucleic acid molecule by this method can be referred to as the "mutation characteristics" introduced. The term "type of mutation" refers to the nature of the mutation, i.e., whether it is a substitution mutation, an addition mutation, or a deletion mutation, and if it is a substitution mutation, what is the starting nucleotide and what does the starting nucleotide mutate to (e.g., an A-to-G mutation, with an A starting nucleotide that mutates to G)?

[0281] A user can determine the "mutation profile" introduced by a given method by copying a test target nucleic acid molecule and then subjecting some of the copies to the method of the present invention that involves introducing mutations, while retaining some of the copies (without mutating them). Then, the user can sequence the copies that have been subjected to the method of the present invention that involves introducing mutations and the retained copies. Finally, the user can align the sequences of the copies that have been subjected to the method of the present invention that involves introducing mutations and the sequences of the retained copies to determine the number of mutations introduced, the type of mutations, and the location of each mutation. Alternatively, the user can use a test target nucleic acid molecule with a known sequence. Then, the user only needs to subject the test target nucleic acid molecule to the method of the present invention that involves introducing mutations and then sequence the resulting mutated target nucleic acid molecule to see what mutation profile has been introduced.

[0282] A user may wish to alter the mutation profile in a variety of ways. For example, as described above, it is advantageous to be able to reduce mutation bias. Thus, in one embodiment, unequal concentrations of dNTPs are used to reduce the bias in the introduced mutation profile. In another embodiment, the method is a method for introducing mutations with a low-bias mutation profile.

[0283] This application demonstrates that the use of unequal concentrations of dNTPs can be used to reduce the bias in the introduced mutation profile. For example, if a DNA polymerase (such as the low-bias DNA polymerase described above) is used to mutate a target nucleic acid molecule and a higher number of G-to-A mutations are introduced compared to other mutations, the user can lower the concentration of dATP relative to the other dNTPs, and this can reduce the frequency of incorporation of an A nucleotide in place of dGTP, thereby reducing the number of G-to-A mutations.

[0284] Similarly, if nucleotide analogs are used when introducing mutations into a target nucleic acid molecule, the relative concentrations of dNTPs can be varied to alter the mutagenic properties. For example, dPTP can be used to introduce G-to-A, C-to-T, A-to-G, and T-to-C mutations. As described in more detail above, dPTP can substitute for a T nucleotide or a C nucleotide, and depending on whether dPTP is in its amino or imino form, it can then pair with an A nucleotide or a G nucleotide. This results in two scenarios. In the first scenario, dPTP substitutes for T (mutation step) in, for example, the sense strand, and it can then pair with A (no mutation) or G (A-to-G mutation) in the antisense strand. If dPTP substitutes for T and pairs with G in the antisense strand, the mutant G will pair with C to introduce a T-to-C mutation in the replica of the sense strand (recovery step). Conversely, dPTP can substitute for T in the antisense strand, which can result in an A-to-G mutation in the sense strand and a T-to-C mutation in the replica of the antisense strand. In the second scenario, dPTP substitutes for C (e.g., in the sense strand), and it can then pair with A (G-to-A mutation) or G (no mutation) in the antisense strand (mutation step). If dPTP substitutes for C and pairs with A in the antisense strand, the mutant A will pair with T to introduce a C-to-T mutation in the replica of the sense strand (recovery step). Conversely, dPTP can substitute for C in the antisense strand, which can result in a G-to-A mutation in the sense strand and a C-to-T mutation in the replica of the antisense strand.

[0285] This application demonstrates that if the ratio of G-to-A and C-to-T mutations is higher than the ratio of A-to-G and T-to-C mutations, decreasing the concentration of dTTP relative to other dNTPs (and preferably relative to the concentration of dCTP) will encourage the incorporation of dPTP in place of dTTP, increasing the instances of the first scenario relative to the second scenario, which means that the A-to-G and T-to-C mutations introduced in the first scenario will be increased. Similarly, this application demonstrates that if the level of dATP is decreased during the recovery step, the levels of G-to-A and C-to-T mutations increase. This is because in scenario 2 above, if dATP is present at a lower concentration compared to other dNTPs (and preferably relative to the concentration of dGTP), this means that dPTP incorporated in place of a C nucleotide will pair with G more frequently and fewer G-to-A or C-to-T mutations will be introduced. These two scenarios are as Figure 7 shown.

[0286] Even the low-bias DNA polymerases disclosed herein introduce mutations into the target nucleic acid molecule with a small bias. This application demonstrates that using unequal concentrations of dNTPs and a low-bias DNA polymerase can actually eliminate any mutational bias.

[0287] Based on the information provided in the present application, a person skilled in the art is capable of determining how changing the concentrations of various dNTPs will affect the mutagenic properties depending on whether a nucleotide analogue is used and, if so, which nucleotide analogue is used. Thus, in some embodiments, the method of using unequal concentrations of dNTPs includes the step of identifying the dNTPs whose levels should be increased or decreased to reduce the bias in the introduced mutagenic properties.

[0288] Optionally, the unequal concentrations of dNTPs include dTTP at a concentration lower than the concentrations of the other dNTPs. As described above, this can increase the rates of T-to-C and A-to-G mutations introduced when dPTP is used as a nucleotide analogue. Optionally, the unequal concentrations of dNTPs include dTTP at a concentration less than 75%, less than 70%, less than 60%, less than 55%, 25% to 75%, 25% to 70%, 25% to 60%, or approximately 50% of the concentration of dATP, dCTP, or dGTP. Optionally, the unequal concentrations of dNTPs include dTTP at a concentration less than 60% of the concentration of dCTP. Optionally, the unequal concentrations of dNTPs include dTTP at a concentration of 25% to 60% of the concentration of dCTP.

[0289] Optionally, the unequal concentrations of dNTPs include dATP at a concentration lower than the concentrations of the other dNTPs. As described above, this can increase the rates of G-to-A and C-to-T mutations introduced when dPTP is used as a nucleotide analogue. Optionally, the unequal concentrations of dNTPs include dATP at a concentration less than 75%, less than 70%, less than 60%, less than 55%, 25% to 75%, 25% to 70%, 25% to 60%, or approximately 50% of the concentration of dTTP, dCTP, or dGTP. Optionally, the unequal concentrations of dNTPs include dATP at a concentration less than 75%, less than 70%, less than 60%, less than 55%, 25% to 75%, 25% to 70%, 25% to 60%, or approximately 50% of the concentration of dGTP. Optionally, the unequal concentrations of dNTPs include dATP at a concentration less than 60% of the concentration of dGTP. Optionally, the unequal concentrations of dNTPs include dATP at a concentration of 25% to 60% of the concentration of dGTP.

[0290] As described in the above two cases, when using dPTP as a nucleotide analogue, reducing dTTPs increases T-to-C and A-to-G mutations by encouraging replacement of T nucleotides in the target nucleic acid molecule with dPTP. Thus, it is preferred to use such unequal concentrations of dNTPs in the mutagenesis step (e.g., the step of performing PCR in the presence of dPTP): dTTP having a concentration lower than the concentrations of the other dNTPs. Similarly, when using dPTP as a nucleotide analogue, reducing dATP reduces the number of dPTPs that replace C nucleotides and pair with dATP, thereby increasing G-to-A and C-to-T mutations. Since dPTP pairing with dATP tends to occur during the recovery step, reducing dATP during the recovery step increases the number of G-to-A and C-to-T mutations. Optionally, therefore, the step of amplifying at least one target nucleic acid molecule comprising a nucleotide analogue or at least one mutated target nucleic acid molecule in the absence of a nucleotide analogue is performed using unequal concentrations of dNTPs, and the unequal concentrations of dNTPs comprise dATP having a concentration lower than the concentrations of the other dNTPs.

[0291] Example

[0292] Example 1 - Mutating Nucleic Acid Molecules Using PrimeStar GXL Other Polymerase

[0293] Fragment the DNA molecule to a suitable size (e.g., 10 kb) using tag fragmentation and attach defined sequence primer sites (adapters) to each end.

[0294] The first step is a tag fragmentation reaction to fragment the DNA. Under the following conditions, perform tag fragmentation on 50 ng of high molecular weight genomic DNA in one or more bacterial strains in a volume of 4 μl or less. Mix 50 ng of DNA with 4 μl of Nextera transposase (diluted 1:50) and 8 μl of 2X tag fragmentation buffer (20 mM Tris [pH 7.6], 20 mM MgCl, 20% (v / v) dimethylformamide) for a total volume of 16 μl. Incubate the reaction at 55 °C for 5 minutes, add 4 μl of NT buffer (or 0.2% SDS) to the reaction, and incubate the reaction at room temperature for 5 minutes.

[0295] Clean the tag fragmentation reaction using SPRIselect beads (Beckman Coulter) according to the manufacturer's instructions, perform left size selection using 0.6 volume of beads, and elute the DNA in molecular grade water.

[0296] Subsequently, PCR was performed with standard dNTPs and dPTPs in 6 limited cycles. Using Primestar GXL, 12.5 ng of the tagged fragmented and purified DNA was added to a total reaction volume of 25 μl, which contained 1× GXL buffer, 200 μM each of dATP, dTTP, dGTP, and dCTP, 0.5 mM dPTP, and 0.4 μM of custom primers (Table 2).

[0297] Table 2:

[0298]

[0299] Table 2. Custom primers were used for mutagenic PCR of a 10 kbp template. XXXXXX is a defined, sample-specific 6-8 nt barcode sequence. NNNNNN is a 6 nt random nucleotide region.

[0300] The reaction was carried out in the presence of Primestar GXL with the following thermal cycling. The initial gap extension was carried out at 68 °C for 3 minutes, followed by 6 cycles of 10 seconds at 98 °C, 15 seconds at 55 °C, and 10 minutes at 68 °C.

[0301] The next stage was PCR without dPTP to remove dPTP from the template and replace it with transition mutations ("recovery PCR"). The PCR reaction was cleaned with SPRIselect beads to remove excess dPTP and primers, and then another 10 rounds (minimum 1 round, maximum 20 rounds) of amplification were carried out using primers annealing to the fragment ends introduced during the dPTP incorporation cycles (Table 3).

[0302] Table 3

[0303] i7 Flowcell primer CAAGCAGAAGACGGCATACGA i5 Flowcell primer AATGATACGGCGACCACCGA

[0304] Subsequently, a gel extraction step was carried out to size-select the amplified and mutated fragments in the desired size range, e.g., 7 kb - 10 kb. Gel extraction can be done manually or by an automated system (e.g., BluePippin). Subsequently, another round of PCR, 16 - 20 cycles ("enrichment PCR"), was carried out.

[0305] After amplifying a defined number of long mutant templates, the templates were randomly fragmented to generate a set of overlapping shorter fragments for sequencing. Fragmentation was carried out by tag fragmentation.

[0306] Subject the long DNA fragments from the previous step to a standard tag fragmentation reaction (e.g., Nextera XT or Nextera Flex), with the difference that the reaction is divided into three pools for PCR amplification. This enables selective amplification of fragments derived from each end of the original template (including sample barcodes), as well as internal fragments from long templates that have been re-tag fragmented at both ends. This effectively creates three pools for sequencing on an Illumina instrument (e.g., MiSeq or HiSeq).

[0307] Repeat this method using standard Taq (Jena Biosciences) and a mixture of Taq and the proofreading polymerase (DeepVent), called LongAmp (New England Biolabs).

[0308] The data obtained from this experiment are as Figure 1 depicted. dPTP was not used as a control. The reads were mapped relative to the E. coli genome, resulting in a median mutation rate of ~8%.

[0309] Example 2 - Comparison of Mutation Frequencies for Different DNA Polymerases

[0310] Perform mutagenesis on a series of different DNA polymerases (Table 4). As described in the method of Example 1, genomic DNA from E. coli strain MG1655 is tag fragmented to generate long fragments and the beads are washed. Then, 6 cycles of "mutagenesis PCR" are performed in the presence of 0.5 mM dPTP, SPRIselect bead purification is carried out, and an additional 14 - 16 cycles of "recovery PCR" are performed in the absence of dPTP. The resulting long mutant templates are then subjected to a standard tag fragmentation reaction (see Example 1), and the "internal" fragments are amplified and sequenced on an Illumina MiSeq instrument.

[0311] The mutation rates are described in Table 4, where the frequencies of base substitutions are normalized by the dPTP mutagenesis reaction, as measured by Illumina sequencing of DNA from a known reference genome. For Taq polymerase, even when used in a buffer optimized for Thermococcus polymerase, only ~12% of the mutations occur at template G + C sites. Thermococcus-like polymerases produce 58% - 69% mutations at template G + C sites, while polymerases derived from Pyrococcus produce 88% mutations at template G + C sites.

[0312] The enzymes were obtained from Jena Biosciences (Taq), Takara (Primestar variant), Merck Millipore (KOD DNA polymerase), and New England Biolabs (Phusion).

[0313] Taq was tested with the supplied buffer and also with Primestar GXL buffer (Takara) for this experiment. All other reactions were carried out using the standard buffer supplied for each polymerase.

[0314] Table 4

[0315]

[0316] Example 3 - Determination of the dPTP mutagenesis rate

[0317] We performed dPTP mutagenesis on a series of genomic DNA samples with different G + C content (33% - 66%) levels using the Thermococcus polymerase (Primestar GXL; Takara) under a single set of reaction conditions. Mutagenesis and sequencing were carried out as described in the method of Example 3, except that 10 cycles of "rescue PCR" were performed. As predicted, despite the diversity in G + C content, the mutation rates were roughly similar between samples (median rate of 7% - 8%) ( Figure 2 ).

[0318] Example 4 - Measurement of template amplification bias

[0319] The template amplification bias of two polymerases was measured: Kapa HiFi, which is a proofreading polymerase commonly used in Illumina sequencing protocols; and PrimeStar GXL, which is a KOD family polymerase known for its ability to amplify long fragments. In the first experiment, Kapa HiFi was used to amplify a limited number of Escherichia coli genomic DNA templates of approximately 2 kbp in size. The ends of these amplified fragments were then sequenced. A similar experiment was performed on fragments of approximately 7 kbp - 10 kbp from Escherichia coli using PrimeStar GXL. The position of each end sequence read was determined by mapping relative to the Escherichia coli reference genome. The distances between adjacent fragment ends were measured. These distances were compared to a set of distances randomly sampled from a uniform distribution. D was compared by the non-parametric Kolmolgorov-Smirnov test. When the two samples are from the same distribution, the value of D approaches zero. For the low-bias PrimeStar polymerase, we observed D = 0.07 when measuring 50,000 fragment ends compared to a uniform random sample of 50,000 genomic positions. For the Kapa HiFi polymerase, we observed D = 0.14 at 50,000 fragment ends.

[0320] Example 5 - Preferential Amplification of Longer Templates Using Two Identical Primer Binding Sites and a Single Primer Sequence

[0321] As described above, tag fragmentation can be used to fragment DNA molecules and simultaneously introduce primer binding sites (adaptors) to the ends of the fragments. The Nextera tag fragmentation system (Illumina) utilizes a transposase loaded with one of two unique adaptors (herein referred to as X and Y). This generates a random mixture of products, some of which have the same end sequences (X-X, Y-Y), while others have unique ends (X-Y). The standard Nextera protocol uses two different primer sequences to selectively amplify the "X-Y" products that contain different adaptors on each end (required for sequencing using Illumina technology). However, it is also possible to use a single primer sequence to amplify the "X-X" or "Y-Y" fragments that have the same end adaptors.

[0322] To generate long mutant templates containing the same end adaptors, 50 ng of high molecular weight genomic DNA (Escherichia coli strain MG1655) was first tag fragmented as described in Example 1 and then cleaned with SPRIselect beads. Subsequently, 5 cycles of "mutagenic PCR" were performed as detailed in Example 1, combining standard dNTPs and dPTPs, with the difference that a single primer sequence (Table 5) was used.

[0323] The PCR reaction was cleaned with SPRIselect beads to remove excess dPTP and primers, and then 10 additional cycles of "recursive PCR" were performed in the absence of dPTP to replace dPTP in the template with conversion mutations. Recursive PCR was performed with a single primer that annealed to the fragment ends introduced during the dPTP incorporation cycles, enabling selective amplification of the mutant templates generated in the previous PCR step.

[0324] Table 5:

[0325]

[0326] Table 5. Primers were used to generate mutant templates with the same basic adapter structure at both ends. The primer "single_mut" was used for mutagenic PCR of DNA fragments generated by Nextera tag fragmentation. This primer contains a 5' portion that introduces an additional primer binding site at the fragment ends. The primer "single_rec" was able to anneal to this site and was used during recursive PCR to selectively amplify the mutant templates generated with the single_mut primer. XXXXXXXXXXX is a defined, sample-specific 13nt barcode sequence. NNN is a 3nt random nucleotide region.

[0327] As a control, mutant templates with different adapters at both ends were generated using the same protocol as above, except that two different primer sequences were used during both mutagenic PCR (see Table 2) and recursive PCR (see Table 3). The final PCR products were cleaned with SPRIselect beads and analyzed on a high-sensitivity DNA chip using the 2100 Bioanalyzer system (Agilent). As shown in Figure xxx, the templates generated with the same terminal adapters were on average significantly longer than the control samples containing dual adapters. The smallest size of the detectable control templates was ~800bp, while no templates below 2000bp were observed in the single adapter samples.

[0328] Mutant templates with the same terminal adapter (blue) and control templates with dual adapters were run on an Agilent 2100 Bioanalyzer (high-sensitivity DNA kit) to compare size characteristics. The use of the same terminal adapter suppressed the amplification of templates <2kbp. Data are shown in Figure 6 in.

[0329] Example 8 - Further Reduction of the Mutation Bias of Thermococcus Polymerase by Altering the Native dNTP Levels during PCR

[0330] Although Thermococcus polymerase generates a more balanced mutational profile compared to other DNA polymerases, it does exhibit a small bias towards mutations at G and C sites (see Table 4). To eliminate this residual bias, we tested the effect of changing the concentration of natural dNTPs during the mutagenesis and recovery PCR steps on the relative incorporation rates of different nucleotides.

[0331] First, long mutant templates were prepared from bacterial genomic DNA (E. coli strain MG1655) using the method outlined in Example 5, except that the concentration of each nucleotide in the PCR reaction was variable. This was achieved by adding separate solutions of the four natural nucleotides (purchased from New England Biolabs, at standard final concentrations of 200 μM or minimum concentrations of 160 μM (80% relative to standard) or 100 μM (50%)) to the PCR mixture separately. Only one nucleotide was changed per reaction. As a control, an equimolar dNTP mixture provided with Primestar GXL polymerase (Takara) was used, with all natural nucleotides added to the same final concentration of 200 μM. Five mutagenic PCR cycles and twelve recovery cycles were performed using the primers shown in Table 5. The resulting long mutant templates were then subjected to a standard tag fragmentation reaction (see Example 1), and the "internal" fragments were amplified and sequenced on an Illumina MiSeq instrument. The mutation frequencies were determined by alignment with a known reference sequence.

[0332] As shown in Table 6, changes in the concentration of individual dNTPs during mutagenesis and / or recovery PCR altered the observed mutational profile. Importantly, it was found that limiting the amount of dTTP to 50% during mutagenesis resulted in nearly equal mutation frequencies for each nucleotide (Table 3). This confirmed that the residual mutational bias of Thermococcus polymerase can be eliminated by changing the dNTP levels.

[0333] Table 6.

[0334]

[0335] SEQUENCE LISTING <110> LONGAS TECHNOLOGIES PTY LTD <120> ENZYME <130> N411620WO <140> PCT / GB2019 / 050443 <141> 2019-02-19 <150> GB 1802744.1 <151> 2018-02-20 <160> 142 <170> PatentIn version 3.5 <210> 1 <211> 2325 <212> DNA <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus sp. KS-1 <400> 1 atgatcctcg acactgacta cataactgag aatggaaaac ccgtcataag gattttcaag 60 aaggagaacg gcgagtttaa gattgagtac gataggactt ttgaacccta catttacgcc 120 ctcctgaagg acgattctgc cattgaggag gtcaagaaga taaccgccga gaggcacgga 180 acggttgtaa cggttaagcg ggctgaaaag gttcagaaga agttcctcgg gagaccagtt 240 gaggtctgga aactctactt tactcaccct caggacgtcc cagcgataag ggacaagata 300 cgagagcatc cagcagttat tgacatctac gagtacgaca tacccttcgc caagcgctac 360 ctcatagaca agggattagt gccaatggaa ggcgacgagg agctgaaaat gcttgccttt 420 gatatcgaga cgctctacca tgagggcgag gagttcgccg aggggccaat ccttatgata 480 agctacgccg acgaggaagg ggccagggtg ataacgtgga agaacgcgga tctgccctac 540 agctacgccg acgaggaagg ggccagggtg ataacgtgga agaacgcgga tctgccctac 540 gttgacgtcg tctcgacgga gagggagatg ataaagcgct tcctaaaggt ggtcaaagag 600 gttgacgtcg tctcgacgga gagggagatg ataaagcgct tcctaaaggt ggtcaaagag 600 aaagatcctg acgtcctaat aacctacaac ggcgacaact tcgacttcgc ctacctaaaa 660 aaagatcctg acgtcctaat aacctacaac ggcgacaact tcgacttcgc ctacctaaaa 660 aaacgctgtg aaaagcttgg aataaacttc acgctcggaa gggacggaag cgagccgaag 720 aaacgctgtg aaaagcttgg aataaacttc acgctcggaa gggacggaag cgagccgaag 720 attcagagga tgggcgacag gtttgccgtc gaagtgaagg gacggataca cttcgatctc 780 attcagagga tgggcgacag gtttgccgtc gaagtgaagg gacggataca cttcgatctc 780 tatcctgtga taagacggac gataaacctg cccacataca cgcttgaggc cgtttatgaa 840 tatcctgtga taagacggac gataaacctg cccacataca cgcttgaggc cgtttatgaa 840 gccgtcttcg gtcagccgaa ggagaaggtc tacgctgagg agatagctac agcttgggag 900 gccgtcttcg gtcagccgaa ggagaaggtc tacgctgagg agatagctac agcttgggag 900 agcggtgaag gccttgagag agtagccaga tactcgatgg aagatgcgaa ggtcacatac 960 agcggtgaag gccttgagag agtagccaga tactcgatgg aagatgcgaa ggtcacatac 960 gagcttggga aggagttttt ccctatggag gcccagcttt ctcgcttaat cggccagtcc 1020 gagcttggga aggagttttt ccctatggag gcccagcttt ctcgcttaat cggccagtcc 1020 ctctgggacg tctcccgctc cagcactggc aacctcgttg agtggttcct cctcaggaag 1080 ctctgggacg tctcccgctc cagcactggc aacctcgttg agtggttcct cctcaggaag 1080 gcctacgaga ggaatgagct ggccccgaac aagcccgatg aaaaggagct ggccagaaga 1140 gcctacgaga ggaatgagct ggccccgaac aagcccgatg aaaaggagct ggccagaaga 1140 cgacagagct atgaaggagg ctatgtaaaa gagcccgaga gagggttgtg ggagaacata 1200 cgacagagct atgaaggagg ctatgtaaaa gagcccgaga gagggttgtg ggagaacata 1200 gtgtacctag attttagatc tctgtacccc tcaatcatca tcacccacaa cgtctcgccg 1260 gatactctca acagggaagg atgcaaggaa tatgacgttg ccccccaggt cggtcaccgc 1320 ttctgcaagg acttcccagg atttatcccg agcctgcttg gagacctcct agaggagagg 1380 cagaagataa agaagaagat gaaggccacg attgacccga tcgagaggaa gctcctcgat 1440 tacaggcaga gggccatcaa gatcctggcc aacagctact acggttacta cggctatgca 1500 agggcgcgct ggtactgcaa ggagtgtgca gagagcgtaa cggcctgggg aagggagtac 1560 ataacgatga ccatcagaga gatagaggaa aagtacggct ttaaggtaat ctacagcgac 1620 accgacggat tttttgccac aatacctgga gccgatgctg aaaccgtcaa aaagaaggcg 1680 atggagttcc tcaagtatat caacgccaaa ctcccgggcg cgcttgagct cgagtacgag 1740 ggcttctaca aacgcggctt cttcgtcacg aagaagaagt acgcggtgat agacgaggaa 1800 ggcaagataa caacgcgcgg acttgagatt gtgaggcgcg actggagcga gatagcgaaa 1860 gagacgcagg cgagggttct tgaagctttg ctaaaggacg gtgacgtcga gaaggccgtg 1920 aggatagtca aagaagttac cgaaaagctg agcaagtacg aggttccgcc ggagaagctg 1980 gtgatccacg agcagataac gagggattta aaggactaca aggcaaccgg tccccacgtt 2040 gccgttgcca agaggttggc cgcgagagga gtcaaaatac gccctggaac ggtgataagc 2100 tacatcgtgc tcaagggctc tgggaggata ggcgacaggg cgataccgtt cgacgagttc 2160 gacccgacga agcacaagta cgacgccgag tactacattg agaaccaggt tctcccagcc 2220 gttgagagaa ttctgagagc cttcggttac cgcaaggaag acctgcgcta ccagaagacg 2280 agacaggttg gtctgggagc ctggctgaag ccgaagggaa cttga 2325 <210> 2 <211> 774 <212> PRT <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus sp. KS-1 <400> 2 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asn Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Thr 50 55 60 Val Lys Arg Ala Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Glu His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Asp Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Glu Gly Ala Arg Val Ile Thr Trp Lys Asn Ala 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Arg Glu Met Ile Lys 180 185 190 Arg Phe Leu Lys Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Thr Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Val Phe Gly Gln Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Ala Thr Ala Trp Glu Ser Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Tyr 305 310 315 320 Glu Leu Gly Lys Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Ser Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Asp Glu Lys Glu Leu Ala Arg Arg Arg Gln Ser Tyr 370 375 380 Glu Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Lys Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Ala Asn Ser Tyr Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Thr Met Thr Ile Arg Glu Ile 515 520 525 Glu Glu Lys Tyr Gly Phe Lys Val Ile Tyr Ser Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Ala Asp Ala Glu Thr Val Lys Lys Lys Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Lys Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Lys Ala Val 625 630 635 640 Arg Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Arg Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Lys Pro Lys Gly Thr 770 <210> 3 <211> 2325 <212> DNA <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus celer <400> 3 atgatcctcg acgctgacta catcaccgaa gatgggaagc ccgtcgtgag gatattcagg 60 aaggagaagg gcgagttcag aatcgactac gacagggact tcgagcccta catctacgcc 120 ctcctgaagg acgattcggc catcgaggag gtgaagagga taaccgttga gcgccacggg 180 aaggccgtca gggttaagcg ggtggagaag gtcgaaaaga agttcctcaa caggccgata 240 gaggtctgga agctctactt caatcacccg caggacgttc cggcgataag ggacgagata 300 aggaagcatc cggccgtcgt tgatatctac gagtacgaca tccccttcgc caagcgctac 360 ctcatcgata aggggctcgt cccgatggag ggggaggagg agctcaaact gatggccttc 420 gacatcgaga ccctctacca cgagggagac gagttcgggg aggggccgat cctgatgata 480 agctacgccg acggggacgg ggcgagggtc ataacctgga agaagatcga cctcccctac 540 gtcgacgtcg tctcgaccga gaaggagatg ataaagcgct tcctccaggt ggtgaaggag 600 aaggacccgg acgtgctcgt aacttacaac ggcgacaact tcgacttcgc ctacctgaag 660 agacgctccg aggagcttgg attgaagttc atcctcggga gggacgggag cgagcccaag 720 atccagcgca tgggcgaccg cttcgccgtc gaggtgaagg ggaggataca cttcgacctc 780 tacccggtga taaggcgcac cgtgaacctg ccgacctaca cgctcgaggc ggtctacgag 840 gccatcttcg ggaggccaaa ggagaaggtc tacgccgggg agatagtgga ggcctgggaa 900 accggcgagg gtcttgagag ggttgcccgc tactccatgg aggacgcaaa ggttaccttc 960 gagctcggga gggagttctt cccgatggag gcccagctct cgaggctcat cggccagggt 1020 ctctgggacg tctcccgctc gagcaccggc aacctggtcg agtggttcct cctgaggaag 1080 gcctacgaga ggaacgaact ggccccgaac aagccgagcg gccgggaagt ggagatcagg 1140 aggcgtggct acgccggtgg ttacgttaag gagccggaga ggggtttatg ggagaacatc 1200 gtgtacctcg actttcgctc tctttacccc tccatcatca taacccacaa cgtctcgccc 1260 gataccctaa acagggaggg ctgtgagaac tacgacgtcg ccccccaggt ggggcataag 1320 ttctgcaaag attttccggg cttcatcccg agcctgctcg gaggcctgct tgaggagagg 1380 cagaagataa agcggaggat gaaggcctct gtggatcccg ttgagcggaa gctcctcgat 1440 tacaggcaga gggccatcaa gatactggcc aacagcttct acggatacta cggctacgcg 1500 agggcgaggt ggtactgcag ggagtgcgcg gagagcgtta ccgcctgggg cagggagtac 1560 atcgataggg tcatcaggga gctcgaggag aagttcggct tcaaggtgct ctacgcggac 1620 acggacggac tgcacgccac gatccccggg gcggacgccg ggaccgtcaa ggagagggcg 1680 agggggttcc tgagatacat caaccccaag ctccccggcc tcctggagct cgagtacgag 1740 gggttctacc tgaggggttt cttcgtgacg aagaagaagt acgcggtcat agacgaggag 1800 ggcaagataa ccacgcgcgg cctcgagata gtcaggcggg actggagcga ggtggccaag 1860 gagacgcagg cgagggtcct ggaggcgata ctgaggcacg gtgacgtcga ggaggccgtt 1920 agaatcgtca gggaggtaac cgaaaagctg agcaagtacg aggttccgcc ggagaaactg 1980 gtgatccacg agcagataac gagggatttg agggactaca aagccacggg accgcacgtg 2040 gcggtggcga agcgcctggc cgggaggggg gtaaggatac gccccgggac ggtgataagc 2100 tacatcgtcc tcaagggctc cggaaggata ggggacaggg cgattccctt cgacgagttc 2160 gacccgacta agcacaggta cgacgccgac tactacatcg agaaccaggt tctgccagcc 2220 gtcgagagga tcctgaaggc cttcggctac cgcaaggagg acctgaaata ccagaagacg 2280 aggcaggtgg gcctgggtgc gtggctcaac gcggggaagg ggtga 2325 <210> 4 <211> 774 <212> PRT <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus celer <400> 4 Met Ile Leu Asp Ala Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Val 1 5 10 15 Arg Ile Phe Arg Lys Glu Lys Gly Glu Phe Arg Ile Asp Tyr Asp Arg 20 25 30 Asp Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Arg Ile Thr Val Glu Arg His Gly Lys Ala Val Arg 50 55 60 Val Lys Arg Val Glu Lys Val Glu Lys Lys Phe Leu Asn Arg Pro Ile 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Asn His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Glu Ile Arg Lys His Pro Ala Val Val Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Glu Glu Glu Leu Lys Leu Met Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Asp Glu Phe Gly Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Gly Asp Gly Ala Arg Val Ile Thr Trp Lys Lys Ile 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Lys Glu Met Ile Lys 180 185 190 Arg Phe Leu Gln Val Val Lys Glu Lys Asp Pro Asp Val Leu Val Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Arg Arg Ser Glu 210 215 220 Glu Leu Gly Leu Lys Phe Ile Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Val Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Ile Phe Gly Arg Pro Lys Glu 275 280 285 Lys Val Tyr Ala Gly Glu Ile Val Glu Ala Trp Glu Thr Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Phe 305 310 315 320 Glu Leu Gly Arg Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Gly Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Ser Gly Arg Glu Val Glu Ile Arg Arg Arg Gly Tyr 370 375 380 Ala Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Glu Asn Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Lys Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Gly Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Arg Arg Met Lys Ala Ser Val Asp Pro Val Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Ala Asn Ser Phe Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Arg Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Asp Arg Val Ile Arg Glu Leu 515 520 525 Glu Glu Lys Phe Gly Phe Lys Val Leu Tyr Ala Asp Thr Asp Gly Leu 530 535 540 His Ala Thr Ile Pro Gly Ala Asp Ala Gly Thr Val Lys Glu Arg Ala 545 550 555 560 Arg Gly Phe Leu Arg Tyr Ile Asn Pro Lys Leu Pro Gly Leu Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Leu Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Val Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Ile Leu Arg His Gly Asp Val Glu Glu Ala Val 625 630 635 640 Arg Ile Val Arg Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Arg Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Gly 675 680 685 Arg Gly Val Arg Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Arg Tyr Asp Ala Asp Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Lys Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Lys Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Asn Ala Gly Lys Gly 770 <210> 5 <211> 2328 <212> DNA <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus siculi <400> 5 atgatcctcg acacggacta catcacggaa gatgggaaac ccgtcataag gatattcaag 60 aaagagaacg gcgagttcaa gatcgagtac gacaggactt ttgaacccta catctacgcc 120 ctcctgaagg acgactccgc gattgaggat gttaaaaaga taaccgccga gaggcacgga 180 acggtggtga aggtcaagcg cgccgaaaag gtgcagaaga agttcctagg caggccggtt 240 gaagtctgga agctctactt cacccacccc caagatgtcc cggcgataag ggacaagatt 300 aggaagcatc cagctgtaat tgacatctac gagtacgaca taccattcgc caagcgctac 360 ctcatcgaca agggcctgat tccgatggag ggtgaagaag agcttaagat gctcgccttc 420 gacattgaga cgctctacca tgagggtgag gagttcgccg aggggcctat tctgatgata 480 agctacgccg acgagagcga ggcacgcgtc atcacctgga agaaaatcga cctcccctac 540 gttgacgtcg tctcaacgga gaaggagatg ataaagcgct tcctccgcgt tgtgaaggag 600 aaagatcccg atgtcctcat aacctacaac ggcgacaact tcgacttcgc ctacctgaag 660 aagcgctgtg aaaagcttgg aataaacttc ctccttggaa gggacgggag cgagccgaag 720 atccagagaa tgggtgaccg cttcgccgtt gaggtgaagg ggaggataca cttcgacctc 780 tatcctgtaa taaggcgcac gataaacctg ccgacctaca tgcttgaggc agtctacgag 840 gccatctttg ggaagccaaa ggagaaggtt tacgccgagg agatagccac cgcttgggaa 900 accggagagg gccttgagag ggtggctcgc tactctatgg aggacgcgaa ggtcacgttt 960 gagcttggaa aggagttctt cccgatggag gcccaacttt cgaggttggt cggccagagc 1020 ttctgggatg tcgcgcgctc aagcacgggc aatctggtcg agtggttcct cctcaggaag 1080 gcctacgaga ggaacgagct ggctccaaac aagccctctg gaagggaata tgacgagagg 1140 cgcggtggat acgccggcgg ctacgtcaag gaaccggaaa agggcctgtg ggagaacata 1200 gtctacctcg actataaatc tctctacccc tcaatcatca tcacccacaa cgtctcgccc 1260 gataccctca accgcgaggg ctgtaaggag tatgacgtag ctccacaggt cggccaccgc 1320 ttctgcaagg actttccagg cttcatcccg agcctgctcg gggatctcct ggaggagagg 1380 cagaagataa agaggaagat gaaggcaaca attgacccga tcgagagaaa gctccttgat 1440 tacaggcaac gggccatcaa gatccttcta aatagttttt acggctacta cggctacgca 1500 agggctcgct ggtactgcaa ggagtgtgcc gagagcgtta cggcatgggg aagggaatat 1560 atcaccatga caatcaggga aatagaagag aagtatggct ttaaagtact ttatgcggac 1620 actgacggct tcttcgcgac gattcccggg gaagatgccg agaccatcaa aaagagggcg 1680 atggagttcc tcaagtacat aaacgccaaa ctccccggtg cgctcgaact tgagtacgag 1740 gacttctaca ggcgcggctt cttcgtcacc aagaagaaat acgcggttat cgacgaggag 1800 ggcaagataa caacgcgcgg gctggagatc gtcaggcgcg actggagcga gatagccaag 1860 gagacgcagg cgcgggttct ggaggccctt ctgaaggacg gtgacgtcga agaggccgtg 1920 agcatagtca aagaagtgac cgagaagctg agcaagtacg aggttccgcc ggagaagctc 1980 gttatccacg agcagataac gcgcgagctg aaggactaca aggcaacggg accacacgtg 2040 gcgatagcga agaggttagc cgcgagaggc gtcaaaatcc gccccgggac agtcatcagc 2100 tacatcgtgc tcaagggctc cgggaggata ggcgacaggg cgattccctt cgacgagttc 2160 gaccccacga agcacaagta cgatgcagag tactacatcg agaaccaggt tctacctgcc 2220 gtcgagagga ttctgaaggc cttcggctat cgcggtgagg agctcagata ccagaagacg 2280 aggcaggttg gacttggggc gtggctgaag ccgaagggga aggggtga 2328 <210> 6 <211> 775 <212> PRT <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus siculi <400> 6 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Asp Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Lys 50 55 60 Val Lys Arg Ala Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Lys His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Ile Pro 115 120 125 Met Glu Gly Glu Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Ser Glu Ala Arg Val Ile Thr Trp Lys Lys Ile 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Lys Glu Met Ile Lys 180 185 190 Arg Phe Leu Arg Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Leu Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Met Leu Glu Ala Val Tyr Glu Ala Ile Phe Gly Lys Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Ala Thr Ala Trp Glu Thr Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Phe 305 310 315 320 Glu Leu Gly Lys Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Val Gly Gln Ser Phe Trp Asp Val Ala Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Ser Gly Arg Glu Tyr Asp Glu Arg Arg Gly Gly Tyr 370 375 380 Ala Gly Gly Tyr Val Lys Glu Pro Glu Lys Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Tyr Lys Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Arg Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Leu Asn Ser Phe Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Thr Met Thr Ile Arg Glu Ile 515 520 525 Glu Glu Lys Tyr Gly Phe Lys Val Leu Tyr Ala Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Glu Asp Ala Glu Thr Ile Lys Lys Arg Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Asp Phe Tyr Arg Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Glu Ala Val 625 630 635 640 Ser Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Glu Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Ile Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Lys Ala Phe Gly Tyr Arg Gly 740 745 750 Glu Glu Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Lys Pro Lys Gly Lys Gly 770 775 <210> 7 <211> 774 <212> PRT <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus kodakarensis <400> 7 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Phe Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Thr 50 55 60 Val Lys Arg Val Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Glu His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Asp Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr Glu Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Glu Gly Ala Arg Val Ile Thr Trp Lys Asn Val 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Arg Glu Met Ile Lys 180 185 190 Arg Phe Leu Arg Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Ala Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Val Phe Gly Gln Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Thr Thr Ala Trp Glu Thr Gly Glu Asn 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Tyr 305 310 315 320 Glu Leu Gly Lys Glu Phe Leu Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Ser Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Asp Glu Lys Glu Leu Ala Arg Arg Arg Gln Ser Tyr 370 375 380 Glu Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Lys Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Ala Asn Ser Tyr Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Thr Met Thr Ile Lys Glu Ile 515 520 525 Glu Glu Lys Tyr Gly Phe Lys Val Ile Tyr Ser Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Ala Asp Ala Glu Thr Val Lys Lys Lys Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Glu Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Lys Ala Val 625 630 635 640 Arg Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Arg Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Ser Ala Trp 755 760 765 Leu Lys Pro Lys Gly Thr 770 <210> 8 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 8 tagaattgaa gaa 13 <210> 9 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 9 tggccatagc tac 13 <210> 10 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 10 gtcatctgcg acc 13 <210> 11 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 11 ttcgcgcttg gac 13 <210> 12 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 12 cgcgaaccgt tag 13 <210> 13 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 13 ttgcagcctc taa 13 <210> 14 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 14 tctactagta cga 13 <210> 15 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 15 gtaggttcta ctg 13 <210> 16 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 16 gccaatatca agt 13 <210> 17 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 17 ctatcttgct ggt 13 <210> 18 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 18 gttctcatag gta 13 <210> 19 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 19 gtctatgaac caa 13 <210> 20 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 20 cggagcgctt att 13 <210> 21 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 21 tatgccatga gga 13 <210> 22 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 22 atacgactcg gag 13 <210> 23 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 23 gatggaactc agc 13 <210> 24 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 24 ggacctgcat gaa 13 <210> 25 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 25 tagactggaa ctt 13 <210> 26 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 26 gaattacctc gtt 13 <210> 27 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 27 aggatcaggc tac 13 <210> 28 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 28 acgcgtagaa gag 13 <210> 29 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 29 cttcgagact tac 13 <210> 30 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 30 gacggctaac tcc 13 <210> 31 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 31 ttagcattct ctt 13 <210> 32 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 32 gcaaggcata gta 13 <210> 33 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 33 acctagatat gga 13 <210> 34 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 34 acgccaaggc gta 13 <210> 35 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 35 tatgacggat ccg 13 <210> 36 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 36 cctccattag aga 13 <210> 37 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 37 attgaatact ctg 13 <210> 38 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 38 gagatgagaa gaa 13 <210> 39 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 39 tctgagtagc cgg 13 <210> 40 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 40 aataggtagt acg 13 <210> 41 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 41 gtcgaagaag tcc 13 <210> 42 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 42 tactgcatct cgt 13 <210> 43 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 43 gacgtattag agc 13 <210> 44 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 44 cctgcattat tcg 13 <210> 45 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 45 acgaatgatg ctc 13 <210> 46 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 46 tactagcaga gat 13 <210> 47 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 47 ctcctcatct tcc 13 <210> 48 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 48 tcctctgcgc tgc 13 <210> 49 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 49 ccttctcagt ccg 13 <210> 50 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 50 cagcttcata gcg 13 <210> 51 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 51 ttgactctcg cgc 13 <210> 52 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 52 tatcctgagc gat 13 <210> 53 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 53 aacgcctagc cga 13 <210> 54 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 54 ccgaagacgt cat 13 <210> 55 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 55 gagttctcca gat 13 <210> 56 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 56 tgcatccgcg ctt 13 <210> 57 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 57 cctgaactca agt 13 <210> 58 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> sample tag sequence <400> 58 ggtcgtatgc gta 13 <210> 59 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 59 aggcctctct acc 13 <210> 60 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 60 gtactccatc caa 13 <210> 61 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 61 cagcggacgc gct 13 <210> 62 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 62 atctctctta gca 13 <210> 63 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 63 aagcaataat aat 13 <210> 64 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 64 aaggcgactc cga 13 <210> 65 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 65 acgtctctag gag 13 <210> 66 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 66 ccatcagacc tct 13 <210> 67 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 67 acttaatcgt act 13 <210> 68 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 68 tggaattctc caa 13 <210> 69 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 69 ccatacgatc agg 13 <210> 70 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 70 ttatggagca ata 13 <210> 71 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 71 gctcggcgtt cga 13 <210> 72 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 72 ttggccagtc gct 13 <210> 73 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 73 cagatacgta gag 13 <210> 74 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 74 aatgctatta tcc 13 <210> 75 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 75 gcagcatgcc gat 13 <210> 76 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 76 ggagagttac ctc 13 <210> 77 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 77 gagagtccat gat 13 <210> 78 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 78 caatctattc tga 13 <210> 79 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 79 gctcttagta tcc 13 <210> 80 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 80 ccatagttat ggt 13 <210> 81 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 81 tgcgagatcg aag 13 <210> 82 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 82 agagaagtcg agt 13 <210> 83 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 83 ggtaactcca tat 13 <210> 84 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 84 tgctattcca ggc 13 <210> 85 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 85 aaccgcgagg ctc 13 <210> 86 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 86 ttctagagat acc 13 <210> 87 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 87 ttcgctcaag tat 13 <210> 88 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 88 cagagaaggc gca 13 <210> 89 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 89 tagaattggc ctc 13 <210> 90 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 90 ggccattctc cag 13 <210> 91 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 91 tccaacgcgc gtt 13 <210> 92 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 92 gccgcagatt acg 13 <210> 93 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 93 gcagttcgaa cgc 13 <210> 94 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 94 ttctctctgc agg 13 <210> 95 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 95 taagctacca gcg 13 <210> 96 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 96 ctgcatgagg ttg 13 <210> 97 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 97 ttgcctagcg agg 13 <210> 98 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 98 caactgaatt agg 13 <210> 99 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 99 aagcggtcct ctt 13 <210> 100 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 100 aatggaagga ccg 13 <210> 101 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 101 gagttagtaa gtt 13 <210> 102 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 102 ttcctaattc caa 13 <210> 103 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 103 gttctggttc gct 13 <210> 104 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 104 gttcatctct tcc 13 <210> 105 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 105 attccgagga aga 13 <210> 106 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 106 cttagccgag aga 13 <210> 107 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 107 gtctgctacg ctt 13 <210> 108 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 108 atggcgccgc gca 13 <210> 109 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 109 taattggtta tct 13 <210> 110 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 110 tcggttataa gtc 13 <210> 111 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 111 tgcctgagaa cgt 13 <210> 112 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 112 agatgcggtt aac 13 <210> 113 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 113 atggaatagg cga 13 <210> 114 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 114 agagatgcga tcg 13 <210> 115 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 115 ctccaactaa cgt 13 <210> 116 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 116 gccttgctac tgg 13 <210> 117 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 117 cttcgtctct acg 13 <210> 118 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 118 acgctcatag cct 13 <210> 119 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 119 gtcgaagata agg 13 <210> 120 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 120 gccggagtcc tcg 13 <210> 121 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 121 tatacggcga cct 13 <210> 122 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 122 aggtagatat tcg 13 <210> 123 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 123 ttaaggtact gct 13 <210> 124 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 124 cggatctggt ata 13 <210> 125 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 125 gaggtctcgg agg 13 <210> 126 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 126 ggcatcgatg gac 13 <210> 127 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 127 gatctccgat ata 13 <210> 128 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 128 gattcggaat act 13 <210> 129 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 129 ctgcgatccg gcc 13 <210> 130 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 130 gatccggttg caa 13 <210> 131 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 131 cgtcaggctt gac 13 <210> 132 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 132 tcggcaaggc gag 13 <210> 133 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 133 gaacggcgaa cgc 13 <210> 134 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 134 cctcaagcgg act 13 <210> 135 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 135 gaagccagat ggt 13 <210> 136 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 136 tgctcatacc aat 13 <210> 137 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> i7 custom index primer <220> <221> misc_feature <222> (25)..(36) <223> n is a, c, g, or t <400> 137 caagcagaag acggcatacg agatnnnnnn nnnnnngtct cgtgggctcg g 51 <210> 138 <211> 55 <212> DNA <213> Artificial Sequence <220> <223> i5 custom index primer <220> <221> misc_feature <222> (30)..(41) <223> n is a, c, g, or t <400> 138 aatgatacgg cgaccaccga gatctacacn nnnnnnnnnn ntcgtcggca gcgtc 55 <210> 139 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> i7 flow cell primer <400> 139 caagcagaag acggcatacg a 21 <210> 140 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> i5 flow cell primer <400> 140 aatgatacgg cgaccaccga 20 <210> 141 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> single_mut for mutagenesis <220> <221> misc_feature <222> (19)..(34) <223> n is a, c, g, or t <400> 141 tcggtctgcg cctctagcnn nnnnnnnnnn nnnngtctcg tgggctcgga g 51 <210> 142 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> single_rec primer for recovery <400> 142 caagcagaag acggcatacg agattcggtc tgcgcctcta gc 42

Claims

1. A method for introducing substitution mutations into at least one target DNA molecule, the method comprising: a. providing at least one sample comprising at least one target DNA molecule; and b. amplifying the at least one target DNA molecule using a low-bias high-fidelity DNA polymerase, wherein the low bias includes a low template amplification bias such that the DNA polymerase is capable of amplifying a nucleic acid fragment of 7 kbp to 10 kbp with a Kolmolgorov-Smirnov D of less than 0.1; wherein the step of amplifying the at least one target DNA molecule is carried out in the presence of a nucleotide analogue and comprises replicating the at least one target DNA molecule for at least 2 rounds, wherein in the first round of replication, the DNA polymerase incorporates the nucleotide analogue in place of a nucleotide, and in the second round of replication, the nucleotide analogue pairs with a natural nucleotide to introduce a substitution mutation in the complementary strand.

2. Use of a low-bias high-fidelity DNA polymerase in a method for introducing mutations into at least one target DNA molecule, wherein the low bias includes a low template amplification bias such that the low-bias high-fidelity DNA polymerase is capable of amplifying a nucleic acid fragment of 7 kbp to 10 kbp with a Kolmolgorov-Smirnov D of less than 0.1, wherein the method comprises: providing at least one sample comprising at least one target DNA molecule; and using the DNA polymerase to amplify the at least one target DNA molecule; wherein the step of amplifying the at least one target DNA molecule is carried out in the presence of a nucleotide analogue and comprises replicating the at least one target DNA molecule for at least 2 rounds, wherein in the first round of replication, the DNA polymerase incorporates the nucleotide analogue in place of a nucleotide, and in the second round of replication, the nucleotide analogue pairs with a natural nucleotide to introduce a substitution mutation in the complementary strand.

3. The method or use according to claim 1 or 2, characterized in that, The DNA polymerase mutates the adenine, thymine, guanine, and cytosine nucleotides in the at least one target DNA molecule at a ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or 1:1:1:1, respectively.

4. The method or use according to claim 1 or 2, characterized in that The DNA polymerase mutates the adenine, thymine, guanine, and cytosine nucleotides in the at least one target DNA molecule at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively.

5. The method or use according to claim 1 or 2, characterized in that, The DNA polymerase mutates 1% to 15%, 2% to 10%, or 8% of the nucleotides in the at least one target DNA molecule.

6. The method or use according to claim 1 or 2, characterized in that, In each round of replication, the DNA polymerase mutates 0% to 3%, or 0% to 2% of the nucleotides in the at least one target DNA molecule.

7. The method or use according to claim 1 or 2, characterized in that The DNA polymerase uses nucleotide analogs to mutate adenine, thymine, guanine, and / or cytosine in the at least one target DNA molecule.

8. The method or use according to claim 1 or 2, characterized in that The DNA polymerase replaces guanine, cytosine, adenine, and / or thymine with nucleotide analogs.

9. The method or use according to claim 1 or 2, characterized in that, The DNA polymerase uses nucleotide analogs to introduce guanine or adenine nucleotides at a ratio of 0.5 - 1.5:0.5 - 1.5, 0.6 - 1.4:0.6 - 1.4, 0.7 - 1.3:0.7 - 1.3, 0.8 - 1.2:0.8 - 1.2, or 1:1, respectively.

10. The method or use according to claim 1 or 2, characterized in that, The DNA polymerase uses nucleotide analogs to introduce guanine or adenine nucleotides at a ratio of 0.7 - 1.3:0.7 - 1.3, respectively.

11. The method or use according to claim 1 or 2, characterized in that, The method includes the step of amplifying the at least one target DNA molecule using a low - bias high - fidelity DNA polymerase, which is carried out in the presence of the nucleotide analog, and the step of amplifying the at least one target DNA molecule provides at least one target DNA molecule containing the nucleotide analog.

12. The method or use according to claim 1 or 2, characterized in that, The nucleotide analog is dPTP.

13. The method or use according to claim 12, characterized in that, The DNA polymerase introduces guanine - to - adenine substitution mutations, cytosine - to - thymine substitution mutations, adenine - to - guanine substitution mutations, and thymine - to - cytosine substitution mutations.

14. The method or use according to claim 13, wherein The DNA polymerase introduces guanine - to - adenine substitution mutations, cytosine - to - thymine substitution mutations, adenine - to - guanine substitution mutations, and thymine - to - cytosine substitution mutations at a ratio of 0.5 - 1.5:0.5 - 1.5:0.5 - 1.5:0.5 - 1.5, 0.6 - 1.4:0.6 - 1.4:0.6 - 1.4:0.6 - 1.4, 0.7 - 1.3:0.7 - 1.3:0.7 - 1.3:0.7 - 1.3, 0.8 - 1.2:0.8 - 1.2:0.8 - 1.2:0.8 - 1.2, or 1:1:1:1, respectively.

15. The method or use according to claim 13, characterized in that, The DNA polymerase introduces guanine - to - adenine substitution mutations, cytosine - to - thymine substitution mutations, adenine - to - guanine substitution mutations, and thymine - to - cytosine substitution mutations at a ratio of 0.7 - 1.3:0.7 - 1.3:0.7 - 1.3:0.7 - 1.3, respectively.

16. The method or use according to claim 1 or 2, characterized in that, In the absence of nucleotide analogs, the high - fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, 0% to 0.0015%, or 0% to 0.001% mutations per round of replication.

17. The method or use according to claim 11, characterized in that, The method includes a further step of amplifying the at least one target DNA molecule containing the nucleotide analog in the absence of nucleotide analogs.

18. The method or use according to claim 17, characterized in that, The step of amplifying the at least one target DNA molecule containing the nucleotide analog in the absence of nucleotide analogs is carried out using a low - bias DNA polymerase.

19. The method or use according to claim 1 or 2, characterized in that, The method provides at least one target DNA molecule with a mutation, and the method further includes a step of further using the DNA polymerase to amplify the at least one target DNA molecule with the mutation.

20. The method or use according to claim 1 or 2, characterized in that, The DNA polymerase comprises a proofreading domain and / or a domain with enhanced processivity.

21. The method or use according to claim 1 or 2, characterized in that, The DNA polymerase comprises a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids from the following sequences: a. The sequence of SEQ ID NO.2; b. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.2; c. The sequence of SEQ ID NO.4; d. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.4; e. The sequence of SEQ ID NO.6; f. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.6; g. The sequence of SEQ ID NO.7; or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.

7.

22. The method or use according to claim 21, characterized in that, The DNA polymerase includes: The sequence of SEQ ID NO.2; a. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.2; b. The sequence of SEQ ID NO.4; c. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.4; d. The sequence of SEQ ID NO.6; f. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.6; g. The sequence of SEQ ID NO.7; or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.

7.

23. The method or use according to claim 22, wherein The DNA polymerase includes a sequence that is at least 98% identical to SEQ ID NO.

2.

24. The method or use according to claim 22, characterized in that, The DNA polymerase includes a sequence that is at least 98% identical to SEQ ID NO.

4.

25. The method or use according to claim 22, characterized in that, The DNA polymerase includes a sequence that is at least 98% identical to SEQ ID NO.

6.

26. The method or use according to claim 22, characterized in that, The DNA polymerase includes a sequence that is at least 98% identical to SEQ ID NO.

7.

27. The method or use according to claim 1 or 2, characterized in that, The DNA polymerase is a Thermococcus polymerase or a derivative thereof.

28. The method or use according to claim 27, characterized in that, The DNA polymerase is a Thermococcus polymerase.

29. The method or use according to claim 27, characterized in that, The Thermococcus polymerase is derived from a Thermococcus strain selected from the group consisting of Thermococcus kodakarensis, Thermococcus celer, Thermococcus sp. KOD1, and Thermococcus sp. KS-1.

30. The method or use according to claim 1 or 2, characterized in that, It further includes introducing a barcode into the at least one target DNA molecule.

31. The method or use according to claim 1 or 2, characterized in that, It further includes introducing a sample tag into the at least one target DNA molecule.

32. The method or use according to claim 31, characterized in that, Use a set of sample tags, and label the target DNA molecules from different samples with different sample tags from the set.

33. The method or use according to claim 32, characterized in that, Each sample tag differs from substantially all other sample tags in the group by at least 1 low-probability mutation difference or at least 3 high-probability mutation differences, wherein the low-probability mutations are transversion mutations or indel mutations, and the high-probability mutations are transition mutations.

34. The method or use according to claim 33, characterized in that Each sample tag differs from substantially all other sample tags in the group by at least 3 low-probability mutation differences.

35. The method or use according to claim 33, characterized in that, Each sample tag differs from substantially all other sample tags in the group by 3 to 25, or 3 to 10, low-probability mutation differences.

36. The method or use according to claim 33, characterized in that, The low-probability mutation is a transversion mutation.

37. The method or use according to claim 33, characterized in that, The low-probability mutation is an indel mutation.

38. The method or use according to claim 32, characterized in that, Each sample tag differs from substantially all other sample tags in the group by at least 5 high-probability mutation differences.

39. The method or use according to claim 38, characterized in that, Each sample tag differs from substantially all other sample tags in the group by 5 to 25, or 5 to 10, high-probability mutation differences.

40. The method or use according to claim 32, characterized in that, The group of sample tags is obtained by including the following method: a. Analyze a method for introducing mutations into at least one target nucleic acid molecule and determine the average number of low-probability mutations that occur during the method for introducing mutations into at least one target nucleic acid molecule; and b. Determine the sequence of the group of sample tags, wherein each sample tag differs by more low-probability differences from substantially all sample tags in the group compared to the average number of low-probability mutations that occur during the method for introducing mutations into at least one target nucleic acid molecule.

41. The method or use according to claim 1 or 2, characterized in that, It further includes introducing adaptors into each of the at least one target DNA molecule.

42. The method or use according to claim 41, characterized in that, It includes introducing a first adaptor at the 3'-end of the at least one target DNA molecule and a second adaptor at the 5'-end of the at least one target DNA molecule, wherein the first adaptor and the second adaptor can anneal to each other.

43. The method or use according to claim 42, characterized in that, Use primers that are identical to each other and complementary to a part of the first adaptor to amplify the at least one target DNA molecule.

44. The method or use according to claim 42, characterized in that, The first adaptor is complementary to a nucleic acid molecule that is at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to the second adaptor.

45. The method or use according to claim 43, characterized in that, The primer contains a second primer binding site, and the method includes: using the primer to amplify the at least one target DNA molecule, removing the primer, and using a second set of primers that anneal to the second primer binding site to further amplify the at least one target DNA molecule.

46. The method or use according to claim 1 or 2, characterized in that, The method further includes introducing a barcode, a sample tag, and an adaptor into each of the target DNA molecules.

47. The method or use according to claim 1 or 2, characterized in that, The barcode, sample tag, and / or adaptor are introduced by tag fragmentation or by ligation after cleavage.

48. The method or use according to claim 1 or 2, characterized in that, The at least one target DNA molecule is greater than 1 kbp, greater than 1.5 kbp, greater than 2 kbp, greater than 4 kbp, greater than 5 kbp, greater than 7 kbp, or greater than 8 kbp.

49. A method for determining the sequence of at least one target DNA molecule, the method including the method for introducing mutations according to any one of claims 1 or 3 - 48.

50. The method according to claim 49, characterized in that, It includes the steps of: a. Performing the method according to any one of claims 1 or 3 - 48 to provide at least one mutated target DNA molecule; b. Sequencing a region of the at least one mutant target DNA molecule to provide mutant sequence reads; and c. Using the mutant sequence reads to assemble a sequence of at least a portion of the at least one target DNA molecule.

51. The method according to claim 49, characterized in that, Comprising the steps of: a. Performing the method according to any one of claims 1 or 3 - 48 to provide at least one mutant target DNA molecule; b. Fragmenting the at least one mutant target DNA molecule and / or amplifying the at least one mutant target DNA molecule to provide at least one fragmented and / or amplified mutant target DNA molecule; c. Sequencing a region of the at least one fragmented and / or amplified mutant target DNA molecule to provide mutant sequence reads; and d. Using the mutant sequence reads to assemble a sequence of at least a portion of the at least one target DNA molecule.

52. A method for engineering a protein, the method comprising the method for introducing a mutation according to any one of claims 1 or 3 - 48.

53. The method according to claim 52, characterized in that, Comprising the steps of: a. Performing the method according to any one of claims 1 or 3 - 48 to provide at least one mutant target DNA molecule; b. Inserting the at least one mutant target DNA molecule into a vector; and c. Expressing the protein encoded by the at least one mutant target DNA molecule.

54. The method according to claim 53, wherein Comprising the steps of: a. Providing at least one sample, the at least one sample comprising at least one target DNA molecule; and b. Amplifying the at least one target DNA molecule using a low - bias high - fidelity DNA polymerase in the presence of a nucleotide analogue to provide at least one target DNA molecule comprising the nucleotide analogue; c. Amplifying the at least one target DNA molecule comprising the nucleotide analogue in the absence of the nucleotide analogue to provide at least one mutant target DNA molecule; d. Inserting the at least one mutant target DNA molecule into a vector; and e. Expressing the protein encoded by the at least one mutant target DNA molecule.

55. The method according to claim 53 or 54, characterized in that, The method further comprises the step of testing the activity of the protein encoded by the at least one mutant target DNA molecule or evaluating the structure of the protein encoded by the at least one mutant target DNA molecule.

56. The method according to claim 53 or 54, characterized in that The vector is a plasmid, a virus or an artificial chromosome.

57. The method according to claim 53 or 54, characterized in that The step of expressing the protein encoded by the at least one mutant target DNA molecule is achieved by transforming bacterial cells with the vector, transfecting eukaryotic cells or transducing eukaryotic cells.

58. The method according to any one of claims 52 to 54, characterized in that, The step of amplifying the at least one target DNA molecule using a low - bias high - fidelity DNA polymerase is performed using unequal concentrations of dNTPs.

59. The method according to any one of claims 52 to 54, wherein: (i) The method comprises a further step of amplifying the at least one target DNA molecule comprising the nucleotide analogue in the absence of the nucleotide analogue, and the further step of amplifying the at least one target DNA molecule comprising the nucleotide analogue in the absence of the nucleotide analogue is performed using unequal concentrations of dNTPs; or (ii) The method provides at least one target DNA molecule with a mutation, the method comprising a further step of amplifying the at least one target DNA molecule with the mutation using a low-bias DNA polymerase, and the further step of amplifying the at least one target DNA molecule with the mutation using the low-bias DNA polymerase is carried out using unequal concentrations of dNTPs.

60. The method according to claim 58, characterized in that, Unequal concentrations of dNTPs are used to alter the characteristics of the introduced mutation.

61. The method according to claim 60, characterized in that, Unequal concentrations of dNTPs are used to reduce the bias in the characteristics of the introduced mutation.

62. The method according to claim 58, wherein The unequal concentrations of dNTPs comprise dATP, dCTP, dTTP, and dGTP, and the concentration of one or two of dATP, dCTP, dTTP, or dGTP is lower than the concentration of the other dNTPs.

63. The method according to claim 58, wherein, Using unequal concentrations of dNTPs comprises the step of identifying the dNTPs whose levels should be increased or decreased to reduce the bias in the characteristics of the introduced mutation.

64. The method according to claim 58, wherein The unequal concentrations of dNTPs comprise dTTP at a concentration lower than the concentration of the other dNTPs.

65. The method according to claim 64, wherein The unequal concentrations of dNTPs comprise dTTP such that the concentration is less than 75%, less than 70%, less than 60%, less than 55% of the concentration of dATP, dCTP, or dGTP, and is 25% to 75%, 25% to 70%, 25% to 60%, or 50% of the concentration of dATP, dCTP, or dGTP.

66. The method according to claim 65, wherein The unequal concentrations of dNTPs comprise dTTP such that the concentration is less than 75%, less than 70%, less than 60%, less than 55% of the concentration of dCTP, and is 25% to 75%, 25% to 70%, 25% to 60%, or 50% of the concentration of dCTP.

67. The method according to claim 66, wherein The unequal concentrations of dNTPs comprise dTTP at a concentration less than 60% of the concentration of dCTP.

68. The method according to claim 63, wherein The unequal concentrations of dNTPs comprise dTTP at a concentration of 25% to 60% of the concentration of dCTP.

69. The method according to claim 59, characterized in that, The step of amplifying the at least one target DNA molecule comprising a nucleotide analogue or the at least one target DNA molecule with the mutation in the absence of a nucleotide analogue is carried out using unequal concentrations of dNTPs.

70. The method according to claim 69, wherein The unequal concentrations of dNTPs comprise dATP at a concentration lower than the concentration of the other dNTPs.

71. The method according to claim 70, wherein The unequal concentrations of dNTPs comprise dATP such that the concentration is less than 75%, less than 70%, less than 60%, less than 55% of the concentration of dTTP, dCTP, or dGTP, and is 25% to 75%, 25% to 70%, 25% to 60%, or 50% of the concentration of dTTP, dCTP, or dGTP.

72. The method according to claim 71, wherein The dNTPs with unequal concentrations comprise dATP such that the concentration is less than 75%, less than 70%, less than 60%, less than 55% of the concentration of dGTP, and is 25% to 75%, 25% to 70%, 25% to 60%, or 50% of the concentration of dGTP.

73. The method according to claim 72, characterized in that, The unequal concentrations of dNTPs comprise dATP at a concentration less than 60% of the concentration of dGTP.

74. The method according to claim 72, characterized in that, The unequal concentrations of dNTPs comprise dATP at a concentration of 25% to 60% of the concentration of dGTP.

Citation Information

Patent Citations

  • Mutation method for enhancing beta-cyclodextrin production capacity of beta-cyclodextrin glycosyltransferase

    CN103555685A