Method for introducing mutations
Low-bias DNA polymerases with sample tags and adapters address mutation bias and template amplification issues, improving sequencing and protein manipulation by ensuring uniform mutations and efficient amplification of longer nucleic acids.
Patent Information
- Application Number
- JP2023221021
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-02-20
- Filing Date
- 2023-12-27
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2039-02-19
AI Technical Summary
Existing DNA polymerases introduce mutations with high bias, leading to uneven mutation rates and template amplification, complicating sequencing and protein manipulation applications.
Employing a low-bias DNA polymerase that uniformly randomizes mutations and reduces template amplification bias, combined with sample tags and adapters to distinguish sequences and preferentially amplify longer nucleic acids.
Enhances sequencing accuracy by ensuring unique mutation patterns and efficient amplification of longer nucleic acids, facilitating sequence assembly and protein manipulation.
Smart Images

Figure 0007717145000007 
Figure 0007717145000008 
Figure 0007717145000009
Abstract
Description
Technical Field
[0001] The present invention relates to a method for introducing mutations into one or more nucleic acid molecules, the use of a low-bias DNA polymerase in a method for introducing mutations into one or more nucleic acid molecules, a sample tag group, a method for designing the sample tag group, a computer-readable medium, and a method for selectively amplifying a target nucleic acid molecule.
Background Art
[0002] DNA polymerases can be used to introduce mutations into nucleic acid sequences. This can be useful in a plurality of applications. For example, mutagenesis techniques can be useful in applications including sequencing assisted by mutagenesis techniques (SAM), and for introducing mutations into protein sequences to find mutations that affect the activity of the protein.
[0003] Mutations may be introduced using a DNA polymerase that exhibits low fidelity. Low-fidelity DNA polymerases make mistakes during replication, which results in the introduction of mutations. However, many low-fidelity DNA polymerases introduce mutations at a rate of less than 2% per round of mutagenesis reaction (one round of replication), and for some applications, a higher mutagenesis rate is useful. In addition, low-fidelity DNA polymerases may introduce mutations in a biased manner. Such DNA polymerases are sometimes referred to as high-bias DNA polymerases.
[0004] Mutations may be introduced by replicating a sequence using a DNA polymerase in the presence of a nucleotide analog such as dPTP. The DNA polymerase may incorporate a nucleotide analog instead of a natural nucleotide. As a result, in subsequent replication cycles, the nucleotide analog can pair with a natural nucleotide that was not present in the original sequence, thereby introducing a mutation. By utilizing the introduction of mutations by replicating a sequence in the presence of a nucleotide analog, a higher mutation rate can be achieved.
[0005] By using a commonly used DNA polymerase (e.g., Taq polymerase), nucleotide analogs can be incorporated instead of natural nucleotides. However, these polymerases are high-bias polymerases. High-bias DNA polymerases may exhibit two possible biases, namely mutation bias and template amplification bias.
[0006] Some high-bias polymerases have a high mutation bias, but this is because the polymerase does not uniformly and randomly mutate all four natural nucleotides (adenine, cytosine, guanine, and thymine). For example, high-bias DNA polymerases may mutate some nucleotides more frequently than others. The adenine / thymine pair is connected by two hydrogen bonds, while the guanine / cytosine pair is connected by three hydrogen bonds. Therefore, high-bias DNA polymerases may be more likely to introduce mutations into the adenine / thymine pair than the guanine / cytosine pair.
[0007] High-bias polymerases with high mutational bias may not be able to randomly incorporate nucleotide analogs. For example, high-bias polymerases may act preferentially in the replacement of a specific base with a nucleotide analog. dPTP is interconvertible between two different tautomeric forms, the imino form and the amino form. The imino tautomer can form a Watson-Crick base pair with adenine, and the amino form can form a Watson-Crick base pair with guanine (Kong Thoo Lin P, Brown D M (1989). “Synthesis and duplex stability of oligonucleotides containing cytosine-thymine analogues”. Nucleic Acids Research. 17:10373-10383; Stone M J et al (1991). “Molecular basis for methoxyamine-initiated mutagenesis: 1 H nuclear magnetic resonance studies of base-modified oligodeoxynucleotides.” Journal of Molecular Biology. 222:711-723; Nedderman A N R et al (1993). “Molecular basis for methoxyamine initiated mutagenesis: 1"H nuclear magnetic resonance studies of oligonucleotide duplexes containing base-modified cytosine residues". Journal of Molecular Biology. 230:1068-1076; Moore M H et al. (1995). "Direct observation of two base-pairing modes of a cytosine-thymine analogue with guanine in a DNA Z-form duplex. Significance for base analogue mutagenesis". Journal of Molecular Biology. 251:665-673). This essentially means that by utilizing replication in the presence of dPTP, substituents can be introduced into the nucleotide sequence in place of adenine, cytosine, guanine, or thymine. However, in aqueous solution, the ratio of the imino form to the amino form of dPTP has been shown to be approximately 10:1 (Harris V H et al. (2003). "The effect of tautomeric constant on the specificity of nucleotide incorporation during DNA replication: support for the rare tautomer hypothesis of substitution mutagenesis". Journal of Molecular Biology. 326:1389-1401).Thus, when using a polymerase such as Taq polymerase to introduce mutations using dPTP, the polymerase introduces adenine and thymine substitutions much more frequently than guanine and cytosine substitutions (Zaccolo M et al. (1996). “An approach to random mutagenesis of DNA using mixtures of triphosphate derivatives of nucleoside analogues”. Journal of Molecular Biology. 255:589-603; Harris V H et al. (2003). “The effect of tautomeric constant on the specificity of nucleotide incorporation during DNA replication: support for the rare tautomer hypothesis of substitution mutagenesis”. Journal of Molecular Biology. 326:1389-1401).
[0008] Second, high-bias polymerases can exhibit template amplification bias, i.e., they may replicate some template nucleic acid molecules with a higher success rate per PCR cycle than others. Over multiple cycles of PCR, this bias can create extreme differences in copy number between templates. Regions of the template nucleic acid molecule may form secondary structures or may contain some nucleotides (e.g., guanine or cytosine nucleotides) at a higher proportion than others. High-bias polymerases may be more effective, for example, at amplifying template nucleic acid molecules rich in guanine and cytosine compared to those rich in adenine and thymine, or may be more effective at amplifying template nucleic acid molecules that do not form secondary structures.
[0009] Many applications of mutagenesis are more effective when mutagenesis can be performed with low bias (both mutation bias and template amplification).
[0010] Most second-generation sequencing platforms can only sequence short nucleic acid fragments, and since the target nucleic acid sequence needs to be amplified during the sequencing process to provide sufficient nucleic acid molecules for the sequencing step, it has been found that accurate assembly of genomic sequences is difficult. If a user wishes to sequence a larger nucleic acid sequence, this can be achieved by sequencing multiple regions of the target nucleic acid molecule. The user must then computationally assemble the sequence of the entire nucleic acid sequence from the sequences of these regions.
[0011] Assembling a nucleic acid sequence using the sequences of multiple regions can be difficult. In particular, when long regions of the sequence are very similar to each other, it will be difficult to determine whether the sequences of two regions are both sequences of replicas of the same original template nucleic acid molecule or correspond to sequences obtained from two different original template nucleic acid molecules. Similarly, it will be difficult to determine whether the sequences of two regions correspond to sequences of replicas of the same part of the template nucleic acid molecule or actually correspond to two different repetitive sequences within the template nucleic acid molecule. These difficulties can be avoided by introducing mutations into the target nucleic acid molecule before amplification. The user may then confirm that fragments with the same mutation pattern may be derived from the same part of the same original template nucleic acid molecule. This type of sequencing method is sometimes referred to as sequencing assisted by mutagenesis (SAM).
Prior Art Documents
Non-Patent Documents
[0012]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 6
Non-Patent Document 7
Summary of the Invention
[0013] The above sequencing method is more effective when the mutations introduced into the target nucleic acid molecule are uniformly random. If the mutations are uniformly random, then, for example, any given portion of the template nucleic acid molecule is more likely to have a unique mutation pattern. Therefore, there is a need to identify a DNA polymerase that can introduce mutations uniformly randomly (with low mutation bias).
[0014] In addition, sequencing methods using DNA polymerases with high template amplification bias may be limited. A DNA polymerase with high template amplification bias will replicate and / or mutate some target nucleic acid molecules more than others, and for this reason, a sequencing method using such a high-bias DNA polymerase may not be able to sequence some target nucleic acid molecules sufficiently.
[0015] The inventors have identified a polymerase that is a low-bias polymerase (having both low template amplification bias and low mutation bias) and is particularly useful in methods for introducing mutations into at least one target nucleic acid molecule.
[0016] A user may wish to use the method of the present invention on two or more samples at a time. In such a case, it would be advantageous for the user to be able to identify which target nucleic acid molecule is derived from which original sample. Such identification could be achieved by labeling the target nucleic acid molecule with a sample tag. However, since the sample tag itself may mutate during the course of the method, the inventors have determined a method for designing sample tags that can be distinguished from each other even if they mutate.
[0017] The user may also wish to ensure that, by utilizing the method of the present invention, longer target nucleic acid molecules are preferentially mutated and amplified compared to shorter target nucleic acid molecules. The inventors have found that this can be achieved by introducing special primer binding sites at each end of the target nucleic acid molecule.
[0018] Accordingly, in a first aspect of the present invention, there is provided a method for introducing mutations into at least one target nucleic acid molecule, comprising a. and b. below, namely a. providing at least one sample containing at least one target nucleic acid molecule, and b. amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase is provided.
[0019] In a second aspect of the present invention, there is provided the use of a low-bias DNA polymerase in a method for introducing mutations into at least one target nucleic acid molecule.
[0020] In a third aspect of the present invention, there is provided a method for determining the sequence of at least one target nucleic acid molecule, comprising the method for introducing mutations of the present invention.
[0021] In a fourth aspect of the present invention, a method for manipulating a protein is provided, which includes a method for introducing a mutation of the present invention.
[0022] In a fifth aspect of the present invention, a sample tag group is provided, wherein each sample tag is different from substantially all other sample tags of the group by at least one low-probability mutation difference or at least three high-probability mutation differences.
[0023] In a sixth aspect of the present invention, the following a. and b., namely a. Analyzing a method for introducing a mutation into at least one target nucleic acid molecule and determining the average number of low-probability mutations that occur during the method for introducing a mutation into this at least one target nucleic acid molecule, and b. Determining the sequence for the group in which each sample tag is different from substantially all sample tags of the sample tag group by a low-probability difference greater than the average number of low-probability mutations that occur during the method for introducing a mutation into at least one target nucleic acid molecule. A method for designing a sample tag group suitable for use in a method for introducing a mutation into at least one target nucleic acid molecule is provided, which includes the above.
[0024] In a seventh aspect of the present invention, the following a. and b., namely a. Preparing at least one sample containing at least one target nucleic acid molecule, and b. Introducing a mutation into the at least one target nucleic acid molecule by amplifying the at least one target nucleic acid molecule using a DNA polymerase to prepare a mutated at least one target nucleic acid molecule. Here, step b. is to be performed using dNTPs of non-uniform concentration. A method for introducing a mutation into at least one target nucleic acid molecule is provided, which includes the above.
[0025] In an eighth aspect of the present invention, a sample tag group obtainable by the method for designing the sample tag group of the present invention is provided.
[0026] In a ninth aspect of the present invention, there is provided a computer-readable medium configured to implement a method for designing a sample tag group of the present invention.
[0027] In a tenth aspect of the present invention, the following a. to c., that is, a. preparing at least one sample containing a target nucleic acid molecule; b. introducing a first adapter to the 3'-end of the target nucleic acid molecule and introducing a second adapter to the 5'-end of the target nucleic acid molecule; and c. amplifying the target nucleic acid molecule using a primer complementary to a part of the first adapter, wherein the first adapter and the second adapter are capable of annealing to each other, A method for selectively amplifying a target nucleic acid molecule having a length greater than 1 kbp is provided.
Brief Description of the Drawings
[0028]
Figure 1-1
Figure 1-2
Figure 2
Figure 3-1
Figure 3-2
Figure 3-3
Figure 3-4
Figure 3-5
Figure 3-6
Figure 3-7
Figure 4-1
Figure 4-2
Figure 4-3
Figure 4-4
Figure 4-5
Figure 5
Figure 6
Figure 7
Mode for Carrying Out the Invention
[0029] General definitions Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0030] In general, the term "comprising" shall mean including but not limited to. For example, the expression "a method for introducing a mutation into at least one target nucleic acid molecule comprising a certain step" should be construed to mean that the method includes the recited step, but may perform additional steps.
[0031] In some embodiments of the present invention, the word "comprising" is replaced with the expression "consisting of". The term "consisting of" is intended to be limiting. For example, the expression "a method for introducing a mutation into at least one target nucleic acid molecule consisting of a certain step" should be understood to mean that the method includes the recited step and does not perform additional steps.
[0032] For the purposes of the present invention, to determine the percent identity of two sequences (e.g., two polynucleotide sequences), the sequences are aligned for optimal comparison (e.g., gaps can be introduced into the first sequence for optimal alignment with the second sequence). Next, the nucleotides or amino acid residues at each position are compared. If a position in the first sequence is occupied by the same residue as the corresponding position in the second sequence, then the residue is identical at that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., percent identity = number of identical positions / total number of positions × 100). Typically, sequence comparison is performed over the entire length of the reference sequence. For example, to assess whether a test sequence is at least 95% identical to SEQ ID NO:2 (reference sequence), one of ordinary skill in the art would perform an alignment over the entire length of SEQ ID NO:2 and identify how many positions in the test sequence are identical to the positions in SEQ ID NO:2. If at least 80% of those positions are identical, then the test sequence is at least 80% identical to SEQ ID NO:2. If the sequence is shorter than SEQ ID NO:2, gaps should be considered as non-identical positions.
[0033] One of ordinary skill in the art is aware of the various computer programs available for determining homology or identity between two sequences. For example, comparison of sequences and determination of the percent identity between two sequences can be accomplished using mathematical algorithms. In one embodiment, the percent identity between two amino acid or nucleic acid sequences is determined using the Needleman and Wunsch (1970) algorithm incorporated into the GAP program in the Accelrys GCG software package (available at http: / / www.accelrys.com / products / gcg / ), using a Blosum 62 matrix or PAM250 matrix, and gap weight 16, 14, 12, 10, 8, 6, or 4 and length weight 1, 2, 3, 4, 5, or 6.
[0034] Method for introducing mutations into at least one target nucleic acid molecule In one aspect, the present invention provides a method for introducing a mutation into at least one target nucleic acid molecule. In a further aspect, the present invention provides the use of a low-bias DNA polymerase in a method for introducing a mutation into at least one target nucleic acid molecule.
[0035] The mutation may be a substitution mutation, an insertion mutation or a deletion mutation. For the purposes of the present invention, the term "substitution mutation" should be construed to mean that one nucleotide has been replaced by another nucleotide. For example, the conversion of the sequence ATCC to the sequence AGCC is a substitution mutation. For the purposes of the present invention, the term "insertion mutation" should be construed to mean that at least one nucleotide has been added to the sequence. For example, the conversion of the sequence ATCC to the sequence ATTCC is an example of an insertion mutation (an additional T nucleotide has been inserted). For the purposes of the present invention, the term "deletion mutation" should be construed to mean that at least one nucleotide has been removed from the sequence. For example, the conversion of the sequence ATTCC to ATCC is an example of a deletion mutation (the T nucleotide has been removed). Preferably, the mutation is a substitution mutation.
[0036] For the purposes of the present invention, "nucleic acid molecule" refers to a polymeric form of nucleotides of any length. The nucleotides may be deoxyribonucleotides, ribonucleotides or analogs thereof. Preferably, the target nucleic acid molecule is composed of deoxyribonucleotides or ribonucleotides. Even more preferably, the target nucleic acid molecule is composed of deoxyribonucleotides, i.e., the target nucleic acid molecule is a DNA molecule.
[0037] At least one "target nucleic acid molecule" can be any nucleic acid molecule into which the user of the method desires to introduce a mutation. The target nucleic acid molecule may form part of a larger nucleic acid molecule such as a chromosome. The target nucleic acid molecule may contain one gene, a plurality of genes, or a fragment of a gene. The target nucleic acid molecule may have a size greater than 1 kbp, greater than 1.5 kbp, greater than 2 kbp, greater than 4 kbp, greater than 5 kbp, greater than 7 kbp, greater than 8 kbp, from 1 kbp to 50 kbp, or from 1 kbp to 20 kbp.
[0038] The term "at least one target nucleic acid molecule" is considered interchangeable with the term "at least one target nucleic acid molecule(s)".
[0039] "At least one target nucleic acid molecule" can be single-stranded or part of a double-stranded complex. For example, if the at least one target nucleic acid molecule is composed of deoxyribonucleotides, the target nucleic acid molecule may form part of a double-stranded DNA complex. In that case, one strand (e.g., the coding strand) is regarded as the at least one target nucleic acid molecule, and the other strand is a nucleic acid molecule complementary to the at least one target nucleic acid molecule.
[0040] A method for introducing a mutation into at least one target nucleic acid molecule is as follows a. and b., namely a. preparing at least one sample containing at least one target nucleic acid molecule, and b. amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase.
[0041] Preparing at least one sample containing at least one target nucleic acid molecule A method for introducing a mutation into at least one target nucleic acid molecule may include the step of preparing at least one sample containing at least one target nucleic acid molecule.
[0042] The at least one sample may include any sample containing at least one target nucleic acid molecule. The at least one sample may be obtained from any source. For example, the at least one sample may include a sample of nucleic acid derived from a human, such as a sample extracted from a skin specimen of a human patient. Alternatively, the at least one sample may be derived from other sources such as a sample obtained from a water supply. Such a sample may contain billions of template nucleic acid molecules. It would be possible to mutate each of these billions of target nucleic acid molecules simultaneously using the method of the present invention, and for this purpose, there is no upper limit to the number of target nucleic acid molecules that can be used in the method of the present invention.
[0043] In certain embodiments, step a. includes providing two or more samples. For example, step a. may include providing 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 15, 20, 25, 50, 75, or 100 samples. Depending on the situation, step a. includes providing less than 2000, less than 1000, less than 750, or less than 500 samples. In further embodiments, step a. includes providing 2-100, 2-75, 2-50, 2-25, 5-15, or 7-15 samples.
[0044] amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase The method of the present invention may include amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase.
[0045] Amplifying the at least one target nucleic acid molecule refers to providing (preparing) the at least one target nucleic acid molecule and / or at least one nucleic acid molecule complementary to a replication product of the at least one target nucleic acid molecule by replicating the at least one target nucleic acid molecule. By amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase, the number of replication products of the at least one target nucleic acid molecule increases, and mutations are introduced into the at least one target nucleic acid molecule. Because mutations are introduced, the replication products are not necessarily identical to the original at least one target nucleic acid molecule. The original at least one target nucleic acid molecule and the replication products of the at least one target nucleic acid molecule may be collectively referred to as "at least one mutated target nucleic acid molecule".
[0046] For example, amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase may include incubating a sample containing the at least one target nucleic acid molecule with the low-bias DNA polymerase and appropriate primers under conditions suitable for the low-bias DNA polymerase to catalyze the production of replication products of the at least one target nucleic acid molecule.
[0047] Appropriate primers include short nucleic acid molecules complementary to regions adjacent to the at least one target nucleic acid molecule or to regions adjacent to nucleic acid molecules complementary to the at least one target nucleic acid molecule. For example, if the target nucleic acid molecule is part of a chromosome, the primer may be complementary to a nucleic acid molecule complementary to the region of the chromosome immediately 3' to the 3' end of the target nucleic acid molecule and to the region immediately 5' to the 5' end of the target nucleic acid molecule, or the primer may be complementary to a nucleic acid molecule complementary to the region of the chromosome immediately 3' to the 3' end of a nucleic acid molecule complementary to the target nucleic acid molecule and to the region immediately 5' to the 5' end of a nucleic acid molecule complementary to the target nucleic acid molecule. Alternatively, the user may introduce primer binding sites (short nucleic acid sequences) into regions adjacent to the at least one target nucleic acid molecule. This is described in more detail in the section entitled "Barcodes, Samples and Adapters".
[0048] Suitable conditions include a temperature at which a low-bias DNA polymerase can catalyze the production of a replication product of the at least one target nucleic acid molecule. For example, a temperature of 40°C to 90°C, 50°C to 80°C, 60°C to 70°C, or about 68°C may be used.
[0049] The step of amplifying the at least one target nucleic acid molecule may include multiple rounds of replication. For example, the step of amplifying the at least one target nucleic acid molecule preferably includes the following i) and ii), that is, i) a round of replicating the at least one target nucleic acid molecule to provide at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule, and ii) a round of replicating the at least one target nucleic acid molecule to provide a replication product of the at least one target nucleic acid molecule and includes.
[0050] Depending on the situation (optionally), the step of amplifying the at least one target nucleic acid molecule includes replicating this at least one target nucleic acid molecule for at least 2, at least 4, at least 6, at least 8, or at least 10 rounds. Some of these rounds of replicating the at least one target nucleic acid molecule may be performed in the presence of nucleotide analogs. Depending on the situation, the step of amplifying the at least one target nucleic acid molecule includes replication at a temperature of 60°C to 80°C at least 1, at least 2, at least 3, at least 4, at least 5, or at least 6 times.
[0051] Depending on the situation (optionally), the step of amplifying the at least one target nucleic acid molecule is performed using the polymerase chain reaction (PCR). PCR involves multiple rounds of the following steps a) to d) to replicate nucleic acid molecules, that is, a) Melting, b) Annealing, c) Extension, and d) Elongation is a process that requires
[0052] Mix the nucleic acid molecule (e.g., the at least one target nucleic acid molecule) with appropriate primers and a polymerase, such as the low-bias DNA polymerase of the present invention. In the melting step, the nucleic acid molecule is heated to a temperature above 90°C such that the double-stranded nucleic acid molecule denatures (separates into two strands). In the annealing step, the nucleic acid molecule is cooled to a temperature below 75°C, e.g., 55°C - 70°C, about 55°C, or about 68°C, so that the primer can anneal to the nucleic acid molecule. In the extension step, the nucleic acid molecule is heated to a temperature above 60°C so that the DNA polymerase can catalyze primer extension and the addition of nucleotides complementary to the template strand. In the elongation step, the nucleic acid molecule is heated to a temperature at which the DNA polymerase exhibits high activity, e.g., 60°C - 70°C, to catalyze the addition of further complementary nucleic acids to complete the new nucleic acid strand.
[0053] Depending on the situation, the method of the present invention includes multiple rounds of PCR using a low-bias DNA polymerase.
[0054] Low-bias DNA polymerase The method of the present invention may include the step of amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase.
[0055] According to the present invention, a "low-bias DNA polymerase" is a DNA polymerase that (a) exhibits a low mutation bias and / or (b) exhibits a low template amplification bias.
[0056] Low mutation bias A low-bias DNA polymerase that exhibits low mutation bias is a DNA polymerase that can mutate adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine at similar ratios. In certain embodiments, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at similar ratios.
[0057] Depending on the situation, the low-bias DNA polymerase can mutate adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine at a ratio of 0.5 - 1.5:0.5 - 1.5, 0.6 - 1.4:0.6 - 1.4, 0.7 - 1.3:0.7 - 1.3, 0.8 - 1.2:0.8 - 1.2, or approximately 1:1, respectively. Preferably, the low-bias DNA polymerase can mutate guanine and adenine at a ratio of 0.5 - 1.5:0.5 - 1.5, 0.6 - 1.4:0.6 - 1.4, 0.7 - 1.3:0.7 - 1.3, 0.8 - 1.2:0.8 - 1.2, or approximately 1:1, respectively. Preferably, the low-bias DNA polymerase can mutate thymine and cytosine at a ratio of 0.5 - 1.5:0.5 - 1.5, 0.6 - 1.4:0.6 - 1.4, 0.7 - 1.3:0.7 - 1.3, 0.8 - 1.2:0.8 - 1.2, or approximately 1:1, respectively.
[0058] In such an embodiment, in the step of amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase, the DNA polymerase mutates adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine nucleotides in the at least one target nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively. Preferably, the low-bias DNA polymerase mutates guanine and adenine nucleotides in the at least one target nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively. Preferably, the low-bias DNA polymerase mutates thymine and cytosine nucleotides in the at least one target nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively.
[0059] Depending on the situation, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at a ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3.
[0060] In such an embodiment, in the step of amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase, the DNA polymerase may mutate the adenine, thymine, guanine, and cytosine nucleotides in the at least one target nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase mutates the adenine, thymine, guanine, and cytosine nucleotides in the at least one target nucleic acid molecule at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3.
[0061] Adenine, thymine, cytosine, and / or guanine may be substituted with another nucleotide. For example, if the low-bias DNA polymerase can mutate adenine, by amplifying the at least one target nucleic acid molecule in the presence of the low-bias DNA polymerase, at least one adenine nucleotide in the nucleic acid molecule may be substituted with thymine, guanine, or cytosine. Similarly, if the low-bias DNA polymerase can mutate thymine, by amplifying the at least one target nucleic acid molecule in the presence of the low-bias DNA polymerase, at least one thymine nucleotide may be substituted with adenine, guanine, or cytosine. If the low-bias DNA polymerase can mutate guanine, by amplifying at least one target nucleotide in the presence of the low-bias DNA polymerase, at least one guanine nucleotide may be substituted with thymine, adenine, or cytosine. If the low-bias DNA polymerase can mutate cytosine, by amplifying at least one target nucleotide in the presence of the low-bias DNA polymerase, at least one cytosine nucleotide may be substituted with thymine, guanine, or adenine.
[0062] Although low-bias DNA polymerases may not be able to directly substitute nucleotides, they may still be able to mutate a nucleotide by replacing the corresponding nucleotide on the complementary strand. For example, if the target nucleic acid molecule contains thymine, an adenine nucleotide will be present at the corresponding position of the at least one nucleic acid molecule that is complementary to the at least one target nucleic acid molecule. The low-bias DNA polymerase may be able to replace the adenine nucleotide of the at least one nucleic acid molecule that is complementary to the at least one target nucleic acid molecule with guanine, and for this reason, when the at least one nucleic acid molecule that is complementary to the at least one target nucleic acid molecule is replicated, as a result, cytosine will be present in the corresponding replicated at least one target nucleic acid molecule where thymine was originally present (substitution from thymine to cytosine).
[0063] In certain embodiments, the low-bias DNA polymerase mutates 1% to 15%, 2% to 10%, or about 8% of the nucleotides in at least one target nucleic acid. In such embodiments, the step of amplifying the at least one target nucleic acid molecule using the low-bias DNA polymerase is performed in such a way that 1% to 15%, 2% to 10%, or about 8% of the nucleotides in the at least one target nucleic acid are mutated. For example, if the user desires to mutate about 8% of the nucleotides in the target nucleic acid molecule and the low-bias DNA polymerase mutates about 1% of the nucleotides per round of replication, the step of amplifying the at least one target nucleic acid molecule using the low-bias DNA polymerase may include 8 rounds of replication.
[0064] In certain embodiments, the low-bias DNA polymerase can mutate 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5% of the nucleotides in the at least one target nucleic acid molecule per round of replication. In certain embodiments, the low-bias DNA polymerase mutates 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5% of the nucleotides in the at least one target nucleic acid molecule per round of replication. The actual amount of mutations that occur in each round may vary, but on average will be 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5%.
[0065] Whether a DNA polymerase can mutate nucleotides, and if so, at what rate Whether a low-bias DNA polymerase can mutate a certain percentage of the nucleotides in the at least one target nucleic acid molecule per round of replication can be determined by amplifying a nucleic acid molecule of known sequence in the presence of the low-bias DNA polymerase over a defined number of replication rounds. The resulting amplified nucleic acid molecule can then be sequenced to calculate the percentage of nucleotides mutated per round of replication. For example, a nucleic acid molecule of known sequence can be amplified using 10 rounds of PCR in the presence of the low-bias DNA polymerase. The resulting nucleic acid molecule can then be sequenced. If the resulting nucleic acid molecule contains 10% nucleotides that are different from the corresponding nucleotides in the original known sequence, the user will then understand that the low-bias DNA polymerase can mutate on average 1% of the nucleotides in the at least one target nucleic acid molecule per round of replication. Similarly, to examine whether a low-bias DNA polymerase mutates a certain percentage of the nucleotides in the at least one target nucleic acid molecule in a given method, the user can perform the method on a nucleic acid molecule of known sequence and, upon completion of the method, can utilize sequencing methods to determine the percentage of mutated nucleotides.
[0066] A low-bias DNA polymerase can mutate nucleotides such as adenine when used to amplify a nucleic acid molecule, provided that the low-bias DNA polymerase provides a nucleic acid molecule in which some examples of the nucleotide have been substituted or deleted. Preferably, the term "mutate (cause to mutate)" refers to the introduction of a substitution mutation, and in some embodiments, the term "mutate (cause to mutate)" can be replaced with "introduce a substitution of ~".
[0067] A low-bias DNA polymerase mutates a nucleotide such as adenine in at least one target nucleic acid molecule in the method of the present invention when performing the step of amplifying at least one target nucleic acid molecule using the low-bias DNA polymerase, provided that this step results in at least one target nucleic acid molecule that has mutated (some examples of the nucleotide have mutated). For example, when a low-bias DNA polymerase mutates an adenine in the at least one target nucleic acid molecule, performing the step of amplifying the at least one target nucleic acid molecule using the low-bias DNA polymerase results in at least one target nucleic acid molecule that has mutated (at least one adenine has been substituted or deleted).
[0068] To determine whether a particular mutation can be introduced by a DNA polymerase, one of ordinary skill in the art need only test the DNA polymerase using a nucleic acid molecule of a known sequence. Suitable nucleic acid molecules having a known sequence are fragments obtained from bacterial genomes having a known sequence, such as Escherichia coli (E. coli) MG1655. One of ordinary skill in the art can amplify a nucleic acid molecule of a known sequence using PCR in the presence of a low-bias DNA polymerase. One of ordinary skill in the art can then sequence the amplified nucleic acid molecule and further determine whether its sequence is the same as the original known sequence. Even if not up to that point, one of ordinary skill in the art can determine the nature of the mutation. For example, if one of ordinary skill in the art wishes to determine whether a DNA polymerase can mutate adenine using a nucleotide analog, one of ordinary skill in the art can amplify a nucleic acid molecule of a known sequence using PCR in the presence of the nucleotide analog and sequence the resulting amplified nucleic acid molecule. If the amplified DNA has a mutation at a position corresponding to an adenine nucleotide in the known sequence, one of ordinary skill in the art will understand that the DNA polymerase can mutate adenine using the nucleotide analog.
[0069] The rate ratio can be calculated in a similar manner. For example, if one of ordinary skill in the art wishes to determine the rate ratio at which guanine and cytosine nucleotides mutate, one of ordinary skill in the art can amplify a nucleic acid molecule having a known sequence using PCR in the presence of a low-bias DNA polymerase. One of ordinary skill in the art can then sequence the resulting amplified nucleic acid molecule and determine how many guanine nucleotides have been substituted or deleted and how many cytosine nucleotides have been substituted or deleted. The rate ratio is the ratio of the number of guanine nucleotides that have been substituted or deleted to the number of cytosine nucleotides that have been substituted or deleted. For example, if 16 guanine nucleotides have been replaced or deleted and 8 cytosine nucleotides have been replaced or deleted, the guanine and cytosine nucleotides are mutating at a rate ratio of 16:8 or 2:1, respectively.
[0070] Using nucleotide analogs Low-bias DNA polymerases may not be able to directly replace nucleotides with other nucleotides (at least not with high frequency), but they may still be able to mutate nucleic acid molecules if nucleotide analogs are used. Low-bias DNA polymerases may be able to replace nucleotides with other natural nucleotides (i.e., cytosine, guanine, adenine or thymine) or nucleotide analogs.
[0071] For example, the low-bias DNA polymerase may be a high-fidelity DNA polymerase. High-fidelity DNA polymerases generally tend not to introduce mutations because they are highly accurate. However, the inventors have found that some high-fidelity DNA polymerases may still be able to mutate target nucleic acid molecules because they may be able to introduce nucleotide analogs into the target nucleic acid molecule.
[0072] In certain embodiments, in the absence of nucleotide analogs, the high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, from 0% to 0.0015%, or from 0% to 0.001% mutations per round of replication.
[0073] In certain embodiments, the low-bias DNA polymerase can incorporate nucleotide analogs into the at least one target nucleic acid molecule. In certain embodiments, the low-bias DNA polymerase incorporates nucleotide analogs into the at least one target nucleic acid molecule. In certain embodiments, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and / or cytosine using nucleotide analogs. In certain embodiments, the low-bias DNA polymerase uses nucleotide analogs to mutate adenine, thymine, guanine, and / or cytosine in the at least one target nucleic acid molecule. In certain embodiments, the DNA polymerase replaces guanine, cytosine, adenine, and / or thymine with nucleotide analogs. In certain embodiments, the DNA polymerase can replace guanine, cytosine, adenine, and / or thymine with nucleotide analogs.
[0074] The incorporation of nucleotide analogs into the at least one target nucleic acid molecule can be used to mutate nucleotides because the nucleotide analog may be incorporated in place of an existing nucleotide and the nucleotide analog may base pair with a nucleotide in the opposite strand. For example, dPTP can be incorporated into a nucleic acid molecule in place of a pyrimidine nucleotide (which may replace thymine or cytosine) (see Figure 7). Once incorporated into a nucleic acid strand, dPTP can base pair with adenine when it is in the imino tautomeric form. Thus, when a complementary strand is formed, that complementary strand may have an adenine at the position complementary to dPTP. Similarly, once incorporated into a nucleic acid strand, dPTP can base pair with guanine when it is in the amino tautomeric form. Thus, when a complementary strand is formed, that complementary strand may have a guanine at the position complementary to dPTP.
[0075] For example, when introducing dPTP into the at least one target nucleic acid molecule of the present invention, when at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule is formed, the at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule contains adenine or guanine at a position complementary to dPTP in the at least one target nucleic acid molecule (depending on whether dPTP is in its amino form or imino form). When replicating the at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule, the resulting replica of the at least one target nucleic acid molecule contains thymine or cytosine at a position corresponding to dPTP in the at least one target nucleic acid molecule. Therefore, a mutation to thymine or cytosine can be introduced into the at least one target nucleic acid molecule to be mutated.
[0076] Alternatively, when introducing dPTP into at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule, when a replica of the at least one target nucleic acid molecule is formed, the replica of the at least one target nucleic acid molecule contains adenine or guanine at a position complementary to dPTP in the at least one nucleic acid molecule complementary to the at least one target nucleic acid molecule (depending on the tautomeric form of dPTP). Therefore, a mutation to adenine or guanine can be introduced into the at least one target nucleic acid molecule that has been mutated.
[0077] In certain embodiments, the low-bias DNA polymerase can replace cytosine or thymine with nucleotide analogs. In further embodiments, the low-bias DNA polymerase uses nucleotide analogs to introduce guanine or adenine nucleotides at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively. The guanine or adenine nucleotides may be introduced by a low-bias DNA polymerase that pairs them with nucleotide analogs such as dPTP. In further embodiments, the low-bias DNA polymerase uses nucleotide analogs to introduce guanine or adenine nucleotides at a ratio of 0.7-1.3:0.7-1.3, respectively.
[0078] One of ordinary skill in the art can determine, using conventional methods, whether a low-bias DNA polymerase can incorporate nucleotide analogs into the at least one target nucleic acid molecule, or whether the adenine, thymine, guanine, and / or cytosine in the at least one target nucleic acid molecule can be mutated using conventional methods and nucleotide analogs.
[0079] For example, to determine whether a low-bias DNA polymerase can incorporate nucleotide analogs into at least one target nucleic acid molecule, a person skilled in the art can use the low-bias DNA polymerase to amplify a nucleic acid molecule in two rounds of replication. The first round of replication should be performed in the presence of nucleotide analogs, and the second round of replication should be performed in the absence of nucleotide analogs. The resulting amplified nucleic acid molecule can be sequenced to determine whether mutations have been introduced, and if so, how many mutations have been introduced. The user should repeat this experiment without using nucleotide analogs and compare the number of mutations introduced with and without using nucleotide analogs. If the number of mutations introduced with nucleotide analogs is significantly higher than the number of mutations introduced without using nucleotide analogs, the user can conclude that the low-bias DNA polymerase can incorporate nucleotide analogs. Similarly, a person skilled in the art can determine whether a DNA polymerase incorporates nucleotide analogs or uses nucleotide analogs to mutate adenine, thymine, guanine, and / or cytosine. One skilled in the art need only carry out the method in the presence of nucleotide analogues and determine whether the method results in mutations at positions originally occupied by adenine, thymine, guanine, and / or cytosine.
[0080] If the user desires to mutate the at least one target nucleic acid molecule using nucleotide analogues, the method may comprise amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase, wherein amplifying the at least one target nucleic acid molecule using the low-bias DNA polymerase is carried out in the presence of nucleotide analogues, and wherein amplifying the at least one target nucleic acid molecule provides at least one target nucleic acid molecule comprising nucleotide analogues.
[0081] Suitable nucleotide analogs include dPTP (2'-deoxy-P-nucleoside-5'-triphosphate), 8-oxo-dGTP (7,8-dihydro-8-oxoguanine), 5Br-dUTP (5-bromo-2'-deoxy-uridine-5'-triphosphate), 2OH-dATP (2-hydroxy-2'-deoxyadenosine-5'-triphosphate), dKTP (9-(2-deoxy-β-D-ribofuranosyl)-N6-methoxy-2,6,-diaminopurine-5'-triphosphate) and dITP (2'-deoxyinosine 5'-triphosphate). The nucleotide analog may be dPTP. By using the nucleotide analog, substitution mutations described in Table 1 may be introduced.
[0082]
Table 1
[0083] By using various nucleotide analogs alone or in combination, various mutations can be introduced into the at least one target nucleic acid molecule. Therefore, the low-bias DNA polymerase may introduce substitution mutations from guanine to adenine, from cytosine to thymine, from adenine to guanine, and from thymine to cytosine using nucleotide analogs. The low-bias DNA polymerase may, depending on the situation, be able to introduce substitution mutations from guanine to adenine, from cytosine to thymine, from adenine to guanine, and from thymine to cytosine using nucleotide analogs.
[0084] Low-bias DNA polymerases may be able to introduce guanine-to-adenine, cytosine-to-thymine, adenine-to-guanine, and thymine-to-cytosine substitution mutations at ratios of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or approximately 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase is capable of introducing guanine-to-adenine, cytosine-to-thymine, adenine-to-guanine, and thymine-to-cytosine substitution mutations at rate ratios of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively. Suitable methods for determining whether a low-bias DNA polymerase can introduce substitution mutations and at what rate ratios are described under the heading "Whether a DNA Polymerase Can Mutate Nucleotides, and If So, at What Rate?"
[0085] In some methods, low-bias DNA polymerases introduce guanine-to-adenine substitution mutations, cytosine-to-thymine substitution mutations, adenine-to-guanine substitution mutations, and thymine-to-cytosine substitution mutations at ratios of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase introduces guanine-to-adenine substitution mutations, cytosine-to-thymine substitution mutations, adenine-to-guanine substitution mutations, and thymine-to-cytosine substitution mutations at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively. Suitable methods for determining whether substitution mutations have been introduced and at what ratios are described under the heading "Can the DNA polymerase mutate nucleotides, and if so, at what rate?".
[0086] Generally, when a low-bias DNA polymerase uses nucleotide analogs to introduce mutations, this requires two or more rounds of replication. In the first round of replication, the low-bias DNA polymerase introduces a nucleotide analog instead of a nucleotide, and in a further second round of replication, the nucleotide analog pairs with a native nucleotide to introduce a substitution mutation into the complementary strand. The second round of replication may be performed in the presence of the nucleotide analog. However, this method may further include the step of amplifying the at least one target nucleic acid molecule containing the nucleotide analog in the absence of the nucleotide analog. The step of amplifying the at least one target nucleic acid molecule containing the nucleotide analog in the absence of the nucleotide analog may be performed using a low-bias DNA polymerase.
[0087] Depending on the situation, the method provides at least one mutated target nucleic acid molecule, and the method further includes the step of amplifying this at least one mutated target nucleic acid molecule using a low-bias DNA polymerase.
[0088] Low-template amplification bias The low-bias DNA polymerase may have a low-template amplification bias. The low-bias DNA polymerase has a low-template amplification bias if the low-bias DNA polymerase can amplify different target nucleic acid molecules with a similar degree of success per cycle. High-bias DNA polymerases may struggle to amplify template nucleic acid molecules with a high G:C content or a high degree of secondary structure. In certain embodiments, the low-bias DNA polymerase of the present invention has a low-template amplification bias for template nucleic acid molecules that are less than 25,000, less than 10,000, 1 to 15,000, or 1 to 10,000 nucleotides in length.
[0089] In certain embodiments, to determine whether a DNA polymerase has a low-template amplification bias, one of ordinary skill in the art can amplify a wide range of different sequences using the DNA polymerase and sequence the resulting amplified DNA to examine whether the different sequences are amplified at different levels. For example, one of ordinary skill in the art can select a wide range of short (optionally 50 nucleotide) nucleic acid molecules with different characteristics, such as nucleic acid molecules showing a high GC content, nucleic acid molecules showing a low GC content, nucleic acid molecules with a high degree of secondary structure, and nucleic acid molecules with a low degree of secondary structure. The user can then amplify those sequences using the DNA polymerase and quantify the level at which each of the nucleic acid molecules is amplified. In certain embodiments, when the levels are within 25%, 20%, 10%, or 5% of each other, the DNA polymerase in question has a low-template amplification bias.
[0090] Alternatively, in certain embodiments, the DNA polymerase has low amplification bias if it can amplify a 7 - 10 kbp fragment (Kolmogorov - Smirnov D < 0.1, < 0.09, or < 0.08). The Kolmogorov - Smirnov D for a particular low - bias DNA polymerase to amplify a 7 - 10 kbp fragment may be determined using the assay described in Example 4.
[0091] The low - bias DNA polymerase may be a high - fidelity DNA polymerase. A high - fidelity DNA polymerase is not highly error - prone and thus, when used to amplify a target nucleic acid molecule in the absence of nucleotide analogs, is a DNA polymerase that usually does not introduce a large number of mutations. The reason that high - fidelity DNA polymerases are not normally used in methods for introducing mutations is that error - prone DNA polymerases are generally considered to be more effective. However, this application demonstrates that certain high - fidelity polymerases can introduce mutations using nucleotide analogs and that those mutations may be introduced with a lower bias compared to error - prone polymerases such as DNATaq polymerase.
[0092] High - fidelity DNA polymerases have additional advantages. By using a high - fidelity DNA polymerase, mutations can be introduced when used with nucleotide analogs, but in the absence of nucleotide analogs, the high - fidelity DNA polymerase can replicate the target nucleic acid molecule very accurately. This means that the user can very effectively mutate the at least one target nucleic acid molecule and then amplify this mutated at least one target nucleic acid molecule with high accuracy using the same DNA polymerase. When mutating a target nucleic acid molecule by using a low - fidelity DNA polymerase, it may be necessary to remove the low - fidelity DNA polymerase from the reaction mixture before amplifying the target nucleic acid molecule.
[0093] The high-fidelity DNA polymerase may have proofreading activity. The proofreading activity may help the DNA polymerase amplify the target nucleic acid sequence with high accuracy. For example, the low-bias DNA polymerase may contain a proofreading domain. The proofreading domain may check whether the nucleotide added by the polymerase is correct (confirming that the nucleotide forms an exact pair with the corresponding nucleic acid of the complementary strand), and if not, remove the nucleotide from the nucleic acid molecule. The inventors have surprisingly found that in some DNA polymerases, the proofreading domain tolerates the pairing of natural nucleotides with nucleotide analogs. Appropriate proofreading domain structures and sequences are known to those skilled in the art. DNA polymerases containing a proofreading domain include members of DNA polymerase families I, II, and III, such as Pfu polymerase (derived from Pyrococcus furiosus), T4 polymerase (derived from bacteriophage T4), and the Thermococcal polymerases described in more detail below.
[0094] In certain embodiments, in the absence of nucleotide analogs, the high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, from 0% to 0.0015%, or from 0% to 0.001% mutations per round of replication.
[0095] In addition, the low-bias DNA polymerase may contain a processivity-enhancing domain. The processivity-enhancing domain enables the DNA polymerase to amplify the target nucleic acid molecule more rapidly. This is advantageous as it enables the methods of the present invention to be carried out more rapidly.
[0096] Thermococcal polymerase In one embodiment, the low-bias DNA polymerase is a fragment or variant of a polypeptide comprising SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, or SEQ ID NO:7. The polypeptides of SEQ ID NOs:2, 4, 6, and 7 are polymerases from Thermococcales archaea. The polymerases of SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, or SEQ ID NO:7 are low-bias DNA polymerases that exhibit high fidelity and are capable of mutating target nucleic acid molecules by incorporating nucleotide analogs such as dPTP. The polymerases of SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, or SEQ ID NO:7 are particularly advantageous because they have low mutation bias and low template amplification bias. They are also highly processive and high-fidelity polymerases that contain a proofreading domain, meaning that they can rapidly and accurately amplify mutated target nucleic acid molecules in the absence of nucleotide analogs.
[0097] Low-bias DNA polymerases are as follows: a. The sequence of SEQ ID NO: 2; b. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:2; c. The sequence of SEQ ID NO: 4; d. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:4; e. the sequence of SEQ ID NO: 6; f. a sequence at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:6; g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 7 The amino acid sequence may comprise a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of the amino acid sequence.
[0098] Preferably, the low-bias DNA polymerase is one of the following a to h: a. The sequence of SEQ ID NO: 2; b. A sequence that is at least 98% or at least 99% identical to SEQ ID NO: 2, c. The sequence of SEQ ID NO: 4, d. A sequence that is at least 98% or at least 99% identical to SEQ ID NO: 4, e. The sequence of SEQ ID NO: 6, f. A sequence that is at least 98% or at least 99% identical to SEQ ID NO: 6, g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 98% or at least 99% identical to SEQ ID NO: 7 and contains at least 700 consecutive amino acid fragments.
[0099] The low-bias DNA polymerase is as follows a. to h., namely a. The sequence of SEQ ID NO: 2, b. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 2, c. The sequence of SEQ ID NO: 4, d. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 4, e. The sequence of SEQ ID NO: 6, f. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 6, g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 7 and may contain.
[0100] Preferably, the low-bias DNA polymerase is as follows a. to h., namely a. The sequence of SEQ ID NO: 2, b. A sequence that is at least 98% or at least 99% identical to SEQ ID NO: 2, c. The sequence of SEQ ID NO: 4, d. A sequence that is at least 98% or at least 99% identical to SEQ ID NO: 4, e. The sequence of SEQ ID NO: 6, f. A sequence that is at least 98% or at least 99% identical to SEQ ID NO: 6, g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 98% or at least 99% identical to SEQ ID NO: 7 comprising.
[0101] The low-bias DNA polymerase may be a polymerase of an archaeon of the order Thermococcales or a derivative thereof. The DNA polymerases of SEQ ID NOs: 2, 4, 6, and 7 are polymerases of archaea of the order Thermococcales. The polymerases of archaea of the order Thermococcales are advantageous because they are generally high-fidelity polymerases that can introduce mutations using nucleotide analogs with low mutation and template amplification bias.
[0102] The polymerase of an archaeon of the order Thermococcales is a polymerase having the polypeptide sequence of a polymerase isolated from a strain of the genus Thermococcus. A derivative of the polymerase of an archaeon of the order Thermococcales may be a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of the polymerase of an archaeon of the order Thermococcales, or may be at least 95%, at least 98%, at least 99%, or 100% identical to a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of the polymerase of an archaeon of the order Thermococcales. A derivative of the polymerase of an archaeon of the order Thermococcales may be at least 95%, at least 98%, at least 99%, or 100% identical to the polymerase of an archaeon of the order Thermococcales. A derivative of the polymerase of an archaeon of the order Thermococcales may be at least 98% identical to the polymerase of an archaeon of the order Thermococcales.
[0103] Thermococcal archaeal polymerases from any strain may be useful in the context of the present invention. In one embodiment, the Thermococcal archaeal polymerase is derived from a Thermococcales strain selected from the group consisting of T. kodakarensis, T. celer, T. siculi, and T. sp. KS-1. The Thermococcales archaeal polymerases from these strains are set forth as SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, and SEQ ID NO:7.
[0104] Depending on the context, a low-bias DNA polymerase is one that exhibits high catalytic activity at temperatures between 50°C and 90°C, between 60°C and 80°C, or at about 68°C.
[0105] Barcodes, Sample Tags and Adapters The method may further comprise introducing a barcode into the target nucleic acid molecule. For purposes of the present invention, a barcode is a degenerate or randomly generated nucleotide sequence. The term "barcode" is synonymous with the terms "unique molecular identifier" (UMI) or "unique molecular tag" (UMT). The method may comprise introducing one, two, or more barcodes into the target nucleic acid molecule. In a preferred embodiment, the method comprises introducing various barcodes into the target nucleic acid molecule, such that, after the barcodes are introduced, a majority of the original target nucleic acid molecules contain unique barcodes compared to other original target nucleic acid molecules.
[0106] The introduction of barcodes into target nucleic acid molecules can be useful when the method for introducing mutations of the present invention is utilized as part of a method for determining sequences. The use of barcodes can help the user identify from which of the original at least one target nucleic acid molecules each sequence of at least one target nucleic acid molecule (or at least one amplified or fragmented target nucleic acid molecule) is derived. If different barcodes are used for each of the original target nucleic acid molecules, the user can sequence the barcode or the target nucleic acid molecule, and the sequences of the target nucleic acid molecules containing the same barcode are likely to be the sequences of the target nucleic acid molecules generated from the same original target nucleic acid molecule.
[0107] A method for introducing a mutation into at least one target nucleic acid molecule may include introducing a sample tag into the target nucleic acid molecule. The sample tag is a series of short nucleic acids having a known (specific) sequence. For example, the method of the present invention may be performed on a plurality of target nucleic acid molecules collected from different samples. Those samples may be pooled, but the sample tag may be introduced into the target nucleic acid molecule in the sample (labeling the target nucleic acid molecule with the sample tag) before pooling. Target nucleic acid molecules obtained from different samples may be labeled with different sample tags. Depending on the situation, target nucleic acid molecules obtained from the same sample may be labeled with the same sample tag, or a sample tag obtained from the same subgroup of sample tags. For example, if the user decides to use two samples, the target nucleic acid molecules in the first sample may be labeled with a first sample tag having a specific sequence, and the target nucleic acid molecules in the second sample may be tagged with a second sample tag having a second specific sequence. Similarly, if the user decides to use two samples, the target nucleic acid molecules in the first sample may be labeled with a sample tag obtained from a first subgroup of sample tags, and the target nucleic acid molecules in the second sample may be labeled with a sample tag obtained from a second subgroup of sample tags. The user will understand that any target nucleic acid molecule containing the first sample tag or a sample tag obtained from the first subgroup of sample tags originated from the first sample, and any target nucleic acid molecule containing the second sample tag or a sample tag obtained from the second subgroup of sample tags originated from the second sample. Which tag was used to label the target nucleic acid sequence can be determined by sequencing the target nucleic acid sequence. Appropriate sequencing methods are described in more detail below.
[0108] In certain embodiments, prior to the step of amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase, a sample tag is introduced (the target nucleic acid molecule is labeled with the sample tag). This is advantageous because it can pool samples at an early stage of the method, reducing the processing time, the number of reagents required, and the potential for errors during sample processing. However, if a sample tag is introduced prior to the step of amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase, the sample tag may be mutated by the low-bias DNA polymerase. The inventors have designed a group of sample tags that are designed to be distinguishable from each other even if they are mutated.
[0109] In certain embodiments, a group of sample tags is used to label target nucleic acid molecules from different samples with different sample tags from the group. Target nucleic acid molecules obtained from the same sample may be labeled with the same sample tag from the group of sample tags or with sample tags from the same subgroup of sample tags from the group of sample tags. For example, if the group of sample tags includes sample tags named A, B, C, and D, all target nucleic acid molecules in the first sample may be labeled using A or A / B, and all target nucleic acid molecules in the second sample may be labeled using C or C / D. Each sample tag in the group of sample tags may differ from substantially all of the other sample tags in the group by at least one low-probability mutation difference. Each sample tag in the group of sample tags may differ from all of the other sample tags in the group by at least one low-probability mutation difference.
[0110] In one aspect, the present invention provides a group of sample tags, where each sample tag in the group is different from substantially all of the other sample tags in the group by at least one low-probability mutation difference. Each sample tag may be different from all of the other sample tags in the group by at least one low-probability mutation difference.
[0111] By the term "differing by at least one low-probability mutation from substantially all other sample tags of the group", the inventors mean that each tag is designed such that, when the sample tag is mutated by at least one low-probability mutation, the tags still differ from each other by a large degree (substantially all, or all other tags). In certain embodiments, the term "substantially all other sample tags" refers to at least 90%, at least 95%, or at least 98% of the other sample tags. A low-probability mutation is a mutation that rarely occurs in the methods for introducing mutations of the present invention. For example, a low-probability mutation may be a transversion mutation or an indel mutation. Transversion mutations and indel mutations rarely occur when the method for introducing mutations of the present invention is carried out using dPTP as a nucleotide analog. A transversion mutation is the replacement of a purine nucleotide with a pyrimidine nucleotide (adenine to cytosine, adenine to thymine, guanine to cytosine, or guanine to thymine), or the replacement of a pyrimidine nucleotide with a purine nucleotide (cytosine to adenine, cytosine to guanine, thymine to adenine, or thymine to guanine). An indel mutation is a deletion mutation or an insertion mutation. Suitable tags may be designed by a computer using statistical techniques. For example, one skilled in the art would be able to determine which type of mutation is a low-probability mutation in the methods for introducing mutations of the present invention. One skilled in the art would be able to determine the type of mutation introduced by carrying out the method for introducing mutations of the present invention and sequencing the nucleic acid molecular product. The most frequently occurring mutations are high-probability mutations, and the mutations that occur with the lowest frequency are low-probability mutations.
[0112] A user can create a suitable sample tag by utilizing the method for designing a group of sample tags of the present invention.
[0113] Depending on the situation, each sample tag differs from substantially all other sample tags in the sample tag group by at least 2, at least 3, at least 4, at least 5, 3 - 50, 3 - 25, or 3 - 10 low - probability mutation differences. Depending on the situation, each sample tag differs from all other sample tags in the sample tag group by at least 2, at least 3, at least 4, at least 5, 3 - 50, 3 - 25, or 3 - 10 low - probability mutation differences.
[0114] Each sample tag may differ from substantially all other sample tags in the sample tag group by at least 2 high - probability mutation differences. High - probability mutation differences are mutations that frequently occur in the method for introducing mutations of the present invention. For example, the high - probability mutation difference may be a transition mutation. A transition mutation is the replacement of a purine nucleotide with another purine nucleotide (adenine to guanine or guanine to adenine), or the replacement of a pyrimidine nucleotide with another pyrimidine nucleotide (cytosine to thymine or thymine to cytosine).
[0115] Each sample tag may differ from all other sample tags in the sample tag group by at least 2 high - probability mutation differences, that is, each sample tag is designed such that even if the tag mutates by at least 2 high - probability mutations, the tags still differ from each other.
[0116] Depending on the situation, each sample tag differs from substantially all other sample tags in the sample tag group by at least 3, 2 - 50, 3 - 25, or 3 - 10 high - probability mutation differences. Depending on the situation, each sample tag differs from all other sample tags in the sample tag group by at least 3, 2 - 50, 5 - 25, or 5 - 10 high - probability mutation differences.
[0117] In one embodiment, each sample tag is at least 8 nucleotides in length, at least 10 nucleotides in length, at least 12 nucleotides in length, 8 - 50 nucleotides in length, 10 - 50 nucleotides in length, or 10 - 50 nucleotides in length.
[0118] Suitable sample tags are those with SEQ ID NOs: 8 to 136.
[0119] The method may further include introducing an adapter to each of the target nucleic acid molecules. The adapter may contain a primer binding site. For the purposes of the present invention, the primer binding site is a known sequence of nucleotides that is long enough for the primer to specifically hybridize. Depending on the situation, the primer binding site is at least 8, at least 10, at least 12, 8 to 50, or 10 to 25 nucleotides in length.
[0120] The method may include introducing a first adapter to the 3'-end of the at least one target nucleic acid molecule and introducing a second adapter to the 5'-end of the at least one target nucleic acid molecule, where the first adapter and the second adapter are capable of annealing to each other.
[0121] In one aspect, the present invention provides a method for selectively amplifying nucleic acid molecules larger than 1 kbp in length, comprising: a. providing at least one sample containing a target nucleic acid molecule; b. introducing a first adapter to the 3'-end of the target nucleic acid molecule and a second adapter to the 5'-end of the target nucleic acid molecule; and c. amplifying the target nucleic acid molecule using a primer that is complementary to a part of the first adapter, where the first adapter and the second adapter are capable of annealing to each other.
[0122] The second adapter may contain a part that is complementary to the first primer binding site, and the first adapter may contain the first primer binding site.
[0123] The inventors have found that by introducing a first adapter and a second adapter that can anneal to each other to the at least one target nucleic acid molecule, their method of the present invention can selectively amplify and / or mutate long target nucleic acid molecules. When the first adapter can anneal to the second adapter, they may then (as shown in Figure 5) result in at least one self-annealed target nucleic acid molecule in the method of the present invention. Since the self-annealed target nucleic acid molecule is not replicated, it is not amplified and / or mutated by the method of the present invention. The likelihood that the first adapter and the second adapter will anneal to each other during the method of the present invention is higher for shorter nucleic acid molecules than for longer target nucleic acid molecules. For these reasons, the addition of the first adapter and the second adapter to the at least one target nucleic acid molecule of the present invention can be utilized to selectively amplify a larger at least one target nucleic acid molecule.
[0124] A method for selectively amplifying a nucleic acid molecule can be a method for selectively amplifying a target nucleic acid molecule longer than 1.5 kbp. The method may further include a step of sequencing the target nucleic acid molecule. Specific examples of conceivable sequencing methods include sequencing methods including the Maxam-Gilbert method, the Sanger method, nanopore sequencing, or bridge PCR. In a typical embodiment, the sequencing step involves bridge PCR. Depending on the situation, the bridge PCR step is performed using an extension time of more than 5 seconds, more than 10 seconds, more than 15 seconds, or more than 20 seconds. An example of the use of bridge PCR is the Illumina Genome Analyzer Sequencers.
[0125] The user can determine whether the first adapter and the second adapter can anneal to each other. In certain embodiments, the user can confirm whether the first adapter and the second adapter can anneal to each other by preparing a nucleic acid molecule containing the first adapter and further examining whether a primer containing the second adapter can initiate replication of the nucleic acid molecule under PCR conditions.
[0126] Alternatively, in certain embodiments, the first adapter and the second adapter can be considered to anneal to each other if they hybridize under the following conditions: mixing two primers at equimolar concentrations (e.g., 50 μM) and then incubating at a high temperature such as 95° C. for 5 minutes to ensure that the primers are single-stranded. This solution is then slowly cooled to room temperature (25° C.) over about 45 minutes.
[0127] The method may include amplifying a target nucleic acid molecule using primers that are identical to each other or substantially identical to each other. The primer may be complementary to a part of the first adapter. Two primers are "substantially identical" to each other if they have the same sequence or a sequence that differs by 1, 2, or 3 nucleotides. In a preferred embodiment, the method of the present invention includes amplifying a target nucleic acid molecule using primers that have the same sequence or differ by a single nucleotide difference.
[0128] In certain embodiments, the first adapter and the second adapter include sequences that are complementary to each other or substantially complementary to each other. The first adapter may be substantially complementary to the second adapter if the first adapter is complementary to a nucleic acid molecule that is at least 80%, at least 90%, at least 95%, or at least 99% identical to the second adapter.
[0129] The user may use primers that contain a primer binding site, and by using these primers, the replicas of the at least one target nucleic acid molecule generated in the final round of replication may be selectively amplified. For example, a first set of primers containing a third primer binding site may be used for one round of replication. In a further round of replication, a second set of primers that bind to the third primer binding site may be used. The second set of primers only replicates the replicas of the at least one target nucleic acid molecule generated in the previous round of replication using the first set of primers.
[0130] The third and further sets of primers may be used. Selectively replicating the replicas of the previous round of replication is advantageous because it can ensure that each amplified target nucleic acid molecule contains a high level of mutations (since only at least one target nucleic acid molecule that has undergone at least one amplification by a low-bias DNA polymerase is replicated).
[0131] Therefore, the method of the present invention may include the following (a) to (c), that is, (a) introducing a first adapter containing a first primer binding site at the 3'-end of the at least one target nucleic acid molecule or a plurality of target nucleic acid molecules, and introducing a second adapter containing a portion complementary to the first primer binding site at the 5'-end of the at least one target nucleic acid molecule or a plurality of target nucleic acid molecules, where the first adapter and the second adapter are capable of annealing to each other; (b) using a first set of primers that are complementary to the first primer binding site and contain a second primer binding site, and using a low-bias DNA polymerase as appropriate to amplify the target nucleic acid molecule; and (c) using a second set of primers that are complementary to the second primer binding site, and using a low-bias DNA polymerase as appropriate to amplify the target nucleic acid molecule. It may be included.
[0132] The second set of primers may include a third primer binding site, and further amplification steps may be performed using a third or further set of primers that are complementary to the third or further primer binding site.
[0133] Barcodes, sample tags and / or adapters may be introduced using any suitable method, for example, by combining PCR, tagging and physical shearing or restriction digestion of the target nucleic acid followed by adapter ligation (which may be blunt-end ligation). For example, PCR can be performed using a first set of primers capable of hybridizing to the at least one target nucleic acid molecule against at least one target template nucleic acid molecule. Barcodes, sample tags and adapters may be introduced into each of the at least one target nucleic acid molecule by PCR using primers comprising a portion (5' end portion) comprising the barcode, sample tag and / or adapter, and a portion (3' end portion) having a sequence capable of hybridizing (which may be complementary) to the at least one target nucleic acid molecule. Such primers hybridize to the target nucleic acid molecule and then, by PCR primer extension, a nucleic acid molecule comprising the barcode, sample tag and / or adapter is provided. By utilizing further cycles of PCR with these primers, the barcode, sample tag and / or adapter can be added to the other end of the at least one target nucleic acid molecule. The primers may be degenerate, i.e., the 3' end portions of the primers may be similar to each other but may not be identical.
[0134] Barcodes, sample tags, and / or adapters may be introduced using tagmentation. Barcodes, sample tags, and / or adapters may be introduced using direct tagmentation or by two cycles of PCR using a primer that includes a defined sequence introduced by tagmentation, a portion that hybridizes to the defined sequence, and a portion that includes the barcode, sample tag, and / or adapter. Barcodes, sample tags, and / or adapters may be introduced by restriction digestion of the at least one original target nucleic acid molecule followed by ligation of a nucleic acid that includes the barcode, sample tag, and / or adapter. The restriction digestion of the at least one original nucleic acid molecule should be performed such that a nucleic acid molecule (at least one target template nucleic acid molecule) that includes the region to be sequenced by the digestion results. Barcodes, sample tags, and / or adapters may also be introduced by shearing of the at least one target nucleic acid molecule followed by end repair, A-tailing, and subsequent ligation of a nucleic acid that includes the barcode, sample tag, and / or adapter.
[0135] Method for determining the sequence of at least one target nucleic acid molecule One aspect of the invention relates to a method for determining the sequence of at least one target nucleic acid molecule, which includes a method for introducing a mutation of the invention.
[0136] As described above, the method for introducing a mutation of the invention may be useful as part of a method for determining the sequence of at least one target nucleic acid molecule because the mutation may enable a person skilled in the art to assemble the sequence.
[0137] As described in the Background section, the sequencing method can be improved by incorporating a step of introducing mutations into at least one target nucleic acid molecule to be sequenced. The user often amplifies and / or fragments the at least one target nucleic acid molecule before sequencing the at least one target nucleic acid molecule. The user then assembles a consensus sequence for at least one of the target nucleic acid molecules from the sequences of regions of the at least one amplified or fragmented target nucleic acid molecule. Introducing mutations into the at least one target nucleic acid molecule prior to amplification or fragmentation helps the user identify which of the original at least one template nucleic acid molecules each sequence of the regions of the at least one amplified or fragmented target nucleic acid molecule is derived from, and thus may help improve the accuracy of the consensus sequence.
[0138] The more random the introduced mutations are, the easier it becomes to identify which of the original at least one target nucleic acid molecules each sequence of the at least one amplified or fragmented target nucleic acid molecule is derived from. The method for introducing mutations of the present invention uses a low-bias DNA polymerase, and by utilizing this method, mutations can be introduced in a substantially random manner, making it ideal for inclusion in a method for determining the sequence of at least one target nucleic acid molecule.
[0139] A method for determining the sequence of at least one target nucleic acid molecule may include the following steps a. to c., namely a. preparing at least one mutated target nucleic acid molecule by performing the method for introducing mutations into at least one target nucleic acid molecule of the present invention, b. preparing mutated sequence reads by sequencing regions of the at least one mutated target nucleic acid molecule, and c. assembling a sequence for at least a part of the at least one target nucleic acid molecule using the mutated sequence reads may be included.
[0140] Generally, the sequencing step can be performed using any sequencing method. Specific examples of conceivable sequencing methods include the Maxam-Gilbert method, the Sanger method, nanopore sequencing, or a sequencing method including bridge PCR. In a typical embodiment, the sequencing step involves bridge PCR. Depending on the situation, the bridge PCR step is performed using an extension time of more than 5 seconds, more than 10 seconds, more than 15 seconds, or more than 20 seconds. An example of the use of bridge PCR is the Illumina Genome Analyzer Sequencers.
[0141] The method may include preparing mutant sequence reads by sequencing at least one region of the mutant target nucleic acid molecule. The region may correspond to a fragment that may include a substantial portion of at least one mutant target nucleic acid molecule. For some reason, it may not be possible to sequence all of the at least one mutant target nucleic acid molecule, but the user may still find that the sequence of a part of the at least one mutant target nucleic acid molecule is useful. The region of the at least one mutant target nucleic acid molecule may include the entire length of the at least one mutant target nucleic acid molecule.
[0142] The method may include assembling a sequence for at least a part of at least one target nucleic acid molecule from the mutant sequence reads. The sequence may be assembled by aligning the mutant sequence reads and grouping together the reads that share the same mutation pattern. The sequence is assembled from the mutant sequence reads within the same group. The assembly may be performed using software such as Clustal W2, IDBA-UD, or SOAPdenovo.
[0143] A method for determining the sequence of at least one target nucleic acid molecule includes the following steps a. to d., namely a. A step of preparing at least one mutant target nucleic acid molecule by implementing a method for introducing a mutation into at least one target nucleic acid molecule of the present invention, b. preparing at least one fragmented and / or amplified mutated target nucleic acid molecule by fragmenting and / or amplifying the at least one mutated target nucleic acid molecule; c. preparing mutated sequence reads by sequencing regions of the at least one fragmented and / or amplified mutated target nucleic acid molecule; and d. assembling a sequence for at least a portion of the at least one target nucleic acid molecule using the mutated sequence reads may be included.
[0144] The step of amplifying the at least one mutated target nucleic acid molecule can be carried out by any suitable amplification technique such as PCR. Optionally, PCR is performed using a low-bias DNA polymerase under conditions such as those described under the title "amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase".
[0145] The step of fragmenting the at least one mutated target nucleic acid molecule can be carried out using any suitable method. For example, fragmentation can be carried out using restriction digestion or PCR using primers complementary to at least one internal region of the at least one mutated target nucleic acid molecule. Preferably, fragmentation is carried out using a technique that generates random fragments. The term "random fragment" refers to a fragment created randomly, for example, a fragment created by tagmentation. Fragments created using restriction enzymes are not "random" because restriction digestion occurs at specific DNA sequences defined by the restriction enzyme used. Even more preferably, fragmentation is carried out by tagmentation. When fragmentation is carried out by tagmentation, an adapter region is introduced into the at least one mutated target nucleic acid molecule depending on the situation by the tagmentation reaction. This adapter region may be a short DNA sequence that encodes, for example, an adapter that allows the at least one mutated target nucleic acid molecule to be sequenced using Illumina technology.
[0146] The fragmentation step may include a further step of enriching the at least one mutated fragmented target nucleic acid molecule. The step of enriching the at least one mutated fragmented target nucleic acid molecule may be carried out by PCR. Optionally, PCR is carried out using a low-bias DNA polymerase under conditions such as those described under the heading "amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase".
[0147] Method for manipulating a protein The method for introducing a mutation of the present invention may be useful as part of a method for manipulating a protein. For example, protein manipulation may involve searching for mutations that increase or decrease the activity of a protein, or that alter its structure. As part of protein manipulation, a user may wish to randomly mutate a protein to examine how the mutation affects the activity or structure of the protein. The method of the present invention results in very random mutagenesis and, therefore, is a method that can be advantageously used as part of a method for manipulating a protein.
[0148] Accordingly, in one aspect of the present invention, there is provided a method for manipulating a protein, which includes the method for introducing a mutation of the present invention.
[0149] The method may include the following steps a. to c., namely a. a step of preparing at least one mutated target nucleic acid molecule by carrying out the method for introducing a mutation of the present invention, b. a step of inserting the at least one mutated target nucleic acid molecule into a vector, and c. a step of expressing the protein encoded by the at least one mutated target nucleic acid molecule and may be included.
[0150] The method may include the following steps a. to e., namely a. Preparing at least one mutated target nucleic acid molecule by implementing a method for introducing mutations of the present invention; b. Preparing a target nucleic acid molecule containing a nucleotide analog by amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase in the presence of the nucleotide analog; c. Preparing at least one mutated target nucleic acid molecule by amplifying the target nucleic acid molecule containing the nucleotide analog in the absence of the nucleotide analog; d. Inserting the at least one mutated target nucleic acid molecule into a vector, and e. Expressing the protein encoded by the at least one mutated target nucleic acid molecule may be included.
[0151] Any suitable vector can be used. Depending on the situation, the vector is a plasmid, virus, cosmid or artificial chromosome. Typically, the vector further includes control sequences operably linked to the inserted sequence, resulting in the ability to express the polypeptide. Preferably, the vector of the present invention further includes appropriate initiators, promoters, enhancers and other elements that may be required and are arranged in the correct orientation to enable the expression of the polypeptide.
[0152] Depending on the situation, the step of expressing the at least one mutated target nucleic acid molecule is achieved by transforming bacterial cells, transfecting eukaryotic cells or transducing eukaryotic cells using the vector. Depending on the situation, the bacterial cells are Escherichia coli (E. coli) cells.
[0153] For example, the step of expressing the at least one mutated target nucleic acid molecule may include inserting the at least one mutated target nucleic acid molecule into a plasmid vector and transforming Escherichia coli with the plasmid. The plasmid may contain control elements suitable for expression in Escherichia coli, such as the lac or T7 promoter (Dubendorff JW, Studier FW (1991). "Controlling basal expression in an inducible T7 expression system by blocking the target T7 promoter with lac repressor". Journal of Molecular Biology. 219 (1):45-59.)). Appropriate expression techniques are described in Sambrook, J. et al., (1989) Molecular Cloning:A Laboratory Manual Second Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York.
[0154] Alternatively, the step of expressing the at least one mutated target nucleic acid molecule may include expressing a fragment directly generated by the step of amplifying the target nucleic acid molecule using an in vitro method.
[0155] The method may further include the step of testing the activity of the protein encoded by the at least one mutated target nucleic acid molecule or the step of evaluating the structure of the protein.
[0156] The step of testing the activity of the protein encoded by the at least one mutated target nucleic acid molecule or the step of evaluating the structure of the protein can be carried out using any well-known technique. For example, a person skilled in the art would recognize appropriate techniques for evaluating the structure of a protein, such as nuclear magnetic resonance (NMR) techniques, microscopy techniques such as cryo-electron microscopy, small-angle X-ray scattering techniques, or X-ray crystallography.
[0157] Similarly, one of ordinary skill in the art would recognize techniques that can be used to evaluate the activity of a protein. The method used will depend on the protein encoded by the at least one mutated target nucleic acid molecule. For example, if the protein encoded by the at least one mutated target nucleic acid molecule is a blood coagulation factor, one of ordinary skill in the art would utilize, for example, a chromogenic clotting assay to test the protein for clotting activity. Alternatively, if the protein encoded by the at least one mutated target nucleic acid molecule is an enzyme, one of ordinary skill in the art can test the activity of the enzyme by measuring the rate at which the enzyme catalyzes its reaction, for example, by measuring the decrease in the concentration of the starting material or the increase in the concentration of the end product of the reaction catalyzed by the enzyme.
[0158] Method for designing a sample tag group In one aspect, the present invention is as follows a. and b., namely a. analyzing a method for introducing a mutation into at least one target nucleic acid molecule and determining the average number of low-probability mutations that occur during the method for introducing a mutation into the at least one target nucleic acid molecule, and b. determining the sequence for a group in which each sample tag differs from substantially all of the sample tags in the sample tag group by a low-probability mutation difference greater than the average number of low-probability mutations that occur during the method for introducing a mutation into the at least one target nucleic acid molecule The present invention further provides a method for designing a sample tag group suitable for use in a method for introducing a mutation into at least one target nucleic acid molecule, comprising the above.
[0159] For example, a user may create a first estimated sample tag using a computer program for creating a random array. Add the first estimated sample tag to the sample tag group. The user then creates a second estimated sample tag in the same way and further compares the sequence of the second estimated sample tag with that of the first estimated sample tag to check whether the second sample tag is different from the first sample tag such that even if a reasonable number of low-probability mutations are introduced into the second estimated sample tag, the second estimated sample tag still differs from the first estimated sample tag. If different, add the second estimated sample tag to the sample tag group. If not different, discard this second estimated sample tag. This may be repeated for the third and subsequent estimated sample tags.
[0160] As described above, with respect to sample tags, it is advantageous to add them to at least one target nucleic acid molecule in a method for introducing mutations into at least one target nucleic acid molecule. However, adding a sample tag before introducing a mutation may mean that the sample tag mutates and as a result cannot be used to identify target nucleic acid molecules arising from the same sample or different samples. This can be avoided by designing the sample tags such that they are sufficiently different from each other so that the user can identify them even if they mutate.
[0161] The method may further comprise the following a.(i) and (ii), namely a.(i) analyzing a method for introducing mutations into at least one target nucleic acid molecule and determining the average number of high-probability mutations occurring during the method for introducing mutations into at least one target nucleic acid molecule, and (ii) determining the sequences for a group in which each sample tag is different from substantially all of the sample tags in the sample tag group by a high-probability mutation difference greater than the average number of high-probability mutations occurring during the method for introducing mutations into at least one target nucleic acid molecule may further be included.
[0162] Low-probability mutations may be transversion mutations or indel mutations. High-probability mutations may be transition mutations.
[0163] The method may be a method implemented by a computer.
[0164] In a further aspect of the present invention, there is provided a computer-readable medium configured to implement a method for designing a set of sample tags suitable for use in a method for introducing mutations into at least one target nucleic acid molecule. In a further aspect of the present invention, there is provided a set of sample tags obtainable by the method for designing the sample tags of the present invention. Depending on the situation, the set of sample tags is obtained by the method for designing the sample tags of the present invention.
[0165] Using dNTPs at non-uniform concentrations The step of amplifying the at least one target nucleic acid using a low-bias DNA polymerase may be performed using dNTPs at non-uniform concentrations.
[0166] In one aspect of the present invention, the following a. and b., namely a. preparing at least one sample containing at least one target nucleic acid molecule, and b. introducing a mutation into the at least one target nucleic acid molecule by amplifying the at least one target nucleic acid molecule using a DNA polymerase to prepare at least one mutated target nucleic acid molecule, where step b. is to be performed using dNTPs at non-uniform concentrations, A method for introducing a mutation into at least one target nucleic acid molecule is provided, which includes the above steps.
[0167] To enable amplification of the at least one target nucleic acid using a DNA polymerase (e.g., a low-bias DNA polymerase), the target nucleic acid may be exposed to the DNA polymerase and dNTPs under conditions suitable for DNA replication, for example, within a PCR apparatus. When the step of amplifying the at least one target nucleic acid is performed using dNTPs at non-uniform concentrations, the target nucleic acid is exposed to a DNA polymerase (e.g., a low-bias DNA polymerase) and dNTPs, where the concentrations of the dNTPs are different relative to each other.
[0168] The term dNTP shall refer to deoxynucleotides. Specifically, in the context of the present application, the term "dNTP" shall refer to a solution containing dTTP (deoxythymidine triphosphate) or dUTP (deoxyuridine), dGTP (deoxyguanosine triphosphate), dCTP (deoxycytidine triphosphate), and dATP (deoxyadenosine triphosphate). Depending on the situation, "dNTP" shall refer to a solution containing dTTP (deoxythymidine triphosphate), dGTP (deoxyguanosine triphosphate), dCTP (deoxycytidine triphosphate), and dATP (deoxyadenosine triphosphate).
[0169] The expression "non-uniform concentrations of dNTPs" means that the four types of dNTPs are present in the solution at different concentrations relative to each other. For example, one type of dNTP may be present at a (higher) concentration compared to the other three types of dNTPs, two types of dNTPs may be present at a (higher) concentration compared to the other two types of dNTPs, or three types of dNTPs may be present at a (higher) concentration compared to the other one type of dNTP.
[0170] dGTP may be present at a (higher) concentration compared to dCTP, dTTP, and dATP, dGTP may be present at a (higher) concentration compared to dTTP and dATP, dGTP may be present at a (higher) concentration compared to dATP, dGTP may be present at a (higher) concentration compared to dTTP, dCTP may be present at a (higher) concentration compared to dGTP, dTTP, and dATP, dCTP may be present at a (higher) concentration compared to dTTP and dATP, dCTP may be present at a (higher) concentration compared to dATP, dCTP may be present at a (higher) concentration compared to dTTP, dTTP may be present at a (higher) concentration compared to dGTP, dCTP, and dATP, dTTP may be present at a (higher) concentration compared to dGTP and dCTP, dTTP may be present at a (higher) concentration compared to dCTP, dTTP may be present at a (higher) concentration compared to dGTP, dATP may be present at a (higher) concentration compared to dGTP, dTTP, and dCTP, dATP may be present at a (higher) concentration compared to dGTP and dCTP, dATP may be present at a (higher) concentration compared to dGTP, dATP may be present at a higher concentration compared to dGTP, dCTP and dATP may be present at a (higher) concentration compared to dGTP and dCTP, or dGTP and dCTP may be present at a (higher) concentration compared to dATP and dTTP.
[0171] The user may prepare a solution of dNTPs at non-uniform concentrations in any convenient manner. Since solutions of dATP, dTTP, dGTP, and dTTP are readily available commercially, the user only needs to mix them in an appropriate ratio.
[0172] Depending on the situation, the method is either (i) or (ii) below, namely (i) further comprising the step of amplifying the at least one target nucleic acid molecule comprising a nucleotide analog in the absence of the nucleotide analog, and performing the further step of amplifying the at least one target nucleic acid molecule comprising a nucleotide analog in the absence of the nucleotide analog using non-uniform concentrations of dNTPs, or (ii) preparing at least one mutated target nucleic acid molecule, and further comprising the step of amplifying the at least one mutated target nucleic acid molecule using a low-bias DNA polymerase, and performing the further step of amplifying the at least one mutated target nucleic acid molecule using a low-bias DNA polymerase using non-uniform concentrations of dNTPs.
[0173] Depending on the situation, introducing a mutation into the at least one target nucleic acid molecule by amplifying the at least one target nucleic acid molecule using a DNA polymerase to prepare the at least one mutated target nucleic acid molecule is performed in the presence of a nucleotide analog. Depending on the situation, the method for introducing a mutation into the at least one target nucleic acid molecule comprises the step of amplifying the at least one mutated target nucleic acid molecule in the absence of the nucleotide analog, and depending on the situation, this step is performed using non-uniform concentrations of dNTPs.
[0174] When introducing a mutation into at least one target nucleic acid molecule by using a nucleotide analog, this generally involves two amplification steps. In the first amplification step, the nucleotide analog is incorporated into the target nucleic acid molecule (mutation step). In the second amplification step, the nucleotide analog pairs with a natural nucleotide, thereby introducing the mutation into one strand of the target nucleic acid molecule (recovery step). When the target nucleic acid molecule is further amplified, this mutation is transmitted to both strands of the target nucleic acid molecule. Depending on the situation, both the first (mutation) amplification step and the second (recovery) amplification step may be performed using dNTPs at non-uniform concentrations. Depending on the situation, the non-uniform concentrations of dNTPs are different in the first (mutation) amplification step and the second (recovery) amplification step. For example, the non-uniform concentration of dNTPs may contain a lower concentration of dTTP than the other dNTPs in the first (mutation) amplification step, and the non-uniform concentration of dNTPs may contain a lower concentration of dATP than the other dNTPs in the second (recovery) amplification step. The step of amplifying at least one target nucleic acid molecule using a low-bias DNA polymerase, or the step of preparing at least one target nucleic acid molecule that has been mutated, may correspond to one or more "mutation steps". A further step of amplifying at least one target nucleic acid molecule containing a nucleotide analog in the absence of the nucleotide analog, or a further step of amplifying at least one target nucleic acid molecule that has been mutated, may correspond to one or more "recovery steps".
[0175] Depending on the situation, the nucleotide analog is dPTP.
[0176] In certain embodiments, the profile of mutations introduced is altered by using dNTPs at non-uniform concentrations. The non-uniform concentrations of dNTPs are used in methods that include introducing mutations into at least one target nucleic acid molecule. Thus, the method results in target nucleic acid molecules that contain mutations (e.g., the mutated target nucleic acid molecules described herein). The number of mutations introduced into a given target nucleic acid molecule by the method, the type of mutation, and the position of each mutation may be referred to as the “mutation profile” introduced. The term “type of mutation” refers to the nature of the mutation, i.e., whether it is a substitution, addition, or deletion mutation, and if it is a substitution mutation, what the starting nucleotide was and what it was mutated to (e.g., a mutation from A to G has an A starting nucleotide that was mutated to G).
[0177] A user may determine the “mutation profile” introduced by a given method by replicating a test target nucleic acid molecule and then subjecting a portion of the replication to a method that includes introducing the mutations of the invention (while keeping a portion of the replication intact (without mutating them)). The user may then sequence the replication that was subjected to the method that includes introducing the mutations of the invention, and the replication that was kept intact. Finally, the user can determine the number of mutations introduced, the type of mutation, and the position of each mutation by aligning the sequence of the replication that was subjected to the method that includes introducing the mutations of the invention with the sequence of the replication that was kept intact. Alternatively, the user may use a test target nucleic acid molecule of known sequence. In that case, the user may subject the test target nucleic acid molecule to a method that includes introducing the mutations of the invention and then sequence the resulting mutated target nucleic acid molecule to simply determine what mutation profile was introduced.
[0178] A user may wish to modify the mutation profile in many ways. For example, as described above, it is advantageous to be able to reduce the mutation bias. Thus, in one embodiment, the bias in the profile of mutations introduced is reduced by using dNTPs at non-uniform concentrations. In a further embodiment, this method is a method for introducing mutations into a low-bias mutation profile.
[0179] This application demonstrates that the bias in the profile of mutations introduced can be reduced by taking advantage of the use of dNTPs at non-uniform concentrations. For example, when mutating a target nucleic acid molecule by using a DNA polymerase (e.g., the low-bias DNA polymerase described above) to introduce a greater number of G-to-A mutations compared to other mutations, the user can reduce the concentration of dATP relative to the other dNTPs, and this may result in a decrease in the frequency with which an A nucleotide is incorporated instead of a dGTP, and thus a decrease in the number of G-to-A mutations.
[0180] Similarly, when using nucleotide analogs to introduce mutations into a target nucleic acid molecule, the mutation profile can be altered by changing the relative concentrations of dNTPs. For example, by using dPTP, mutations can be introduced from G to A, C to T, A to G, and T to C. As described in more detail previously, dPTP can replace a T nucleotide or a C nucleotide, and depending on whether the dPTP is in its amino or imino form, the dPTP can later pair with an A nucleotide or a G nucleotide. This leads to two scenarios. In the first scenario, dPTP replaces T (mutation step) in the sense strand (for example), and then the dPTP can pair with an A (no mutation) or a G (mutation from A to G) in the antisense strand. When dPTP replaces T and pairs with a G in the antisense strand, this mutant G introduces a mutation from T to C in the replica of the sense strand by pairing with a C (recovery step). Conversely, dPTP may replace T in the antisense strand, which may result in a mutation from A to G in the sense strand and a mutation from T to C in the replica of the antisense strand. In the second scenario, dPTP replaces C (mutation step) in the sense strand (for example), and then the dPTP can pair with an A (mutation from G to A) or a G (no mutation) in the antisense strand. When dPTP replaces C and pairs with an A in the antisense strand, this mutant A introduces a mutation from C to T in the replica of the sense strand by pairing with a T (recovery step). Conversely, dPTP may replace C in the antisense strand, which may result in a mutation from G to A in the sense strand and a mutation from C to T in the replica of the antisense strand.
[0181] This application demonstrates that when the rate of mutations from G to A and from C to T is higher than the rate of mutations from A to G and from T to C, decreasing the concentration of dTTP compared to other dNTPs (and preferably compared to the concentration of dCTP) at that time promotes the incorporation of dPTP instead of dTTP and increases the cases of the first scenario relative to the second scenario (meaning an increase in the mutations from A to G and from T to C introduced in the first scenario). Similarly, this application demonstrates that when the level of dATP decreases during the recovery step, the level of mutations from G to A and from C to T increases at that time. This is because in Scenario 2 above, when dATP is present at a lower concentration compared to other dNTPs (and preferably compared to the concentration of dGTP), this means that dPTP incorporated instead of the C nucleotide pairs more frequently with G and fewer mutations from G to A or from C to T are introduced. The two scenarios described above are shown in Figure 7.
[0182] Even the low - bias DNA polymerases disclosed herein introduce mutations into target nucleic acid molecules with a small bias. This application demonstrates that using non - uniform concentrations of dNTPs with a low - bias DNA polymerase can substantially eliminate any mutation bias.
[0183] Based on the information provided in this application, it is within the ability of one of ordinary skill in the art to determine (decide) how changes in the concentrations of various dNTPs affect the mutation profile depending on whether a particular nucleotide analog is used and, if so, which dNTPs are used. Thus, in some embodiments, the method of using non - uniform concentrations of dNTPs includes the step of identifying the dNTPs whose levels should be increased or decreased to reduce the bias of the profile of mutations to be introduced.
[0184] Depending on the situation, the dNTP with non-uniform concentration contains dTTP at a concentration lower than that of other dNTPs. As described above, this can increase the ratio of mutations from T to C and from A to G introduced when using dPTP as a nucleotide analog. Depending on the situation, the dNTP with non-uniform concentration contains dTTP at a concentration less than 75%, less than 70%, less than 60%, less than 55%, 25% - 75%, 25% - 70%, 25% - 60%, or about 50% of the concentration of dATP, dCTP, or dGTP. Depending on the situation, the dNTP with non-uniform concentration contains dTTP at a concentration less than 60% of the concentration of dCTP. Depending on the situation, the dNTP with non-uniform concentration contains dTTP at a concentration of 25% - 60% of the concentration of dCTP.
[0185] Depending on the situation, the dNTP with non-uniform concentration contains dATP at a concentration lower than that of other dNTPs. As described above, this can increase the ratio of mutations from G to A or from C to T introduced when using dPTP as a nucleotide analog. Depending on the situation, the dNTP with non-uniform concentration contains dATP at a concentration less than 75%, less than 70%, less than 60%, less than 55%, 25% - 75%, 25% - 70%, 25% - 60%, or about 50% of the concentration of dTTP, dCTP, or dGTP. Depending on the situation, the dNTP with non-uniform concentration contains dATP at a concentration less than 75%, less than 70%, less than 60%, less than 55%, 25% - 75%, 25% - 70%, 25% - 60%, or about 50% of the concentration of dGTP. Depending on the situation, the dNTP with non-uniform concentration contains dATP at a concentration less than 60% of the concentration of dGTP. Depending on the situation, the dNTP with non-uniform concentration contains dATP at a concentration of 25% - 60% of the concentration of dGTP.
[0186] As described in the above two scenarios, when using dPTP as a nucleotide analog, reducing dTTP increases the mutations from T to C and from A to G by promoting the replacement of T nucleotides in the target nucleic acid molecule with dPTP. Therefore, a non-uniform concentration of dNTPs containing a lower concentration of dTTP than other dNTPs is preferably used in the mutagenesis step (e.g., the PCR step in the presence of dPTP). Similarly, when using dPTP as a nucleotide analog, reducing dATP increases the mutations from G to A and from C to T because it reduces the number of dPTPs that pair with dATP and replace C nucleotides. Since the pairing of dPTP with dATP tends to occur during the recovery step, reducing dATP during the recovery step increases the number of mutations from G to A and from C to T. Therefore, depending on the situation, the step of amplifying at least one target nucleic acid molecule containing a nucleotide analog in the absence of the nucleotide analog, or the step of amplifying at least one mutated target nucleic acid molecule in the absence of the nucleotide analog, is performed using a non-uniform concentration of dNTPs, and the non-uniform concentration of dNTPs contains a lower concentration of dATP compared to other dNTPs.
Example
[0187] Example 1 - Mutating nucleic acid molecules using PrimeStar GXL of other polymerases The DNA molecule was fragmented to an appropriate size (e.g., 10 kb), and defined sequence priming sites (adapters) were ligated to each end using tagmentation.
[0188] The first step was a tagmentation reaction to fragment DNA. 50 ng of high-molecular-weight genomic DNA from one or more strains in a volume of 4 μL or less was subjected to tagmentation under the following conditions: 50 ng of DNA was combined with 4 μL of Nextera Transposase (diluted 1:50) and 8 μL of 2X tagmentation buffer (20 mM Tris [pH 7.6], 20 mM MgCl, 20% (v / v) dimethylformamide) for a total volume of 16 μL. The reaction was incubated at 55°C for 5 minutes, and 4 μL of NT buffer (or 0.2% SDS) was added to the reaction, after which the reaction was incubated at room temperature for 5 minutes.
[0189] The tagmentation reaction was purified using SPRIselect beads (Beckman Coulter) according to the manufacturer's instructions using 0.6 volumes of beads for left-side size selection, and then DNA was eluted in molecular biology-grade water.
[0190] Following this, PCR was performed for six cycles using standard dNTP and dPTP combinations. Primestar GXL was used, and 12.5 ng of tagmented and purified DNA was added to a total reaction volume of 25 μL containing 1× GXL buffer, 200 μM each of dATP, dTTP, dGTP, and dCTP, as well as 0.5 mM dPTP, and 0.4 μM of the custom primers (Table 2).
[0191] [Table 2]
[0192] The reactions were subjected to thermal cycling in the presence of Primestar GXL as follows: an initial gap extension at 68°C for 3 minutes, followed by six cycles of 98°C for 10 seconds, 55°C for 15 seconds, and 68°C for 10 minutes.
[0193] The next step is PCR without dPTP ( "recovery PCR") to remove dPTPs derived from the template and replace them with transition mutations. Excess dPTP and primers were removed by purifying the PCR reaction with SPRIselect beads, and then subjected to an additional 10 rounds (minimum 1 round, maximum 20) of amplification using primers (Table 3) that anneal to the fragment ends introduced during the dPTP incorporation cycle.
[0194]
Table 3
[0195] After this, a gel extraction step was performed to size select the amplified and mutated fragments within the desired size range (e.g., 7 - 10 kb). Gel extraction can be performed manually or via an automated system such as BluePippin. After this, an additional 16 - 20 cycles of PCR were performed ( "enrichment PCR").
[0196] After amplifying a defined number of long mutant templates, overlapping shorter fragment pools for sequencing were created by performing random fragmentation of the templates. Fragmentation was performed by tagmentation.
[0197] The long DNA fragments obtained in the previous step were subjected to a standard tagmentation reaction (e.g., Nextera XT or Nextera Flex), except that the reaction was divided into three pools for PCR amplification. This enables selective amplification of fragments derived from each end of the original template (e.g., containing the barcode of the sample), as well as internally fragmented fragments newly tagged at both ends from said long template. This effectively creates three pools for sequencing on an Illumina device (e.g., MiSeq or HiSeq).
[0198] The method was repeated using standard Taq (Jena Biosciences) and a blend of Taq and a proofreading polymerase (DeepVent) called LongAmp (New England Biolabs).
[0199] The data obtained in this experiment are shown in Figure 1. dPTP was not used as a control. Reads were mapped to the E. coli genome, achieving a median mutation rate of approximately 8%.
[0200] Example 2 - Comparison of mutation frequencies of various DNA polymerases Mutagenesis was performed using a wide variety of DNA polymerases (Table 4). Genomic DNA from E. coli strain MG1655 was tagmented to generate long fragments and bead-purified as described in the methods of Example 1. This was followed by six cycles of "mutagenesis PCR" in the presence of 0.5 mM dPTP, SPRIselect bead purification, and an additional 14-16 cycles of "recovery PCR" in the absence of dPTP. The resulting long mutant templates were then subjected to a standard tagmentation reaction (see Example 1), and the "internal" fragments were amplified and sequenced on a MiSeq instrument.
[0201] Mutation rates are listed in Table 4 and are normalized to the frequency of base substitutions via dPTP mutagenesis reactions, as measured using Illumina sequencing of DNA from known reference genomes. For Taq polymerase, even when used in a buffer optimized for Thermococcus archaeal polymerases, less than 12% of mutations occurred at template G+C sites. Thermococcus-like polymerases generated 58–69% of mutations at template G+C sites, while polymerases from Pyrococcus archaea contributed 88% of mutations to template G+C sites.
[0202] The enzymes were obtained from Jena Biosciences (Taq), Takara (various Primestar), Merck Millipore (KOD DNA polymerase), and New England Biolabs (Phusion).
[0203] Taq was tested for this experiment using the supplied buffer and additionally Primestar GXL Buffer (Takara). All other reactions were carried out using the standard supplied buffer for each polymerase.
[0204]
Table 4
[0205] Example 3 - Determining the dPTP mutagenesis rate The inventors performed dPTP mutagenesis on a wide range of genomic DNA samples showing various levels of G+C content (33 - 66%) using the polymerase of Thermococcus archaea (Primestar GXL; Takara) under a single set of reaction conditions. Mutagenesis and sequencing were carried out as described in the method of Example 3, except that 10 cycles of "recovery PCR" were performed. As expected, the mutation rates were approximately the same among the samples despite the diversity of G+C content (median mutation rate 7 - 8%) (Figure 2).
[0206] Example 4 - Measuring template amplification bias The template amplification bias was measured for two polymerases, namely Kapa HiFi, a proofreading polymerase commonly used in the Illumina sequencing protocol, and PrimeStar GXL, a polymerase of the KOD family known for its ability to amplify long fragments. In the first experiment, a limited number of E. coli genomic DNA templates having a size of about 2 kbp were amplified by using Kapa HiFi. The ends of these amplified fragments were then sequenced. A similar experiment was conducted using PrimeStar GXL on fragments of about 7 - 10 kbp obtained from E. coli. The position of each end sequence read was determined by mapping to the E. coli reference genome. The distance between adjacent fragment ends was measured. These distances were compared to a series of distances randomly drawn from a uniform distribution. The comparison was performed via the D of the non - parametric Kolmogorov - Smirnov test. When two samples are derived from the same distribution, the value of D approaches zero. For the low - bias PrimeStar polymerase, the inventors observed D = 0.07 when measured compared to a uniform random sample of 50,000 genomic positions for 50,000 fragment ends. For the Kapa HiFi polymerase, the inventors observed D = 0.14 for 50,000 fragment ends.
[0207] Example 5 - Use of two identical primer binding sites and a single primer sequence for selective amplification of longer templates As described above, by using tagmentation, DNA molecules can be fragmented while introducing primer binding sites (adapters) at the ends of the fragments. The Nextera tagmentation system (Illumina) utilizes a transposase enzyme that has one of two unique adapters (referred to herein as X and Y). This results in the generation of a random mixture of amplification products that have some identical end sequences (X-X, Y-Y) and some unique ends (X-Y). In the standard Nextera protocol, two separate primer sequences are used to selectively amplify "X-Y" products that contain different adapters at each end (if required for sequencing by Illumina technology). However, it is also possible to amplify "X-X" or "Y-Y" fragments that have identical end adapters by using a single primer sequence.
[0208] To create long mutant templates containing identical end adapters, 50 ng of high molecular weight genomic DNA (E. coli strain MG1655) was first subjected to tagmentation as described in Example 1 and then purified using SPRIselect beads. Thereafter, 5 cycles of "mutagenic PCR" were performed using a standard combination of dNTPs and dPTPs, which was carried out as detailed in Example 1 except that a single primer sequence was used (Table 5).
[0209] The excess dPTP and primers were removed by purifying the above PCR reaction with SPRIselect beads, and then the reaction was subjected to an additional 10 cycles of "recovery PCR" in the absence of dPTP to replace the dPTP in the template by transition mutations. The recovery PCR was carried out using a single primer that annealed to the fragment ends introduced during the dPTP incorporation cycles, thereby enabling selective amplification of the mutant templates generated in the previous PCR step.
[0210]
Table 5
[0211] As a control, a mutant template with different adapters at each end was prepared using the same protocol as above, except that it was used in both mutagenic PCR (shown in Table 2) and recovery PCR (Table 3) with two separate primer arrays. The final PCR product was purified with SPRIselect beads and further analyzed with a high-sensitivity DNA chip using the 2100 Bioanalzyer System (Agilent). As shown in Figure xxx, the templates prepared with the same end adapter were on average significantly longer than the control samples containing double adapters. The control template could be detected down to a minimum size of about 800 bp, while for the single adapter samples, no templates less than 2000 bp were observed.
[0212] The size profiles of mutant templates with the same end adapter (blue) and control templates with double adapters were compared by using them on an Agilent 2100 Bioanalyzer (high-sensitivity DNA kit). The use of the same end adapter inhibits the amplification of templates less than 2 kbp. This data is shown in Figure 6.
[0213] Example 8 - Further reducing the mutation bias of Thermococcus archaeal polymerase by changing the natural dNTP levels during PCR The polymerases of Thermococcus archaea generate a much more balanced mutation profile compared to other DNA polymerases, although they exhibit a slight bias towards mutations at G and C sites (see Table 4). To eliminate this residual bias, the inventors verified the effect of changing the concentration of natural dNTPs during the mutagenic and recovery PCR steps to affect the relative incorporation rates of various nucleotides.
[0214] First, long mutant templates were prepared from bacterial genomic DNA (E. coli strain MG1655) using the procedure outlined in Example 5, except that the concentrations of the individual nucleotides in the PCR reaction were varied. This was achieved by separately adding individual solutions of the four natural nucleotides (purchased from New England Biolabs) to the PCR mixture at a standard final concentration of 200 μM, or at lower concentrations of 160 μM (80% of the standard) or 100 μM (50%). Only one nucleotide was changed per reaction. As a control, equimolar dNTP mixtures supplemented with Primestar GXL polymerase (Takara) were used to add all natural nucleotides to the same final concentration of 200 μM. Five mutagenic PCR cycles and 12 recovery cycles were performed using the primers shown in Table 5. The resulting long mutant templates were then subjected to a standard tagging reaction (see Example 1), and the "internal" fragments were further amplified before sequencing on an Illumina MiSeq instrument. Mutation frequencies were determined by comparison to a known reference sequence.
[0215] As shown in Table 6, changes in the concentration of individual dNTPs during mutagenic and / or recovery PCR altered the profile of the mutations observed. Importantly, restricting the amount of dTTP to 50% during mutagenesis was found to result in substantially the same mutation frequency for each nucleotide (Table 3). This supports the conclusion that the residual mutational bias of Thermococcus polymerases can be eliminated through changes in dNTP levels.
[0216]
Table 6
[10] The method or use according to any one of Embodiments 1 to 9, wherein the low-bias DNA polymerase uses a nucleotide analog to mutate adenine, thymine, guanine, and / or cytosine in the at least one target nucleic acid molecule.
[11] The method or use according to any one of Embodiments 1 to 10, wherein the low-bias DNA polymerase replaces guanine, cytosine, adenine, and / or thymine with a nucleotide analog.
[12] The method or use according to any one of Embodiments 1 to 11, wherein the low-bias DNA polymerase uses a nucleotide analog to introduce guanine or adenine nucleotides at a ratio of 0.5 to 1.5:0.5 to 1.5, 0.6 to 1.4:0.6 to 1.4, 0.7 to 1.3:0.7 to 1.3, 0.8 to 1.2:0.8 to 1.2, or about 1:1, respectively.
[13] The method or use according to any one of Embodiments 1 to 12, wherein the low-bias DNA polymerase uses a nucleotide analog to introduce guanine or adenine nucleotides at a ratio of 0.7 to 1.3:0.7 to 1.3, respectively.
[14] The method or use according to any one of Embodiments 9 to 13, wherein the method includes a step of amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase, the step of amplifying the at least one target nucleic acid molecule using the low-bias DNA polymerase is performed in the presence of a nucleotide analog, and the step of amplifying the at least one target nucleic acid molecule provides at least one target nucleic acid molecule containing the nucleotide analog.
[15] The method or use according to any one of Embodiments 9 to 14, wherein the nucleotide analog is dPTP.
[16] The method or use according to embodiment 15, wherein the low-bias DNA polymerase introduces substitution mutations from guanine to adenine, substitution mutations from cytosine to thymine, substitution mutations from adenine to guanine, and substitution mutations from thymine to cytosine.
[17] The method or use according to embodiment 16, wherein the low-bias DNA polymerase introduces substitution mutations from guanine to adenine, substitution mutations from cytosine to thymine, substitution mutations from adenine to guanine, and substitution mutations from thymine to cytosine at a ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1, respectively.
[18] The method or use according to embodiment 16 or 17, wherein the low-bias DNA polymerase introduces substitution mutations from guanine to adenine, substitution mutations from cytosine to thymine, substitution mutations from adenine to guanine, and substitution mutations from thymine to cytosine at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively.
[19] The method or use according to any one of embodiments 1 to 18, wherein the low-bias DNA polymerase is a high-fidelity DNA polymerase.
[20] The method or use according to embodiment 19, wherein in the absence of nucleotide analogs, the high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, 0% to 0.0015%, or 0% to 0.001% mutations per round of replication.
[21] The method or use according to embodiment 14 or 15, wherein the method further comprises the step of amplifying the at least one target nucleic acid molecule containing nucleotide analogs in the absence of nucleotide analogs.
[22] The method or use according to embodiment 21, wherein the step of amplifying the at least one target nucleic acid molecule containing nucleotide analogs in the absence of nucleotide analogs is performed using a low-bias DNA polymerase.
[23] The method according to any one of embodiments 1 to 22, or use, which provides at least one target nucleic acid molecule in which the method has mutated, and further includes a further step of amplifying the at least one target nucleic acid molecule in which the method has mutated using a low-bias DNA polymerase.
[24] The method according to any one of embodiments 1 to 23, or use, wherein the low-bias DNA polymerase has a low template amplification bias.
[25] The method according to any one of embodiments 1 to 24, or use, wherein the low-bias DNA polymerase includes a proofreading domain and / or a processivity-enhancing domain.
[26] The low-bias DNA polymerase is as follows a. to h., that is a. The sequence of SEQ ID NO: 2, b. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 2, c. The sequence of SEQ ID NO: 4, d. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 4, e. The sequence of SEQ ID NO: 6, f. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 6, g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 7 The method according to any one of embodiments 1 to 25, or use, which includes at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acid fragments.
[27] The low-bias DNA polymerase is as follows a. to h., that is a. The sequence of SEQ ID NO: 2, b. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 2, c. The sequence of SEQ ID NO: 4, d. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 4, e. The sequence of SEQ ID NO: 6, f. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 6, g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 7 The method or use according to embodiment 26, which includes the same.
[28] The method according to embodiment 27, or use, wherein the low-bias DNA polymerase includes a sequence that is at least 98% identical to SEQ ID NO: 2.
[29] The method or use according to embodiment 27, wherein the low-bias DNA polymerase comprises a sequence that is at least 98% identical to SEQ ID NO: 4.
[30] The method or use according to embodiment 27, wherein the low-bias DNA polymerase comprises a sequence that is at least 98% identical to SEQ ID NO: 6.
[31] The method or use according to embodiment 27, wherein the low-bias DNA polymerase comprises a sequence that is at least 98% identical to SEQ ID NO: 7.
[32] The method or use according to any one of embodiments 1 to 31, wherein the low-bias DNA polymerase is a thermococcal polymerase of archaea of the order Thermococcales, or a derivative thereof.
[33] The method or use according to embodiment 32, wherein the low-bias DNA polymerase is a thermococcal polymerase of archaea of the order Thermococcales.
[34] The method or use according to embodiment 32 or 33, wherein the thermococcal polymerase of archaea of the order Thermococcales is derived from a strain of archaea of the order Thermococcales selected from the group consisting of T. kodakarensis, T. siculi, T. celer, and T. sp KS-1.
[35] The method or use according to any one of embodiments 1 to 34, further comprising introducing a barcode into the at least one target nucleic acid molecule.
[36] The method or use according to any one of embodiments 1 to 35, further comprising introducing a sample tag into the at least one target nucleic acid molecule.
[37] The method or use according to embodiment 36, using a sample tag group and labeling target nucleic acid molecules from different samples with different sample tags from the group.
[38] The method or use according to embodiment 37, wherein each sample tag is different from substantially all other sample tags of the sample tag group by at least one low-probability mutation difference or at least three high-probability mutation differences.
[39] The method or use according to embodiment 38, wherein each sample tag is different from substantially all other sample tags of the sample tag group by at least three low-probability mutation differences.
[40] The method or use according to embodiment 38 or 39, wherein each sample tag is different from substantially all other sample tags of the sample tag group by 3 to 25, or 3 to 10 low-probability mutation differences.
[41] The method or use according to any one of embodiments 38 to 40, wherein the low-probability mutation is a transversion mutation or an indel mutation.
[42] The method or use according to any one of embodiments 37 to 41, wherein each sample tag differs from substantially all other sample tags in the sample tag group by at least 5 high-probability mutation differences.
[43] The method or use according to embodiment 42, wherein each sample tag differs from substantially all other sample tags in the sample tag group by a high-probability mutation difference of 5 to 25, or 5 to 10.
[44] The method or use according to any one of embodiments 38 to 43, wherein the high-probability mutation is a transition mutation.
[45] The method or use according to any one of embodiments 37 to 44, wherein the sample tag group can be obtained by the method according to any one of embodiments 71 to 75.
[46] The method or use according to any one of embodiments 1 to 45, further comprising introducing an adapter into each of the at least one target nucleic acid molecule.
[47] The method or use according to embodiment 46, comprising introducing a first adapter at the 3' end of the at least one target nucleic acid molecule and introducing a second adapter at the 5' end of the at least one target nucleic acid molecule, wherein the first adapter and the second adapter are capable of annealing to each other.
[48] The method or use according to embodiment 47, amplifying the at least one target nucleic acid molecule using a primer that is identical to each other and complementary to a part of the first adapter.
[49] The method or use according to embodiment 47 or 48, wherein the first adapter is complementary to a nucleic acid molecule that is at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to the second adapter.
[50] The method or use according to embodiment 48 or 49, wherein the primer contains a second primer binding site, and the method comprises amplifying the at least one target nucleic acid molecule using the primer, removing the primer, and further amplifying the at least one target nucleic acid molecule using a second set of primers that anneal to the second primer binding site.
[51] The method or use according to any one of embodiments 1 to 50, further comprising introducing a barcode, a sample tag, and an adapter into each of the target nucleic acid molecules.
[52] The method or use according to any one of embodiments 1 to 51, wherein the barcode, the sample tag, and / or the adapter are introduced by tagmentation or by cleavage and ligation.
[53] The method or use according to any one of embodiments 1 to 52, wherein the at least one target nucleic acid molecule is greater than 1 kbp, greater than 1.5 kbp, greater than 2 kbp, greater than 4 kbp, greater than 5 kbp, greater than 7 kbp, or greater than 8 kbp.
[54] A method for determining the sequence of at least one target nucleic acid molecule, comprising a method for introducing a mutation according to embodiment 1 or any one of embodiments 3 to 53.
[55] The steps of a. to c. below, namely a. preparing at least one mutated target nucleic acid molecule by performing the method according to embodiment 1 or any one of embodiments 3 to 53; b. preparing mutated sequence reads by sequencing a region of the at least one mutated target nucleic acid molecule; and c. assembling a sequence for at least a portion of the at least one target nucleic acid molecule using the mutated sequence reads The method according to embodiment 54.
[56] The steps of a. to d. below, namely a. preparing at least one mutated target nucleic acid molecule by performing the method according to embodiment 1 or any one of embodiments 3 to 53; b. preparing at least one fragmented and / or amplified mutated target nucleic acid molecule by fragmenting and / or amplifying the at least one mutated target nucleic acid molecule; c. preparing mutated sequence reads by sequencing a region of the at least one fragmented and / or amplified mutated target nucleic acid molecule; and d. assembling a sequence for at least a portion of the at least one target nucleic acid molecule using the mutated sequence reads The method according to embodiment 54.
[57] A method for manipulating a protein, comprising a method for introducing a mutation according to embodiment 1 or any one of embodiments 3 to 53.
[58] The steps of a. to c. below, namely a. Preparing at least one mutated target nucleic acid molecule by implementing the method according to any one of Embodiment 1 or Embodiments 3 to 53; b. Inserting the at least one mutated target nucleic acid molecule into a vector; and c. Expressing the protein encoded by the at least one mutated target nucleic acid molecule The method according to Embodiment 57, comprising the above steps.
[59] The following steps a. to e., namely a. Preparing at least one sample containing at least one target nucleic acid molecule; and b. Preparing at least one target nucleic acid molecule containing the nucleotide analog by amplifying the at least one target nucleic acid molecule using a low-bias DNA polymerase in the presence of the nucleotide analog; c. Preparing at least one mutated target nucleic acid molecule by amplifying the at least one target nucleic acid molecule containing the nucleotide analog in the absence of the nucleotide analog; d. Inserting the at least one mutated target nucleic acid molecule into a vector; and e. Expressing the protein encoded by the at least one mutated target nucleic acid molecule The method according to Embodiment 58, comprising the above steps.
[60] The method according to Embodiment 58 or 59, further comprising testing the activity of the protein encoded by the at least one mutated target nucleic acid molecule or evaluating the structure of the protein.
[61] The method according to any one of Embodiments 58 to 60, wherein the vector is a plasmid, virus, cosmid, or artificial chromosome.
[62] The method according to any one of Embodiments 58 to 61, wherein the step of expressing the protein encoded by the at least one mutated target nucleic acid molecule is achieved by transforming bacterial cells, transfecting eukaryotic cells, or transducing eukaryotic cells using the vector.
[63] A sample tag group, wherein each sample tag is different from substantially all other sample tags in the group by at least one low-probability mutation difference or at least three high-probability mutation differences.
[64] The sample tag group according to Embodiment 63, wherein each sample tag is different from substantially all other sample tags in the sample tag group by at least three low-probability mutation differences.
[65] The sample tag group according to embodiment 63 or 64, wherein each sample tag differs from substantially all other sample tags in the sample tag group by a low probability mutation difference of 3 to 25, or 3 to 10.
[66] The sample tag group according to any one of embodiments 63 to 65, wherein the low probability mutation is a transversion mutation or an indel mutation.
[67] The sample tag group according to any one of embodiments 63 to 66, wherein each sample tag differs from substantially all other sample tags in the sample tag group by a high probability mutation difference of at least 5.
[68] The sample tag group according to any one of embodiments 63 to 67, wherein each sample tag differs from substantially all other sample tags in the sample tag group by a high probability mutation difference of 5 to 25, or 5 to 10.
[69] The sample tag group according to any one of embodiments 63 to 68, wherein the high probability mutation is a transition mutation.
[70] The sample tag group according to any one of embodiments 63 to 69, wherein each sample tag has a length of at least 8 nucleotides, at least 10 nucleotides, at least 12 nucleotides, 8 to 50 nucleotides, 10 to 50 nucleotides, or 10 to 50 nucleotides.
[71] The following a. and b., that is a. Analyzing a method for introducing a mutation into at least one target nucleic acid molecule and determining the average number of low probability mutations that occur during the method for introducing a mutation into the at least one target nucleic acid molecule, and b. Determining the sequence for a group in which each sample tag differs from substantially all sample tags in the sample tag group by a low probability difference greater than the average number of low probability mutations that occur during the method for introducing a mutation into the at least one target nucleic acid molecule A method for designing a sample tag group suitable for use in a method for introducing a mutation into at least one target nucleic acid molecule, comprising.
[72] The following a.(i) and (ii), that is a.(i) Analyzing a method for introducing a mutation into the at least one target nucleic acid molecule and determining the average number of high probability mutations that occur during the method for introducing a mutation into the at least one target nucleic acid molecule, and (ii) determining the sequence for a group in which each sample tag is different from substantially all sample tags of the sample tag group by a high-probability difference greater than the average number of high-probability mutations that occur during the method for introducing mutations into the at least one target nucleic acid molecule The method according to embodiment 71, further comprising this.
[73] The method according to embodiment 71 or 72, wherein the low-probability mutation is a transversion mutation or an indel mutation.
[74] The method according to any one of embodiments 72 to 73, wherein the high-probability mutation is a transition mutation.
[75] The method according to any one of embodiments 71 to 74, which is a method implemented by a computer.
[76] The method or use according to any one of embodiments 1 to 75, wherein the step of amplifying the at least one target nucleic acid molecule using the low-bias DNA polymerase is performed using dNTPs at non-uniform concentrations.
[77] The following (i) or (ii), that is (i) The method includes a further step of amplifying the at least one target nucleic acid molecule containing a nucleotide analog in the absence of the nucleotide analog, and in the absence of this nucleotide analog, the further step of amplifying the at least one target nucleic acid molecule containing a nucleotide analog is performed using dNTPs at non-uniform concentrations, or (ii) The method provides at least one target nucleic acid molecule that has been mutated, and the method includes a further step of amplifying the at least one mutated target nucleic acid molecule using the low-bias DNA polymerase, and the further step of amplifying the at least one mutated target nucleic acid molecule using the low-bias DNA polymerase is performed using dNTPs at non-uniform concentrations. The method according to any one of embodiments 1 to 76.
[78] The following a. and b., that is a. preparing at least one sample containing at least one target nucleic acid molecule, and b. introducing a mutation into the at least one target nucleic acid molecule by amplifying the at least one target nucleic acid molecule using a DNA polymerase to prepare at least one mutated target nucleic acid molecule. Here, it is assumed that step b. is performed using dNTPs at non-uniform concentrations. A method for introducing a mutation into at least one target nucleic acid molecule, including this.
[79] The method according to embodiment 78, wherein step b. is carried out in the presence of a nucleotide analog.
[80] The method according to embodiment 79, wherein the nucleotide analog is dPTP.
[81] The method according to embodiment 79 or 80, further comprising step c. of amplifying the at least one mutated target nucleic acid molecule in the absence of a nucleotide analog.
[82] The method according to embodiment 81, wherein step c. is carried out using dNTPs at non-uniform concentrations.
[83] The method according to any one of embodiments 76 to 82, wherein by using dNTPs at non-uniform concentrations, the profile of the mutations to be introduced is altered.
[84] The method according to embodiment 83, wherein by using dNTPs at non-uniform concentrations, the bias in the profile of the mutations to be introduced is reduced.
[85] The method according to any one of embodiments 1 to 84, wherein the method is for introducing mutations with a low-bias mutation profile.
[86] The method according to any one of embodiments 76 to 85, wherein the non-uniform concentration dNTPs comprise dATP, dCTP, dTTP and dGTP, and one or two of dATP, dCTP, dTTP or dGTP are at a lower concentration compared to the other dNTPs.
[87] The method according to any one of embodiments 76 to 86, wherein the use of non-uniform concentration dNTPs comprises the step of identifying the dNTP for which the level should be increased or decreased to reduce the bias in the profile of the mutations to be introduced.
[88] The method according to any one of embodiments 76 to 87, wherein the non-uniform concentration dNTPs comprise dTTP at a lower concentration than the other dNTPs.
[89] The method according to embodiment 88, wherein the non-uniform concentration dNTPs comprise dTTP at a concentration of less than 75%, less than 70%, less than 60%, less than 55%, 25% - 75%, 25% - 70%, 25% - 60%, or about 50% of the concentration of dATP, dCTP or dGTP.
[90] The method according to embodiment 89, wherein the non-uniform concentration dNTPs comprise dTTP at a concentration of less than 75%, less than 70%, less than 60%, less than 55%, 25% - 75%, 25% - 70%, 25% - 60%, or about 50% of the concentration of dCTP.
[91] The method according to embodiment 90, wherein the non-uniform concentration dNTPs comprise dTTP at a concentration of less than 60% of the concentration of dCTP.
[92] The method according to embodiment 87, wherein the dNTP with the non-uniform concentration contains dTTP at a concentration of 25% to 60% of the concentration of dCTP.
[93] The method according to any one of embodiments 77 or 81 to 92, wherein the step of amplifying at least one target nucleic acid molecule containing a nucleotide analog or the step of amplifying at least one mutated target nucleic acid molecule in the absence of a nucleotide analog is performed using dNTP with a non-uniform concentration.
[94] The method according to embodiment 93, wherein the dNTP with the non-uniform concentration contains dATP at a lower concentration compared to other dNTPs.
[95] The method according to embodiment 94, wherein the dNTP with the non-uniform concentration contains dATP at a concentration of less than 75%, less than 70%, less than 60%, less than 55%, 25% to 75%, 25% to 70%, 25% to 60%, or about 50% of the concentration of dTTP, dCTP or dGTP.
[96] The method according to embodiment 95, wherein the dNTP with the non-uniform concentration contains dATP at a concentration of less than 75%, less than 70%, less than 60%, less than 55%, 25% to 75%, 25% to 70%, 25% to 60%, or about 50% of the concentration of dGTP.
[97] The method according to embodiment 96, wherein the dNTP with the non-uniform concentration contains dATP at a concentration of less than 60% of the concentration of dGTP.
[98] The method according to embodiment 96 or 97, wherein the dNTP with the non-uniform concentration contains dATP at a concentration of 25% to 60% of the concentration of dGTP.
[99] A sample tag group obtainable by the method according to any one of embodiments 71 to 74.
[0100] A computer-readable medium configured to implement the method according to any one of embodiments 71 to 74.
[0101] The following a. to c., that is a. preparing at least one sample containing a target nucleic acid molecule, b. introducing a first adapter to the 3' end of the target nucleic acid molecule and introducing a second adapter to the 5' end of the target nucleic acid molecule, and c. amplifying the target nucleic acid molecule using a primer complementary to a part of the first adapter, wherein the first adapter and the second adapter are capable of annealing to each other, A method for selectively amplifying a target nucleic acid molecule larger than 1 kbp in length.
[0102] The method according to embodiment 101, wherein the primers are identical to each other. The method according to embodiment 101 or 102, wherein the first adapter is complementary to a nucleic acid molecule that is at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to the second adapter. The method according to any one of embodiments 101 to 103, wherein the method is for selectively amplifying a target nucleic acid molecule larger than 1.5 kbp in length. The method according to any one of embodiments 101 to 104, further comprising the step of sequencing the target nucleic acid molecule.
[0217] This disclosure includes the following sequence information. SEQUENCE LISTING <110> ILLUMINA SINGAPORE PTE. LTD. <120> Method for Introducing Mutations <130> PA23-680 <140> <141> 2019-02-19 <150> GB 1802744.1 <151> 2018-02-20 <160> 142 <170> PatentIn version 3.5 <210> 1 <211> 2325 <212> DNA <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus sp. KS-1 <400> 1 atgatcctcg acactgacta cataactgag aatggaaaac ccgtcataag gattttcaag 60 aaggagaacg gcgagtttaa gattgagtac gataggactt ttgaacccta catttacgcc 120 ctcctgaagg acgattctgc cattgaggag gtcaagaaga taaccgccga gaggcacgga 180 acggttgtaa cggttaagcg ggctgaaaag gttcagaaga agttcctcgg gagaccagtt 240 gaggtctgga aactctactt tactcaccct caggacgtcc cagcgataag ggacaagata 300 cgagagcatc cagcagttat tgacatctac gagtacgaca tacccttcgc caagcgctac 360 ctcatagaca agggattagt gccaatggaa ggcgacgagg agctgaaaat gcttgccttt 420 gatatcgaga cgctctacca tgagggcgag gagttcgccg aggggccaat ccttatgata 480 agctacgccg acgaggaagg ggccagggtg ataacgtgga agaacgcgga tctgccctac 540 gttgacgtcg tctcgacgga gagggagatg ataaagcgct tcctaaaggt ggtcaaagag 600 aaagatcctg acgtcctaat aacctacaac ggcgacaact tcgacttcgc ctacctaaaa 660 aaacgctgtg aaaagcttgg aataaacttc acgctcggaa gggacggaag cgagccgaag 720 attcagagga tgggcgacag gtttgccgtc gaagtgaagg gacggataca cttcgatctc 780 tatcctgtga taagacggac gataaacctg cccacataca cgcttgaggc cgtttatgaa 840 gccgtcttcg gtcagccgaa ggagaaggtc tacgctgagg agatagctac agcttgggag 900 agcggtgaag gccttgagag agtagccaga tactcgatgg aagatgcgaa ggtcacatac 960 gagcttggga aggagttttt ccctatggag gcccagcttt ctcgcttaat cggccagtcc 1020 ctctgggacg tctcccgctc cagcactggc aacctcgttg agtggttcct cctcaggaag 1080 gcctacgaga ggaatgagct ggccccgaac aagcccgatg aaaaggagct ggccagaaga 1140 cgacagagct atgaaggagg ctatgtaaaa gagcccgaga gagggttgtg ggagaacata 1200 gtgtacctag attttagatc tctgtacccc tcaatcatca tcacccacaa cgtctcgccg 1260 gatactctca acagggaagg atgcaaggaa tatgacgttg ccccccaggt cggtcaccgc 1320 ttctgcaagg acttcccagg atttatcccg agcctgcttg gagacctcct agaggagagg 1380 cagaagataa agaagaagat gaaggccacg attgacccga tcgagaggaa gctcctcgat 1440 tacaggcaga gggccatcaa gatcctggcc aacagctact acggttacta cggctatgca 1500 agggcgcgct ggtactgcaa ggagtgtgca gagagcgtaa cggcctgggg aagggagtac 1560 ataacgatga ccatcagaga gatagaggaa aagtacggct ttaaggtaat ctacagcgac 1620 accgacggat tttttgccac aatacctgga gccgatgctg aaaccgtcaa aaagaaggcg 1680 atggagttcc tcaagtatat caacgccaaa ctcccgggcg cgcttgagct cgagtacgag 1740 ggcttctaca aacgcggctt cttcgtcacg aagaagaagt acgcggtgat agacgaggaa 1800 ggcaagataa caacgcgcgg acttgagatt gtgaggcgcg actggagcga gatagcgaaa 1860 gagacgcagg cgagggttct tgaagctttg ctaaaggacg gtgacgtcga gaaggccgtg 1920 aggatagtca aagaagttac cgaaaagctg agcaagtacg aggttccgcc ggagaagctg 1980 gtgatccacg agcagataac gagggattta aaggactaca aggcaaccgg tccccacgtt 2040 gccgttgcca agaggttggc cgcgagagga gtcaaaatac gccctggaac ggtgataagc 2100 tacatcgtgc tcaagggctc tgggaggata ggcgacaggg cgataccgtt cgacgagttc 2160 gacccgacga agcacaagta cgacgccgag tactacattg agaaccaggt tctcccagcc 2220 gttgagagaa ttctgagagc cttcggttac cgcaaggaag acctgcgcta ccagaagacg 2280 agacaggttg gtctgggagc ctggctgaag ccgaagggaa cttga 2325 <210> 2 <211> 774 <212> PRT <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus sp. KS-1 <400> 2 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asn Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Thr 50 55 60 Val Lys Arg Ala Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Glu His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Asp Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Glu Gly Ala Arg Val Ile Thr Trp Lys Asn Ala 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Arg Glu Met Ile Lys 180 185 190 Arg Phe Leu Lys Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Thr Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Val Phe Gly Gln Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Ala Thr Ala Trp Glu Ser Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Tyr 305 310 315 320 Glu Leu Gly Lys Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Ser Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Asp Glu Lys Glu Leu Ala Arg Arg Arg Gln Ser Tyr 370 375 380 Glu Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Lys Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Ala Asn Ser Tyr Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Thr Met Thr Ile Arg Glu Ile 515 520 525 Glu Glu Lys Tyr Gly Phe Lys Val Ile Tyr Ser Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Ala Asp Ala Glu Thr Val Lys Lys Lys Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Lys Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Lys Ala Val 625 630 635 640 Arg Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Arg Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Lys Pro Lys Gly Thr 770 <210> 3 <211> 2325 <212> DNA <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus celer <400> 3 atgatcctcg acgctgacta catcaccgaa gatgggaagc ccgtcgtgag gatattcagg 60 aaggagaagg gcgagttcag aatcgactac gacagggact tcgagcccta catctacgcc 120 ctcctgaagg acgattcggc catcgaggag gtgaagagga taaccgttga gcgccacggg 180 aaggccgtca gggttaagcg ggtggagaag gtcgaaaaga agttcctcaa caggccgata 240 gaggtctgga agctctactt caatcacccg caggacgttc cggcgataag ggacgagata 300 aggaagcatc cggccgtcgt tgatatctac gagtacgaca tccccttcgc caagcgctac 360 ctcatcgata aggggctcgt cccgatggag ggggaggagg agctcaaact gatggccttc 420 gacatcgaga ccctctacca cgagggagac gagttcgggg aggggccgat cctgatgata 480 agctacgccg acggggacgg ggcgagggtc ataacctgga agaagatcga cctcccctac 540 gtcgacgtcg tctcgaccga gaaggagatg ataaagcgct tcctccaggt ggtgaaggag 600 aaggacccgg acgtgctcgt aacttacaac ggcgacaact tcgacttcgc ctacctgaag 660 agacgctccg aggagcttgg attgaagttc atcctcggga gggacgggag cgagcccaag 720 atccagcgca tgggcgaccg cttcgccgtc gaggtgaagg ggaggataca cttcgacctc 780 tacccggtga taaggcgcac cgtgaacctg ccgacctaca cgctcgaggc ggtctacgag 840 gccatcttcg ggaggccaaa ggagaaggtc tacgccgggg agatagtgga ggcctgggaa 900 accggcgagg gtcttgagag ggttgcccgc tactccatgg aggacgcaaa ggttaccttc 960 gagctcggga gggagttctt cccgatggag gcccagctct cgaggctcat cggccagggt 1020 ctctgggacg tctcccgctc gagcaccggc aacctggtcg agtggttcct cctgaggaag 1080 gcctacgaga ggaacgaact ggccccgaac aagccgagcg gccgggaagt ggagatcagg 1140 aggcgtggct acgccggtgg ttacgttaag gagccggaga ggggtttatg ggagaacatc 1200 gtgtacctcg actttcgctc tctttacccc tccatcatca taacccacaa cgtctcgccc 1260 gataccctaa acagggaggg ctgtgagaac tacgacgtcg ccccccaggt ggggcataag 1320 ttctgcaaag attttccggg cttcatcccg agcctgctcg gaggcctgct tgaggagagg 1380 cagaagataa agcggaggat gaaggcctct gtggatcccg ttgagcggaa gctcctcgat 1440 tacaggcaga gggccatcaa gatactggcc aacagcttct acggatacta cggctacgcg 1500 agggcgaggt ggtactgcag ggagtgcgcg gagagcgtta ccgcctgggg cagggagtac 1560 atcgataggg tcatcaggga gctcgaggag aagttcggct tcaaggtgct ctacgcggac 1620 acggacggac tgcacgccac gatccccggg gcggacgccg ggaccgtcaa ggagagggcg 1680 agggggttcc tgagatacat caaccccaag ctccccggcc tcctggagct cgagtacgag 1740 gggttctacc tgaggggttt cttcgtgacg aagaagaagt acgcggtcat agacgaggag 1800 ggcaagataa ccacgcgcgg cctcgagata gtcaggcggg actggagcga ggtggccaag 1860 gagacgcagg cgagggtcct ggaggcgata ctgaggcacg gtgacgtcga ggaggccgtt 1920 agaatcgtca gggaggtaac cgaaaagctg agcaagtacg aggttccgcc ggagaaactg 1980 gtgatccacg agcagataac gagggatttg agggactaca aagccacggg accgcacgtg 2040 gcggtggcga agcgcctggc cgggaggggg gtaaggatac gccccgggac ggtgataagc 2100 tacatcgtcc tcaagggctc cggaaggata ggggacaggg cgattccctt cgacgagttc 2160 gacccgacta agcacaggta cgacgccgac tactacatcg agaaccaggt tctgccagcc 2220 gtcgagagga tcctgaaggc cttcggctac cgcaaggagg acctgaaata ccagaagacg 2280 aggcaggtgg gcctgggtgc gtggctcaac gcggggaagg ggtga 2325 <210> 4 <211> 774 <212> PRT <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus celer <400> 4 Met Ile Leu Asp Ala Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Val 1 5 10 15 Arg Ile Phe Arg Lys Glu Lys Gly Glu Phe Arg Ile Asp Tyr Asp Arg 20 25 30 Asp Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Arg Ile Thr Val Glu Arg His Gly Lys Ala Val Arg 50 55 60 Val Lys Arg Val Glu Lys Val Glu Lys Lys Phe Leu Asn Arg Pro Ile 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Asn His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Glu Ile Arg Lys His Pro Ala Val Val Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Glu Glu Glu Leu Lys Leu Met Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Asp Glu Phe Gly Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Gly Asp Gly Ala Arg Val Ile Thr Trp Lys Lys Ile 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Lys Glu Met Ile Lys 180 185 190 Arg Phe Leu Gln Val Val Lys Glu Lys Asp Pro Asp Val Leu Val Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Arg Arg Ser Glu 210 215 220 Glu Leu Gly Leu Lys Phe Ile Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Val Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Ile Phe Gly Arg Pro Lys Glu 275 280 285 Lys Val Tyr Ala Gly Glu Ile Val Glu Ala Trp Glu Thr Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Phe 305 310 315 320 Glu Leu Gly Arg Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Gly Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Ser Gly Arg Glu Val Glu Ile Arg Arg Arg Gly Tyr 370 375 380 Ala Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Glu Asn Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Lys Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Gly Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Arg Arg Met Lys Ala Ser Val Asp Pro Val Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Ala Asn Ser Phe Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Arg Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Asp Arg Val Ile Arg Glu Leu 515 520 525 Glu Glu Lys Phe Gly Phe Lys Val Leu Tyr Ala Asp Thr Asp Gly Leu 530 535 540 His Ala Thr Ile Pro Gly Ala Asp Ala Gly Thr Val Lys Glu Arg Ala 545 550 555 560 Arg Gly Phe Leu Arg Tyr Ile Asn Pro Lys Leu Pro Gly Leu Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Leu Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Val Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Ile Leu Arg His Gly Asp Val Glu Glu Ala Val 625 630 635 640 Arg Ile Val Arg Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Arg Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Gly 675 680 685 Arg Gly Val Arg Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Arg Tyr Asp Ala Asp Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Lys Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Lys Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Asn Ala Gly Lys Gly 770 <210> 5 <211> 2328 <212> DNA <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus siculi <400> 5 atgatcctcg acacggacta catcacggaa gatgggaaac ccgtcataag gatattcaag 60 aaagagaacg gcgagttcaa gatcgagtac gacaggactt ttgaacccta catctacgcc 120 ctcctgaagg acgactccgc gattgaggat gttaaaaaga taaccgccga gaggcacgga 180 acggtggtga aggtcaagcg cgccgaaaag gtgcagaaga agttcctagg caggccggtt 240 gaagtctgga agctctactt cacccacccc caagatgtcc cggcgataag ggacaagatt 300 aggaagcatc cagctgtaat tgacatctac gagtacgaca taccattcgc caagcgctac 360 ctcatcgaca agggcctgat tccgatggag ggtgaagaag agcttaagat gctcgccttc 420 gacattgaga cgctctacca tgagggtgag gagttcgccg aggggcctat tctgatgata 480 agctacgccg acgagagcga ggcacgcgtc atcacctgga agaaaatcga cctcccctac 540 gttgacgtcg tctcaacgga gaaggagatg ataaagcgct tcctccgcgt tgtgaaggag 600 aaagatcccg atgtcctcat aacctacaac ggcgacaact tcgacttcgc ctacctgaag 660 aagcgctgtg aaaagcttgg aataaacttc ctccttggaa gggacgggag cgagccgaag 720 atccagagaa tgggtgaccg cttcgccgtt gaggtgaagg ggaggataca cttcgacctc 780 tatcctgtaa taaggcgcac gataaacctg ccgacctaca tgcttgaggc agtctacgag 840 gccatctttg ggaagccaaa ggagaaggtt tacgccgagg agatagccac cgcttgggaa 900 accggagagg gccttgagag ggtggctcgc tactctatgg aggacgcgaa ggtcacgttt 960 gagcttggaa aggagttctt cccgatggag gcccaacttt cgaggttggt cggccagagc 1020 ttctgggatg tcgcgcgctc aagcacgggc aatctggtcg agtggttcct cctcaggaag 1080 gcctacgaga ggaacgagct ggctccaaac aagccctctg gaagggaata tgacgagagg 1140 cgcggtggat acgccggcgg ctacgtcaag gaaccggaaa agggcctgtg ggagaacata 1200 gtctacctcg actataaatc tctctacccc tcaatcatca tcacccacaa cgtctcgccc 1260 gataccctca accgcgaggg ctgtaaggag tatgacgtag ctccacaggt cggccaccgc 1320 ttctgcaagg actttccagg cttcatcccg agcctgctcg gggatctcct ggaggagagg 1380 cagaagataa agaggaagat gaaggcaaca attgacccga tcgagagaaa gctccttgat 1440 tacaggcaac gggccatcaa gatccttcta aatagttttt acggctacta cggctacgca 1500 agggctcgct ggtactgcaa ggagtgtgcc gagagcgtta cggcatgggg aagggaatat 1560 atcaccatga caatcaggga aatagaagag aagtatggct ttaaagtact ttatgcggac 1620 actgacggct tcttcgcgac gattcccggg gaagatgccg agaccatcaa aaagagggcg 1680 atggagttcc tcaagtacat aaacgccaaa ctccccggtg cgctcgaact tgagtacgag 1740 gacttctaca ggcgcggctt cttcgtcacc aagaagaaat acgcggttat cgacgaggag 1800 ggcaagataa caacgcgcgg gctggagatc gtcaggcgcg actggagcga gatagccaag 1860 gagacgcagg cgcgggttct ggaggccctt ctgaaggacg gtgacgtcga agaggccgtg 1920 agcatagtca aagaagtgac cgagaagctg agcaagtacg aggttccgcc ggagaagctc 1980 gttatccacg agcagataac gcgcgagctg aaggactaca aggcaacggg accacacgtg 2040 gcgatagcga agaggttagc cgcgagaggc gtcaaaatcc gccccgggac agtcatcagc 2100 tacatcgtgc tcaagggctc cgggaggata ggcgacaggg cgattccctt cgacgagttc 2160 gaccccacga agcacaagta cgatgcagag tactacatcg agaaccaggt tctacctgcc 2220 gtcgagagga ttctgaaggc cttcggctat cgcggtgagg agctcagata ccagaagacg 2280 aggcaggttg gacttggggc gtggctgaag ccgaagggga aggggtga 2328 <210> 6 <211> 775 <212> PRT <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus siculi <400> 6 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Asp Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Lys 50 55 60 Val Lys Arg Ala Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Lys His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Ile Pro 115 120 125 Met Glu Gly Glu Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Ser Glu Ala Arg Val Ile Thr Trp Lys Lys Ile 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Lys Glu Met Ile Lys 180 185 190 Arg Phe Leu Arg Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Leu Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Met Leu Glu Ala Val Tyr Glu Ala Ile Phe Gly Lys Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Ala Thr Ala Trp Glu Thr Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Phe 305 310 315 320 Glu Leu Gly Lys Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Val Gly Gln Ser Phe Trp Asp Val Ala Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Ser Gly Arg Glu Tyr Asp Glu Arg Arg Gly Gly Tyr 370 375 380 Ala Gly Gly Tyr Val Lys Glu Pro Glu Lys Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Tyr Lys Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Arg Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Leu Asn Ser Phe Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Thr Met Thr Ile Arg Glu Ile 515 520 525 Glu Glu Lys Tyr Gly Phe Lys Val Leu Tyr Ala Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Glu Asp Ala Glu Thr Ile Lys Lys Arg Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Asp Phe Tyr Arg Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Glu Ala Val 625 630 635 640 Ser Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Glu Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Ile Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Lys Ala Phe Gly Tyr Arg Gly 740 745 750 Glu Glu Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Lys Pro Lys Gly Lys Gly 770 775 <210> 7 <211> 774 <212> PRT <213> Artificial Sequence <220> <223> DNA polymerase from Thermococcus kodakarensis <400> 7 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Phe Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Thr 50 55 60 Val Lys Arg Val Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Glu His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Asp Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr Glu Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Glu Gly Ala Arg Val Ile Thr Trp Lys Asn Val 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Arg Glu Met Ile Lys 180 185 190 Arg Phe Leu Arg Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Ala Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Val Phe Gly Gln Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Thr Thr Ala Trp Glu Thr Gly Glu Asn 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Tyr 305 310 315 320 Glu Leu Gly Lys Glu Phe Leu Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Ser Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Asp Glu Lys Glu Leu Ala Arg Arg Arg Gln Ser Tyr 370 375 380 Glu Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Lys Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Ala Asn Ser Tyr Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Thr Met Thr Ile Lys Glu Ile 515 520 525 Glu Glu Lys Tyr Gly Phe Lys Val Ile Tyr Ser Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Ala Asp Ala Glu Thr Val Lys Lys Lys Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Glu Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Lys Ala Val 625 630 635 640 Arg Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Arg Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Ser Ala Trp 755 760 765 Leu Lys Pro Lys Gly Thr 770 <210> 8 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 8 tagaattgaa gaa 13 <210> 9 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 9 tggccatagc tac 13 <210> 10 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 10 gtcatctgcg acc 13 <210> 11 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 11 ttcgcgcttg gac 13 <210> 12 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 12 cgcgaaccgt tag 13 <210> 13 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 13 ttgcagcctc taa 13 <210> 14 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 14 tctactagta cga 13 <210> 15 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 15 gtaggttcta ctg 13 <210> 16 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 16 gccaatatca agt 13 <210> 17 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 17 ctatcttgct ggt 13 <210> 18 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 18 gttctcatag gta 13 <210> 19 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 19 gtctatgaac caa 13 <210> 20 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 20 cggagcgctt att 13 <210> 21 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 21 tatgccatga gga 13 <210> 22 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 22 atacgactcg gag 13 <210> 23 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 23 gatggaactc agc 13 <210> 24 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 24 ggacctgcat gaa 13 <210> 25 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 25 tagactggaa ctt 13 <210> 26 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 26 gaattacctc gtt 13 <210> 27 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 27 aggatcaggc tac 13 <210> 28 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 28 acgcgtagaa gag 13 <210> 29 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 29 cttcgagact tac 13 <210> 30 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 30 gacggctaac tcc 13 <210> 31 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 31 ttagcattct ctt 13 <210> 32 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 32 gcaaggcata gta 13 <210> 33 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 33 acctagatat gga 13 <210> 34 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 34 acgccaaggc gta 13 <210> 35 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 35 tatgacggat ccg 13 <210> 36 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 36 cctccattag aga 13 <210> 37 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 37 attgaatact ctg 13 <210> 38 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 38 gagatgagaa gaa 13 <210> 39 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 39 tctgagtagc cgg 13 <210> 40 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 40 aataggtagt acg 13 <210> 41 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 41 gtcgaagaag tcc 13 <210> 42 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 42 tactgcatct cgt 13 <210> 43 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 43 gacgtattag agc 13 <210> 44 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 44 cctgcattat tcg 13 <210> 45 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 45 acgaatgatg ctc 13 <210> 46 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 46 tactagcaga gat 13 <210> 47 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 47 ctcctcatct tcc 13 <210> 48 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 48 tcctctgcgc tgc 13 <210> 49 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 49 ccttctcagt ccg 13 <210> 50 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 50 cagcttcata gcg 13 <210> 51 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 51 ttgactctcg cgc 13 <210> 52 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 52 tatcctgagc gat 13 <210> 53 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 53 aacgcctagc cga 13 <210> 54 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 54 ccgaagacgt cat 13 <210> 55 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 55 gagttctcca gat 13 <210> 56 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 56 tgcatccgcg ctt 13 <210> 57 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 57 cctgaactca agt 13 <210> 58 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> sample tag sequence <400> 58 ggtcgtatgc gta 13 <210> 59 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 59 aggcctctct acc 13 <210> 60 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 60 gtactccatc caa 13 <210> 61 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 61 cagcggacgc gct 13 <210> 62 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 62 atctctctta gca 13 <210> 63 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 63 aagcaataat aat 13 <210> 64 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 64 aaggcgactc cga 13 <210> 65 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 65 acgtctctag gag 13 <210> 66 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 66 ccatcagacc tct 13 <210> 67 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 67 acttaatcgt act 13 <210> 68 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 68 tggaattctc caa 13 <210> 69 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 69 ccatacgatc agg 13 <210> 70 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 70 ttatggagca ata 13 <210> 71 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 71 gctcggcgtt cga 13 <210> 72 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 72 ttggccagtc gct 13 <210> 73 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 73 cagatacgta gag 13 <210> 74 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 74 aatgctatta tcc 13 <210> 75 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 75 gcagcatgcc gat 13 <210> 76 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 76 ggagagttac ctc 13 <210> 77 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 77 gagagtccat gat 13 <210> 78 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 78 caatctattc tga 13 <210> 79 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 79 gctcttagta tcc 13 <210> 80 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 80 ccatagttat ggt 13 <210> 81 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 81 tgcgagatcg aag 13 <210> 82 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 82 agagaagtcg agt 13 <210> 83 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 83 ggtaactcca tat 13 <210> 84 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 84 tgctattcca ggc 13 <210> 85 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 85 aaccgcgagg ctc 13 <210> 86 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 86 ttctagagat acc 13 <210> 87 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 87 ttcgctcaag tat 13 <210> 88 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 88 cagagaaggc gca 13 <210> 89 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 89 tagaattggc ctc 13 <210> 90 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 90 ggccattctc cag 13 <210> 91 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 91 tccaacgcgc gtt 13 <210> 92 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 92 gccgcagatt acg 13 <210> 93 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 93 gcagttcgaa cgc 13 <210> 94 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 94 ttctctctgc agg 13 <210> 95 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 95 taagctacca gcg 13 <210> 96 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 96 ctgcatgagg ttg 13 <210> 97 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 97 ttgcctagcg agg 13 <210> 98 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 98 caactgaatt agg 13 <210> 99 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 99 aagcggtcct ctt 13 <210> 100 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 100 aatggaagga ccg 13 <210> 101 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 101 gagttagtaa gtt 13 <210> 102 <21{1}> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 102 ttcctaattc caa 13 <210> 103 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 103 gttctggttc gct 13 <210> 104 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 104 Note: There seems to be a small error in the original text where the "1" in "<211> 13" was likely a typo and should be "11" as per the pattern. It has been corrected in the translation for better readability. If this is not an error, please let me know and I'll adjust the translation accordingly. gttcatctct tcc 13 <210> 105 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 105 attccgagga aga 13 <210> 106 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 106 cttagccgag aga 13 <210> 107 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 107 gtctgctacg ctt 13 <210> 108 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 108 atggcgccgc gca 13 <210> 109 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 109 taattggtta tct 13 <210> 110 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 110 tcggttataa gtc 13 <210> 111 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 111 tgcctgagaa cgt 13 <210> 112 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 112 agatgcggtt aac 13 <210> 113 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 113 atggaatagg cga 13 <210> 114 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 114 agagatgcga tcg 13 <210> 115 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 115 ctccaactaa cgt 13 <210> 116 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 116 gccttgctac tgg 13 <210> 117 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 117 cttcgtctct acg 13 <210> 118 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 118 acgctcatag cct 13 <210> 119 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 119 gtcgaagata agg 13 <210> 120 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 120 gccggagtcc tcg 13 <210> 121 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 121[[ID=:39]] tatacggcga cct 13 <210> 122 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 122 aggtagatat tcg 13 <210> 123 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 123 ttaaggtact gct 13 <210> 124 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 124 cggatctggt ata 13 <210> 125 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 125 gaggtctcgg agg 13 <210> 126 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 126 ggcatcgatg gac 13 <210> 127 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 127 gatctccgat ata 13 <210> 128 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 128 gattcggaat act 13 <210> 129 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 129 ctgcgatccg gcc 13 <210> 130 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 130 gatccggttg caa 13 <210> 131 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 131 cgtcaggctt gac 13 <210> 132 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 132 tcggcaaggc gag 13 <210> 133 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 133 gaacggcgaa cgc 13 <210> 134 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 134 cctcaagcgg act 13 <210> 135 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 135 gaagccagat ggt 13 <210> 136 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 136 tgctcatacc aat 13 <210> 137 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> i7 custom index primer (Table 2) <220> <221> misc_feature <222> (25)..(36) <223> n is a, c, g, or t <400> 137 caagcagaag acggcatacg agatnnnnnn nnnnnngtct cgtgggctcg g 51 <210> 138 <211> 55 <212> DNA <213> Artificial Sequence <220> <223> i5 custom index primer (Table 2) <220> <221> misc_feature <222> (30)..(41) <223> n is a, c, g, or t <400> 138 aatgatacgg cgaccaccga gatctacacn nnnnnnnnnn ntcgtcggca gcgtc 55 <210> 139 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> i7 flow cell primer (Table 3) <400> 139 caagcagaag acggcatacg a 21 <210> 140 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> i5 flow cell primer (Table 3) <400> 140 aatgatacgg cgaccaccga 20 <210> 141 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> single_mut for mutagenesis (Table 5) <220> <221> misc_feature <222> (19)..(34) <223> n is a, c, g, or t <400> 141 tcggtctgcg cctctagcnn nnnnnnnnnn nnnngtctcg tgggctcgga g 51 <210> 142 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> single_rec primer for recovery (Table 5) <400> 142 caagcagaag acggcatacg agattcggtc tgcgcctcta gc 42
Claims
**Claim 1** The following a. and b., namely a. providing at least one target nucleic acid molecule; b. using a DNA polymerase to amplify the at least one target nucleic acid molecule in the presence of nucleotide analogs and using non-uniform concentrations of dNTPs to introduce mutations into the at least one target nucleic acid molecule; and c. amplifying the at least one target nucleic acid molecule with mutations in the absence of the nucleotide analogs, A method for producing at least one target nucleic acid molecule into which a mutation has been introduced, comprising the steps of: **Claim 2** The method according to claim 1, wherein the nucleotide analog is dPTP. **Claim 3** The method according to claim 1, wherein step c. is performed using non-uniform concentrations of dNTPs. **Claim 4** The method according to any one of claims 1 to 3, wherein the non-uniform concentrations of dNTPs comprise dATP, dCTP, dTTP, and dGTP, and one or more of dATP, dCTP, dTTP, or dGTP are at a lower concentration compared to the other dNTPs. **Claim 5** The method according to any one of claims 1 to 4, wherein the use of non-uniform concentrations of dNTPs comprises identifying a dNTP whose level should be increased or decreased to reduce bias in the profile of mutations introduced into the at least one target nucleic acid molecule with mutations. **Claim 6** The method according to any one of claims 1 to 5, wherein the non-uniform concentrations of dNTPs comprise dTTP at a lower concentration than the other dNTPs. **Claim 7** The method according to claim 6, wherein the non-uniform concentrations of dNTPs comprise dTTP at a concentration less than 75% of the concentration of dATP, dCTP, or dGTP. **Claim 8** The method according to claim 6, wherein the non-uniform concentrations of dNTPs comprise dTTP at a concentration of 25% to 60% of the concentration of dCTP. **Claim 9** The method according to any one of claims 1 to 5, wherein the non-uniform concentrations of dNTPs comprise dATP at a lower concentration than the other dNTPs. **Claim 10** The method according to claim 9, wherein the non-uniform concentrations of dNTPs comprise dATP at a concentration less than 75% of the concentration of dTTP, dCTP, or dGTP. **Claim 11** The method according to claim 9, wherein the non-uniform concentrations of dNTPs comprise dATP at a concentration of 25% to 60% of the concentration of dGTP. **Claim 12** The method according to any one of claims 1 to 11, wherein the at least one target nucleic acid molecule comprises a plurality of target nucleic acid molecules. **Claim 13** The method according to any one of claims 1 to 12, wherein the amplification comprises amplification by polymerase chain reaction. **Claim 14** A method for determining the sequence of at least one target DNA molecule, comprising the following steps: a. preparing at least one mutated target DNA molecule by performing the method according to any one of claims 1 to 13; b. sequencing a region of the at least one mutated target DNA molecule to prepare a mutated sequence read; and c. using the mutated sequence read to assemble the sequence of at least a part of the at least one mutated target DNA molecule to determine the sequence of the at least one target DNA molecule. The method as described above.
Citation Information
Patent Citations
Compositions and methods for random nucleic acid mutagenesis
JP2005501527A
Structure for presenting desired peptide sequences
US20030215914A1
Polypeptides having DNA polymerase activity
WO2005118815A1
Amplification methods to minimise sequence specific bias
WO2011106368A2