Sequencing Algorithm
The method sequences nucleic acid molecules by analyzing non-mutated and mutated reads to resolve repetitive regions and structural mutations, enhancing sequencing accuracy and efficiency.
Patent Information
- Application Number
- JP2024018627
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-20
- Filing Date
- 2024-02-09
- Publication Date
- 2026-01-07
- Estimated Expiration
- 2039-08-12
AI Technical Summary
Next-generation sequencing technologies struggle with accurately sequencing nucleic acid molecules containing repetitive regions and resolving structural mutations, often leading to ambiguity in distinguishing between repeats and similar sequences, and require complex, processing-intensive methods like sequencing assisted mutagenesis (SAM) to construct sequences.
A method involving sequencing a pair of samples, one non-mutated and one mutated, to generate non-mutated and mutant sequence reads, analyzing mutant reads to construct the sequence of the target template nucleic acid molecule, and controlling the number of template molecules in the sample to optimize sequencing efficiency.
Enables accurate, efficient sequencing of nucleic acid molecules by distinguishing repetitive regions and resolving structural mutations, reducing artifacts from mutations, and optimizing data coverage for sequence construction.
Smart Images

Figure 0007795566000019 
Figure 0007795566000020 
Figure 0007795566000021
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for sequencing at least one target template nucleic acid molecule using non-mutated sequence reads and mutant sequence reads. The present invention also relates to a method for sequencing at least one target template nucleic acid molecule in a sample, comprising controlling or normalizing the number of target template nucleic acid molecules in the sample. The present invention also relates to a computer program adapted to perform the method, a computer-readable medium comprising the computer program, and a computer-implemented method. [Background technology]
[0002] The ability to sequence nucleic acid molecules is a very useful tool in a wide variety of applications. However, it can be difficult to determine the exact sequence of nucleic acid molecules that contain difficult structures, such as nucleic acid molecules that contain repetitive regions. It can also be difficult to resolve structural mutations, such as the haplotype structure of diploid and polyploid organisms.
[0003] Many of the more modern technologies (so-called next-generation sequencing technologies) can only accurately sequence short nucleic acid molecules. Next-generation sequencing technologies can be used to sequence longer nucleic acid sequences, but this is often difficult. Next-generation sequencing technologies can be used to generate short sequence reads corresponding to the sequence of a portion of a nucleic acid molecule, and a complete sequence can be constructed from the short sequence reads. When a nucleic acid molecule contains a repetitive region, it may be unclear to the user whether two sequence reads with similar sequences correspond to the sequences of two repeats in a longer sequence or two copies of the same sequence. Similarly, a user may want to simultaneously sequence two similar nucleic acid molecules, and it may be difficult to determine whether two sequence reads with similar sequences correspond to the sequences of the same original nucleic acid molecule or the sequences of two different original nucleic acid molecules.
[0004] The construction of sequences from short sequence reads can be facilitated by using sequencing assisted mutagenesis (SAM) technology. Generally, SAM involves introducing mutations into target template nucleic acid sequences. The introduced mutation patterns can help the user of this method to construct sequences of nucleic acid molecules from short sequence reads.
[0005] For example, if the template nucleic acid molecule contains repetitive regions, the repeats can be distinguished from one another by different mutation patterns, allowing the repetitive regions to be correctly separated and assembled.
[0006] Generally, SAM technology involves mutating copies of a target template nucleic acid molecule and then assembling the sequences of the mutated copies based on their mutation patterns. The user can then generate a consensus sequence from the sequences of the mutated copies. Because the mutated copies thereby contain mutations at different positions, the consensus sequence may be representative of the original template nucleic acid molecule. However, the consensus sequence may contain artifacts from the mutation process. Furthermore, generating a consensus sequence involves the use of a computer program, which is complex and processing-intensive.
[0007] Thus, there remains a need for methods for sequencing at least one target template nucleic acid molecule from which sequence reads can be constructed accurately, quickly, and efficiently. Summary of the Invention
[0008] Overview of the invention The present inventors have developed a new and improved method for determining the sequence of at least one target template nucleic acid molecule. Accordingly, in a first aspect of the present invention, (a) providing a pair of samples, each sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule in a first of the pair of samples to obtain a non-mutated sequence read; (c) introducing a mutation into at least one target template nucleic acid molecule in a second of the pair of samples to obtain at least one mutated target template nucleic acid molecule; (d) sequencing a region of at least one mutation target template nucleic acid molecule to obtain a mutation sequence read; (e) analyzing the mutant sequence reads and assembling a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads using information obtained from the analysis of the mutant sequence reads. The present invention provides a method for sequencing at least one target template nucleic acid molecule, comprising:
[0009] In a second aspect of the present invention, (a)(i) non-mutated sequence reads, and (ii) Mutation sequence reads obtaining data including (b) analyzing the mutant sequence reads and constructing a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads using information obtained from the analysis of the mutant sequence reads. The present invention provides a method for generating the sequence of at least one target template nucleic acid molecule, comprising:
[0010] In a third aspect of the invention, there is provided a computer program adapted to carry out the method of the invention.
[0011] In a fourth aspect of the present invention, there is provided a computer readable medium comprising the computer program of the present invention.
[0012] In a fifth aspect of the present invention, there is provided a computer-implemented method comprising the method of the present invention.
[0013] In a sixth aspect of the present invention, (a) providing at least one sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule; and (c) constructing the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule. The present invention provides a method for sequencing at least one target template nucleic acid molecule, comprising: (i) providing at least one sample containing at least one target template nucleic acid molecule includes controlling the number of target template nucleic acid molecules in the at least one sample; and / or (ii) preparing at least one sample by pooling two or more subsamples (partial samples), and normalizing the number of target template nucleic acid molecules in each of the subsamples;
[0014] In a sixth aspect of the present invention, (a) providing at least one sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule; and (c) constructing the sequence of at least a portion of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule. The present invention provides a method for sequencing at least one target template nucleic acid molecule, comprising: (i) providing at least one sample containing at least one target template nucleic acid molecule includes controlling the number of target template nucleic acid molecules in the at least one sample; and / or (ii) preparing at least one sample by pooling two or more subsamples, and normalizing the number of target template nucleic acid molecules in each of the subsamples; [Brief explanation of the drawings]
[0015] [Figure 1-1]Figure 1-1 shows the levels of mutations obtained with three different polymerases in the presence or absence of dPTP. Panel A shows data obtained using Taq (Jena Biosciences), panel B shows data obtained using LongAmp (New England Biolabs), and panel C shows data obtained using Primestar GXL (Takara). Dark gray bars show results obtained in the absence of dPTP, and light gray bars show results obtained in the presence of 0.5 mM dPTP. [Figure 1-2] This is a continuation of Figure 1-1. [Figure 2] Figure 2 shows the mutation rates obtained by dPTP mutagenesis using Thermococcus polymerase (Primestar GXL; Takara) on templates of various G + C content. The median mutation rate observed was approximately 7% for the S. aureus low-GC template (33% GC), while the median for other templates was approximately 8%. [Figure 3-1] Figure 3-1 is the sequence listing. [Figure 3-2] This is a continuation of Figure 3-1. [Figure 3-3] This is a continuation of Figure 3-1. [Figure 3-4] This is a continuation of Figure 3-1. [Figure 3-5] This is a continuation of Figure 3-1. [Figure 3-6] This is a continuation of Figure 3-1. [Figure 3-7] This is a continuation of Figure 3-1. [Figure 3-8] This is a continuation of Figure 3-1. [Figure 3-9] This is a continuation of Figure 3-1. [Figure 3-10] This is a continuation of Figure 3-1. [Figure 3-11] This is a continuation of Figure 3-1. [Figure 3-12] This is a continuation of Figure 3-1. [Figure 4] FIG. 4 shows the lengths of the fragments obtained using the method described in Example 5. [Figure 5]Figure 5 shows the distribution of values using variational inference on the simulated data. Panel A shows the values of M estimated using variational inference on the simulated data. The true value is 0.895 for identities ([1,1], [2,2], [3,3], [4,4]), 0.1 for transitions ([1,3], [2,4], [3,1], [4,2]), and 0.005 for transversions (all other entries). Panel B shows the values of z estimated using variational inference on the simulated data. The true value of z is 1 for the same [1:5] and 0 for the same [91:95]. [Figure 6] Figure 6 shows precision-recall plots for simulated data using cutoff values ranging from 100 to 10,000 in steps of 100. 2,000 trials were performed for each threshold, including 1,000 read pairs that originated from the same template and 1,000 that did not. [Figure 7] FIG. 7 is a flow diagram illustrating a method of sequencing at least one target template nucleic acid molecule of the present invention. [Figure 8] FIG. 8 is a flow diagram illustrating a method for generating the sequence of at least one target template nucleic acid molecule of the present invention. [Figure 9] FIG. 9 shows an assembly graph in panel A and, in panel B, mapping of mutant sequence reads to the assembly graph. [Figure 10] FIG. 10 shows the size of target nucleic acid molecules amplified using adapters that anneal to each other (right line) or standard adapters (left line). [Figure 11] 11 is a graph showing the linear relationship between sample dilution factor and the observed number of unique templates. A starting sample of target template nucleic acid molecules was serially diluted and end-sequenced to identify and quantify the number of unique templates at each dilution. [Figure 12]Figure 12 shows graphs demonstrating normalization of template counts among individual samples within a pool. (A) shows the number of unique templates for 66 barcoded bacterial genomes determined from pooled samples before normalization. (B) shows the number of templates (expressed per megabase (Mb) of genome content) for the same samples after normalization, demonstrating much lower variation. [Figure 13] FIG. 13 shows the workflow for the construction (assembly) of a bacterial genome according to the present invention. [Figure 14] FIG. 14 shows comparative assembly statistics from 65 bacterial genomes of a standard read assembly compared to our assembly (Morphoseq assembly). [Figure 15] FIG. 15 shows exemplary assembly metrics for the assembly of a bacterial genome for short read assemblies compared to the assembly of the present invention. [Figure 16]Figure 16 shows an exemplary workflow of the present invention for generating synthetic long reads. (a) Preparation of long mutant templates. First, target genomic DNA is tagged to obtain long templates containing end adapters. The templates are then amplified in the presence of the mutagenic nucleotide analog dPTP, which is randomly incorporated opposite A and G residues on both product strands (mutagenesis PCR). This step also introduces (i) a sample tag and (ii) an additional adapter sequence at the template end to facilitate downstream amplification of products containing P bases. Further amplification (recovery PCR) is performed in the absence of dPTP, during which template P residues are replaced with natural nucleotides, generating transition mutations (shown in red). The sample is then size-selected (8-10 kb), restricting it to a certain number of unique templates and selectively enriching to obtain multiple copies of each unique molecule. (b) Preparation, sequencing, and analysis of short-read libraries. The long mutant templates are then processed for short-read sequencing by further tagging and library amplification. During this step, fragments derived from both ends of the full-length template are amplified and barcoded separately from the random "internal" fragments using different primers targeting the original template end adapters (dark gray) and internally tagged adapters (light gray). Both libraries are sequenced alongside a non-mutated reference library generated in parallel, and a custom algorithm is used to reconstruct synthetic long reads. This involves generating an assembly graph from the reference data, mapping mutated reads against it and linking them together due to different patterns of overlapping mutations. The final synthetic long reads correspond to the identified pathways through the unmatched all-mutation assembly graph. DETAILED DESCRIPTION OF THE INVENTION
[0016] Detailed Description of the Invention General definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0017] In general, the word "comprising" is intended to mean inclusive without limitation. For example, the phrase "a method for sequencing at least one target template nucleic acid molecule, comprising [a step]" should be interpreted to mean that the method includes the recited steps, and that additional steps may be performed.
[0018] In some embodiments of the invention, the word "comprising" is replaced with the word "consisting." The word "consisting" is intended to be limiting. For example, the phrase "a method for sequencing at least one target template nucleic acid molecule, comprising [a step]" should be interpreted to mean that the method includes the recited steps and that no additional steps are performed.
[0019] Method for sequencing at least one target template nucleic acid molecule In some aspects, the present invention provides methods for sequencing at least one target template nucleic acid molecule or for generating the sequence of at least one target template nucleic acid molecule.
[0020] For purposes of the present invention, the terms "determining" and "generating" may be used interchangeably. However, sequence "determining" methods generally include steps such as sequencing steps, while sequence "generating" methods may be limited to steps that can be implemented on a computer.
[0021] The method can be used to determine or generate the complete sequence of at least one target template nucleic acid molecule. Alternatively, the method can be used to determine or generate a partial sequence (i.e., a partial sequence of the molecule) of at least one target template nucleic acid molecule. For example, when it is impossible or difficult to determine the complete sequence, the user may find that a partial sequence of at least one target template nucleic acid molecule is useful or even sufficient for their purpose.
[0022] For purposes of the present invention, a "nucleic acid molecule" refers to a polymeric form of nucleotides of any length. Nucleotides can be deoxyribonucleotides, ribonucleotides, or their analogs. Preferably, at least one target template nucleic acid molecule is composed of deoxyribonucleotides or ribonucleotides. Even more preferably, at least one target template nucleic acid molecule is composed of deoxyribonucleotides, i.e., at least one target template nucleic acid molecule is a DNA molecule. At least one "target template nucleic acid molecule" can be any nucleic acid molecule that a user wishes to sequence. At least one "target template nucleic acid molecule" can be single-stranded or can be part of a double-stranded complex. If at least one target template nucleic acid molecule is composed of deoxyribonucleotides, it can form part of a double-stranded DNA complex. In this case, one strand (e.g., the coding strand) is considered to be at least one target template nucleic acid molecule, and the other strand is a nucleic acid molecule complementary to at least one target template nucleic acid molecule. The at least one target template nucleic acid molecule may be a DNA molecule corresponding to a gene, may contain introns, may be an intergenic region, may be an intragenic region, may be a genomic region spanning multiple genes, or may actually be the entire genome of an organism.
[0023] The terms "at least one target template nucleic acid molecule" and "at least one of the target template nucleic acid molecules" are considered synonymous and may be used interchangeably herein.
[0024] In the methods of the present invention, any number of at least one target template nucleic acid molecule can be simultaneously sequenced. Thus, in one embodiment of the present invention, the at least one target template nucleic acid molecule comprises a plurality of target template nucleic acid molecules. Optionally, the at least one target template nucleic acid molecule comprises at least 10, at least 20, at least 50, at least 100, or at least 250 target template nucleic acid molecules. Optionally, the at least one target template nucleic acid molecule comprises 10-1000, 20-500, or 50-100 target template nucleic acid molecules.
[0025] A method for sequencing at least one target template nucleic acid molecule, comprising: (a) providing a pair of samples, each sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule in a first of the pair of samples to obtain a non-mutated sequence read; (c) introducing a mutation into at least one target template nucleic acid molecule in a second of the pair of samples to obtain at least one mutated target template nucleic acid molecule; (d) sequencing a region of at least one mutation target template nucleic acid molecule to obtain a mutation sequence read; (e) analyzing the mutant sequence reads and assembling a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads using information obtained from the analysis of the mutant sequence reads. Includes:
[0026] A method for generating a sequence of at least one target template nucleic acid molecule, comprising: (a)(i) non-mutated sequence reads, and (ii) Mutation sequence reads obtaining data including (b) analyzing the mutant sequence reads and constructing a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads using information obtained from the analysis of the mutant sequence reads. Includes:
[0027] Preparing pairs of samples, each containing at least one target template nucleic acid molecule. A method of sequencing at least one target template nucleic acid molecule can include providing a pair of samples, each sample comprising at least one target template nucleic acid molecule.
[0028] The method of the present invention uses information obtained from the analysis of the mutant sequence reads to construct the sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads. The method of the present invention may include introducing a mutation into at least one target template nucleic acid molecule in a second sample pair. Thus, a mutant sequence read can be obtained by sequencing a region of at least one mutant target template nucleic acid molecule in the second sample pair, and a non-mutated sequence read can be obtained by sequencing a region of at least one non-mutated target template nucleic acid molecule in the first sample pair.
[0029] A portion of the mutant sequence reads and a portion of the non-mutated sequence reads correspond to the same original target template nucleic acid molecule so that a user can use information obtained by analyzing mutant sequence reads from a second sample to construct a sequence that primarily comprises non-mutated sequence from a first sample.
[0030] For example, if a user wishes to sequence target template nucleic acid molecules A and B, a first sample contains template nucleic acid molecules A and B, and a second sample contains template nucleic acid molecules A and B. A and B in the first sample can be sequenced to obtain unmutated sequence reads for A and B, and A and B in the second sample can be mutated and sequenced to obtain mutant sequence reads for A and B.
[0031] Because the first of the sample pair and the second of the sample pair both contain at least one target template nucleic acid molecule, the sample pair can be derived from the same target organism or taken from the same original sample.
[0032] For example, if a user intends to sequence at least one target template nucleic acid molecule in a sample, the user can collect a pair of samples from the same original sample. Optionally, the user can replicate at least one target template nucleic acid molecule in the original sample and then collect a pair of samples from it. The user may intend to sequence various nucleic acid molecules derived from a particular organism, such as E. coli. In this case, the first of the pair of samples can be a sample of E. coli from one source, and the second of the pair of samples can be a sample of E. coli from a second source.
[0033] The sample pair can be from any source that contains or is suspected of containing at least one target template nucleic acid molecule. The sample pair can include a sample of nucleic acid molecules from a human, for example, a sample extracted from a skin swab of a human patient. Alternatively, the sample pair can be from another source, such as a water supply. Such a sample can contain billions of template nucleic acid molecules. Using the methods of the present invention, it is possible to simultaneously sequence each of these billions of target template nucleic acid molecules, and therefore there is no upper limit to the number of target template nucleic acid molecules that can be used in the methods of the present invention.
[0034] In one embodiment, multiple pairs of samples may be prepared. For example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 15, 20, 25, 50, 75, or 100 pairs of samples may be prepared. Optionally, fewer than 100, 75, 50, 25, 20, 15, 11, 10, 9, 8, 7, 6, 5, or 4 samples may be prepared. Optionally, 2 to 100, 2 to 75, 2 to 50, 2 to 25, 5 to 15, or 7 to 15 pairs of samples may be prepared.
[0035] When preparing multiple pairs of samples, at least one target template nucleic acid molecule in different pairs of samples can be labeled with different sample tags.For example, when user intends to prepare two pairs of samples, all or substantially all of at least one target template nucleic acid molecule in the first pair of samples can be labeled with sample tag A, and all or substantially all of at least one target template nucleic acid molecule in the second pair of samples can be labeled with sample tag B.Sample tag is described in more detail in the section " sample tag and barcode ".
[0036] Controlling the number of target template nucleic acid molecules in a sample As described above, the sequencing method of the present invention involves constructing the sequence of at least a portion of at least one target template nucleic acid molecule from non-mutated reads using information obtained from the analysis of corresponding mutant sequence reads. Typically, target template nucleic acid molecules in a sample can be combined to generate the sequence of one or more larger nucleic acid molecules present in the sample. In a representative embodiment, target template nucleic acid molecules can be combined to generate the sequence of a genome. Sequencing operations generate a finite amount of data in the form of resulting sequencing reads. To construct the sequence of a target template nucleic acid molecule from the resulting sequencing reads (and thus to combine target template nucleic acid molecules to generate the sequence of one or more larger target template nucleic acid molecules), it is preferable to ensure that the coverage of the target template nucleic acid molecule between the sequencing reads is adequate (i.e., sufficient to construct a sequence) and that excessively redundant (i.e., overlapping) sequencing reads are not generated for each target template nucleic acid molecule. For example, if a sample contains too many target template nucleic acid molecules and not enough sequence reads are generated from each target template nucleic acid molecule, it may be impossible to construct a sequence for each target template nucleic acid molecule (i.e., there may not be enough data for each template). On the other hand, if a sample contains too few target template nucleic acid molecules, it may be possible to construct each target template nucleic acid molecule, but it may not be possible to combine the target template nucleic acid molecules to generate a sequence for a larger nucleic acid molecule, such as a genome (i.e., there may be too much data for each template, and therefore insufficient data for the entire sample).
[0037] With these considerations in mind, it is convenient for the user to control the number of unique target template nucleic acid molecules present in the first sample pair and / or the second sample pair.The user can then select the optimal number of unique target template nucleic acid molecules present in the first sample pair and / or the second sample pair.The optimal number of unique target template nucleic acid molecules can depend on several different factors that are understood by the user.For example, if the target template nucleic acid molecules are longer, their sequencing will be more difficult, and the user will want to select fewer unique target template nucleic acid molecules.
[0038] Thus, the methods of the invention can include a step of providing pairs of samples, each sample containing at least one target template nucleic acid molecule, which step includes controlling the number of target template nucleic acid molecules in the first and / or second of the sample pair.
[0039] It may be useful to control the number of target template nucleic acid molecules in the first sample pair.However, it is particularly preferred that the number of target template nucleic acid molecules in the second sample pair be controlled for the second sample pair (i.e., the sample that contains at least one target template nucleic acid molecule to be introduced mutation).In the method of the present invention, at least one target template nucleic acid molecule in the second sample pair is mutated, and is used to reconstruct the sequence of target template nucleic acid molecule.In this case, the number of target template nucleic acid molecules in the second sample pair may be crucial.Therefore, it may be particularly advantageous to control the number of target template nucleic acid molecules in the second sample pair.
[0040] Similarly, in one aspect of the invention, (a) providing at least one sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule; and (c) constructing the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule. The present invention provides a method for sequencing at least one target template nucleic acid molecule, comprising:
[0041] Similarly, in one aspect of the invention, (a) providing at least one sample containing at least one target template nucleic acid molecule; (b) sequencing at least a portion of at least one target template nucleic acid molecule; and (c) constructing the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule. The present invention provides a method for sequencing at least one target template nucleic acid molecule, comprising:
[0042] For purposes of this application, the phrase "controlling the number of target template nucleic acid molecules" in a sample refers to obtaining a desired number of target template nucleic acid molecules in a sample. In certain embodiments, this may involve manipulating or adjusting the sample to contain a desired number of target template nucleic acid molecules (e.g., by diluting the sample or pooling the sample with another sample that also contains target template nucleic acid molecules).
[0043] It is understood that "controlling the number of target template nucleic acid molecules" does not have to be completely precise. This is because, for example, it is difficult to obtain the exact number of template nucleic acid molecules by diluting a sample using conventional techniques. However, if a user finds that a sample contains about twice the desired number of target template nucleic acid molecules, the user can dilute the sample to obtain a diluted sample containing about half the number of target template nucleic acid molecules present in the original sample (e.g., 45% to 55% of the number of target template nucleic acid molecules present in the original sample).
[0044] Controlling the number of target template nucleic acid molecules can include measuring the number of target template nucleic acid molecules in a sample (e.g., a user can measure the number of target template nucleic acid molecules in the first of a pair of samples, the second of a pair of samples, or at least one sample). The term "measuring" can be substituted for the term "estimating" herein. Generally, measuring the number of target template nucleic acid molecules in a sample is performed as part of a process for controlling the number of target template nucleic acid molecules in a sample, and controlling the number of target template nucleic acid molecules in a sample can be performed to help a user confirm that the sample contains a number of target template nucleic acid molecules appropriate for use in a particular sequencing method (i.e., within a desired range). However, such a process for controlling the number of target template nucleic acid molecules need not be completely rigorous. A method for approximately controlling the number of target template nucleic acid molecules in a sample would be useful for improving methods for sequencing target template nucleic acid molecules. In one embodiment, "measuring the number of target template nucleic acid molecules" refers to determining the number of target template nucleic acid molecules in a sample to at least within the correct order of magnitude, i.e., within 10-fold, more preferably within 5-fold, 4-fold, 3-fold, or 2-fold of the true number. More preferably, the number of target template nucleic acid molecules in a sample can be determined to within at least 50%, or within 40%, or within 30%, or within 25%, or within 20%, or within 15%, or within 10% of the true number. Any method can be used to measure the number of target template nucleic acid molecules in a sample.
[0045] A sample (for example, the first of a pair of samples, the second of a pair of samples, or at least one sample) can be diluted before or during the measurement of the number of target template nucleic acid molecules in the sample.For example, if a user believes that a sample contains a large number of target template nucleic acid molecules, he may want to dilute the sample to obtain a sample containing an appropriate number of target template nucleic acid molecules, for example, for accurate measurement by sequencing.Therefore, a diluted sample can be prepared.Therefore, the number of target template nucleic acid molecules can be measured in the diluted sample, thereby determining the number of target template nucleic acid molecules in the sample.
[0046] In certain embodiments, it may be advantageous to prepare multiple diluted samples, each with a different dilution factor. For example, if the user is unsure of the number of target template nucleic acid molecules present in a sample, he or she may wish to generate a dilution series and measure the number of target template nucleic acid molecules in each dilution (i.e., each diluted sample). Thus, measuring the number of target template nucleic acid molecules may involve preparing a dilution series in the first of a pair of samples, the second of a pair of samples, or at least one sample to obtain a dilution series including diluted samples. The dilution series may include 1-50, 1-25, 1-20, 1-15, 1-10, 1-5 diluted samples, 5-25, 5-20, 5-15, or 5-10 diluted samples.
[0047] Such a dilution series can be prepared by performing serial dilutions. As desired, a sample can be diluted 2-fold to 20-fold, 5-fold to 15-fold, or about 10-fold. For example, to obtain a dilution series of 10 samples, each diluted 10-fold, the user would prepare a 10-fold dilution of the sample, then isolate a portion of the diluted sample and further dilute it 10-fold, and so on, until 10 diluted samples are obtained.
[0048] A user may prepare 10 diluted samples and only determine the number of target template nucleic acid molecules in fewer than 10 of the diluted samples. For example, if a user determines the number of target template nucleic acid molecules in five diluted samples and accurately determines the number of target template nucleic acid molecules in the fifth diluted sample, there is no need to further determine the number of target template nucleic acid molecules in any of the other diluted samples. In yet other embodiments, a user may correlate results from multiple diluted samples to further increase the reliability of the results. Advantageously, this may also provide the user with information about the dynamic range over which the number of target template nucleic acid molecules in a sample can be accurately determined under a given set of conditions. However, a user may only need to perform a single dilution to accurately determine the number of target template nucleic acid molecules in a sample.
[0049] In certain embodiments, the number of target template nucleic acid molecules in a sample (or diluted sample) can be measured by determining the molar concentration of the target template nucleic acid molecules in the sample. This can be done, for example, by electrophoresis. In certain embodiments, the number of target template nucleic acid molecules in a sample can be determined by high-resolution microfluidic electrophoresis, in which the sample can be loaded into a microchannel, and the target template nucleic acid molecules can be separated by electrophoresis and detected by their fluorescence. Systems suitable for measuring the number of target template nucleic acid molecules in this way include the Agilent 2100 Bioanalyzer and the Agilent 4200 Tapestation.
[0050] In another embodiment, the number of target template nucleic acid molecules can be determined by sequencing the target template nucleic acid molecules in one or more of the first of the pair of samples, the second of the pair of samples, at least one sample, or a diluted sample.
[0051] In certain embodiments, the method may include determining the number of target template nucleic acid molecules by sequencing the target template nucleic acid molecules in one or more of the diluted samples.
[0052] The target template nucleic acid can be sequenced using any sequencing method. Examples of possible sequencing methods include Maxam-Gilbert sequencing, Sanger sequencing, sequencing involving bridge amplification (e.g., bridge PCR), or any high-throughput sequencing (HTS) method, such as those described in Maxam AM, Gilbert W (February 1977), "A new method for sequencing DNA", Proc. Natl. Acad. Sci. USA 74 (2): 560-4; Sanger F, Coulson AR (May 1975), "A rapid method for determining sequences in DNA by primed synthesis with DNA polymerase", J. Mol. Biol. 94 (3): 441-8; and Bentley DR, Balasubramanian S et al. (2008), "Accurate whole human genome sequencing using reversible terminator chemistry", Nature 456 (7218): 53-59. Determining the number of target template nucleic acid molecules can involve amplifying the target template nucleic acid molecules in one or more of the first of a pair of samples, the second of a pair of samples, at least one sample, or a diluted sample, and then sequencing the target template nucleic acid molecules (or, alternatively, the amplified target template nucleic acid molecules). Amplifying the target template nucleic acid molecules provides the user with multiple copies of the target template nucleic acid molecules, allowing the user to more accurately sequence the target template nucleic acid molecules (since sequencing techniques are not perfectly accurate, sequencing multiple copies of the target template nucleic acid sequence and then calculating a consensus sequence from the sequences of the copies improves accuracy). Producing multiple copies of a certain number of unique target template nucleic acid molecules in a sample and sequencing a portion of the entire (amplified) sample allows for obtaining sequence information from all of the target template nucleic acid molecules.
[0053] Suitable methods for amplifying at least one target template nucleic acid molecule are known in the art. For example, PCR is commonly used. PCR is described in more detail in the section "Introduction of mutations into at least one target template nucleic acid molecule."
[0054] In a typical embodiment, the sequencing step can include bridge amplification. Optionally, the bridge amplification step is performed with an extension time of more than 5 seconds, more than 10 seconds, more than 15 seconds, or more than 20 seconds. An example of the use of bridge amplification is in the Illumina Genome Analyzer Sequencer. Preferably, paired-end sequencing is used.
[0055] Determining the number of target template nucleic acid molecules may include fragmenting the target template nucleic acid molecules in one or more of the first sample pair, the second sample pair, at least one sample, or diluted samples. This may be particularly advantageous, for example, when the sequencing platform precludes the use of long nucleic acid molecules as templates. Fragmentation may be performed using any suitable technique. For example, fragmentation may be performed using restriction digestion or PCR using primers complementary to at least one internal region of at least one mutant target nucleic acid molecule. Preferably, fragmentation is performed using a technique that generates random fragments. The term "random fragments" refers to randomly generated fragments, such as fragments generated by tagging. Fragments generated using restriction enzymes are not "random" because restriction digestion occurs at a specific DNA sequence defined by the restriction enzyme used. Even more preferably, fragmentation is performed by tagging. When fragmentation is performed by tagging, the tagging reaction may optionally introduce an adapter region into the target template nucleic acid molecule. The adapter region is a short DNA sequence that may encode an adapter that allows at least one target nucleic acid molecule to be sequenced using, for example, Illumina technology.
[0056] In certain embodiments, determining the number of target template nucleic acid molecules comprises amplifying and fragmenting the target template nucleic acid molecules in one or more of the first sample pair, the second sample pair, at least one sample, or a diluted sample, and then sequencing the target template nucleic acid molecules (or, alternatively, the amplified and fragmented target template nucleic acid molecules). Amplification and fragmentation can be performed in any order prior to sequencing. In one embodiment, determining the number of target template nucleic acid molecules can comprise amplifying, then fragmenting, and then sequencing the target template nucleic acid molecules in one or more of the first sample pair, the second sample pair, at least one sample, or a diluted sample. Alternatively, determining the number of target template nucleic acid molecules can comprise fragmenting, then amplifying, and then sequencing the target template nucleic acid molecules in one or more of the first sample pair, the second sample pair, at least one sample, or a diluted sample. Alternatively, amplification and fragmentation can be performed simultaneously, i.e., in a single step. If the target template nucleic acid molecule is very long (e.g., too long to be sequenced using conventional techniques), it may be useful for the method to include fragmenting and then amplifying the target template nucleic acid molecule.
[0057] Measuring the number of target template nucleic acid molecules can include determining the total number of target template nucleic acid molecules in sample.However, preferably, measuring the number of target template nucleic acid molecules includes determining the number of unique target template nucleic acid molecule sequences in one or more of the first pair of samples, the second pair of samples, at least one sample or diluted sample.As mentioned above, when at least one target template nucleic acid sequence is part of a sample that contains a large number of different target template nucleic acid sequences, it is more difficult to sequence the at least one target template nucleic acid sequence.Therefore, by reducing the number of unique target template nucleic acid molecules, the method for sequencing at least one target template nucleic acid molecule becomes more convenient.
[0058] As described elsewhere herein, introducing mutations into a target template nucleic acid sequence can facilitate the construction of at least a portion of the sequence of the target template nucleic acid. Mutation of a target template nucleic acid molecule can be particularly useful, for example, in determining whether sequence reads are likely to be derived from the same target template nucleic acid molecule, or whether sequence reads are likely to be derived from different target template nucleic acid molecules. Thus, in certain embodiments of this aspect of the present invention, when measuring the number of target template nucleic acid molecules by sequencing, it can be useful to introduce mutations into the target template nucleic acid molecule. Thus, in certain such embodiments, measuring the number of target template nucleic acid molecules can include mutating the target template nucleic acid molecule.
[0059] Mutating the target template nucleic acid molecule can be performed by any convenient means. In particular, mutating the target template nucleic acid molecule can be performed as described elsewhere herein. In a particularly preferred embodiment, mutations can be introduced by using a low-bias DNA polymerase. In additional or alternative embodiments, mutating the target template nucleic acid molecule can include amplifying the target template nucleic acid molecule in the presence of a nucleotide analog, such as dPTP.
[0060] In a preferred embodiment, determining the number of target template nucleic acid molecules comprises: (i) mutating a target template nucleic acid molecule to obtain a mutant target template nucleic acid molecule; (ii) sequencing a region of the mutant target template nucleic acid molecule; and (iii) identifying the number of unique mutant target template nucleic acid molecules based on the number of unique mutant target template nucleic acid molecule sequences; may include:
[0061] To quantify the number of target template nucleic acid molecules in a sample, the user does not need the complete sequence of each target template nucleic acid molecule. Rather, what is needed is sufficient information about the sequences of the different target template nucleic acid molecules (or, if applicable, the amplified and fragmented target template nucleic acid molecules) in the sample to allow the user to estimate the total number of target template nucleic acid molecules and / or the number of unique target template nucleic acid molecules. For this reason, the user may choose to sequence only a region of each target template nucleic acid molecule. For example, in certain embodiments, the user may choose to sequence the terminal regions of each unique target template nucleic acid molecule or fragmented target template nucleic acid molecule as part of the process of determining the number of unique target template nucleic acid molecules. Thus, the user may sequence the 3' and / or 5' terminal regions of the target template nucleic acid molecule or fragmented target template nucleic acid molecule as part of the process of determining the number of target template nucleic acid molecules. The terminal region of a target template nucleic acid molecule includes the terminal (e.g., 5'- or 3'-terminal) nucleotide in the target template nucleic acid molecule (i.e., the 5'-most or 3'-most nucleotide in the target template nucleic acid molecule) and a contiguous nucleotide stretch of a desired length adjacent thereto.
[0062] In certain exemplary embodiments, determining the number of target template nucleic acid molecules may include introducing a barcode (also referred to herein as a unique molecular tag or unique molecular identifier, as described below) or a pair of barcodes into the target template nucleic acid molecule (in other words, labeling the target template nucleic acid molecule with a barcode or a pair of barcodes) to obtain a barcoded target template nucleic acid molecule. As described elsewhere herein, the barcodes are suitably degenerate, such that substantially each target template nucleic acid molecule contains a unique or substantially unique sequence, and each (or substantially each) target template nucleic acid molecule may be labeled with a different barcode sequence. Introduction of barcodes into target template nucleic acid molecules may be performed as described elsewhere herein. In certain embodiments, the barcode sequence may be introduced at the end of the target template nucleic acid molecule, i.e., as an additional sequence 5' to the 5'-terminal (or 5'-most) nucleotide or 3' to the 3'-terminal (or 3'-most) nucleotide of the target template nucleic acid molecule.
[0063] In a preferred embodiment, the target template nucleic acid molecule labeled with a barcode sequence can be sequenced to determine the number of target template nucleic acid molecules in a sample. More specifically, the region of the target template nucleic acid molecule that contains the barcode sequence can be sequenced to determine the number of target template nucleic acid molecules in a sample. The barcode sequence is substantially unique, and therefore, labeling the target template nucleic acid molecule with the barcode sequence introduces a substantially unique (and therefore quantifiable) sequence into the target template nucleic acid molecule. Therefore, in such an embodiment, the number of unique barcodes identified by sequencing can determine the number of unique target template nucleic acid molecules in a sample.
[0064] Thus, in certain embodiments, determining the number of target template nucleic acid molecules is performed by: (i) sequencing a region of the barcoded target template nucleic acid molecule that includes the barcode or pair of barcodes; and (ii) determining the number of uniquely barcoded target template nucleic acid molecules based on the number of unique barcodes or barcode pairs; may include:
[0065] In yet another embodiment, it may not be necessary to use one or more barcodes to determine the number of target template nucleic acid molecules present in a sample. In certain exemplary embodiments, the number of target template nucleic acid molecules can be determined by sequencing the terminal regions of the target template nucleic acid molecules. Optionally, the user can then identify the number of unique terminal sequences present and / or the user can then map the sequences of the terminal regions to a reference sequence, such as a reference genome. While not wishing to be bound by theory, it is believed that such an approach can enable the determination of the number of target template nucleic acid molecules, since the sequence of each target template nucleic acid molecule can start from a different position in the reference sequence.
[0066] Furthermore, the sequencing step in this aspect of the invention may be a "rough" sequencing step, in that the user may not require precise sequence information in order to be able to determine the number of target template nucleic acid molecules in a sample. Typically, the sequencing step can be performed on a set of molecules that is not fully amplified, which may allow the step to be performed more quickly and / or at lower cost.
[0067] Optionally, determining the number of unique target template nucleic acid molecules in a sample can include sequencing a terminal region of the barcoded target template nucleic acid molecule that comprises a barcode or pair of barcodes. Thus, reference to sequencing a terminal region of a target template nucleic acid molecule can include sequencing a terminal region of a barcoded target template nucleic acid molecule that can comprise a barcode or pair of barcodes.
[0068] Once the number of unique target template nucleic acid molecules in a sample has been determined, the sample can be adjusted to control the number of target template nucleic acid molecules in the sample so that the sample contains a desired number of unique target template nucleic acid molecules. In certain embodiments, this can include diluting the sample. Thus, controlling the number of target template nucleic acid molecules in a sample can include determining the number of target template nucleic acid molecules in the sample and diluting the sample so that the sample contains a desired number of target template nucleic acid molecules.
[0069] As noted above, the sample according to this aspect of the invention can be any sample, and in particular can be the first or second sample in the methods of the invention. Thus, in certain embodiments, controlling the number of target template nucleic acid molecules in the first of the pair of samples and / or the second of the pair of samples comprises measuring the number of target template nucleic acid molecules and diluting the first of the pair of samples and / or the second of the pair of samples so that the first of the pair of samples and / or the second of the pair of samples contains the desired number of target template nucleic acid molecules.
[0070] Pooling of subsamples to obtain a sample A sample can be obtained by pooling several subsamples, which allows target template nucleic acid molecules from multiple samples (e.g., multiple sources) to be sequenced simultaneously, and this can enable greater sample throughput to be achieved, reducing the cost and time required to sequence the target template nucleic acid molecules.
[0071] Thus, the methods of the present invention can be performed on a sample obtained by pooling two or more subsamples. In certain embodiments, the first of a pair of samples can be obtained by pooling two or more subsamples. In further embodiments, the second of a pair of samples can be obtained by pooling two or more subsamples. Thus, the first and / or second sample can be obtained by pooling two or more subsamples. Alternatively, the first and second samples can be taken from the pooled sample and subjected to the methods of the present invention.
[0072] Thus, this aspect of the invention allows for the sequencing of at least one target template nucleic acid molecule from each of two or more smaller samples pooled to obtain the sample.
[0073] One problem associated with pooling samples for sequencing is that each sample may contain different number of target nucleic acid molecules.Therefore, it may be beneficial for pooled sample to contain the target template nucleic acid molecules from each of its constituent subsamples in desired amount, more particularly in desired ratio.In other words, it may be beneficial for pooled sample to contain an appropriate number (i.e., within desired range) of unique target template nucleic acid molecules from each of its subsamples, so that individual sequencing methods can be used to sequence the target template nucleic acid molecules from each of the subsamples in pooled sample.
[0074] As a representative example, two separate subsamples (sample Y and sample Z) may be obtained. If the total number of target template nucleic acid molecules in sample Y is 100 times the total number of target template nucleic acid molecules in sample Z, then if sample Y and sample Z are pooled in equal amounts and the pooled sample is subjected to a sequencing method, the number of sequencing reads obtained from the target template nucleic acid molecules in sample Y would be expected to be 100 times the number of sequencing reads obtained from the target template nucleic acid molecules in sample Z. Thus, pooling samples in this manner not only results in insufficient sequencing reads from sample Z to enable the performance of a sequence assembly step using the sequence reads obtained from sample Z, but may also complicate the performance of the sequence assembly step on the sequencing reads obtained from sample Y.
[0075] Thus, the methods of the invention may include normalizing the number of target template nucleic acid molecules in each of the subsamples pooled to obtain the first of the sample pair and / or the second of the sample pair.
[0076] More generally, however, the present invention provides: (a) providing at least one sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule; and (c) constructing the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule. wherein the at least one sample is obtained by pooling two or more subsamples, and the number of target template nucleic acid molecules in each of the subsamples is normalized.
[0077] For purposes of this application, the phrases "normalizing the number of target template nucleic acid molecules in each of the subsamples" and "normalizing the number of target template nucleic acid molecules in each of the pooled subsamples" refer to pooling the subsamples so that the total number of target template nucleic acid molecules in the pooled sample derived from each of the subsamples is a desired amount. In some embodiments, the number of unique target template nucleic acid molecules is normalized. A "unique target template nucleic acid molecule" is a target template nucleic acid molecule that comprises a different nucleic acid sequence. Optionally, each of at least one target template nucleic acid molecule can be a unique target template nucleic acid molecule. Unique target template nucleic acid molecules can differ by only a single nucleotide in their sequence, or can be substantially different from each other.
[0078] The normalization step can advantageously allow the number of target template nucleic acid molecules from each subsample to be obtained in a desired ratio. In certain embodiments, this can involve manipulating or adjusting each subsample so that, when pooled, the pooled sample contains the desired number of target template nucleic acid molecules from each of the subsamples. In other words, this step can be understood as allowing the number of target template nucleic acid molecules in a pooled sample from each of two or more subsamples to be controlled, or allowing the number of target template nucleic acid molecules in at least one sample from each of two or more subsamples to be controlled.
[0079] Therefore, from another perspective, the present invention is (a) providing at least one sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule; and (c) constructing the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule. The present invention provides a method for sequencing at least one target template nucleic acid molecule, comprising: preparing at least one sample comprising at least one target template nucleic acid molecule; and pooling two or more subsamples; and controlling the number of target template nucleic acid molecules in at least one sample from each of the two or more subsamples.
[0080] In certain embodiments, normalizing the number of target template nucleic acid molecules in each of the subsamples may involve obtaining a similar number of target template nucleic acid molecules in the pooled sample from each of the subsamples (i.e., about a 1:1 ratio). Such an embodiment may be particularly useful, for example, when each subsample is derived from a sample containing genomes of similar size. However, in other embodiments, the number of target template nucleic acid molecules may be obtained in different amounts, i.e., the number of target template nucleic acid molecules from a first subsample may be greater than the number of target template nucleic acid molecules from a second subsample. Such an embodiment may be desirable, for example, when the first subsample is derived from a larger genome and the second subsample is derived from a sample containing a smaller genome.
[0081] It will be understood that "normalizing the number of target template nucleic acid molecules in each of the pooled subsamples" does not have to be completely precise. This is because, for example, it may be difficult to measure the number of target template nucleic acid molecules in each of the subsamples. However, if a user finds that a subsample contains approximately twice the desired number of target template nucleic acid molecules, the user can normalize the number of target template nucleic acid molecules in the subsamples so that the number of target template nucleic acid molecules in the pooled sample is approximately half the number of target template nucleic acid molecules present in the subsample (e.g., 45% to 55% of the number of target template nucleic acid molecules present in the subsample).
[0082] Normalization of the number of target template nucleic acid molecules in each of the subsamples can be considered, in its broadest sense, to correspond to the control of the number of target template nucleic acid molecules from each of the subsamples obtained in the pooled sample. Thus, normalization of the number of target template nucleic acid molecules can include measuring the number of target template nucleic acid molecules in each of the subsamples.
[0083] In certain embodiments, the number of target template nucleic acid molecules in a subsample may be measured as described elsewhere herein, particularly in the context of methods for controlling the number of target template nucleic acid molecules in a sample.
[0084] In a preferred embodiment, normalizing the number of target template nucleic acid molecules in each subsample can include labeling target template nucleic acid molecules from different subsamples with different sample tags. A sample tag is a tag used to label a substantial portion or all of at least one target template nucleic acid molecule in a sample. Labeling target template nucleic acid molecules in different subsamples with different sample tags can allow template target nucleic acid molecules from different subsamples to be distinguished. Thus, sample tags can be particularly useful in this aspect of the present invention, because their use can allow the number of target template nucleic acid molecules in each of two or more subsamples to be measured simultaneously. In particular, sample tags can allow the number of target template nucleic acid molecules in each of two or more subsamples to be measured in a single sample. Preferably, target template nucleic acid molecules can be labeled with sample tags before pooling the subsamples. Thus, in certain embodiments, this aspect of the invention may involve preparing preparatory pools of subsamples, each containing target template nucleic acid molecules labeled with a sample tag, and determining the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pools.
[0085] From another perspective, the present invention is (a) labeling target template nucleic acid molecules from two or more different subsamples with different sample tags; (b) pooling two or more subsamples to obtain a preliminary pool of subsamples; and (c) determining the number of target template nucleic acid molecules in the preliminary pool labeled with each sample tag; The present invention provides a method for determining the number of target template nucleic acid molecules in two or more subsamples, comprising:
[0086] If desired, two or more preliminary pools can be prepared, for example, each containing subsamples obtained in different amounts or proportions and / or each composed of different subsamples (eg, different combinations of subsamples).
[0087] In certain embodiments, the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool can be measured using techniques described elsewhere herein for measuring the number of target template nucleic acid molecules in a sample (particularly in the context of controlling the number of target template nucleic acid molecules in a sample). In this regard, those skilled in the art will understand that the target template nucleic acid molecules from each sample are distinguishable based on the sample tag they contain, and therefore, measuring the number of target template nucleic acid molecules in a preliminary pool labeled with any given sample tag can be performed by adapting methods for measuring the total number of target template nucleic acid molecules present in an individual sample.
[0088] In this regard, in certain embodiments, the reserve pool can be diluted before or during the measurement of the number of target template nucleic acid molecules labeled with each sample tag. The dilution can be performed as described elsewhere herein. For example, in certain embodiments, a serial dilution can be performed on the reserve pool to obtain a serial dilution comprising the diluted reserve pool.
[0089] As described elsewhere, two or more different preparatory pools can be prepared, each of which can be diluted to a different degree, for example, by serial dilution.
[0090] In a particularly preferred embodiment, the number of target template nucleic acid molecules labeled with each sample tag in the reserve pool can be determined by sequencing the labeled (sample-tagged) target template nucleic acid molecules in the reserve pool or diluted reserve pool. Sequencing can be performed according to any convenient sequencing method, such as those described elsewhere herein. Preferably, sequencing the labeled target template nucleic acid molecules can include sequencing the sample tags of the labeled target template nucleic acid molecules.
[0091] In certain embodiments, determining the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool may include an amplification step. Suitable methods for amplifying labeled target template nucleic acid molecules are known in the art, and amplification may be performed, for example, as described elsewhere herein. In certain embodiments, determining the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool may include amplifying the target template nucleic acid molecules and then sequencing them.
[0092] In certain embodiments, the target template nucleic acid molecules in the subsamples can be amplified, i.e., amplified before pooling two or more subsamples to obtain a preliminary pooled sample. Amplification can be performed before labeling the target template nucleic acid molecules in the subsamples with sample tags, or in certain preferred embodiments, can be performed simultaneously with labeling the target template nucleic acid molecules in the subsamples with sample tags (e.g., using PCR primers containing sample barcodes). In other embodiments, the target template nucleic acid molecules labeled with sample tags can be amplified before obtaining a preliminary pooled sample.
[0093] In yet another embodiment, determining the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool can include amplifying the target template nucleic acid molecules labeled with the sample tag in the preliminary pool, i.e., after pooling two or more subsamples.
[0094] If desired, more than one amplification step can be performed, for example, a first amplification can be performed before or simultaneously with labeling the target template nucleic acid molecules in the subsample with sample tags, and a second amplification can be performed to amplify the target template nucleic acid molecules labeled with sample tags (this second amplification can be performed on a subsample or pre-pooled sample, as outlined above).
[0095] After amplification, determining the number of target template nucleic acid molecules labeled with each sample tag in the reserve pool can include sequencing the target template nucleic acid molecules labeled with each sample tag (i.e., sample tag-labeled target template nucleic acid molecules) in the reserve pool or diluted reserve pool. Thus, in a preferred embodiment, determining the number of target template nucleic acid molecules labeled with each sample tag in the reserve pool can include amplifying and then sequencing the target template nucleic acid molecules labeled with each sample tag in the reserve pool or diluted reserve pool.
[0096] Determining the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool can include a fragmentation step. Preferably, the target template nucleic acid molecules in the pooled sample are fragmented, i.e., fragmented after preparation of the pooled sample. Fragmentation can be performed using any suitable technique, including any of the techniques described elsewhere herein.
[0097] In certain embodiments, measuring the number of target template nucleic acid molecules labeled with each sample tag can include both amplification and fragmentation steps before sequencing the target template nucleic acid molecules in the preliminary pool or diluted preliminary pool. Thus, in a preferred embodiment, the target nucleic acid molecules in a subsample can be amplified, fragmented, and labeled with a sample tag, and then two or more subsamples can be pooled to obtain a preliminary pooled sample, and the target template nucleic acid molecules can be sequenced. Amplification and fragmentation can be performed in any order. In one embodiment, the target template nucleic acid molecules in the subsample can be amplified and then fragmented, or fragmented and then amplified, and then labeled with a sample tag. In other embodiments, the amplification, fragmentation, and labeling of the target template nucleic acid molecules can be performed simultaneously, i.e., in a single step. A particularly preferred method for amplifying, fragmenting, and labeling target template nucleic acid molecules in a single step can be performed using tagging and PCR, particularly using PCR primers containing sample tags. Thus, after such a step, the amplified and fragmented target nucleic acid molecules are labeled with a sample tag and can be identified as originating from a particular subsample when pooled, for example, as a pre-pooled sample during sequencing.
[0098] Determining the number of target template nucleic acid molecules labeled with each sample tag in the reserve pool can include determining the number of target template nucleic acid molecules (optionally, unique target template nucleic acid molecules) containing each sample tag (i.e., labeled with each sample tag) in the reserve pool (or diluted reserve pool). Preferably, however, determining the number of target template nucleic acid molecules containing each sample tag includes determining the number of unique target template nucleic acid sequences containing each sample tag in the reserve pool (or diluted reserve pool).
[0099] As described elsewhere, mutating target template nucleic acid molecules can be particularly useful, for example, in determining whether sequence reads are likely to originate from the same or different target template nucleic acid molecules, and thus can be useful in determining the number of target template nucleic acid molecules in a preliminary pool that originate from a particular subsample.
[0100] Thus, in certain embodiments, determining the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool (or diluted preliminary pool) can include mutating the target template nucleic acid molecule. In certain embodiments, it is possible to mutate the target template nucleic acid molecule in the preliminary pooled sample. However, mutating the target template nucleic acid molecule can preferably be performed in a subsample, i.e., before pooling two or more samples to obtain a pooled sample. In particularly preferred embodiments, it is possible to mutate the target template nucleic acid molecule before or simultaneously with labeling the target template nucleic acid molecule with a sample tag. It may be preferable not to mutate the sample tag sequence used to label the target template nucleic acid molecule. Mutagenizing the target template nucleic acid molecule can be performed by any convenient means, including any of the means described elsewhere herein. Thus, in one embodiment, mutations can be introduced by using a low-bias DNA polymerase. In other embodiments, mutating the target template nucleic acid molecule can include amplifying the target template nucleic acid molecule in the presence of a nucleotide analog, such as dPTP.
[0101] In a preferred embodiment, determining the number of target template nucleic acid molecules labeled with each sample tag in the reserve pool comprises: (i) mutating a target template nucleic acid molecule to obtain a mutant target template nucleic acid molecule; (ii) sequencing a region of the mutant target template nucleic acid molecule; and (iii) determining the number of unique mutant target template nucleic acid molecules containing each sample tag based on the number of unique mutant target template nucleic acid molecules labeled with each sample tag.
[0102] As outlined in more detail above, to quantify target template nucleic acid molecules, it may not be necessary to obtain the complete sequence of each target template nucleic acid molecule; it may be sufficient to simply sequence the terminal region of each labeled target template nucleic acid molecule as part of the process of measuring the number of target template nucleic acid molecules in a reserve pool labeled with each sample tag.Therefore, the user may choose to sequence only the terminal region of each target template nucleic acid molecule.As outlined above, the sample tag is preferably sequenced.
[0103] In certain exemplary embodiments, determining the number of target template nucleic acid molecules may include introducing a barcode or pair of barcodes into the target template nucleic acid molecules to obtain barcoded sample-tagged target template nucleic acid molecules. Barcodes suitable for use in such processes, and methods for introducing them into target template nucleic acid molecules, are described in more detail elsewhere herein.
[0104] Preferably, the barcode can be introduced into the target template nucleic acid molecule before pooling the subsamples, i.e., before pooling the subsamples to obtain a preliminary pooled sample. The barcode and sample tag can be introduced into the target template nucleic acid molecule in any order. For example, in one embodiment, the barcode can be introduced into the target template nucleic acid molecule, followed by the sample tag. In another embodiment, the sample tag can be introduced into the target template nucleic acid molecule, followed by the barcode. In yet another embodiment, the sample tag and the barcode tag can be introduced simultaneously. In either case, in certain embodiments, the target template nucleic acid molecule from the subsample can be labeled with both the sample tag and the barcode. In this regard, it should be noted that the sample tag can be particularly useful for identifying a specific target template nucleic acid molecule in the preliminary sample as originating from a specific subsample, while the barcode can be particularly useful for enabling the determination of the number of unique target template nucleic acids from each subsample.
[0105] Thus, in a particularly preferred embodiment, determining the number of target template nucleic acid molecules labeled with each sample tag is performed by: (i) sequencing a region of a barcoded sample-tagged target template nucleic acid molecule; and (ii) determining the number of unique barcoded target template nucleic acid molecules containing each sample tag based on the number of unique barcode or barcode pair sequences associated with each sample tag; may include:
[0106] As described elsewhere herein, the sequencing step in determining the number of target template nucleic acid molecules may be a "rough" sequencing step, in that the user may not require precise sequence information to be able to determine the number of target template nucleic acid molecules in a sample. Instead, sequencing that allows identification of sample tags, barcodes, and / or target template nucleic acid molecules may be sufficient.
[0107] In certain exemplary embodiments, once the number of target template nucleic acid molecules containing different sample tags has been determined, the ratio of the number of target template nucleic acid molecules containing different sample tags can be calculated. In other exemplary embodiments, once the number of target template nucleic acid molecules containing different sample tags has been determined, it can be possible to determine the number of target template nucleic acid molecules (in the pre-pooled sample) resulting from each subsample, thereby calculating the number of target template nucleic acid molecules present in each subsample.
[0108] The pooled sample used in the method of the present invention can be prepared using information about the ratio of target template nucleic acid molecules that contain different sample tags and / or the number of target template nucleic acid molecules that come from each subsample.In particular, such information can be used in the normalization step to normalize the number of target template nucleic acid molecules that come from each of two or more subsamples in the pooled sample, thereby obtaining target template nucleic acid molecules from each subsample at a desired ratio in the pooled sample.
[0109] Therefore, the present invention provides (a) providing at least one sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule; and (c) constructing (assembling) the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule; It is understood that the present invention provides a method for sequencing at least one target template nucleic acid molecule, wherein at least one sample comprises: (i) preparing a preliminary pooled sample by pooling two or more of the subsamples; (ii) determining the number of target template nucleic acid molecules in the preliminary pooled samples resulting from each of the two or more subsamples; and (iii) Pooling two or more subsamples where the number of target template nucleic acid molecules in the sample from each of the subsamples is normalized.
[0110] As mentioned above, normalizing the number of target template nucleic acid molecules in a sample obtained by pooling two or more subsamples can include obtaining target template nucleic acid molecules at a desired ratio from each of the subsamples.In certain embodiments, the sample generated by pooling two or more subsamples is considered to be a repooled sample, where the target template nucleic acid molecules in each of the subsamples are obtained at a desired ratio (i.e., after obtaining a preliminary pool and measuring the number of target template nucleic acid molecules in the preliminary pool resulting from each of two or more subsamples).Therefore, measuring the number of target template nucleic acid molecules in a subsample allows the number of target template nucleic acid molecules in the sample from each of the subsamples to be normalized when repooling the subsamples.
[0111] A sample may be obtained by pooling two or more subsamples according to this aspect of the invention. Thus, two or more, preferably 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000 or more subsamples may be pooled to obtain a sample for use in the methods of the invention (i.e., a pooled sample). In certain embodiments, 2 to 5000, 10 to 1000, or 25 to 150 subsamples may be pooled.
[0112] The phrase "pooling two or more subsamples" does not require the entire subsamples to be combined with another subsample to obtain a sample, but instead preferably refers to obtaining an aliquot of each of the subsamples and combining the aliquots to obtain the sample. Similarly, references to introducing a barcode or tag into a target template nucleic acid molecule in a subsample, or mutating a target template nucleic acid molecule in a subsample, may be understood to mean performing such a step on an aliquot or on a portion of the subsamples.
[0113] In certain embodiments, "pooling two or more subsamples" can include diluting a subsample and combining the diluted subsamples to obtain a sample. In other embodiments, the term can include obtaining aliquots of a sample, diluting the aliquots, and combining the diluted aliquots of the subsamples to obtain a sample. Dilution of a subsample (or aliquot) can include a separate dilution step performed before pooling the subsamples (or aliquots) to obtain a sample. However, it will be understood that pooling two or more subsamples (or aliquots) to obtain a sample can effectively reduce the concentration of target template nucleic acid molecules from each of the subsamples provided in the sample and thus can amount to a dilution step. One skilled in the art will be able to determine the degree to which dilution of each subsample may be required, including dilution that may occur as a result of pooling two or more subsamples (or aliquots).
[0114] Sequencing a region of at least one target template nucleic acid molecule or at least one mutant target template nucleic acid molecule. A method for sequencing at least one target template nucleic acid molecule may include sequencing a region of at least one target template nucleic acid molecule in a first of a pair of samples to obtain a non-mutated sequence read, and / or sequencing a region of at least one mutant target template nucleic acid molecule to obtain a mutant sequence read.
[0115] The sequencing step can be carried out using any sequencing method. Examples of possible sequencing methods include Maxam-Gilbert sequencing, Sanger sequencing, sequencing involving bridge amplification (e.g., bridge PCR), or any high-throughput sequencing (HTS) method, such as those described in Maxam AM, Gilbert W (February 1977), "A new method for sequencing DNA", Proc. Natl. Acad. Sci. USA 74 (2): 560-4; Sanger F, Coulson AR (May 1975), "A rapid method for determining sequences in DNA by primed synthesis with DNA polymerase", J. Mol. Biol. 94 (3): 441-8; and Bentley DR, Balasubramanian S et al. (2008), "Accurate whole human genome sequencing using reversible terminator chemistry", Nature 456 (7218): 53-59. In typical embodiments, at least one, or preferably both, of the sequencing steps can include bridge amplification. Optionally, the bridge amplification step is performed with an extended time of greater than 5 seconds, greater than 10 seconds, greater than 15 seconds, or greater than 20 seconds. One example of the use of bridge amplification is in an Illumina Genome Analyzer Sequencer.
[0116] Optionally, (i) sequencing a region of at least one target template nucleic acid molecule in a first of a pair of samples to obtain a non-mutated sequence read and (ii) sequencing a region of at least one mutant target template nucleic acid molecule to obtain a mutant sequence read can be performed using the same sequencing method. Optionally, (i) sequencing a region of at least one target template nucleic acid molecule in a first of a pair of samples to obtain a non-mutated sequence read and (ii) sequencing a region of at least one mutant target template nucleic acid molecule to obtain a mutant sequence read can be performed using different sequencing methods.
[0117] Optionally, (i) sequencing a region of at least one target template nucleic acid molecule in the first pair of samples to obtain non-mutated sequence reads, and (ii) sequencing a region of at least one mutant target template nucleic acid molecule to obtain mutant sequence reads can be performed using multiple sequencing methods.For example, a portion of at least one target template nucleic acid molecule in the first pair of samples can be sequenced using a first sequencing method, and a portion of at least one target template nucleic acid molecule in the first pair of samples can be sequenced using a second sequencing method.Similarly, a portion of at least one mutant target template nucleic acid molecule can be sequenced using a first sequencing method, and a portion of at least one mutant target template nucleic acid molecule can be sequenced using a second sequencing method.
[0118] If desired, (i) sequencing a region of at least one target template nucleic acid molecule in the first sample pair to obtain a non-mutated sequence read, and (ii) sequencing a region of at least one mutant target template nucleic acid molecule to obtain a mutant sequence read, can be performed at different times. Alternatively, steps (i) and (ii) can be performed during a substantially similar period, for example, within one year of each other. The first sample pair and the second sample pair do not need to be collected at the same time. If the two samples are from the same organism, they can be collected at substantially different times, even several years apart, and therefore the two sequencing steps can also be several years apart. Furthermore, even if the first sample pair and the second sample pair are from the same original sample, the biological sample can be stored for some time, and therefore the sequencing steps do not need to be performed simultaneously.
[0119] The mutated and / or non-mutated sequence reads can be single-end or paired-end sequence reads.
[0120] Optionally, the mutant sequence reads and / or non-mutated sequence reads are greater than 50 bp (i.e., greater than 50 bp), greater than 100 bp, greater than 500 bp, less than 200,000 bp, less than 15,000 bp, less than 1,000 bp, between 50 and 200,000 bp, between 50 and 15,000 bp, or between 50 and 1,000 bp. The longer the read length, the easier it is to use information obtained from analyzing the mutant sequence reads to construct (assemble) the sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads. For example, when using an assembly graph, using longer sequence reads makes it easier to identify valid pathways using the assembly graph. For example, as described in more detail below, identifying valid pathways using the assembly graph can include identifying signature k-mers, and longer read lengths can allow for longer k-mers.
[0121] Optionally, the sequencing step is performed using a sequencing depth of 0.1 to 500 reads, 0.2 to 300 reads, or 0.5 to 150 reads per nucleotide per at least one target template nucleic acid molecule. The greater the sequencing depth, the more accurate the sequence determined / generated, but the more difficult assembly may be.
[0122] Introducing a mutation into at least one target template nucleic acid molecule The method can include introducing a mutation into at least one target template nucleic acid molecule in a second of the pair of samples to obtain at least one mutated target template nucleic acid molecule.
[0123] A mutation can be a substitution mutation, an insertion mutation, or a deletion mutation. For the purposes of the present invention, the term "substitution mutation" should be interpreted to mean that a nucleotide is replaced with another nucleotide. For example, converting the sequence ATCC to the sequence AGCC introduces a single substitution mutation. For the purposes of the present invention, the term "insertion mutation" should be interpreted to mean that at least one nucleotide is added to a sequence. For example, converting the sequence ATCC to the sequence ATTCC is an example of an insertion mutation (an additional T nucleotide is inserted). For the purposes of the present invention, the term "deletion mutation" should be interpreted to mean that at least one nucleotide is removed from a sequence. For example, converting the sequence ATTCC to ATCC is an example of a deletion mutation (a T nucleotide is removed). Preferably, the mutation is a substitution mutation.
[0124] The phrase "introducing mutations into at least one target template nucleic acid molecule" refers to subjecting at least one target template nucleic acid molecule in the second of a pair of samples to conditions that mutate the at least one target template nucleic acid molecule. This can be accomplished using any suitable method. For example, mutations can be introduced by chemical mutagenesis and / or enzymatic mutagenesis.
[0125] Optionally, the step of introducing mutations into the at least one target template nucleic acid molecule mutates 1% to 50%, 3% to 25%, 5% to 20%, or about 8% of the nucleotides of the at least one target template nucleic acid molecule. Optionally, the at least one mutated target template nucleic acid molecule contains 1% to 50%, 3% to 25%, 5% to 20%, or about 8% mutations.
[0126] By performing the steps of introducing mutations into a nucleic acid molecule of known sequence, sequencing the resulting nucleic acid molecule, and determining the percentage of the total number of nucleotides that are changed compared to the original sequence, the user can determine the number of mutations contained in at least one mutated target template nucleic acid molecule and / or the degree to which the step of introducing mutations into at least one target template nucleic acid molecule mutates at least one target template nucleic acid molecule.
[0127] Optionally, the step of introducing mutations into the at least one target template nucleic acid molecule mutates the at least one target template nucleic acid molecule in a substantially random manner. Optionally, the at least one mutated target template nucleic acid molecule comprises a substantially random mutation pattern.
[0128] At least one mutant target template nucleic acid molecule is said to contain a substantially random mutation pattern if it contains mutations at substantially the same level throughout its length. For example, a user can determine whether at least one mutant target template nucleic acid molecule contains a substantially random mutation pattern by mutating a test nucleic acid molecule of known sequence to obtain a mutant test nucleic acid molecule. The sequence of the mutant test nucleic acid molecule can be compared with the test nucleic acid molecule to determine the respective positions of the mutations. The user can then: (i) calculating the distance between each of the mutations; (ii) calculating the average distance; (iii) subsampling distances without replacement to smaller numbers such as 500 or 1000; (iv) constructing a simulated set of 500 or 1000 distances from a geometric distribution (where the mean is obtained by the method of moments consistent with that previously calculated for the observed distances); and (v) Computing Kolmorgorov-Smirnov on two distributions This allows one to determine whether mutations occur at substantially similar levels throughout the length of the mutation test nucleic acid molecule.
[0129] At least one mutant target template nucleic acid molecule can be considered to contain a substantially random mutation pattern if D<0.15, D<0.2, D<0.25, or D<0.3, depending on the length of the non-mutated reads.
[0130] Similarly, a step of introducing mutations into at least one target template nucleic acid molecule can be said to mutate the at least one target template nucleic acid molecule in a substantially random manner if the resulting at least one mutated target template nucleic acid molecule contains a substantially random mutation pattern. Whether a step of introducing mutations into at least one target template nucleic acid molecule mutates the at least one target template nucleic acid molecule in a substantially random manner can be determined by performing a step of introducing mutations into at least one target template nucleic acid molecule in a test nucleic acid molecule of known sequence to obtain a mutant test nucleic acid molecule. The user can then sequence the mutant test nucleic acid molecule to identify which mutations have been introduced and determine whether the mutant test nucleic acid molecule contains a substantially random mutation pattern.
[0131] Optionally, at least one mutation target template nucleic acid molecule comprises an unbiased mutation pattern. Optionally, the step of introducing mutations into at least one target template nucleic acid molecule introduces mutations in an unbiased manner. If the type of mutation introduced is random, at least one mutation target template nucleic acid molecule comprises an unbiased mutation pattern. If the mutation introduced is a substitution mutation, the mutation introduced is random if similar proportions of A (adenosine), T (thymine), C (cytosine), and G (guanine) nucleotides are introduced. The phrase "similar proportions of A (adenosine), T (thymine), C (cytosine), and G (guanine) nucleotides are introduced" means that the number of adenosine nucleotides, thymine nucleotides, cytosine nucleotides, and guanine nucleotides introduced are within 20% of each other (for example, 20 A nucleotides, 18 T nucleotides, 24 C nucleotides, and 22 G nucleotides could be introduced).
[0132] Whether the step of introducing mutations into at least one target template nucleic acid molecule mutates at least one target template nucleic acid molecule in an unbiased manner can be determined by performing the step of introducing mutations into at least one target template nucleic acid molecule in a test nucleic acid molecule of known sequence to obtain a mutant test nucleic acid molecule. The user can then sequence the mutant test nucleic acid molecule to identify which mutations were introduced and determine whether the mutant test nucleic acid molecule contains an unbiased mutation pattern.
[0133] Usefully, the method for generating a sequence of at least one target template nucleic acid molecule can be used even when the step of introducing mutations into at least one target template nucleic acid molecule introduces unevenly distributed (non-uniformly distributed) mutations. Thus, in one embodiment, at least one mutant target template nucleic acid molecule contains unevenly distributed mutations. Optionally, the step of introducing mutations into at least one mutant target template nucleic acid molecule introduces unevenly distributed mutations. Mutations are considered "unevenly distributed" if they are introduced in a biased manner, i.e., if the numbers of adenosine nucleotides, thymine nucleotides, cytosine nucleotides, and guanine nucleotides introduced are not within 20% of each other. Whether at least one mutant target template nucleic acid molecule contains unevenly distributed mutations or whether the step of introducing mutations into at least one target template nucleic acid molecule introduces unevenly distributed mutations can be determined in a manner similar to that described above for determining whether the step of introducing mutations into at least one target template nucleic acid molecule introduces mutations in an unbiased manner.
[0134] Similarly, the method for generating a sequence of at least one target template nucleic acid molecule can be used even when the mutant sequence reads and / or the non-mutated sequence reads contain ubiquitous sequencing errors. Thus, in one embodiment, the mutant sequence reads and / or the non-mutated sequence reads contain ubiquitous sequencing errors. Similarly, in one embodiment, the step of sequencing a region of at least one target template nucleic acid molecule and / or the step of sequencing a region of at least one mutant target template nucleic acid molecule introduces ubiquitous sequence errors.
[0135] Whether a particular step of sequencing a region of at least one target template nucleic acid molecule and / or a region of at least one mutant target template nucleic acid molecule introduces widespread sequence errors will likely depend on the accuracy of the sequencing device and will likely be known to the user. However, a user can determine whether a step of sequencing a region of at least one target template nucleic acid molecule and / or a region of at least one mutant target template nucleic acid molecule introduces widespread sequence errors by performing a sequencing method on a nucleic acid molecule of known sequence and comparing the resulting sequence reads with those of the original nucleic acid molecule of known sequence. The user can then apply the probability function described in Example 6 to determine the values of M and E. If the values of E and the matrix model are not equal or substantially unequal (within 10% of each other), the step of sequencing a region of at least one target template nucleic acid molecule introduces widespread sequence errors.
[0136] Introducing mutations into at least one target template nucleic acid molecule by chemical mutagenesis can be achieved by exposing at least one target template nucleic acid to a chemical mutagen, including mitomycin C (MMC), N-methyl-N-nitrosourea (MNU), nitrous acid (NA), diepoxybutane (DEB), 1,2,7,8-diepoxyoctane (DEO), ethyl methanesulfonate (EMS), methyl methanesulfonate (MMS), N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), 4-nitroquinoline 1-oxide (4-NQO), 2-methyl-6-chloro-9-(3-[ethyl-2-chloroethyl]-aminopropylamino)-acridine dihydrochloride (ICR-170), 2-aminopurine (2A), bisulfite (bisulfite), and hydroxylamine (HA). For example, when a nucleic acid molecule is exposed to bisulfite, the bisulfite deaminates cytosine to form uracil, effectively introducing a CT substitution mutation.
[0137] As described above, the step of introducing mutations into at least one target template nucleic acid molecule can be carried out by enzymatic mutagenesis. Optionally, enzymatic mutagenesis is carried out using a DNA polymerase. For example, some DNA polymerases are error-prone (low-fidelity polymerases), and replication of at least one target template nucleic acid molecule using an error-prone DNA polymerase will introduce mutations. Taq polymerase is an example of a low-fidelity polymerase, and the step of introducing mutations into at least one target template nucleic acid molecule can be carried out by replicating at least one target template nucleic acid molecule using Taq polymerase, for example, by PCR.
[0138] The DNA polymerase can be a low-bias DNA polymerase, which is described in more detail below.
[0139] When the step of introducing mutations into at least one target template nucleic acid molecule is carried out using a DNA polymerase, the at least one target template nucleic acid molecule can be incubated with the DNA polymerase and suitable primers under conditions suitable for the DNA polymerase to catalyze the production of at least one mutated target template nucleic acid molecule.
[0140] Suitable primers comprise short nucleic acid molecules that are complementary to the adjacent regions of at least one target template nucleic acid molecule, or complementary to the adjacent regions of the nucleic acid molecule that is complementary to at least one target template nucleic acid molecule.For example, if at least one target template nucleic acid molecule is a part of a chromosome, the primer is complementary to the region of the chromosome that is immediately 3' to the 3' end of at least one target template nucleic acid molecule and immediately 5' to the 5' end of at least one target template nucleic acid molecule; or the primer is complementary to the region of the chromosome that is immediately 3' to the 3' end of the nucleic acid molecule that is complementary to at least one target template nucleic acid molecule and immediately 5' to the 5' end of the nucleic acid molecule that is complementary to at least one target template nucleic acid molecule.
[0141] Suitable conditions include temperatures that allow a DNA polymerase to replicate at least one target template nucleic acid molecule, such as temperatures between 40°C and 90°C, between 50°C and 80°C, between 60°C and 70°C, or about 68°C.
[0142] Introducing mutations into at least one template nucleic acid molecule can include multiple rounds of replication. For example, introducing mutations into at least one target template nucleic acid molecule preferably includes: i) replicating at least one target template nucleic acid molecule to obtain at least one nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule; and ii) a round of replicating at least one target template nucleic acid molecule to obtain at least one replica of the target template nucleic acid molecule; Includes:
[0143] Optionally, the step of introducing mutations into the at least one target template nucleic acid molecule comprises replicating the at least one target template nucleic acid molecule for at least 2, at least 4, at least 6, at least 8, at least 10, less than 10, less than 8, about 6, 2 to 8, or 1 to 7 rounds. The user may choose a low number of replication rounds to reduce the possibility of introducing amplification bias.
[0144] Optionally, the step of introducing mutations into at least one target template nucleic acid molecule comprises at least 2, at least 4, at least 6, at least 8, at least 10, less than 10, less than 8, about 6, 2 to 8, or 1 to 7 rounds of replication at a temperature between 60°C and 80°C.
[0145] Optionally, the step of introducing mutations into at least one target template nucleic acid molecule is carried out using polymerase chain reaction (PCR), which comprises the following steps to replicate nucleic acid molecules: a) Melting, b) annealing, and c) Lengthening and elongation It is a process that involves multiple rounds of
[0146] A nucleic acid molecule (e.g., at least one target template nucleic acid molecule) is mixed with an appropriate primer and polymerase. In the melting step, the nucleic acid molecule is heated to a temperature above 90°C to denature the double-stranded nucleic acid molecule (separate the two strands). In the annealing step, the nucleic acid molecule is cooled to a temperature below 75°C, e.g., 55-70°C, about 55°C, or about 68°C, to allow the primers to anneal to the nucleic acid molecule. In the extension and elongation step, the nucleic acid molecule is heated to a temperature above 60°C to allow DNA polymerase to catalyze primer extension and add complementary nucleotides to the template strand.
[0147] Optionally, the step of introducing a mutation into the at least one target template nucleic acid molecule comprises replicating the at least one target template nucleic acid molecule using Taq polymerase under highly error-prone reaction conditions. For example, the step of introducing a mutation into the at least one target template nucleic acid molecule comprises replicating the at least one target template nucleic acid molecule using Mn 2+ , Mg 2+ Alternatively, it may involve PCR using Taq polymerase in the presence of unequal dNTP concentrations (eg, excess cytosine, guanine, adenine, or thymine).
[0148] Acquisition of data containing non-mutated and mutated sequence reads The methods of the present invention may include obtaining data comprising non-mutated and mutated sequence reads. The non-mutated and mutated sequence reads may be obtained from any source.
[0149] Optionally, the non-mutated sequence reads are obtained by sequencing a region of at least one target template nucleic acid molecule in a first one of the pair of samples. Optionally, the mutant sequence reads are obtained by introducing a mutation into at least one target template nucleic acid molecule in a second one of the pair of samples to obtain at least one mutant target template nucleic acid molecule, and sequencing a region of the at least one mutant target template nucleic acid molecule.
[0150] Optionally, the non-mutated sequence read comprises the sequence of a region of at least one target template nucleic acid molecule in a first of the pair of samples, and the mutant sequence read comprises the sequence of a region of at least one mutant target template nucleic acid molecule in a second of the pair of samples, wherein the pair of samples are obtained from the same original sample or are derived from the same organism.
[0151] analyzing the mutant sequence reads and constructing a sequence using information obtained from the analysis of the mutant sequence reads. As described above, the first sample and the second sample contain at least one target template nucleic acid molecule, and therefore, the mutation pattern present in the mutation sequence reads can help the user construct the sequence of at least a portion of the at least one target template nucleic acid molecule.
[0152] As mentioned above, for example, when regions of sequences are similar to each other or when sequences contain repetitive portions, assembling a sequence can be difficult. However, a user may be able to more effectively assemble a sequence from a non-mutated sequence read by using information obtained from a mutated sequence read corresponding to the non-mutated sequence read. For example, a mutated sequence read can be used to identify nodes calculated from a non-mutated sequence read that form part of a valid path through a sequence assembly graph.
[0153] In certain embodiments, a sequence can be assembled using information from multiple mutation reads. As described in more detail below, mutation sequence reads that are likely to originate from the same mutation target template nucleic acid molecule can be identified. In certain embodiments, mutation sequence reads can be assembled and / or a consensus sequence can be generated from multiple mutation sequence reads. In certain embodiments, to obtain information for constructing a sequence, long mutation reads (i.e., long composite mutation reads) can be reconstructed from multiple partially overlapping mutation reads that originate from the same mutation target template nucleic acid molecule. Such long composite reads can correspond to identified paths through a non-mutation assembly graph, as described elsewhere herein.
[0154] Preparation of assembly graph The step of analyzing mutant sequence reads and constructing the sequence of at least a portion of at least one target template nucleic acid molecule from non-mutated sequence reads using information obtained from the analysis of the mutant sequence reads may include preparing an assembly graph.
[0155] For purposes of the present invention, an "assembly graph" is a graph that includes nodes calculated from non-mutated sequence reads and paths that (if valid) may correspond to portions of at least one target template nucleic acid molecule. For example, a node may represent a consensus sequence calculated from assembled non-mutated sequence reads.
[0156] Nodes can be calculated from non-mutated sequence reads. However, if some of at least one target template nucleic acid molecule is not correctly sequenced, insufficient non-mutated sequence reads may be available to construct the complete sequence of at least one target template nucleic acid molecule. In this case, nodes can be calculated from a combination of non-mutated sequence reads and mutated sequence reads, where the mutated sequence reads are used to fill in the region of the assembly graph representing the missing non-mutated sequence reads. Optionally, nodes are calculated from non-mutated sequence reads and mutated sequence reads. Because the non-mutated sequence reads closely correspond to the original target template nucleic acid molecule, it is beneficial to use nodes calculated only from non-mutated sequence reads. Therefore, using an assembly graph consisting of nodes calculated from non-mutated sequence reads can avoid artifacts introduced by the mutation process.
[0157] A representation of a suitable assembly graph is shown in panel A of Figure 9.
[0158] Optionally, a node of the assembly graph is a unitig. For purposes of the present invention, the term "unitig" is intended to refer to a portion of at least one target template nucleic acid molecule having a sequence that can be determined with a high level of confidence. For example, a node of the assembly graph may include a unitig corresponding to a consensus sequence of all or a portion of one or more non-mutated sequence reads and / or a consensus sequence of all or a portion of one or more mutant sequence reads. Preferably, a node of the assembly graph includes a unitig corresponding to a consensus sequence of all or a portion of one or more non-mutated sequence reads.
[0159] The assembly graph can be a contig graph, a unitig graph, or a weighted graph. For example, the assembly graph can be a de Bruijn graph.
[0160] Identifying nodes that form part of a valid path through an assembly graph Using information obtained from the analysis of mutant sequence reads to construct a sequence of at least a portion of at least one target template nucleic acid molecule from non-mutated sequence reads may include using information obtained from the analysis of mutant sequence reads to identify nodes calculated from non-mutated sequence reads that form part of valid paths through an assembly graph. Each valid path through the assembly graph may represent (or represent) the sequence of a portion of at least one target template nucleic acid molecule. When the assembly graph includes multiple putative paths from node to node, information obtained from the analysis of mutant sequence reads can be used to obtain an order of the nodes. In other embodiments, information obtained from the analysis of mutant sequence reads can be used to determine the copy number of a given sequence in a genome.
[0161] Optionally, analyzing the mutant sequence reads includes identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule. The method of the present invention can provide a plurality of mutant sequence reads containing mutant sequences corresponding to the same region, i.e., a group of mutant sequence reads corresponding to the same region. Some of the mutant sequence reads in the group may overlap, and some of the mutant sequence reads in the group may repeat. Once the group of mutant sequence reads is mapped to the assembly graph, they can connect nodes calculated from non-mutated sequence reads, and can be used to identify valid paths through the assembly graph, as shown in Figure 9B.
[0162] Thus, optionally, analyzing the mutant sequence reads includes identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule. Optionally, identifying nodes that form part of valid paths through the assembly graph using information obtained by analyzing the mutant sequence reads includes: (i) Computing nodes from non-mutated sequence reads; (ii) mapping mutant sequence reads against the assembly graph; (iii) identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule; and (iv) identifying nodes connected by mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule; wherein nodes connected by mutant sequence reads are likely to originate from the same at least one mutant target template nucleic acid molecule and form part of a valid path through the assembly graph.
[0163] Optionally, mutant sequence reads that are likely to originate from the same mutant target template nucleic acid molecule are assigned to groups.
[0164] Identifying mutant sequence reads that are likely to originate from the same mutant target template nucleic acid molecule As described above, analyzing mutant sequence reads can include identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule.
[0165] Optionally, if mutation sequence reads share a common mutation pattern, they are likely to be derived from the same mutation target template nucleic acid molecule.Optionally, the mutation sequence reads that share a common mutation pattern comprise a common signature k-mer or a common signature mutation.Preferably, the mutation sequence reads that share a common mutation pattern comprise at least 1, at least 2, at least 3, at least 4, at least 5 or at least k common signature k-mers and / or common signature mutations.
[0166] Identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule can be particularly useful when a sample is obtained by pooling two or more subsamples. In certain embodiments, such a process can be used when determining the sequence of at least one target template nucleic acid molecule in a sample obtained by pooling two or more subsamples. More specifically, such a process can be used when determining the sequence of at least one target template nucleic acid molecule from each of two or more subsamples pooled to obtain the sample. Such a process can also be particularly useful when measuring the number of target template nucleic acid molecules in a sample that are from each of two or more subsamples when the target template nucleic acid molecules in the subsamples are mutated.
[0167] Signature k-mer or signature mutation Mutant sequence reads that share a common mutational pattern may include a common signature k-mer and / or a common signature mutation. Preferably, mutant sequence reads that share a common mutational pattern include at least 1, at least 2, at least 3, at least 4, at least 5, or at least k common signature k-mers and / or common signature mutations.
[0168] In the context of the present invention, a "k-mer" refers to a nucleic acid sequence of length k contained within a sequence read. A "signature k-mer" can be a k-mer that does not appear in a non-mutated sequence read but appears at least twice in a mutant sequence read. In one embodiment, a signature k-mer is a k-mer that appears at least n times more frequently in a mutant sequence read than in a non-mutated sequence read, where n is any integer, such as 2, 3, 4, or 5. Optionally, a signature k-mer is a k-mer that appears at least 2, at least 3, at least 4, at least 5, or at least 10 times in a mutant sequence read. Thus, a user can determine whether a mutant sequence read contains a common signature k-mer by dividing the mutant sequence read into k-mers and the non-mutated sequence read into k-mers. The user can then compare the mutant and non-mutated sequence read k-mers and determine which k-mers appear in the mutant sequence read k-mers but not in the non-mutated sequence read k-mers (or which k-mers appear more frequently in the mutant sequence read k-mers than in the non-mutated sequence read k-mers). The user can then evaluate and count k-mers that appear in the mutant sequence read k-mers but not in the non-mutated sequence read k-mers (or that appear less frequently). Any k-mer that appears at least 2, at least 3, at least 4, at least 5, or at least 10 times in the mutant sequence read k-mers but not in the non-mutated sequence read k-mers is a signature k-mer. Any k-mer that appears less than k, less than 5, less than 4, less than 3, or once in the mutant sequence read k-mer but does not appear (or appears less frequently) in the non-mutated sequence read k-mer may be the result of a sequencing error and should therefore be ignored.
[0169] The value of k is user-selectable and can be any value. Optionally, the value of k is at least 5, at least 10, at least 15, less than 100, less than 50, less than 25, between 5 and 100, between 10 and 50, or between 15 and 25. Generally, the user selects a value of k that is as long as possible while ensuring that the proportion of k-mers in the reads containing one or more sequence errors is low. Preferably, the proportion of k-mers in the reads containing sequencing errors is less than 50%, less than 40%, less than 30%, 0% to 50%, 0% to 40%, or 0% to 30%.
[0170] A "signature mutation" can be a nucleotide that occurs at least twice in a mutant sequence read and that appears at the corresponding position in a non-mutated sequence read. In one embodiment, a signature mutation is a mutation that occurs at least n times more frequently in a mutant sequence read than in a non-mutated sequence read, where n is any integer, e.g., 2, 3, 4, or 5. Optionally, a signature mutation is a mutation that occurs at least 2 times, at least 3 times, at least 4 times, at least 5 times, or at least 10 times in a mutant read and that does not appear (or appears less frequently) at the corresponding position in a non-mutated read.
[0171] Optionally, signature mutation is a co-occurring mutation. " Co-occurring mutation " refers to two or more signature mutations that occur in the same mutation sequence read.For example, if a mutation sequence read contains three signature mutations, it contains three co-occurring mutation pairs or one co-occurring mutation triad.If it contains four signature mutations, it contains six co-occurring mutation pairs, four co-occurring mutation triads, and one co-occurring mutation tetraad.
[0172] Optionally, an identified signature mutation can be disregarded if it does not meet certain criteria that suggest that the signature mutation is spurious or not useful in constructing the sequence of at least a portion of at least one target template nucleic acid molecule.
[0173] Optionally, if the corresponding position in the mutation sequence reads that share signature mutation is different from each other at least 1, at least 2, at least 3 or at least 5 nucleotides, signature mutation will be ignored.For example, when two mutation sequence reads overlap and share a common signature mutation in this overlap, the nucleotides in this overlap should be identical.If they have low level of identity, it is likely that an error has occurred, and therefore the mutation sequence read should be ignored.For example, the difference of one nucleotide can be tolerated, because it can be a simple sequence error.
[0174] Optionally, if signature mutation is unexpected mutation, this signature mutation will be ignored.Term " unexpected mutation " refers to the mutation that is unlikely to occur when using the individual step of introducing mutation into at least one target template nucleic acid molecule.For example, if the step of introducing mutation into at least one target template nucleic acid molecule is carried out using the chemical mutagen that only introduces adenine to guanine, any substitution of cytosine will be unexpected, and the mutant sequence read containing this mutation should be ignored.
[0175] Optionally, identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule includes identifying mutant sequence reads that correspond to specific regions of the at least one target template nucleic acid molecule. For example, a user may only be interested in identifying mutant sequence reads that contain signature mutations in regions that overlap with other mutant sequence reads, and signature mutations that occur in other regions may be ignored.
[0176] Generally, mutant sequence reads having a set of signature mutations with a larger intersection and a smaller symmetric difference are more likely to be derived from the same at least one mutant target template nucleic acid molecule. For two mutant sequence reads A and B having signature mutations SM(A) and SM(B), Intersection(SM(A),SM(B))>= C and Symmetric difference(SM(A),SM(B))< Intersection(SM(A),SM(B)) [where C is greater than 4, greater than 5, less than 20, or less than 10, and SM(X) is a set of signature mutations of mutant sequence read X, which may be a subset of the signature mutations of X], then it can be assumed that A and B are derived from the same at least one mutant target template nucleic acid molecule.
[0177] If desired, a set of co-occurring mutations can be used in place of the signature mutation in the following formula: Intersection(SM(A),SM(B))>= C and Symmetric difference (SM(A),SM(B))< C2 * Intersection (SM(A),SM(B)) [where C2 is less than 3, less than 2, or less than or equal to 1.5, and SM(X) is the set of co-occurring mutations in mutant sequence read X, which may be a subset of the signature mutations of X].
[0178] Mutant sequence reads that share a common signature k-mer or a common signature mutation can be grouped together. Preferably, mutant sequence reads are grouped together if they share at least 1, at least 2, at least 3, at least 4, at least 5, or at least k common signature k-mers and / or common signature mutations. In such embodiments, "k" is the length of the k-mer used.
[0179] Determining the probability that two mutant sequence reads are derived from the same mutant target template nucleic acid molecule Mutant sequence reads that are likely to be derived from the same mutant target template nucleic acid molecule can be identified by calculating the odds ratio (the probability that the mutant sequence read is derived from the same mutant target template nucleic acid molecule: the probability that the mutant sequence read is not derived from the same mutant target template nucleic acid molecule).
[0180] If the odds ratio exceeds a certain threshold, the mutation sequence read is likely to be derived from the same at least one mutation target template nucleic acid molecule.Similarly, if the odds ratio of the first mutation sequence read and the second mutation sequence read is higher than that of the first mutation sequence read and other mutation sequence reads that map to the same region of the assembly graph, the first mutation sequence read is likely to be derived from the same at least one target template nucleic acid molecule as the second mutation sequence read.
[0181] The applied threshold can be at any level, and in practice the user decides the threshold for any given alignment method depending on the requirements.
[0182] For example, the user can determine what level of stringency (strictness) is required.When the user uses a method for sequencing or generating a sequence of at least one target template nucleic acid, where accuracy is not important, the selected threshold value can be significantly lower than when the user uses a method for sequencing or generating a sequence of at least one target template nucleic acid, where accuracy is important.For example, when the user uses a method for sequencing or generating a sequence of the target template nucleic acid in a sample to determine whether the sample contains multiple or only one bacterial strain, a lower level of accuracy may be required than when the user uses a method for sequencing or generating a sequence of a specific mutant gene to determine how the mutant gene differs from the native gene.Therefore, the threshold value can be changed (determined) based on the required stringency.
[0183] Similarly, the user can change the threshold depending on the mutation rate used in the step of introducing mutations into at least one target template nucleic acid molecule. When the mutation rate is high, a higher probability threshold can be used because it is easier to determine whether two mutant sequence reads are derived from the same mutant target template nucleic acid molecule.
[0184] Similarly, the user can change the threshold value depending on the size of the at least one target template nucleic acid molecule. The larger the size of the at least one target template nucleic acid molecule, the more difficult it is to sequence the entire length without introducing sequencing errors, and therefore the user may wish to use a higher threshold value for a longer at least one target template nucleic acid molecule.
[0185] Similarly, the user can change the threshold depending on time and resource constraints: if these constraints are higher, the user may settle for a lower threshold, resulting in less accurate alignment.
[0186] In addition, the user can change the threshold value according to the error rate of the step of sequencing at least one region of the mutation target template to obtain the mutation sequence read.When the error rate is high, the user can set a higher threshold value than when the error rate is low.This is because when the error rate is high, the data may have less information value regarding whether two mutation sequence reads are derived from the same mutation target template nucleic acid molecule, especially when the error has a bias in the same manner as the introduced mutation.
[0187] Optionally, identifying mutant sequence reads that are likely to originate from the same mutant target template nucleic acid molecule includes using a probability function based on the following parameters: a. A matrix of nucleotides (N) at each position in the mutated sequence reads and assembly graph; b. The probability (M) that a given nucleotide (i) is mutated to a lead nucleotide (j); c. The probability (E) that a given nucleotide (i) is misread by a lead nucleotide (j) given that the nucleotide is misread; and Probability (Q) that the nucleotide at position dY was misread.
[0188] Using the probability function, the odds ratio (probability that a mutant sequence read originates from the same mutant target template nucleic acid molecule:probability that a mutant sequence read does not originate from the same mutant target template nucleic acid molecule) can be determined.
[0189] If desired, the value of Q can be obtained by performing statistical analysis on the mutation sequence reads and non-mutation sequence reads, or based on prior knowledge of the accuracy of the sequencing method. For example, Q depends on the accuracy of the sequencing method used. Therefore, the user can determine the value of Q by sequencing nucleic acid molecules of known sequence and determining the average number of misread nucleotides. Alternatively, the user can select a subgroup of mutation sequence reads and non-mutation sequence reads and compare them. The difference between the mutation sequence reads and non-mutation sequence reads will be due to either sequencing errors or the introduction of mutations. The user can use statistical analysis to estimate the number of differences due to sequencing errors.
[0190] Optionally, the values of M and E are estimated based on statistical analysis performed on a subset of mutant and non-mutated sequence reads, where the subset includes mutant and non-mutated sequence reads selected as mapping to the same region of the reference assembly graph. An example of a method for determining M and E is provided in Example 6. Briefly, a user can perform statistical analysis on a subset of mutant and non-mutated sequence reads to obtain best-fit values for M and E (through unsupervised learning). Because unsupervised learning can be a computationally expensive process, it is advantageous to perform this step on a subset of mutant and non-mutated sequence reads, and then later apply the values of M and E to the complete set of mutant and non-mutated sequence reads.
[0191] If desired, statistical analysis is performed using Bayesian inference, Monte Carlo methods such as Hamiltonian Monte Carlo, variational inference, or maximum likelihood analogs of Bayesian inference.
[0192] Optionally, identifying mutant sequence reads that are likely to originate from the same mutant target template nucleic acid molecule includes using machine learning or neural nets, such as those described in detail in Russell & Norvig "Artificial Intelligence, a modern approach."
[0193] Pre-clustering Optionally, the method includes a pre-clustering step. For example, the user can perform an initial calculation to assign mutant sequence reads to groups, where each member of the same group has a reasonable likelihood of originating from the same at least one mutation target template nucleic acid molecule. The mutant sequence reads in each group can be mapped to a common position on the assembly graph and / or share a common mutation pattern. Two mutant sequence reads in a group will be mapped to a common position on the assembly graph if they map to the same region and if they overlap in the assembly graph. The likelihood threshold applied in the pre-clustering step can be lower than that applied in the step of identifying mutant sequence reads that are likely to originate from the same at least one mutation target template nucleic acid molecule. In other words, the pre-clustering step can be a step with lower stringency than the step of identifying mutant sequence reads that are likely to originate from the same at least one mutation target template nucleic acid molecule.
[0194] Optionally, the identification of mutant sequence reads that are likely to be derived from the same mutant target template nucleic acid molecule is constrained by the results of the pre-clustering step. For example, the user can apply a lower stringency pre-clustering step to group mutant sequence reads that map to a common region of the assembly graph and have a reasonable likelihood of being derived from the same at least one mutant target template nucleic acid molecule. The user can then apply a higher stringency step to each member of the group, which identifies mutant sequence reads that are likely to be derived from the same at least one mutant target template nucleic acid molecule, to determine which of them are actually likely to be derived from the same at least one mutant target template nucleic acid molecule. The advantage of using a pre-clustering step is that a higher stringency step uses more processing power than a lower stringency step; in this example, the higher stringency step only needs to be applied to mutant sequence reads that are assigned to the same group by the lower stringency step, thereby reducing the overall processing power required.
[0195] Optionally, the pre-clustering step includes Markov clustering or Louvain clustering (https: / / micans.org / mcl / and https: / / arxiv.org / abs / 0803.0476).
[0196] Optionally, the pre-clustering step is performed by assigning mutant sequence reads that share at least one, at least two, at least three, at least five, or at least k signature k-mers or at least one, at least two, at least three, or at least five signature mutations to the same group, as described above. Optionally, mutant sequence reads are reasonably likely to originate from the same at least one mutant target template nucleic acid molecule if they share a common mutation pattern, and the mutant sequence reads that share a common mutation pattern are mutant sequence reads that contain at least one, at least two, at least three, at least five, or at least k common signature k-mers or common signature mutations.
[0197] Optionally, as described under the heading "signature k-mer or signature mutation," a signature k-mer is a k-mer that does not occur (or occurs less frequently) in non-mutated sequence reads but occurs at least twice (optionally, at least three, at least four, at least five, or at least ten times) in mutant sequence reads. Optionally, a signature mutation is a nucleotide that occurs at least twice (optionally, at least three, at least four, at least five, or at least ten times) in mutant sequence reads and does not occur (or occurs less frequently) at the corresponding position in non-mutated sequence reads.
[0198] Ignoring inferred paths by assembly graph In some embodiments of the invention, identifying nodes that form part of a valid path through the assembly graph comprises ignoring a putative path through the assembly graph.
[0199] For example, the estimated path from the assembly graph is (i) if they have ends that do not match those present in the library of end sequences; (ii) if they are the result of mold collisions, (iii) they are longer or shorter than expected; and / or (iv) if they have atypical coverage depth; can be ignored.
[0200] The term "template collision" refers to a situation in which two putative pathways are identified by assembly graphs that correspond to one or more of the same mutant sequence reads or to one or more mutant sequence reads with the same mutation pattern (the two putative pathways collide).
[0201] Ignoring putative paths due to assembly graphs with unmatched ends The method can include preparing a library of paired sequences of the ends of at least one mutant target template nucleic acid molecule. For example, the library can specify that a first at least one target template nucleic acid molecule has end sequences of A and B, and a second at least one target template nucleic acid molecule has end sequences of C and D. The library can be prepared by performing paired-end sequencing of at least one target template nucleic acid molecule. Optionally, the method can include sequencing the ends of at least one target template nucleic acid molecule using mate pair sequencing.
[0202] In such embodiments, identifying nodes that form part of valid pathways through the assembly graph includes ignoring putative pathways that have mismatched ends, i.e., the sequences of the ends of the putative pathway do not correspond to one of the pairs in the library. For example, if a library specifies that a first at least one target template nucleic acid molecule has end sequences of A and B and a second at least one target template nucleic acid molecule has end sequences of C and D, then a putative pathway that pairs end A with end D is a false pathway and should be ignored.
[0203] To ignore putative pathways with mismatched ends, the user can map the sequence of the ends of at least one target template nucleic acid molecule to the assembly graph. Optionally, the user may also wish to map the sequence of the ends of at least one target template nucleic acid molecule to the assembly graph to identify where each at least one target template nucleic acid molecule begins and ends on the assembly graph to aid the user in constructing the sequence of at least a portion of the at least one target template nucleic acid molecule from the non-mutated sequence reads.
[0204] Optionally, at least one target template nucleic acid molecule comprises at least one barcode. Optionally, at least one target template nucleic acid molecule comprises a barcode at each end. The term "at each end" means that a barcode is present fairly close to both ends of at least one target template nucleic acid molecule, for example, within 50 base pairs, 25 base pairs, or 10 base pairs from the end of at least one target template nucleic acid molecule. When at least one target template nucleic acid molecule comprises at least one barcode, it is easier for the user to determine whether a putative pathway has mismatched ends. This is because the end sequences are more distinctive, and it is easier to determine whether the sequences at the two ends that appear to be mismatched are actually mismatched, or whether a sequence error has been introduced into the sequence at one end.
[0205] Barcodes and Sample Tags For purposes of the present invention, a barcode (also referred to herein as a "unique molecular tag" or "unique molecular identifier") is a degenerate or randomly generated nucleotide sequence. A target template nucleic acid molecule may contain one, two, or three barcodes. In certain embodiments, each barcode may have a sequence that is different from all other generated barcodes. However, in other embodiments, two or more barcode sequences may be the same, i.e., a barcode sequence may be present multiple times. For example, at least 90% of the barcode sequences may be different from the sequences of all other barcode sequences. It is merely required that the barcodes be appropriately degenerate so that each target template nucleic acid molecule contains a barcode with a unique or substantially unique sequence compared to each other target template nucleic acid molecule in a sample pair. Thus, labeling (or tagging) target template nucleic acid molecules with barcodes allows target template nucleic acid molecules to be distinguished from one another, thereby facilitating the methods described elsewhere herein. Thus, barcodes may be considered unique molecular tags (UMTs). Barcodes can be 5, 6, 7, 8, 5-25, 6-20 or more nucleotides in length.
[0206] Optionally, as described above, at least one target template nucleic acid molecule in different sample pairs can be labeled with a different sample tag.
[0207] For purposes of the present invention, a sample tag is a tag used to label a substantial portion of at least one target template nucleic acid molecule in a sample. Different sample tags can be used in additional samples to distinguish which of the at least one target template nucleic acid molecule originates from which sample. A sample tag is a known sequence of nucleotides. A sample tag can be 5, 6, 7, 8, 5-25, 6-20, or more nucleotides in length.
[0208] Optionally, the method of the present invention includes a step of introducing at least one barcode or sample tag into at least one target template nucleic acid molecule. The at least one barcode or sample tag can be introduced using any suitable method, including PCR, tagging, and a combination of physical shearing or restriction digestion of the target nucleic acid followed by adapter ligation (optionally, sticky end ligation). For example, PCR can be performed on at least one target template nucleic acid molecule using a first set of primers that can hybridize to at least one target nucleic acid molecule. At least one barcode or sample tag can be introduced into each of the at least one target template nucleic acid molecule by PCR using primers that include a portion (5'-terminal portion) containing a barcode, sample tag, and / or adapter and a portion (3'-terminal portion) having a sequence that can hybridize to (and optionally be complementary to) at least one target nucleic acid molecule. Such primers hybridize to at least one target template nucleic acid molecule, and then PCR primer extension provides at least one target template nucleic acid molecule containing a barcode and / or sample tag. Additional cycles of PCR using these primers can be performed to optionally add additional barcodes or sample tags to the other end of at least one target template nucleic acid molecule. The primers can be degenerate, i.e., the 3'-terminal portions of the primers can be similar to each other but not identical.
[0209] At least one barcode or sample tag can be introduced using tagging. At least one barcode or sample tag can be introduced using direct tagging. Alternatively, at least one barcode or sample tag can be introduced by performing tagging and then performing two cycles of PCR using primers including a portion capable of hybridizing to a predetermined sequence and a portion including a barcode, sample tag, and / or adapter to introduce a predetermined sequence. At least one barcode or sample tag can be introduced by restriction digestion of at least one original target template nucleic acid molecule and subsequent ligation of a nucleic acid including a barcode and / or sample tag. Restriction digestion of at least one original nucleic acid molecule should be performed so that the digestion yields a nucleic acid molecule containing the region to be sequenced (at least one target template nucleic acid molecule). At least one barcode or sample tag can be introduced by shearing at least one target template nucleic acid molecule, followed by end repair, A-tailing, and then ligation of a nucleic acid including a barcode and / or sample tag.
[0210] Ignore predicted paths that are the result of template collisions The method may include ignoring putative pathways that are the result of template collisions. As mentioned above, the term "template collisions" refers to a situation in which two putative pathways are identified (two putative pathways collide) by assembly graphs that correspond to one or more of the same mutant sequence reads or one or more of the mutant sequence reads that have the same mutation pattern. Because each valid pathway should contain a unique set of mutant sequence reads, it is highly likely that at least one of the two colliding putative pathways is false. For these reasons, ignoring putative pathways that are the result of template collisions may reduce the number of false pathways identified.
[0211] Similarly, two different at least one mutant target template nucleic acid molecules may have similar or identical mutation patterns because they did not undergo many mutations during the step of introducing mutations into the at least one target template nucleic acid molecule, or because the mutations they did undergo were coincidentally the same. In such a situation, template collision occurs again. In such a situation, it is virtually impossible to construct the sequence of at least a portion of the at least one target template nucleic acid molecule from non-mutated sequence reads using information obtained by analyzing these insufficiently mutated at least one mutant target template nucleic acid molecule, and the putative pathways corresponding to the nodes calculated from the non-mutated sequence reads derived from such insufficiently mutated at least one mutant target template nucleic acid molecule should be ignored.
[0212] Ignoring estimated routes that are longer or shorter than expected At least one target template nucleic acid molecule can be of a known or predictable length.
[0213] Length can be determined by analyzing the length of at least one target template nucleic acid molecule in a laboratory setting. For example, a user can use gel electrophoresis to isolate a sample of at least one target template nucleic acid molecule and use that sample in the methods of the present invention. In such cases, all of the at least one target template nucleic acid molecule to be sequenced or generated will be within a known size range. For example, a user can extract bands corresponding to at least one target template nucleic acid molecule of 6,000 to 14,000 or 18,000 to 12,000 bp in length from a gel subjected to gel electrophoresis. Alternatively or additionally, the size of at least one target template nucleic acid molecule can be quantified using various methods for determining the size of nucleic acid molecules, including gel electrophoresis. For example, a user can use an instrument such as an Agilent Bioanalyzer or FemtoPulse instrument.
[0214] If the size of at least one target template nucleic acid molecule is known or predictable, predicted paths longer and shorter than a defined length are likely to be inaccurate and should be ignored.
[0215] Ignore estimated paths with atypical coverage depth The method of the present invention can include amplifying at least one mutant target template nucleic acid molecule, i.e., replicating at least one mutated target nucleic acid molecule to obtain copies of at least one mutant target template nucleic acid molecule.For example, the method can include using PCR to amplify at least one mutant target template nucleic acid molecule.The amplification can cause at least some of the mutant target template nucleic acid molecules to be replicated more times than other molecules.If some of the at least one mutant target template nucleic acid molecule are amplified to a higher degree (have a higher coverage depth) than other at least one mutant target template nucleic acid molecule, the putative pathway corresponding to these at least one mutant target template nucleic acid molecule will have a larger number of mutation sequence reads associated with them than others.Similarly, the coverage depth is expected to be consistent throughout the length of at least one template nucleic acid molecule.Therefore, different effective pathway parts are expected to have a similar number of mutation sequence reads associated with them (similar coverage depth). If the estimated path contains a portion with low depth of coverage and a portion with high depth of coverage, the two portions likely do not correspond to the same valid path, and the estimated path is erroneous and should be ignored.
[0216] Assembling at least a portion of the sequence of at least one target template nucleic acid molecule Optionally, a sequence is assembled for at least a portion of at least one target template nucleic acid molecule from non-mutated sequence reads that form part of a validated path according to the assembly graph.
[0217] Optionally, the method does not include generating a consensus sequence from the mutant sequence reads. Optionally, the method does not include constructing the sequence of at least one mutant target template nucleic acid molecule, or a majority of at least one mutant target template nucleic acid molecule.
[0218] A "consensus sequence" is intended to refer to a sequence that includes likely nucleotides at each position, e.g., the most frequently occurring nucleotides at each position in a group of sequence reads that are aligned with each other, as determined by analyzing a group of sequence reads that are aligned with each other.
[0219] The method includes constructing a sequence of at least a portion of at least one target template nucleic acid molecule from nodes that form a valid path through an assembly graph. Optionally, constructing a sequence of at least a portion of at least one target template nucleic acid molecule includes constructing a sequence of at least a portion of at least one target template nucleic acid molecule from nodes that form part of a valid path through the assembly graph.
[0220] Optionally, constructing the sequence of at least a portion of at least one target template nucleic acid molecule includes identifying an "end wall." An end wall is a location on the assembly graph corresponding to a plurality of "end + internal reads" (an end read corresponds to one of the ends of at least one target template nucleic acid molecule, and an internal read corresponds to an internal sequence (i.e., a sequence not present at an end of at least one target template nucleic acid molecule)). The end reads may be generated, for example, using paired-end sequencing. Optionally, an end wall is identified as a location on the assembly graph to which at least five end reads are mapped. Optionally, an end wall is identified as a location on the assembly graph to which two to four end reads are mapped and at least five end reads or internal reads are mapped. Optionally, constructing the sequence of at least a portion of at least one target template nucleic acid molecule includes constructing the sequence of at least a portion of at least one target template nucleic acid molecule from nodes forming part of a valid path through the assembly graph, the construction process starting from the end wall.
[0221] As described above, a valid path through an assembly graph can include connecting nodes. When a series of connecting nodes forms a single path through an assembly graph consisting of one or more nodes (e.g., where the nodes of the graph can be unitary), the sequence covered by the connecting nodes represents at least a portion of at least one target template nucleic acid molecule. These portions can then be constructed by connecting the nodes using standard techniques, such as canu (https: / / github.com / marbl / canu) or miniasm (https: / / github.com / lh3 / miniasm). For example, a user can prepare a consensus sequence from the nodes that form the valid path.
[0222] Optionally, the assembled sequence comprises nodes calculated primarily from non-mutated sequence reads. A constructed sequence is said to comprise nodes calculated primarily from non-mutated sequence reads if the sequence is constructed from nodes calculated from more than 50% of non-mutated sequence reads. Constructing a sequence from nodes calculated primarily from non-mutated sequence reads is advantageous because the assembled sequence most likely corresponds exactly to the original sequence of at least one target template nucleic acid molecule. However, if it is not possible to map non-mutated sequence reads to a portion of the predicted path by the assembly graph, missing portions of the sequence can be constructed from nodes calculated from mutant sequence reads. Preferably, the assembled sequence comprises nodes calculated from more than 50% (i.e., more than 50%), more than 60%, more than 70%, more than 80%, more than 90%, more than 98%, 50% to 100%, 60% to 100%, 70% to 100%, or 80% to 100% of non-mutated sequence reads.
[0223] Amplification of at least one target template nucleic acid molecule The method may include amplifying at least one target template nucleic acid molecule in a first of a pair of samples prior to sequencing a region of the at least one target template nucleic acid molecule.The method may include amplifying at least one target template nucleic acid molecule in a second of a pair of samples prior to sequencing a region of the at least one mutant target template nucleic acid molecule.
[0224] Suitable methods for amplifying at least one target template nucleic acid molecule are known in the art. For example, PCR is commonly used. PCR is described in more detail above in the section "Introduction of mutations into at least one target template nucleic acid molecule."
[0225] Fragmentation of at least one target template nucleic acid molecule The method can include fragmenting at least one target template nucleic acid molecule in a first sample pair prior to sequencing a region of the at least one target template nucleic acid molecule. Optionally, the method includes fragmenting at least one target template nucleic acid molecule in a second sample pair prior to sequencing a region of the at least one mutant target template nucleic acid molecule.
[0226] At least one target template nucleic acid molecule can be fragmented using any suitable technique. For example, fragmentation can be performed using restriction digestion or PCR using primers complementary to at least one internal region of at least one mutant target nucleic acid molecule. Preferably, fragmentation is performed using a technique that generates random fragments. The term "random fragment" refers to randomly generated fragments, such as fragments generated by tagging. Fragments generated using restriction enzymes are not "random" because restriction digestion occurs at a specific DNA sequence defined by the restriction enzyme used. Even more preferably, fragmentation is performed by tagging. When fragmentation is performed by tagging, the tagging reaction optionally introduces an adapter region into at least one mutant target nucleic acid molecule. This adapter region is a short DNA sequence that can encode an adapter that allows at least one mutant target nucleic acid molecule to be sequenced using, for example, Illumina technology.
[0227] Low-bias DNA polymerase As mentioned above, mutations can be introduced using a low-bias DNA polymerase. Low-bias DNA polymerases can introduce mutations in a uniformly random manner, which can be beneficial in the methods of the present invention. This is because when mutations are introduced in a uniformly random manner, any given portion of the template nucleic acid is more likely to have a unique mutation pattern. As mentioned above, unique mutation patterns can be useful for identifying effective pathways through assembly graphs.
[0228] In addition, methods using DNA polymerases with high template amplification bias may be limited: DNA polymerases with high template amplification bias may replicate and / or mutate some target template nucleic acid molecules better than others, and therefore sequencing methods using such highly biased DNA polymerases may not be able to successfully sequence some target template nucleic acid molecules.
[0229] A low-bias DNA polymerase can have a low template amplification bias and / or a low mutation bias.
[0230] Low mutation bias A low-bias DNA polymerase that exhibits low mutation bias is a DNA polymerase that can mutate adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine at similar rates. In one embodiment, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at similar rates.
[0231] Depending on the circumstances, the low-bias DNA polymerase can mutate adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1: 1. Preferably, the low-bias DNA polymerase can mutate guanine and adenine at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1: 1. Preferably, the low-bias DNA polymerase can mutate thymine and cytosine at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively.
[0232] In such embodiments, in the step of introducing mutations into the plurality of target template nucleic acid molecules, the low-bias DNA polymerase mutates adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine nucleotides in the at least one target template nucleic acid molecule at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively. Preferably, the low-bias DNA polymerase mutates guanine and adenine nucleotides in the at least one target template nucleic acid molecule at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1: 1. Preferably, the low-bias DNA polymerase mutates thymine and cytosine nucleotides in the at least one target template nucleic acid molecule at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1: 1.
[0233] Depending on the circumstances, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at a rate ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1. Preferably, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at a rate ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3.
[0234] In such embodiments, in the step of introducing mutations into the at least one target template nucleic acid molecule in a second of the pair of samples, the low-bias DNA polymerase may mutate adenine, thymine, guanine, and cytosine nucleotides in the at least one target template nucleic acid molecule at a rate ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase mutates adenine, thymine, guanine, and cytosine nucleotides in the at least one target template nucleic acid molecule at a rate ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3.
[0235] Adenine, thymine, cytosine, and / or guanine may be replaced with another nucleotide.For example, if a low-bias DNA polymerase can mutate adenine, enzymatic mutagenesis using the low-bias DNA polymerase can replace at least one adenine nucleotide in the nucleic acid molecule with thymine, guanine, or cytosine.Similarly, if a low-bias DNA polymerase can mutate thymine, enzymatic mutagenesis using the low-bias DNA polymerase can replace at least one thymine nucleotide with adenine, guanine, or cytosine.If a low-bias DNA polymerase can mutate guanine, enzymatic mutagenesis using the low-bias DNA polymerase can replace at least one adenine nucleotide with thymine, guanine, or cytosine.If a low-bias DNA polymerase can mutate cytosine, enzymatic mutagenesis using the low-bias DNA polymerase can replace at least one cytosine nucleotide with thymine, guanine, or adenine.
[0236] Although low-bias DNA polymerases may not be able to directly substitute a nucleotide, they may still be able to mutate that nucleotide by replacing the corresponding nucleotide on the complementary strand. For example, if a target template nucleic acid molecule contains thymine, an adenine nucleotide will be present at the corresponding position of at least one nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule. Low-bias DNA polymerases may be able to replace the adenine nucleotide of the at least one nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule with guanine, and thus, when the at least one nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule is replicated, a cytosine will be present in the corresponding replicated at least one target template nucleic acid molecule where the thymine was originally present (thymine-to-cytosine substitution).
[0237] In some embodiments, the low-bias DNA polymerase mutates 1% to 15%, 2% to 10%, or about 8% of the nucleotides in at least one target template nucleic acid. In such embodiments, enzymatic mutagenesis using the low-bias DNA polymerase is carried out in a manner that mutates 1% to 15%, 2% to 10%, or about 8% of the nucleotides in the at least one target template nucleic acid. For example, if a user desires to mutate about 8% of the nucleotides in a target template nucleic acid molecule and the low-bias DNA polymerase mutates about 1% of the nucleotides per round of replication, the step of introducing mutations into the plurality of target template nucleic acid molecules by enzymatic mutagenesis may include eight rounds of replication in the presence of the low-bias DNA polymerase.
[0238] In some embodiments, the low-bias DNA polymerase is capable of mutating 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5% of the nucleotides in the at least one target template nucleic acid molecule per round of replication. In some embodiments, the low-bias DNA polymerase mutates 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5% of the nucleotides in the at least one target template nucleic acid molecule per round of replication. The actual amount of mutations occurring in each round may vary, but will average 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5%.
[0239] Whether DNA polymerase can mutate nucleotides, and if so, how often Whether a low-bias DNA polymerase can mutate a certain percentage of nucleotides in the at least one target template nucleic acid molecule per round of replication can be determined by amplifying a nucleic acid molecule of known sequence in the presence of a low-bias DNA polymerase for a certain number of replication rounds. The resulting amplified nucleic acid molecule can then be sequenced to calculate the percentage of nucleotides mutated per round of replication. For example, a nucleic acid molecule of known sequence can be amplified using 10 rounds of PCR in the presence of a low-bias DNA polymerase. The resulting nucleic acid molecule can then be sequenced. If the resulting nucleic acid molecule contains 10% nucleotides that are different from the corresponding nucleotides in the original known sequence, the user will understand that the low-bias DNA polymerase can mutate, on average, 1% of the nucleotides in the at least one target template nucleic acid molecule per round of replication. Similarly, to determine whether a low-bias DNA polymerase mutates a certain percentage of nucleotides in the at least one target template nucleic acid molecule in a given manner, a user could perform the method on a nucleic acid molecule of known sequence and, once the method is complete, use sequencing to determine the percentage of mutated nucleotides.
[0240] A low-bias DNA polymerase can mutate a nucleotide, such as adenine, when used to amplify a nucleic acid molecule, providing a nucleic acid molecule in which some instances of the nucleotide are substituted or deleted. Preferably, the term "mutate" refers to the introduction of a substitution mutation, and in some embodiments, the term "mutate" can be interchangeable with "introduce a substitution of."
[0241] A low-bias DNA polymerase mutates a nucleotide such as adenine in at least one target template nucleic acid molecule when the step of introducing mutations into a plurality of target template nucleic acid molecules using the low-bias DNA polymerase results in at least one mutated target template nucleic acid molecule (some instances of the nucleotide are mutated). For example, if a low-bias DNA polymerase mutates adenine in the at least one target template nucleic acid molecule, when the step of introducing mutations into a plurality of target template nucleic acid molecules using the low-bias DNA polymerase results in at least one mutated target template nucleic acid molecule (at least one adenine is substituted or deleted).
[0242] To determine whether a DNA polymerase can introduce a particular mutation, one skilled in the art need only test the DNA polymerase using a nucleic acid molecule of known sequence. A suitable nucleic acid molecule with a known sequence is a fragment obtained from a bacterial genome with a known sequence, such as E. coli MG1655. One skilled in the art can use PCR in the presence of a low-bias DNA polymerase to amplify a nucleic acid molecule of known sequence. One skilled in the art can then sequence the amplified nucleic acid molecule and determine whether its sequence is identical to the original known sequence. Even if this is not sufficient, one skilled in the art can determine the nature of the mutation. For example, if one skilled in the art wants to determine whether a DNA polymerase can mutate adenine using a nucleotide analog, one skilled in the art can use PCR to amplify a nucleic acid molecule of known sequence in the presence of a nucleotide analog and sequence the resulting amplified nucleic acid molecule. If the amplified DNA contains a mutation at a position corresponding to the adenine nucleotide in the known sequence, one skilled in the art will understand that a DNA polymerase can mutate adenine using a nucleotide analog.
[0243] The rate ratio can be calculated in a similar manner. For example, if a person skilled in the art wants to determine the rate ratio of guanine and cytosine nucleotide mutations, he or she can use PCR in the presence of a low-bias DNA polymerase to amplify a nucleic acid molecule with a known sequence. He or she can then sequence the resulting amplified nucleic acid molecule to determine how many guanine nucleotides have been substituted or deleted, and how many cytosine nucleotides have been substituted or deleted. The rate ratio is the ratio of the number of substituted or deleted guanine nucleotides to the number of substituted or deleted cytosine nucleotides. For example, if 16 guanine nucleotides have been substituted or deleted, and 8 cytosine nucleotides have been substituted or deleted, the guanine and cytosine nucleotides are mutated at a rate ratio of 16:8 or 2:1, respectively.
[0244] Use nucleotide analogs Although a low-bias DNA polymerase may not be able to directly replace nucleotides with other nucleotides (at least not frequently), the low-bias DNA polymerase may still be able to mutate nucleic acid molecules using nucleotide analogs. A low-bias DNA polymerase may be able to replace nucleotides with other natural nucleotides (i.e., cytosine, guanine, adenine, or thymine) or nucleotide analogs.
[0245] For example, low-bias DNA polymerase can be high-fidelity DNA polymerase.High-fidelity DNA polymerases are generally prone to introduce few mutations because they are highly accurate.However, the present inventors have found that some high-fidelity DNA polymerases can still mutate target template nucleic acid molecules because they can introduce nucleotide analogues into target template nucleic acid molecules.
[0246] In certain embodiments, in the absence of nucleotide analogs, the high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, 0% to 0.0015%, or 0% to 0.001% mutations per round of replication.
[0247] In some embodiments, the low-bias DNA polymerase can incorporate a nucleotide analog into the at least one target template nucleic acid molecule. In some embodiments, the low-bias DNA polymerase incorporates a nucleotide analog into the at least one target template nucleic acid molecule. In some embodiments, the low-bias DNA polymerase can use a nucleotide analog to mutate adenine, thymine, guanine, and / or cytosine. In some embodiments, the low-bias DNA polymerase uses a nucleotide analog to mutate adenine, thymine, guanine, and / or cytosine in the at least one target template nucleic acid molecule. In some embodiments, the DNA polymerase replaces guanine, cytosine, adenine, and / or thymine with a nucleotide analog. In some embodiments, the DNA polymerase can replace guanine, cytosine, adenine, and / or thymine with a nucleotide analog.
[0248] Incorporation of a nucleotide analog into the at least one target template nucleic acid molecule can be used to mutate nucleotides, since the nucleotide analog may be incorporated in place of an existing nucleotide or may pair with a nucleotide in the opposite strand. For example, dPTP can be incorporated into a nucleic acid molecule in place of a pyrimidine nucleotide (may replace thymine or cytosine). When incorporated into a nucleic acid strand, dPTP can pair with adenine when in the imino tautomeric form. Thus, when a complementary strand is formed, the complementary strand may have an adenine present in a position complementary to dPTP. Similarly, when incorporated into a nucleic acid strand, dPTP can pair with guanine when in the amino tautomeric form. Thus, when a complementary strand is formed, the complementary strand may have a guanine present in a position complementary to dPTP.
[0249] For example, when dPTP is introduced into the at least one target template nucleic acid molecule of the present invention, when at least one nucleic acid molecule complementary to the at least one target template nucleic acid molecule is formed, the at least one nucleic acid molecule complementary to this at least one target template nucleic acid molecule contains adenine or guanine at the position complementary to dPTP in the at least one target template nucleic acid molecule (depending on whether dPTP is its amino or imino form).When the at least one nucleic acid molecule complementary to the at least one target template nucleic acid molecule is replicated, the resulting copy of the at least one target template nucleic acid molecule contains thymine or cytosine at the position corresponding to dPTP in the at least one target template nucleic acid molecule.Therefore, mutation to thymine or cytosine can be introduced into the at least one target template nucleic acid molecule to be mutated.
[0250] Alternatively, when dPTP is introduced into at least one nucleic acid molecule that is complementary to said at least one target template nucleic acid molecule, when said at least one target template nucleic acid molecule is replicated, said at least one target template nucleic acid molecule replicates will contain adenine or guanine (depending on the tautomeric form of dPTP) at the position that is complementary to dPTP in said at least one nucleic acid molecule that is complementary to said at least one target template nucleic acid molecule.Therefore, the mutation of adenine or guanine can be introduced into said mutated at least one target template nucleic acid molecule.
[0251] In certain embodiments, the low-bias DNA polymerase can replace cytosine or thymine with a nucleotide analog. In further embodiments, the low-bias DNA polymerase uses the nucleotide analog to introduce guanine or adenine nucleotides at a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively. The guanine or adenine nucleotides may be introduced by the low-bias DNA polymerase by pairing them face-to-face with a nucleotide analog, such as dPTP. In further embodiments, the low-bias DNA polymerase uses the nucleotide analog to introduce guanine or adenine nucleotides at a rate ratio of 0.7-1.3:0.7-1.3, respectively.
[0252] One skilled in the art can use conventional methods to determine whether a low-bias DNA polymerase can incorporate a nucleotide analog into the at least one target template nucleic acid molecule, or whether a nucleotide analog can be used to mutate adenine, thymine, guanine, and / or cytosine in the at least one target template nucleic acid molecule.
[0253] For example, to determine whether a low-bias DNA polymerase can incorporate nucleotide analogs into at least one target template nucleic acid molecule, a person skilled in the art can use the low-bias DNA polymerase to amplify a nucleic acid molecule in two rounds of replication. The first round of replication should be performed in the presence of nucleotide analogs, and the second round of replication should be performed in the absence of nucleotide analogs. The resulting amplified nucleic acid molecule can be sequenced to determine whether mutations have been introduced, and if so, how many mutations have been introduced. The user should repeat this experiment without using nucleotide analogs and compare the number of mutations introduced with and without using nucleotide analogs. If the number of mutations introduced using nucleotide analogs is significantly higher than the number of mutations introduced without using nucleotide analogs, the user can conclude that the low-bias DNA polymerase can incorporate nucleotide analogs. Similarly, a person skilled in the art can determine whether a DNA polymerase incorporates nucleotide analogs or uses nucleotide analogs to mutate adenine, thymine, guanine, and / or cytosine. One skilled in the art need only carry out the method in the presence of nucleotide analogues and determine whether the method results in mutations at positions originally occupied by adenine, thymine, guanine, and / or cytosine.
[0254] If the user desires to mutate the at least one target template nucleic acid molecule using nucleotide analogs, the method may comprise amplifying the at least one target template nucleic acid molecule using a low-bias DNA polymerase, wherein amplifying the at least one target template nucleic acid molecule using the low-bias DNA polymerase is carried out in the presence of nucleotide analogs, and wherein amplifying the at least one target template nucleic acid molecule provides at least one target template nucleic acid molecule comprising nucleotide analogs.
[0255] Suitable nucleotide analogs include dPTP (2'-deoxy-P-nucleoside-5'-triphosphate), 8-oxo-dGTP (7,8-dihydro-8-oxoguanine), 5Br-dUTP (5-bromo-2'-deoxy-uridine-5'-triphosphate), 2OH-dATP (2-hydroxy-2'-deoxyadenosine-5'-triphosphate), dKTP (9-(2-deoxy-β-D-ribofuranosyl)-N6-methoxy-2,6-diaminopurine-5'-triphosphate), and dITP (2'-deoxyinosine 5'-triphosphate). The nucleotide analog may be dPTP. Nucleotide analogs may be used to introduce substitution mutations as described in Table 1.
[0256] [Table 1]
[0257] By using different nucleotide analogs alone or in combination, various mutations can be introduced into the at least one target template nucleic acid molecule.Therefore, low-bias DNA polymerase can use nucleotide analogs to introduce guanine to adenine substitution mutation, cytosine to thymine substitution mutation, adenine to guanine substitution mutation, and thymine to cytosine substitution mutation.Depending on the situation, low-bias DNA polymerase may use nucleotide analogs to introduce guanine to adenine substitution mutation, cytosine to thymine substitution mutation, adenine to guanine substitution mutation, and thymine to cytosine substitution mutation.
[0258] Low-bias DNA polymerases may be able to introduce guanine-to-adenine, cytosine-to-thymine, adenine-to-guanine, and thymine-to-cytosine substitution mutations at ratios of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or approximately 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase is capable of introducing guanine-to-adenine, cytosine-to-thymine, adenine-to-guanine, and thymine-to-cytosine substitution mutations at rate ratios of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively. Suitable methods for determining whether a low-bias DNA polymerase can introduce substitution mutations and at what rate ratios are described under the heading "Whether a DNA Polymerase Can Mutate Nucleotides, and If So, at What Rate?"
[0259] In some methods, the low-bias DNA polymerase introduces guanine-to-adenine, cytosine-to-thymine, adenine-to-guanine, and thymine-to-cytosine substitution mutations at a rate ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1, respectively. Preferably, the low-bias DNA polymerase introduces guanine-to-adenine, cytosine-to-thymine, adenine-to-guanine, and thymine-to-cytosine substitution mutations at a rate ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively. Suitable methods for determining whether substitution mutations have been introduced and at what rate ratios are described under the heading "Whether a DNA Polymerase Can Mutate Nucleotides, and If So, at What Rate?"
[0260] Generally, when a low-bias DNA polymerase uses a nucleotide analog to introduce a mutation, this requires two or more rounds of replication. In the first round of replication, the low-bias DNA polymerase introduces a nucleotide analog instead of a nucleotide, and in the second round of replication, the nucleotide analog pairs with a natural nucleotide, thereby introducing a substitution mutation into the complementary strand. The second round of replication may be carried out in the presence of a nucleotide analog. However, this method may further include a step of amplifying the at least one target template nucleic acid molecule in the second pair of samples containing a nucleotide analog in the absence of a nucleotide analog. The step of amplifying the at least one target template nucleic acid molecule containing a nucleotide analog in the absence of a nucleotide analog may be carried out using a low-bias DNA polymerase.
[0261] Low template amplification bias A low-bias DNA polymerase may have a low template amplification bias. A low-bias DNA polymerase has a low template amplification bias if it can amplify different target template nucleic acid molecules with similar success per cycle. A high-bias DNA polymerase may have difficulty amplifying template nucleic acid molecules with a high G:C content or a large degree of secondary structure. In some embodiments, the low-bias DNA polymerase has a low template amplification bias for template nucleic acid molecules that are less than 25,000, less than 10,000, 1 to 15,000, or 1 to 10,000 nucleotides in length.
[0262] In some embodiments, to determine whether a DNA polymerase has low template amplification bias, a person skilled in the art can use the DNA polymerase to amplify a wide range of different sequences, and then sequence the resulting amplified DNA to determine whether the different sequences are amplified at different levels.For example, a person skilled in the art can select a wide range of short (possibly 50 nucleotides) nucleic acid molecules with different characteristics, such as nucleic acid molecules that exhibit high GC content, nucleic acid molecules that exhibit low GC content, nucleic acid molecules with a high degree of secondary structure, and nucleic acid molecules with a low degree of secondary structure.The user can then use DNA polymerase to amplify these sequences and quantify the level at which each of the nucleic acid molecules is amplified.In some embodiments, if the levels are within 25%, 20%, 10%, or 5% of each other, then the DNA polymerase has low template amplification bias.
[0263] Alternatively, in certain embodiments, a DNA polymerase has low template amplification bias (Kolmogorov-Smirnov D = less than 0.1, less than 0.09, or less than 0.08) if the DNA polymerase can amplify a 7-10 kbp fragment. The Kolmogorov-Smirnov D at which a particular low-bias DNA polymerase can amplify a 7-10 kbp fragment may be determined using the assay described in Example 4.
[0264] The low-bias DNA polymerase may be a high-fidelity DNA polymerase. A high-fidelity DNA polymerase is a DNA polymerase that is not highly error-prone and therefore does not typically introduce many mutations when used to amplify a target template nucleic acid molecule in the absence of nucleotide analogs. High-fidelity DNA polymerases are not typically used in methods for introducing mutations because error-prone DNA polymerases are generally considered more efficient. However, this application demonstrates that certain high-fidelity polymerases can introduce mutations using nucleotide analogs, and that these mutations may be introduced with less bias compared to error-prone polymerases such as DNATaq polymerase.
[0265] High-fidelity DNA polymerase has additional advantages.High-fidelity DNA polymerase can introduce mutations when used with nucleotide analogs, but in the absence of nucleotide analogs, high-fidelity DNA polymerase can replicate target template nucleic acid molecules very accurately.This means that users can mutate at least one target template nucleic acid molecule very effectively, and then use the same DNA polymerase to amplify this mutated at least one target template nucleic acid molecule with high accuracy.When using low-fidelity DNA polymerase to mutate target template nucleic acid molecules, it may be necessary to remove the low-fidelity DNA polymerase from reaction mixture before amplifying target template nucleic acid molecules.
[0266] High-fidelity DNA polymerases may have proofreading activity. Proofreading activity can help DNA polymerases amplify target template nucleic acid sequences with high accuracy. For example, low-bias DNA polymerases may contain a proofreading domain. The proofreading domain may verify whether the nucleotide added by the polymerase is correct (ensuring that the nucleotide pairs correctly with the corresponding nucleic acid in the complementary strand), and if incorrect, delete the nucleotide from the nucleic acid molecule. The present inventors surprisingly found that in some DNA polymerases, the proofreading domain allows for the pairing of natural nucleotides with nucleotide analogs. The structure and sequence of suitable proofreading domains are known to those skilled in the art. DNA polymerases that contain a proofreading domain include members of DNA polymerase families I, II, and III, such as Pfu polymerase (from Pyrococcus furiosus), T4 polymerase (from bacteriophage T4), and polymerases from Thermococcales archaea, which are described in more detail below.
[0267] In certain embodiments, in the absence of nucleotide analogs, the high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, 0% to 0.0015%, or 0% to 0.001% mutations per round of replication.
[0268] In addition, the low-bias DNA polymerase may contain a processivity-enhancing domain, which allows the DNA polymerase to amplify a target template nucleic acid molecule more quickly, which is advantageous because it allows the methods of the invention to be performed more quickly.
[0269] Polymerases from Thermococcal Archaea In one embodiment, the low-bias DNA polymerase is a fragment or variant of a polypeptide comprising SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, or SEQ ID NO:7. The polypeptides of SEQ ID NOs:2, 4, 6, and 7 are polymerases from Thermococcales archaea. The polymerases of SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, or SEQ ID NO:7 are low-bias DNA polymerases that exhibit high fidelity and can mutate target template nucleic acid molecules by incorporating nucleotide analogs such as dPTP. The polymerases of SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, or SEQ ID NO:7 are particularly advantageous because they have low mutation bias and low template amplification bias. They are also highly processive and high-fidelity polymerases that contain a proofreading domain, meaning that they can rapidly and accurately amplify mutated target template nucleic acid molecules in the absence of nucleotide analogs.
[0270] Low-bias DNA polymerases are as follows: a. The sequence of SEQ ID NO: 2; b. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:2; c. The sequence of SEQ ID NO: 4; d. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:4; e. the sequence of SEQ ID NO: 6; f. a sequence at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:6; g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 7 The amino acid sequence may comprise a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of the amino acid sequence.
[0271] Preferably, the low-bias DNA polymerase is one of the following a to h: a. The sequence of SEQ ID NO: 2; b. a sequence that is at least 98%, or at least 99% identical to SEQ ID NO:2; c. The sequence of SEQ ID NO: 4; d. a sequence that is at least 98%, or at least 99% identical to SEQ ID NO:4; e. the sequence of SEQ ID NO: 6; f. a sequence at least 98%, or at least 99% identical to SEQ ID NO:6; g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 98%, or at least 99% identical to SEQ ID NO: 7 and a fragment of at least 700 consecutive amino acids of
[0272] Low-bias DNA polymerases are as follows: a. The sequence of SEQ ID NO: 2; b. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:2; c. The sequence of SEQ ID NO: 4; d. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:4; e. the sequence of SEQ ID NO: 6; f. a sequence at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:6; g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 7 may also include:
[0273] Preferably, the low-bias DNA polymerase is one of the following a. to h., a. The sequence of SEQ ID NO: 2; b. a sequence that is at least 98%, or at least 99% identical to SEQ ID NO:2; c. The sequence of SEQ ID NO: 4; d. a sequence that is at least 98%, or at least 99% identical to SEQ ID NO:4; e. the sequence of SEQ ID NO: 6; f. a sequence at least 98%, or at least 99% identical to SEQ ID NO:6; g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 98%, or at least 99% identical to SEQ ID NO: 7 Includes:
[0274] The low-bias DNA polymerase may be a polymerase from a Thermococcales archaea or a derivative thereof. The DNA polymerases of SEQ ID NOs: 2, 4, 6, and 7 are polymerases from Thermococcales archaea. Thermococcales archaeal polymerases are advantageous because they are generally high-fidelity polymerases with low mutation and template amplification bias and are capable of introducing mutations using nucleotide analogs.
[0275] A Thermococcal archaeal polymerase is a polymerase having the polypeptide sequence of a polymerase isolated from a strain of the genus Thermococcal. A derivative of a Thermococcal archaeal polymerase may be a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of a Thermococcal archaeal polymerase, or may be at least 95%, at least 98%, at least 99%, or 100% identical to a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of a Thermococcal archaeal polymerase. A derivative of a Thermococcal archaeal polymerase may be at least 95%, at least 98%, at least 99%, or 100% identical to a Thermococcal archaeal polymerase. The derivative of a Thermococcal archaeal polymerase may be at least 98% identical to the Thermococcal archaeal polymerase.
[0276] Thermococcal archaeal polymerases from any strain may be useful in the context of the present invention. In one embodiment, the Thermococcal archaeal polymerase is derived from a Thermococcales strain selected from the group consisting of T. kodakarensis, T. celer, T. siculi, and T. sp. KS-1. The Thermococcales archaeal polymerases from these strains are set forth as SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, and SEQ ID NO:7.
[0277] Depending on the context, a low-bias DNA polymerase is one that exhibits high catalytic activity at temperatures between 50°C and 90°C, between 60°C and 80°C, or at about 68°C. [Example]
[0278] Example 1 - Mutating Nucleic Acid Molecules Using Other Polymerases or PrimeStar GXL DNA molecules were fragmented to an appropriate size (e.g., 10 kb) and then tagmentation was used to attach defined sequence priming sites (adapters) to each end.
[0279] The first step was a tagmentation reaction to fragment DNA. 50 ng of high-molecular-weight genomic DNA from one or more strains in a volume of 4 μL or less was subjected to tagmentation under the following conditions: 50 ng of DNA was combined with 4 μL of Nextera Transposase (diluted 1:50) and 8 μL of 2X tagmentation buffer (20 mM Tris [pH 7.6], 20 mM MgCl, 20% (v / v) dimethylformamide) for a total volume of 16 μL. The reaction was incubated at 55°C for 5 minutes, and 4 μL of NT buffer (or 0.2% SDS) was added to the reaction, after which the reaction was incubated at room temperature for 5 minutes.
[0280] The tagmentation reaction was cleaned using SPRIselect beads (Beckman Coulter) according to the manufacturer's instructions for left-hand size selection using 0.6 volume beads, and DNA was eluted with molecular-grade water.
[0281] Following this, PCR was performed for six cycles using standard dNTP and dPTP combinations. Primestar GXL was used, and 12.5 ng of tagmented and purified DNA was added to a total reaction volume of 25 μL containing 1× GXL buffer, 200 μM each of dATP, dTTP, dGTP, and dCTP, as well as 0.5 mM dPTP, and 0.4 μM of the custom primers (Table 2).
[0282] [Table 2]
[0283] The reactions were subjected to thermal cycling in the presence of Primestar GXL as follows: an initial gap extension at 68°C for 3 minutes, followed by six cycles of 98°C for 10 seconds, 55°C for 15 seconds, and 68°C for 10 minutes.
[0284] The next step was a dPTP-free PCR ("recovery PCR") to remove template-derived dPTPs and replace them with transition mutations. The PCR reactions were purified with SPRIselect beads to remove excess dPTP and primers, and then subjected to 10 additional rounds (minimum 1 round, maximum 20) of amplification using primers (Table 3) that anneal to the fragment ends introduced during the dPTP incorporation cycles.
[0285] [Table 3]
[0286] This was followed by a gel extraction step to size-select the amplified and mutated fragments within the desired size range (e.g., 7-10 kb). Gel extraction can be performed manually or via an automated system such as BluePippin. This was followed by an additional round of PCR ("enrichment PCR") with 16-20 cycles.
[0287] After amplifying a defined number of long mutant templates, random fragmentation of the templates was performed to generate overlapping shorter fragments for sequencing. Fragmentation was performed by tagmentation.
[0288] The long DNA fragments obtained in the previous step were subjected to a standard tagmentation reaction (e.g., Nextera XT or Nextera Flex), except that the reaction was divided into three pools for PCR amplification. This allows for selective amplification of fragments from each end of the original template (including the sample tag) as well as internal fragments newly tagmented at both ends from the long template. This effectively creates three pools for sequencing on an Illumina instrument (e.g., MiSeq or HiSeq).
[0289] The method was repeated using standard Taq (Jena Biosciences) and a blend of Taq and a proofreading polymerase (DeepVent) called LongAmp (New England Biolabs).
[0290] The data obtained in this experiment are shown in Figure 1. dPTP was not used as a control. Reads were mapped to the E. coli genome, achieving a median mutation rate of approximately 8%.
[0291] Example 2 - Comparison of mutation frequencies of various DNA polymerases Mutagenesis was performed using a wide variety of DNA polymerases (Table 4). Genomic DNA from E. coli strain MG1655 was tagmented to generate long fragments and bead-purified as described in the methods of Example 1. This was followed by six cycles of "mutagenesis PCR" in the presence of 0.5 mM dPTP, SPRIselect bead purification, and an additional 14-16 cycles of "recovery PCR" in the absence of dPTP. The resulting long mutant templates were then subjected to a standard tagmentation reaction (see Example 1), and the "internal" fragments were amplified and sequenced on a MiSeq instrument.
[0292] Mutation rates are listed in Table 4 and are normalized to the frequency of base substitutions via dPTP mutagenesis reactions, as measured using Illumina sequencing of DNA from known reference genomes. For Taq polymerase, even when used in a buffer optimized for Thermococcus archaeal polymerases, less than 12% of mutations occurred at template G+C sites. Thermococcus-like polymerases generated 58–69% of mutations at template G+C sites, while polymerases from Pyrococcus archaea contributed 88% of mutations to template G+C sites.
[0293] Enzymes were obtained from Jena Biosciences (Taq), Takara (Primestar varieties), Merck Millipore (KOD DNA polymerase), and New England Biolabs (Phusion).
[0294] For this experiment, Taq was tested using the supplied buffer and also Primestar GXL Buffer (Takara). All other reactions were performed using the standard supplied buffer for each polymerase.
[0295] [Table 4]
[0296] Example 3 Determining dPTP Mutagenesis Rate We performed dPTP mutagenesis on a wide range of genomic DNA samples exhibiting various levels of G+C content (33-66%) using a Thermococcus archaeal polymerase (Primestar GXL; Takara) under a single set of reaction conditions. Mutagenesis and sequencing were performed as described in the methods of Example 1, except that 10 cycles of "recovery PCR" were performed. As expected, mutation rates were similar among samples despite the diversity of G+C content (median mutation rates of 7-8%) (Figure 2).
[0297] Example 4 - Measuring template amplification bias Template amplification bias was measured for two polymerases: Kapa HiFi, a proofreading polymerase commonly used in Illumina sequencing protocols, and PrimeStar GXL, a polymerase from the KOD family known for its ability to amplify long fragments. In the first experiment, Kapa HiFi was used to amplify a limited number of E. coli genomic DNA templates approximately 2 kbp in size. The ends of these amplified fragments were then sequenced. A similar experiment was performed using PrimeStar GXL on approximately 7-10 kbp fragments obtained from E. coli. The location of each end sequence read was determined by mapping to the E. coli reference genome. The distances between adjacent fragment ends were measured. These distances were compared to a series of distances randomly drawn from a uniform distribution. The comparison was performed via the nonparametric Kolmogorov-Smirnov test, D. If two samples come from the same distribution, the value of D approaches zero. For the low-bias PrimeStar polymerase, we observed D=0.07 when measured against a uniform random sample of 50,000 genomic locations for 50,000 fragment ends. For Kapa HiFi polymerase, we observed D=0.14 for 50,000 fragment ends.
[0298] Example 5 - Determining the size range of reconstructions Mutant and non-mutant sequence reads were generated, and the sequence of the non-mutant sequence reads was determined using computer-implemented method steps.
[0299] To generate mutant sequence reads, mutant target template nucleic acid molecule fragments were generated using the method described in Example 1, except that the fragment size range was limited to 1-2 kb. The mutant target template nucleic acid molecule fragments were sequenced using an Illumnia MiSeq with a V2 500-cycle flow cell.
[0300] To generate non-mutated sequence reads, the following steps were performed. The first step was a tagging reaction to fragment the DNA. 50 ng of high-molecular-weight genomic DNA in one or more bacterial strains in a volume of 4 μL or less was subjected to a tagging reaction under the following conditions: 50 ng of DNA was combined with 4 μL of Nextera transposase (diluted 1:50) and 8 μL of 2× tagging buffer (20 mM Tris [pH 7.6], 20 mM MgCl, 20% (v / v) dimethylformamide) in a total volume of 16 μL. The reaction was incubated at 55°C for 5 minutes, and 4 μL of NT buffer (or 0.2% SDS) was added to the reaction, and the reaction was incubated at room temperature for 5 minutes.
[0301] SPRIselect beads (Beckman Coulter) were used according to the manufacturer's instructions for left-side size selection using 0.6 volumes of beads. The tagging reactions were washed and the DNA was eluted in molecular-grade water. The long DNA fragments from the previous step were subjected to a standard tagging reaction (e.g., Nextera XT or Nextera Flex) except that the reaction was split into three pools for PCR amplification. This allows for selective amplification of fragments from each end of the original template (containing the sample tag) and internal fragments from the newly tagged long template at both ends. This effectively generates three pools for sequencing on an Illumina instrument (e.g., MiSeq or HiSeq).
[0302] The target template nucleic acid molecule was sequenced by pre-clustering the mutant sequence reads into read groups, and then each group of mutant reads was subjected to de novo assembly using steps 1 and 2 of the A5-miseq assembly pipeline (Coil et al., 2015 Bioinformatics). The analysis yielded 53,053 hypothetical fragments with lengths distributed as shown in Figure 4.
[0303] Example 6 - Testing of Probabilistic Algorithms A probabilistic algorithm was used to determine whether two mutant sequence reads originated from at least one original template nucleic acid molecule. Details of the probabilistic algorithm are as follows:
[0304] Given two non-mutated sequence reads S1 and S2 in a set of mutant sequence reads aligned to a non-mutated reference sequence R, the model described herein seeks to determine whether S1 and S2 were sequenced from the same at least one mutant template nucleic acid molecule or from different templates. The alignment of these three sequences can be represented as a 3×N matrix N of aligned positions (e.g., individual nucleotides s 1,i : s 2,j : r k where the alignment nucleotides are in the same column y of N. For example, n .,y ). For convenience, we define a mapping from nucleotides A, C, G, and T to integers 1, 2, 3, and 4, such that A maps to 1, C maps to 2, and so on. This mapping is implied in the remainder of the description below. Next, we define two 4x4 probability matrices M and E. Each entry m i,j denotes the probability that nucleotide i is mutated to nucleotide j by the mutagenesis process, for i,j∈{A,C,G,T}. Similarly, entry e i,jdenotes the conditional probability that nucleotide i is misread as nucleotide j given that nucleotide j is misread, for i, j ∈ {A, C, G, T}. Furthermore, define a 2 × N matrix Q, where the entries q 1,y and q 2,y indicates the probability that the nucleotide at alignment position y was misread with respect to sequences S1 and S2, respectively, as reported by the sequencer. Finally, we use z∈{0,1} as an indicator value for whether two sequence reads are derived from the same mutant template, where z=1 indicates that S1 and S2 were sequenced from the same template fragment, and z=0 indicates that S1 and S2 were sequenced from different template fragments.
[0305] While the values of Q and N are provided / determined by sequencing and the subsequent read mapping process, the values of M, E, and z are generally unknown. Fortunately, these values (and any other unknown parameters) can be estimated from the data using any of a wide variety of techniques. Based on knowledge of the mutation process, prior distributions can be imposed on the values of the unknown parameters. A Dirichlet distribution is imposed on the rows of M, and m 1, .~Dirichlet(α+β,1-β,1-α,1-β), where the entries correspond to the events A→A (no mutation), A→C (transversion), A→G (transition), and A→T (transversion). Here, α is the unknown transition rate hyperparameter and β is the unknown transversion rate hyperparameter. The full prior for M is defined as follows: m 1, .~Dirichlet(α+β,1-β,1-α,1-β) m 2, .~Dirichlet(1-β,α+β,1-β,1-α) m 3, .~Dirichlet(1-α,1-β,α+β,1-β) m 4, .~Dirichlet(1-β,1-α,1-β,α+β)
[0306] Prior knowledge of the mutation process (e.g., knowledge of the properties of the polymerase or other mutagen) is generally available to the experimenter and may allow hyperpriors to be applied to the α and β terms. More general constructions for the priors on M are possible: uniform priors are applied to the matrices E and z.
[0307] Given the above representation, the likelihood of the data given the model can be expressed as:
[0308]
number
[0309] Here, the central dot in a matrix subscript refers to all members of a row or column, and vector multiplication refers to the dot product. 1 {} is the characteristic function, which takes the value 1 if the expression in the subscript is true, and 0 otherwise.
[0310] Combining the likelihood with the prior probability provides the necessary elements for performing Bayesian inference on unknown values. Numerous methods exist for performing Bayesian inference, including analytically tractable exact methods for posterior probability distributions, as well as a range of Monte Carlo and related methods for approximating posterior distributions. In this case, the model is implemented in the Stan modeling language (see Code Table X1), which can facilitate inference using Hamiltonian Monte Carlo and variational inference using mean-field and full-rank approximations. The variational inference approximation method used relies on stochastic gradient descent to maximize the variational lower bound (ELBO) (Kucukelbir et al., 2015 https: / / arxiv.org / abs / 1506.03431), which requires the probabilistic model to be continuous and differentiable. To meet this requirement, z is implemented as a continuous parameter with support [0,1], and a beta(0.1,0.1) distribution is used as a sparsifying prior to center the posterior density of z around 0 and 1. This approach, which uses continuous relaxation of discrete random variables, is called "concrete distributions" and is described at https: / / arxiv.org / abs / 1611.00712. Fitting the model to a collection of approximately 100 simulated sequence alignments of at least 100 bases in length using variational inference takes only a few minutes of CPU time on a laptop to estimate posterior probabilities for the unknown parameters, giving the posterior distribution of the model parameters shown in Figure 5.
[0311] Although variational inference is faster than many Monte Carlo methods, it is not fast enough to analyze the millions of sequence reads generated in a typical sequencing run. Therefore, we developed a faster method to calculate the probability that two reads (r0 and r1) do or do not originate from the same at least one mutant target template nucleic acid molecule. Given the mutagenesis process and sequencing errors, these probabilities can be expressed as follows:
[0312]
number
[0313] Here, the values of M and E are fixed to the maximum posterior probability or a similar value with a high posterior probability determined by Bayesian (or maximum likelihood) inference using a small subset of the entire data set. The values of N and Q are taken to correspond to the alignment of r0 and r1 to the reference sequence. The log-odds score for two reads originating from a common template can then be simply calculated as follows:
[0314]
number
[0315] Mutant sequence reads are considered to originate from the same at least one target template nucleic acid molecule if their pairwise score is higher than some predetermined cutoff, which in this case is set to 1,000. Tests on simulated data show that this log-odds score can distinguish with high precision and recall whether two mutant reads originate from at least one common target template nucleic acid molecule (Figure 6).
[0316] Example 7 - Using two identical primer binding sites and a single primer sequence for selective amplification of longer templates As described above, tagmentation can be used to fragment DNA molecules while simultaneously introducing primer binding sites (adapters) at the ends of the fragments. The Nextera tagmentation system (Illumina) utilizes a transposase enzyme with one of two unique adapters (referred to herein as X and Y). This generates a random mixture of amplification products, some with identical end sequences (XX, YY) and some with unique ends (XY). The standard Nextera protocol utilizes two distinct primer sequences to selectively amplify "XY" products containing different adapters at each end (as required for sequencing with Illumina technology). However, it is also possible to amplify "XX" or "YY" fragments with identical end adapters using a single primer sequence.
[0317] To generate long mutant templates containing identical terminal adapters, 50 ng of high molecular weight genomic DNA (E. coli strain MG1655) was first subjected to tagmentation and then purified with SPRIselect beads as described in Example 1. This was followed by five cycles of "mutagenesis PCR" using standard dNTP and dPTP combinations, performed as detailed in Example 1 except using a single primer sequence (Table 5).
[0318] The PCR reaction was purified with SPRIselect beads to remove excess dPTP and primers, and then subjected to an additional 10 cycles of "recovery PCR" in the absence of dPTP to replace dPTP in the template with the transition mutation. Recovery PCR was performed using a single primer that anneals to the fragment ends introduced during the dPTP incorporation cycles, thereby allowing selective amplification of the mutant template generated in the previous PCR step.
[0319] [Table 5]
[0320] As a control, mutant templates with different adapters at each end were generated using the same protocol as above, except that two distinct primer sequences were used in both the mutagenesis PCR (shown in Table 2) and the restoration PCR (Table 3). The final PCR products were purified with SPRIselect beads and further analyzed on a high-sensitivity DNA chip using the 2100 Bioanalyzer System (Agilent). As shown in Figure 10, templates generated with identical-end adapters were significantly longer, on average, than control samples containing dual adapters. Control templates could be detected down to a minimum size of approximately 800 bp, whereas no templates shorter than 2000 bp were observed for the single-adapter samples.
[0321] Size profiles were compared using a mutant template with identical-end adapters (blue) and a control template with dual adapters on an Agilent 2100 Bioanalyzer (high-sensitivity DNA kit). The use of identical-end adapters inhibits amplification of templates less than 2 kbp. The data are shown in Figure 10.
[0322] Example 8 – Sample dilution and end sequencing to quantify DNA templates The initial sample of long mutant template for analysis was diluted to a defined number of unique template molecules in preparation for downstream processing, sequencing, and analysis to ensure sufficient sequence data was generated per template for effective template construction.
[0323] First, long mutant templates were prepared from human genomic DNA (genomic NA12878) using the approach outlined in Example 7. Five mutagenic PCR cycles and six recovery cycles were performed, followed by gel extraction, to select templates spanning the 8-10 kb size range. The primers shown in Table 5 were used to obtain templates flanked by identical adapter sequences.
[0324] The size-selected template samples were then serially diluted in 10-fold increments, and DNA sequencing was used to determine the number of unique templates present in each dilution. This involved first amplifying the diluted samples to obtain multiple copies of each unique template. PCR was performed using a single primer (5'-CAAGCAGAAGACGGCATACGA-3') that anneals to the ends of the fragments introduced during the previous recovery PCR step, thereby selectively amplifying templates that had completed the process of dPTP incorporation and substitution, generating transition mutations. A total of 16–30 PCR cycles (depending on the sample dilution factor) were required to generate sufficient material for downstream processing.
[0325] Each PCR product was then fragmented using a standard tagging reaction (see Example 1), and fragments derived from the template ends (containing the sample tag and unique molecular tag) were selectively amplified in preparation for Illumina sequencing. This was achieved using a pair of primers: one that specifically anneals to the original template end (5'-CAAGCAGAAGACGGCATACGA-3') and one that anneals to the adapter introduced during tagging (i5 custom index primer; Table 2). After sequencing the samples on an Illumina MiSeq instrument, unique templates were identified based on sequence information corresponding to both ends of the original template molecule. To achieve this, a clustering algorithm (e.g., vsearch) was used to group together reads with identical sequences that were likely derived from the same original unique template. Other types of sequence information, such as unique molecular tags, can also be used for this purpose. As shown in Figure 11, a clear linear relationship was observed between the sample dilution factor and the number of observed unique templates. This information can be used to determine the exact dilution factor required to control the number of mutant target template nucleic acid molecules in the second sample to the desired number of unique templates in preparation for subsequent sequencing and template construction.
[0326] Example 9 – Dilution and end sequencing to normalize pooled template samples Using the sample dilution and end-sequencing approach described above, multiple template libraries were quantified in pre-pooled samples, and this information was then used to normalize the number of templates between individual samples in the pooled samples.
[0327] First, genomic DNA samples from 96 different bacterial strains were subjected to tagmentation and five cycles of mutagenesis PCR as outlined in Example 5, using a single primer with a unique sample tag for each reaction (design single_mut; Table 5). Equal volumes of each sample-tagged mutagenesis product were then pooled, and the pooled sample was washed with SPRIselect beads to remove excess dPTP and primers. This was followed by six cycles of recovery PCR using the single_rec primer (Table 5) and gel extraction to select templates in the 8-10 kb size range. The pooled template samples were then diluted 1:1000, and end sequencing was performed to determine the number of unique templates present in each bacterial strain within the diluted pool. This was achieved using the approach outlined in Example 7.
[0328] The number of templates was found to vary greatly among strains within the diluted pools, ranging from no detectable templates in some strains to over 1,000 unique templates in others. Sixty-six strains with non-zero template numbers were selected for normalization. To obtain a consistent number of unique templates per unit of genome content (e.g., per Mb) for each strain, normalization pools were prepared by combining various volumes of sample-tagged mutagenesis PCR products based on the observed template numbers and the known genome size of each strain. The normalization pools were then processed for end sequencing as described above to determine the number of unique templates per strain. As expected, after normalization, the number of templates varied much less among strains (Figure 12).
[0329] Example 10 – Use of assembly algorithms to construct bacterial genome sequences Bacterial strains and DNA preparation DNA from 62 bacterial strains was obtained from the BEI resource. These strains were isolates sequenced as part of the Human Microbiome Project. They exhibit a range of GC content (25%–69%) and further details are shown in Table 6. [Table 6] JPEG0007795566000010.jpg247164JPEG0007795566000011.jpg87165
[0330] Three additional strains with well-characterized genomes encompassing a wide range of GC content were included as controls (Escherichia coli K12 MG1655, Staphylococcus aureus ATCC 25923, and Haloferax volcanii DS2). DNA was prepared from these strains using the Qiagen DNeasy UltraClean Microbial Kit according to the manufacturer's instructions (with the following modifications). Overnight cultures (20 mL per strain) were centrifuged at 3200 g for 5 minutes to obtain cell pellets, and each pellet was washed with 5 mL of sterile 0.9% sodium chloride solution. Each pellet was resuspended in 300 μl of PowerBead solution before continuing with the manufacturer's protocol. DNA was eluted in 50 μL of elution buffer preheated to 42°C for E. coli and S. aureus, and 35 μL of elution buffer for H. volcanii.
[0331] DNA concentrations in all samples were measured using the Quant-iT PicoGreen dsDNA kit (Thermo Scientific). For a subset of species, DNA purity and molecular weight were also assessed by Nanodrop (Thermo Scientific) spectrophotometry and agarose gel electrophoresis.
[0332] Morphoseq library preparation Tagging to generate longer fragments DNA from each bacterial genome was arrayed in a 96-well plate, and concentrations were normalized to 10 ng / μl. E. coli MG1655 DNA was included in two separate wells to provide an internal control for sample processing and downstream data analysis. Tagging was performed using Nextera DNA Tagment Enzyme (TDE1; Illumina) diluted 1:50 in storage buffer (5 mM Tris-HCl [pH 8.0], 0.5 mM EDTA, 50% (v / v) glycerol). For each sample, 16 μL of tagmentation reactions were prepared containing 50 ng of DNA and 4 μl of diluted TDE1 in 1× tagmentation buffer (10 mM Tris-HCl [pH 7.6], 10 mM MgCl, 10% (v / v) dimethylformamide). Each reaction was incubated at 55°C for 5 minutes and then cooled to 10°C. SDS was added to a final concentration of 0.04%, and the reactions were incubated at 25°C for an additional 15 minutes. The reaction was subjected to a left-side clean up using SPRIselect magnetic beads (Beckman Coulter) with 0.6 volumes of beads and eluted in 20 μl of molecular grade water (performed according to the manufacturer's instructions).
[0333] Mutagenesis of long DNA fragments PCR to incorporate the mutagenic nucleotide analog dPTP was performed as follows: 5 μl of each of the above washed tagging reactions was used as template in a 25 μl PCR reaction containing 0.625 U PrimeStar GXL polymerase, 1× Primestar GXL buffer, and 0.2 mM dNTPs (all from Takara) with 0.5 mM dPTP (TriLink Biotechnologies) and 0.4 mM Morphoseq index primer (see Table 7; unique index for each sample). A single primer was used during mutagenic PCR to amplify templates containing the same Nextera tagging adapter sequence at both ends. Reactions were subjected to the following cycling conditions: 68°C for 3 minutes, followed by five cycles of 98°C for 10 seconds, 55°C for 15 seconds, and 68°C for 10 minutes.
[0334] At this point, equal volumes (4 μL) of each reaction were combined into a single pool, and the pool was subjected to an additional SPRIselect left-hand bead wash using 0.6x the volume of beads. The purified pool was eluted in 45 μL of molecular-grade water and quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific).
[0335] The pooled sample of dPTP-containing template was then further amplified in the absence of dPTP, thereby substituting the natural dNTP for the nucleotide analog and generating transition mutations due to the ambivalent base-pairing properties of dPTP. This "recovery" PCR contained 1.25 U PrimeStar GXL polymerase, 1x Primestar GXL buffer, and 0.2 mM dNTPs (Takara) along with 0.4 µM recovery primer (see Table 7) and 10 ng of pooled template sample in a total volume of 50 µl. The reaction was subjected to six cycles of 98°C for 10 seconds, 55°C for 15 seconds, and 68°C for 10 minutes.
[0336] Selection of long template size Using a DNA gel electrophoresis approach, recovery PCR products were size-selected to remove unwanted short fragments. 25 μl of the recovery PCR reaction, along with DNA size standards, was loaded onto a 0.9% agarose gel and run overnight (900 min) at 18 V in 1x TBE buffer. Gel slices corresponding to the 8-10 kb size range were excised, and DNA was extracted using the Wizard SV Gel and PCR Cleanup Kit (Promega) according to the manufacturer's instructions. Size-selected DNA was quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific), and the size range was confirmed using a Bioanalyzer High Sensitivity DNA Chip (Agilent).
[0337] Template normalization and quantification To assess template abundance among individual sample-tagged samples in the pooled, size-selected products, the following approach was used. First, size-selected DNA was diluted to 0.1 pg / μL, and 2 μl of the dilution (0.2 pg) was used as input for enrichment PCR to obtain multiple copies of each unique template. Preliminary experiments indicated that this level of dilution limited the diversity of unique templates sufficient to allow accurate template quantification from the sequence output of a single Illumina MiSeq run. The 50 μl enrichment PCR also contained 1.25 U PrimeStar GXL polymerase, 1× Primestar GXL buffer, and 0.2 mM dNTPs (Takara) along with 0.4 μM enrichment primers (see Table 7). The enrichment primers were designed to anneal to the fragment end adapters introduced during the previous recovery PCR step, thereby selectively amplifying templates that had completed the process of dPTP incorporation and replacement. Reactions were subjected to 22 cycles of 98°C for 10 seconds, 55°C for 15 seconds, and 68°C for 10 minutes, then purified by SPRIselect left-hand bead wash using 0.6x beads and eluted in 20 μl of molecular-grade water. Samples were then quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific), and the size range was confirmed using a Bioanalyzer High Sensitivity DNA Chip (Agilent).
[0338] The full-length enrichment products were then fragmented by a second tagging reaction, and fragments derived from the original template ends (including the sample barcodes) were amplified for Illumina sequencing. Tagmentation was performed as described above for long template generation, except that 2 ng of starting DNA was used instead of 50 ng. After SDS treatment, end library PCR reactions were prepared by adding KAPA HiFi HotStart ReadyMix (Kapa Biosystems) to a final concentration of 1×, along with 0.23 μM enrichment primer (which anneals to the Illumina p7 flow cell adapter located at the extreme end of the full-length template) and 0.23 μM custom i5 index primer (which anneals to the internal adapter introduced during the second round of tagging; see Table 7). The reactions were cycled as follows: 72°C for 3 minutes, 98°C for 30 seconds; 12 cycles of 98°C for 15 seconds, 55°C for 30 seconds, and 72°C for 30 seconds; followed by a final extension at 72°C for 5 minutes. The final library was then purified and quantified as described above for the full-length enriched products.
[0339] Illumina sequencing was performed on a MiSeq using V3 chemistry, yielding 2 × 75 nt paired-end reads. The number of unique templates was determined for each individual bacterial genome sample within the dilution pool by first demultiplexing the end-read data based on the index 1 (i7) read sequence and then mapping the read 2 sequence (corresponding to the extreme end of the original genome insert) to the publicly available reference genome for each strain. The number of unique templates was calculated by counting the number of unique mapping start sites (corresponding to the beginning or end of the template), keeping in mind that two sites are expected per template.
[0340] The observed template numbers varied among individual genomes within the diluted pools, ranging from no detectable templates in some samples to over 1,000 unique templates in others. For simplicity, 66 samples with non-zero template numbers were selected for further processing, sequencing, and assembly. A normalized pool was prepared by combining various volumes of the original barcoded mutagenesis PCR products based on the observed template numbers and known genome sizes for each of these samples to obtain a consistent number of unique templates per unit of genome content (e.g., per Mb) for each strain. To confirm successful normalization, the normalized pool was further processed for template quantification by repeating all subsequent steps of library preparation and sequencing described above (recovery PCR, size selection, template dilution and enrichment, end library preparation, Illumina sequencing, and analysis). As expected, after normalization, template numbers were much less variable between strains (Figure 11).
[0341] Template bottlenecking, enrichment and short-read library processing Based on template quantification data from the normalized sample pool and the known sizes of long fragments, a total of 1.5 million unique template targets were selected for processing in Morphoseq sequencing and assembly. This would ensure at least 20x (maximum 90x) theoretical long template coverage per individual genome. To this end, the final long template sample was prepared by diluting the size-selection recovery PCR product from the previous step to 750,000 templates per μl and using 2 μl of the dilution as input for enrichment PCR to obtain multiple copies of each unique template. Enrichment PCR was performed as described above, except that 16 amplification cycles were performed instead of 22.
[0342] To process the final long template samples for short-read (Illumina) sequencing, we first prepared, purified, and quantified barcoded end libraries using the method outlined in the previous section. A second library containing randomly generated internal fragments from the long template was also prepared using the Nextera DNA Flex Library Prep Kit (Illumina) according to the manufacturer's protocol (with some modifications). Specifically, we diluted the BLT (Bead-Linked Transposomes) reagent 1:50 in molecular-grade water, and used 10 μl of this dilution in a tagging reaction with 10 ng of long template DNA. We performed 12 cycles of library amplification using custom i5 and i7 index primers (Table 7) rather than standard Illumina adapters.
[0343] Preparation of the non-mutated reference library Reference libraries were generated for all 66 genomes included in the final Morphoseq pool. Using 10 ng of genomic DNA as input, library preparation was performed according to the procedure described above for the internal Morphoseq library (with additional modifications for the Nextera DNA Flex method). Specifically, Illumina TB1 buffer was replaced with custom tagging buffer (see above), KAPA HiFi HotStart ReadyMix (1x final concentration; Kapa Biosystems) was used instead of the kit polymerase, and Illumina sample purification beads (SPB) were replaced with SPRIselect magnetic beads (Beckman Coulter). Thermal cycling conditions for reference library amplification were as follows: 72°C for 3 minutes, 98°C for 30 seconds; 12 cycles of 98°C for 15 seconds, 55°C for 30 seconds, and 72°C for 30 seconds; followed by a final extension at 72°C for 5 minutes.
[0344] To normalize the reference library, equal volumes of each sample were first pooled together, and the pooled library was sequenced using the MiSeq Reagent Nano Kit (Illumina) to obtain 2 × 150-nt paired-end reads using the MiSeqV2 chemistry. The resulting sequence data was demultiplexed to determine the read counts for each individual genome. These counts were then used to prepare a normalized pool by combining various volumes of each original reference library to achieve equal coverage per genome.
[0345] Illumina sequencing The final sample for Illumina sequencing was prepared by combining the normalized reference pool, the morphoseq end library, and the morphoseq internal library at a molar ratio of 1:1:20, respectively. Sequencing was performed at the Ramaciotti Centre for Genomics at the University of New South Wales (Sydney, Australia) using a NovaSeq 6000 instrument and an S1 flow cell to generate 2 × 150 nt paired-end reads.
[0346] Bacterial genome construction An overview of the workflow for constructing a bacterial genome is shown in Figure 13.
[0347] Non-mutated reference assembly The genome of each bacterial strain was assembled from unmutated paired-end 150-base pair reads. Initial quality filtering to remove low-quality sequences and excise library adapters was performed in bbduk v36.99. Reads were demultiplexed using a custom Python script and assembled using MEGAHIT v1.1.3 with custom parameters (prune-level = 3, low-local-ratio = 0.1, and max-tip-len = 280), which were chosen to reduce the complexity of the resulting genome graph and facilitate better mapping of mutated sequences in the next step (see below). The resulting graphical fragment assembly (gfa file) was used as input to VG(index) v1.14.0 to generate an index suitable for mapping. The resulting graph is referred to as the "indexed unmutated reference assembly graph" or simply the "indexed graph."
[0348] Generation of long synthetic leads (morpholeads) Mutation reads from each end library (end reads) and the pooled internal library (internal reads) were mapped to their corresponding indexed VG bacterial genome assemblies using VG(map) v1.14.0 with default parameters to generate a pair of graphical alignment map (GAM) files for each sample. Data from each sample's GAM pair was combined with information from the corresponding non-mutated reference assembly, processed using custom tools, and stored in an HDF5-formatted database that facilitates parallel processing for many of the remaining steps in reconstructing the original sequence. The morphoread generation process consists of three main stages: "end wall identification," "seeding," and "extension."
[0349] The nature of the processes used to fragment target DNA into long fragments and generate the final short-read library creates a situation in which sequences at the extreme ends of any original template are found only in the second read of the paired Illumina library. When these reads are mapped to the reference genome, they will appear to suddenly pile up at positions corresponding to the ends of the original long DNA template. These positions are called "end walls" and are identified by finding groups of end reads and internal reads that map to the same location in the reference assembly. Any site with at least five end reads mapping in this pattern is designated an end wall. Internal reads are used to increase the mapping count at sites with two to four mapped end reads; if the total increase is at least five, these sites are also designated as end walls.
[0350] The end wall determines the position within the reference assembly from which the algorithm begins constructing long synthetic reads. However, whenever two or more templates have the same start or end position, it is possible to have a single end wall corresponding to multiple original DNA templates. Each DNA template has a unique mutation pattern, and therefore, reads derived from a given template contain a subset of that pattern, which appear as transition mismatches in VG mapping. The "seeding" stage analyzes these mutation patterns in the terminal and internal reads at each end, clustering reads with similar patterns together and generating a single short (400-600 base pairs) morphoread instance per cluster. Each morphoread instance contains a directed acyclic graph-based representation of the mapped mutation reads it contains, called a "consensus graph." The structure of the consensus graph roughly corresponds to a subgraph of the indexed graph, and the position of a read in the consensus graph corresponds to the read's mapping position relative to the indexed graph. The key difference between a consensus graph and its corresponding subgraph of the indexed graph is that edges between nodes in the consensus graph represent paths of mapped reads through the indexed graph; whenever such paths follow a loop in the indexed graph, the nodes in the loop are duplicated, effectively rolling out the loop in the indexed graph and removing the cycle. Thus, individual nodes in the indexed graph correspond to potentially multiple nodes in the consensus graph, and edges in the consensus graph often, but not always, correspond to edges in the indexed graph. The consensus graph stores information about the indexed assembly and mapped mutation reads, and thus it can be used to generate a "consensus sequence" (i.e., containing no mutations) corresponding to a path through the indexed graph, and a "mutation set" containing a consensus of the mutation patterns found in all the internal and terminal reads involved.
[0351] During the "extension" stage, the algorithm walks along the consensus graph starting from the end wall and iteratively adds terminal and internal reads to the morpho read if they match the consensus sequence (more than 90% identity, 100 bp or more overlap), their mutation pattern shares at least three mutations with the mutation set, and contains no more than five mutations that differ from the mutation set. A large number of different mutations is necessary to reduce the impact of errors in individual reads that mimic mutations, and also because reads tested for inclusion in the morpho read may map to nodes that extend beyond the end of the current consensus graph and may contain mutations not yet included in the mutation set of the morpho read. Each time a new read is included in the morpho read, a new node may be added to the consensus graph, potentially making the consensus fragment longer. The algorithm continues walking along the extension consensus graph until a terminal read is incorporated into the morphoread (indicating that the distal end of the original long DNA template has been reached or no reads that can be used to continue extension can be found). The final consensus fragment for each morphoread is written to a FASTA file, and all morphoreads shorter than 500 bp are discarded. The algorithm also generates a BAM file containing the positions of the included terminal and internal reads relative to the consensus sequence, as well as some summary statistics for each morphoread.
[0352] Hybrid genome assembly High-quality morphologic reads were combined with non-mutated reference reads in a hybrid genome assembly using Unicycler v0.4.6 with default parameters.
[0353] result The Morphoseq method consistently produced assemblies with significantly fewer and larger scaffolds (Kruskal Wallis, p < 0.001) than short-read-only assemblies (Figure 14). For Morphoseq and short-read-only assemblies, the median maximum scaffold length as a percentage of genome size was 55.84% vs. 10.15%, and the median number of scaffolds was 17 vs. 192, respectively. Typical assembly metrics for bacterial genomes can be found in Figure 15. [Table 7] JPEG0007795566000013.jpg252166JPEG0007795566000014.jpg250166JPEG00077955660 00015.jpg250166JPEG0007795566000016.jpg250166JPEG0007795566000017.jpg251166
[0354] TIFF0007795566000018.tif95168 The present disclosure includes the following embodiments. [1] (a) providing a pair of samples, each sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule in a first of the pair of samples to obtain a non-mutated sequence read; (c) introducing a mutation into at least one target template nucleic acid molecule in a second of the pair of samples to obtain at least one mutated target template nucleic acid molecule; (d) sequencing a region of the at least one mutant target template nucleic acid molecule to obtain a mutant sequence read; (e) analyzing the mutant sequence reads and constructing a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads using information obtained from the analysis of the mutant sequence reads. A method for sequencing at least one target template nucleic acid molecule, comprising: [2] (a)(i) non-mutated sequence reads, and (ii) Mutation sequence reads obtaining data including (b) analyzing the mutant sequence reads and constructing a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads using information obtained from the analysis of the mutant sequence reads. A method for generating the sequence of at least one target template nucleic acid molecule, comprising: [3] The method of embodiment 1 or 2, wherein analyzing mutant sequence reads and constructing the sequence of at least a portion of at least one target template nucleic acid molecule from non-mutated sequence reads using information obtained from the analysis of the mutant sequence reads comprises preparing an assembly graph. [4] The method of embodiment 3, wherein the assembly graph comprises nodes calculated from non-mutated sequence reads, and each valid path through the assembly graph comprising the nodes represents at least a portion of the sequence of at least one target template nucleic acid molecule. [5] The method of embodiment 4, wherein the node is a unitig. [6] The method of any one of embodiments 3 to 5, wherein constructing a sequence of at least a portion of at least one target template nucleic acid molecule from non-mutated sequence reads using information obtained from analysis of mutant sequence reads comprises identifying nodes that form part of valid paths through an assembly graph using information obtained from analysis of mutant sequence reads. [7] The method of any one of embodiments 4 to 6, wherein a sequence for at least a portion of at least one target template nucleic acid molecule from a node that forms part of a valid path through an assembly graph is constructed. [8] The method of any one of embodiments 1 or 3-7, wherein the pair of samples are taken from the same original sample or derived from the same organism. [9] The method of any one of embodiments 2-7, wherein the non-mutated sequence read comprises the sequence of a region of at least one target template nucleic acid molecule in a first of a pair of samples, and the mutant sequence read comprises the sequence of a region of at least one mutant target template nucleic acid molecule in a second of the pair of samples, and the pair of samples are taken from the same original sample or derived from the same organism.
[10] The method of any one of embodiments 1 to 9, wherein the method does not include constructing a sequence from mutant sequence reads.
[11] The method of any one of embodiments 1 to 10, wherein the method does not include constructing the sequence of at least one mutant target template nucleic acid molecule or the sequence of a majority of at least one mutant target template nucleic acid molecule.
[12] The method of any one of embodiments 1 to 11, wherein analyzing the mutant sequence reads comprises identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule.
[13] Using information obtained from the analysis of mutant sequence reads, the researchers identify nodes that form part of valid paths through an assembly graph. (i) Computing nodes from non-mutated sequence reads; (ii) mapping mutant sequence reads against the assembly graph; (iii) identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule; and (iv) identifying nodes connected by mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule; 7. The method of embodiment 6, wherein nodes connected by mutant sequence reads are likely to originate from the same at least one mutant target template nucleic acid molecule and form part of a valid path through the assembly graph.
[14] The method of embodiment 12 or 13, wherein mutant sequence reads that are likely to originate from the same mutant target template nucleic acid molecule are assigned to groups.
[15] The method of any one of embodiments 12 to 14, wherein mutant sequence reads are likely to originate from the same mutant target template nucleic acid molecule if they share a common mutation pattern.
[16] The method of any one of embodiments 12 to 15, wherein analyzing the mutant sequence reads includes identifying mutant sequence reads that share a common mutation pattern.
[17] The method of embodiment 15 or 16, wherein the mutant sequence reads that share a common mutation pattern include at least 1, at least 2, at least 3, at least 4, at least 5, or at least k common signature k-mers and / or common signature mutations.
[18] The method of embodiment 17, wherein the signature k-mer is a k-mer that does not occur in the non-mutated sequence reads but occurs at least 2 times, at least 3 times, at least 4 times, at least 5 times, or at least 10 times in the mutant sequence reads.
[19] The method of embodiment 17, wherein the signature mutation is a nucleotide that appears at least 2 times, at least 3 times, at least 4 times, at least 5 times, or at least 10 times in the mutant sequence reads but does not appear at the corresponding position in the non-mutated sequence reads.
[20] The method of embodiment 19, wherein the signature mutation is a co-occurring mutation.
[21] The method of embodiment 19 or 20, wherein a signature mutation is ignored if at least one, at least two, at least three, or at least five nucleotides at corresponding positions in mutant sequence reads that share the signature mutation are different from each other.
[22] The method of any one of embodiments 19 to 21, wherein if the signature mutation is an unexpected mutation, the signature mutation is ignored.
[23] The method described in any one of embodiments 19 to 22, wherein the step of identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule comprises identifying mutant sequence reads that correspond to specific regions of the at least one target template nucleic acid molecule.
[24] The method of any one of embodiments 12 to 16 or 23, wherein the mutation sequence reads are likely to be derived from the same mutation target template nucleic acid molecule if the odds ratio between the probability that the mutation sequence reads are derived from the same mutation target template nucleic acid molecule and the probability that the mutation sequence reads are not derived from the same mutation target template nucleic acid molecule exceeds a threshold.
[25] The method of embodiment 24, wherein if the odds ratio of the first mutation sequence read and the second mutation sequence read is higher than the odds ratio of the first mutation sequence read and other mutation sequence reads that map to the same region of the assembly graph, the mutation sequence reads are likely to be derived from the same mutation target template nucleic acid molecule.
[26] The following factors: (i) the stringency required, and / or (ii) the error rate of sequencing a region of at least one mutation target template nucleic acid molecule to obtain mutant sequence reads; and / or (iii) the mutation rate used in the step of introducing mutations into at least one target template nucleic acid molecule; and / or (iv) the size of at least one target template nucleic acid molecule, and / or (v) time constraints, and / or (vi) Resource constraints 26. The method of embodiment 24 or 25, wherein the threshold is determined based on one or more of:
[27] Identifying mutant sequence reads that are likely to originate from the same mutant target template nucleic acid molecule is determined by the following parameters: e. A matrix of nucleotides (N) at each position in the mutant sequence read and assembly graph; f. The probability (M) that a given nucleotide (i) is mutated to a lead nucleotide (j); g. The probability (E) that a given nucleotide (i) is misread by a lead nucleotide (j) given that the nucleotide is misread; and Probability that the nucleotide at position hY was misread (Q) 27. The method of any one of embodiments 12-16 or 23-26, comprising using a probability function based on:
[28] The method of embodiment 27, wherein the value of Q is obtained by performing a statistical analysis on the mutated and non-mutated sequence reads, or is obtained based on prior knowledge of the accuracy of the sequencing method.
[29] The method of embodiment 27 or 28, wherein the values of M and E are estimated based on a statistical analysis performed on a subset of mutated and non-mutated sequence reads, wherein the subset includes mutated and non-mutated sequence reads selected as mapping to the same region of the assembly graph.
[30] The method of embodiment 29, wherein the statistical analysis is performed using Bayesian inference, Monte Carlo methods such as Hamiltonian Monte Carlo, variational inference, or a maximum likelihood analog of Bayesian inference.
[31] The method of any one of embodiments 12 to 16 or 23 to 30, wherein identifying mutant sequence reads that are likely to derive from the same mutant target template nucleic acid molecule comprises using machine learning or neural networks.
[32] The method of any one of embodiments 12 to 31, wherein the method comprises a pre-clustering step.
[33] The method of embodiment 32, wherein the identification of mutant sequence reads that are likely to originate from the same mutant target template nucleic acid molecule is constrained by the results of the pre-clustering step.
[34] The method of embodiment 32 or 33, wherein the pre-clustering step comprises assigning mutant sequence reads to groups, wherein each member of the same group has a reasonable likelihood of originating from the same mutant target template nucleic acid molecule.
[35] The method of any of embodiments 32 to 34, wherein the pre-clustering step comprises Markov clustering or Louvain clustering.
[36] The method of any of embodiments 34 to 35, wherein members of the same group are mapped to a common position on the assembly graph and / or share a common mutation pattern.
[37] The method of embodiment 36, wherein the mutant sequence reads sharing a common mutation pattern are mutant sequence reads that contain at least 1, at least 2, at least 3, at least 4, at least 5, or at least k common signature k-mers and / or common signature mutations.
[38] The method of embodiment 37, wherein the signature k-mer is a k-mer that does not occur in the non-mutated sequence reads but occurs at least 2 times, at least 3 times, at least 4 times, at least 5 times, or at least 10 times in the mutant sequence reads.
[39] The method of embodiment 37, wherein the signature mutation is a nucleotide that appears at least 2 times, at least 3 times, at least 4 times, at least 5 times, or at least 10 times in the mutant sequence reads but does not appear at the corresponding position in the non-mutated sequence reads.
[40] The method of embodiment 39, wherein the signature mutation is a co-occurring mutation.
[41] The method of embodiment 39 or 40, wherein a signature mutation is ignored if at least one, at least two, at least three, or at least five nucleotides at corresponding positions in mutant sequence reads that share the signature mutation differ from each other.
[42] A method according to any one of embodiments 39 to 41, wherein the signature mutation is ignored if the signature mutation is an unexpected mutation.
[43] The method described in any one of embodiments 39 to 42, wherein the step of identifying mutant sequence reads that are likely to originate from the same at least one mutant target template nucleic acid molecule comprises identifying mutant sequence reads that correspond to specific regions of the at least one target template nucleic acid molecule.
[44] The method of any one of embodiments 1 to 43, wherein the method comprises sequencing the ends of at least one target template nucleic acid molecule using paired-end sequencing.
[45] The method of any one of embodiments 1 to 44, wherein the method comprises mapping the sequence of an end of at least one target template nucleic acid molecule to the assembly graph.
[46] The method of any one of embodiments 1 to 45, wherein at least one target template nucleic acid molecule comprises a barcode at each end.
[47] The method of embodiment 46, wherein the method comprises mapping the sequence of an end of at least one target template nucleic acid molecule to an assembly graph, and substantially all ends comprise a barcode.
[48] The method of any one of embodiments 6 to 47, wherein identifying nodes that form part of a valid path through the assembly graph includes ignoring putative paths that have mismatched ends.
[49] The method of any one of embodiments 6 to 48, wherein identifying nodes that form part of a valid path through the assembly graph includes ignoring estimated paths that are the result of template collisions.
[50] A method according to any one of embodiments 6 to 49, wherein identifying nodes that form part of a valid path through the assembly graph includes ignoring estimated paths that are longer or shorter than expected.
[51] A method according to any one of embodiments 6 to 50, wherein identifying nodes that form part of a valid path through the assembly graph includes ignoring estimated paths that have an atypical coverage depth.
[52] The method of any one of embodiments 1 to 51, wherein at least one mutation target template nucleic acid molecule contains 1% to 50%, 3% to 25%, 5% to 20%, or about 8% mutations.
[53] The method of any one of embodiments 1 to 52, wherein at least one mutant target template nucleic acid molecule contains mutations that are unevenly distributed.
[54] The method of any one of embodiments 1 to 53, wherein the mutant sequence reads and / or the non-mutated sequence reads contain unevenly distributed sequencing errors.
[55] The method of any one of embodiments 1 to 54, wherein the step of introducing mutations into at least one mutation target template nucleic acid molecule introduces mutations that are unevenly distributed.
[56] The method of any one of embodiments 1 to 55, wherein sequencing a region of at least one target template nucleic acid molecule and / or sequencing a region of at least one mutant target template nucleic acid molecule introduces non-uniformly distributed sequencing errors.
[57] The method of any one of embodiments 1 to 56, wherein at least one mutation target template nucleic acid molecule comprises a substantially random mutation pattern.
[58] The method of any one of embodiments 1 to 58, wherein multiple pairs of samples are prepared.
[59] The method of embodiment 58, wherein at least one target template nucleic acid molecule in different sample pairs is labeled with a different sample tag.
[60] The method of any one of embodiments 1 or 3-59, further comprising amplifying at least one target template nucleic acid molecule in a first of the pair of samples prior to sequencing a region of the at least one target template nucleic acid molecule.
[61] The method of any one of embodiments 1 or 3-60, further comprising amplifying at least one target template nucleic acid molecule in a second of the pair of samples prior to sequencing a region of the at least one target template nucleic acid molecule.
[62] The method of any one of embodiments 1 or 3-61, further comprising fragmenting at least one target template nucleic acid molecule in a first of the pair of samples prior to sequencing a region of the at least one target template nucleic acid molecule.
[63] The method of any one of embodiments 1 or 3 to 62, further comprising the step of fragmenting at least one target template nucleic acid molecule or at least one mutant target template nucleic acid molecule in a second of the pair of samples prior to the step of sequencing a region of the at least one mutant target template nucleic acid molecule.
[64] The method of any one of embodiments 1-64, wherein at least one target template nucleic acid molecule is greater than 2 kbp, greater than 4 kbp, greater than 5 kbp, greater than 7 kbp, greater than 8 kbp, less than 200 kbp, less than 100 kbp, less than 50 kbp, between 2 kbp and 200 kbp, or between 5 kbp and 100 kbp.
[65] The method of any one of embodiments 1 or 3 to 64, wherein the step of introducing mutations into at least one target template nucleic acid molecule in the second of the pair of samples is carried out by chemical mutagenesis or enzymatic mutagenesis.
[66] The method of embodiment 65, wherein the enzymatic mutagenesis is carried out using a DNA polymerase.
[67] The method of embodiment 66, wherein the DNA polymerase is a low-bias DNA polymerase.
[68] The method of embodiment 67, wherein a low-bias DNA polymerase introduces substitution mutations.
[69] The method of embodiment 67 or 68, wherein the low-bias DNA polymerase mutates adenine, thymine, guanine, and cytosine nucleotides in the at least one target template nucleic acid molecule at a rate ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1, respectively.
[70] The method of any one of embodiments 67-69, wherein the low-bias DNA polymerase mutates adenine, thymine, guanine, and cytosine nucleotides in the at least one target template nucleic acid molecule at a rate ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively.
[71] The method of any one of embodiments 67-70, wherein the low-bias DNA polymerase mutates 1% to 15%, 2% to 10%, or about 8% of the nucleotides in the at least one target template nucleic acid molecule.
[72] The method of any one of embodiments 67 to 71, wherein the low-bias DNA polymerase mutates 0% to 3%, or 0% to 2%, of the nucleotides in the at least one target template nucleic acid molecule per round of replication.
[73] The method of any one of embodiments 67-72, wherein the low-bias DNA polymerase incorporates a nucleotide analog into the at least one target template nucleic acid molecule.
[74] The method of any one of embodiments 67-73, wherein the low-bias DNA polymerase uses nucleotide analogs to mutate adenine, thymine, guanine, and / or cytosine in the at least one target template nucleic acid molecule.
[75] The method of any one of embodiments 67-74, wherein the low-bias DNA polymerase replaces guanine, cytosine, adenine, and / or thymine with nucleotide analogs.
[76] The method of any one of embodiments 67-75, wherein the low-bias DNA polymerase uses nucleotide analogs to introduce guanine or adenine nucleotides at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively.
[77] The method of any one of embodiments 67-76, wherein the low-bias DNA polymerase uses nucleotide analogs to introduce guanine or adenine nucleotides at a rate ratio of 0.7-1.3:0.7-1.3, respectively.
[78] The method of any of embodiments 67-77, wherein the method comprises amplifying the at least one target template nucleic acid molecule in a second one of the pair of samples using a low-bias DNA polymerase, wherein the amplifying the at least one target template nucleic acid molecule using the low-bias DNA polymerase is performed in the presence of nucleotide analogs, and wherein the amplifying the at least one target template nucleic acid molecule provides at least one target template nucleic acid molecule in the second one of the pair of samples that comprises the nucleotide analogs.
[79] The method of any one of embodiments 67 to 78, wherein the nucleotide analog is dPTP.
[80] The method of embodiment 79, wherein the low-bias DNA polymerase introduces guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations.
[81] The method of embodiment 80, wherein the low-bias DNA polymerase introduces guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at a ratio of 0.5 to 1.5:0.5 to 1.5:0.5 to 1.5:0.5 to 1.5, 0.6 to 1.4:0.6 to 1.4:0.6 to 1.4:0.6 to 1.4, 0.7 to 1.3:0.7 to 1.3:0.7 to 1.3:0.7 to 1.3, 0.8 to 1.2:0.8 to 1.2:0.8 to 1.2:0.8 to 1.2, or about 1:1:1:1, respectively.
[82] The method of embodiment 80 or 81, wherein the low-bias DNA polymerase introduces guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at a rate ratio of 0.7 to 1.3:0.7 to 1.3:0.7 to 1.3:0.7 to 1.3, respectively.
[83] The method of any one of embodiments 67 to 82, wherein the low-bias DNA polymerase is a high-fidelity DNA polymerase.
[84] The method of embodiment 83, wherein, in the absence of nucleotide analogs, the high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, 0% to 0.0015%, or 0% to 0.001% mutations per round of replication.
[85] The method of embodiment 83 or 84, wherein the method comprises the further step of amplifying the at least one target template nucleic acid molecule comprising a nucleotide analog in the absence of the nucleotide analog.
[86] The method of embodiment 85, wherein the step of amplifying the at least one target template nucleic acid molecule comprising nucleotide analogs in the absence of the nucleotide analogs is performed using a low-bias DNA polymerase.
[87] The method of any one of embodiments 67-86, wherein the method provides at least one mutated target template nucleic acid molecule, and the method further comprises the additional step of amplifying the at least one mutated target template nucleic acid molecule using a low-bias DNA polymerase.
[88] The method of any one of embodiments 67 to 87, wherein the low-bias DNA polymerase has a low template amplification bias.
[89] The method of any one of embodiments 67 to 88, wherein the low-bias DNA polymerase comprises a proofreading domain and / or a processivity-enhancing domain.
[90] The low-bias DNA polymerase is selected from the following a. to h., a. The sequence of SEQ ID NO: 2; b. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:2; c. The sequence of SEQ ID NO: 4; d. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:4; e. the sequence of SEQ ID NO: 6; f. a sequence at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:6; g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 7 90. The method of any one of embodiments 67 to 89, comprising a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of
[91] The low-bias DNA polymerase is selected from the following a. to h., a. The sequence of SEQ ID NO: 2; b. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:2; c. The sequence of SEQ ID NO: 4; d. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:4; e. the sequence of SEQ ID NO: 6; f. a sequence at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:6; g. The sequence of SEQ ID NO: 7, or h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 7 91. The method of any one of embodiments 67 to 90, comprising:
[92] The method of embodiment 91, wherein the low-bias DNA polymerase comprises a sequence that is at least 98% identical to SEQ ID NO:2.
[93] The method of embodiment 91, wherein the low-bias DNA polymerase comprises a sequence at least 98% identical to SEQ ID NO:4.
[94] The method of embodiment 91, wherein the low-bias DNA polymerase comprises a sequence at least 98% identical to SEQ ID NO:6.
[95] The method of embodiment 91, wherein the low-bias DNA polymerase comprises a sequence at least 98% identical to SEQ ID NO:7.
[96] The method of any one of embodiments 67 to 95, wherein the low-bias DNA polymerase is a thermococcal polymerase from an archaeal order Thermococcales, or a derivative thereof.
[97] The method of embodiment 96, wherein the low-bias DNA polymerase is a polymerase from a Thermococcales archaea.
[98] The method of embodiment 96 or 97, wherein the polymerase of a Thermococcales archaeon is derived from a strain of a Thermococcales archaeon selected from the group consisting of T. kodakarensis, T. siculi, T. celer, and T. sp. KS-1.
[99] A computer program adapted to carry out the method according to any one of embodiments 1 to 98.
[0100] A computer-readable medium containing a computer program described in embodiment 99.
[0101] A computer-implemented method, including the method of any one of embodiments 1 to 98.
[0102] A method described in any one of embodiments 1 or 3 to 98, wherein the step of preparing a pair of samples, each containing at least one target template nucleic acid molecule, includes controlling the number of target template nucleic acid molecules in the first of the pair of samples.
[0103] A method described in any one of embodiments 1, 3 to 98 or 102, wherein the step of preparing a pair of samples, each containing at least one target template nucleic acid molecule, includes controlling the number of target template nucleic acid molecules in the second of the pair of samples.
[0104] A method according to any one of embodiments 1, 3-98 or 102-103, wherein the first of the pair of samples is prepared by pooling two or more subsamples.
[0105] A method according to any one of embodiments 1, 3-98 or 102-104, wherein the second of the pair of samples is prepared by pooling two or more subsamples.
[0106] The method described in embodiment 104 or 105, further comprising a step of normalizing the number of target template nucleic acid molecules in each of the subsamples pooled to prepare the first of the sample pair and / or the second of the sample pair. (a) providing at least one sample containing at least one target template nucleic acid molecule; (b) sequencing a region of at least one target template nucleic acid molecule; and (c) constructing the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule. 1. A method for sequencing at least one target template nucleic acid molecule, comprising: (i) the step of providing at least one sample containing at least one target template nucleic acid molecule includes controlling the number of target template nucleic acid molecules in the at least one sample; and / or (ii) preparing at least one sample by pooling two or more subsamples, wherein the number of target template nucleic acid molecules in each of the subsamples is normalized.
[0108] A method described in any of embodiments 102 to 107, wherein controlling the number of target template nucleic acid molecules includes measuring the number of target template nucleic acid molecules in the first of a pair of samples, the second of a pair of samples, or at least one sample.
[0109] The method described in embodiment 108, wherein measuring the number of target template nucleic acid molecules includes preparing a dilution series of the first of the pair of samples, the second of the pair of samples, or at least one sample to obtain a dilution series including diluted samples.
[0110] A method described in any of embodiments 108 to 190, wherein measuring the number of target template nucleic acid molecules includes sequencing the target template nucleic acid molecules in one or more of the first of the sample pair, the second of the sample pair, at least one sample or a diluted sample.
[0111] The method described in embodiment 110, wherein measuring the number of target template nucleic acid molecules includes amplifying and then sequencing the target template nucleic acid molecules in one or more of the first of the sample pair, the second of the sample pair, at least one sample or a diluted sample.
[0112] The method described in embodiment 110 or 111, wherein measuring the number of target template nucleic acid molecules comprises amplifying and fragmenting the target template nucleic acid molecules in one or more of the first of the sample pair, the second of the sample pair, at least one sample or a diluted sample, and then sequencing the target template nucleic acid molecules.
[0113] A method described in any of embodiments 110 to 112, wherein measuring the number of target template nucleic acid molecules includes identifying the number of unique target template nucleic acid molecule sequences in one or more of the first of the sample pair, the second of the sample pair, at least one sample or a diluted sample.
[0114] A method according to any one of embodiments 110 to 113, wherein measuring the number of target template nucleic acid molecules comprises mutating the target template nucleic acid molecules.
[0115] The method of embodiment 114, wherein mutating the target template nucleic acid molecule comprises amplifying the target template nucleic acid molecule in the presence of nucleotide analogs.
[0116] The method described in embodiment 115, wherein the nucleotide analog is dPTP.
[0117] Measuring the number of target template nucleic acid molecules (i) mutating a target template nucleic acid molecule to obtain a mutated target template nucleic acid molecule; (ii) sequencing a region of the mutated target template nucleic acid molecule; and (iii) identifying the number of unique mutant target template nucleic acid molecules based on the number of unique mutant target template nucleic acid molecule sequences; 117. The method of any one of embodiments 110 to 116, comprising:
[0118] A method described in any of embodiments 108 to 117, wherein measuring the number of target template nucleic acid molecules includes introducing a barcode or a pair of barcodes into the target template nucleic acid molecule to obtain a barcoded target template nucleic acid molecule.
[0119] Measuring the number of target template nucleic acid molecules (i) sequencing a region of the barcoded target template nucleic acid molecule that includes the barcode or pair of barcodes; and (ii) determining the number of uniquely barcoded target template nucleic acid molecules based on the number of unique barcodes or barcode pairs; 119. The method of embodiment 118, comprising:
[0120] A method according to any of embodiments 102 to 119, wherein controlling the number of target template nucleic acid molecules in the first of the sample pairs and / or the second of the sample pairs comprises measuring the number of target template nucleic acid molecules and diluting the first of the sample pairs and / or the second of the sample pairs so that the first of the sample pairs and / or the second of the sample pairs contains the desired number of target template nucleic acid molecules.
[0121] A method according to any of embodiments 106 to 120, wherein normalizing the number of target template nucleic acid molecules in each of the subsamples comprises labeling target template nucleic acid molecules from different subsamples with different sample tags, preferably wherein labeling the target template nucleic acid molecules from different samples is performed before pooling the subsamples.
[0122] The method described in embodiment 121, comprising preparing a preliminary pool of sub-samples constituting the first of the sample pair and / or the second of the sample pair, and measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool.
[0123] The method described in embodiment 122, wherein measuring the number of target template nucleic acid molecules labeled with each sample tag in the reserve pool includes performing a serial dilution in the reserve pool to obtain a serial dilution comprising a diluted reserve pool.
[0124] A method described in any of embodiments 122 to 123, wherein measuring the number of target template nucleic acid molecules labeled with each sample tag in the reserve pool includes sequencing the target template nucleic acid molecules in the reserve pool or diluted reserve pool.
[0125] The method described in embodiment 124, wherein measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool comprises amplifying and then sequencing the target template nucleic acid molecules.
[0126] The method described in embodiment 124 or 125, wherein measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool comprises amplifying, fragmenting, and then sequencing the target template nucleic acid molecules.
[0127] A method described in any of embodiments 122 to 126, wherein measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool includes identifying the number of unique target template nucleic acid molecule sequences having each sample tag.
[0128] A method described in any of embodiments 122 to 127, wherein measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool includes mutating the target template nucleic acid molecules.
[0129] The method of embodiment 128, wherein mutating the target template nucleic acid molecule tag comprises amplifying the target template nucleic acid molecule in the presence of nucleotide analogs.
[0130] The method described in embodiment 129, wherein the nucleotide analog is dPTP.
[0131] Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool (i) mutating a target template nucleic acid molecule to obtain a mutated target template nucleic acid molecule; (ii) sequencing a region of the mutated target template nucleic acid molecule; and (iii) determining the number of unique mutant target template nucleic acid molecules having each sample tag based on the number of unique mutant target template nucleic acid molecules; 131. The method of any one of embodiments 122 to 130, comprising:
[0132] A method described in any of embodiments 122 to 131, wherein measuring the number of target template nucleic acid molecules includes introducing a barcode or a pair of barcodes into the target template nucleic acid molecule to obtain a barcoded sample-tagged target template nucleic acid molecule.
[0133] Measuring the number of target template nucleic acid molecules labeled with each sample tag (i) sequencing a region of the barcoded sample-tagged target template nucleic acid molecule that includes a barcode or pair of barcodes; and (ii) determining the number of unique barcoded target template nucleic acid molecules having each sample tag based on the number of unique barcode or barcode pair sequences associated with each sample tag; 133. The method of embodiment 132, comprising:
[0134] A method according to any one of embodiments 121 to 133, wherein the method comprises calculating the ratio of the number of target template nucleic acid molecules comprising different sample tags.
[0135] A method described in any of embodiments 104 to 134, wherein the first and / or second of a pair of samples is prepared by repooling the subsamples so that the number of target template nucleic acid molecules in each of the subsamples is in the desired ratio.
[0355] The present disclosure includes the following sequence information: SEQUENCE LISTING <110> Illumina Singapore PTE. LTD. <120> Sequencing algorithm <130> PA24-061 <140> <141> 2019-08-12 <150> GB1907101.8 <151> 2019-05-20 <150> GB1813171.4 <151> 2018-08-13 <160> 291 <170> PatentIn version 3.5 <210> 1 <211> 2325 <212> DNA <213> Thermococcus sp. KS-1 <400> 1 atgatcctcg acactgacta cataactgag aatggaaaac ccgtcataag gattttcaag 60 aaggagaacg gcgagtttaa gattgagtac gataggactt ttgaacccta catttacgcc 120 ctcctgaagg acgattctgc cattgaggag gtcaagaaga taaccgccga gaggcacgga 180 acggttgtaa cggttaagcg ggctgaaaag gttcagaaga agttcctcgg gagaccagtt 240 gaggtctgga aactctactt tactcaccct caggacgtcc cagcgataag ggacaagata 300 cgagagcatc cagcagttat tgacatctac gagtacgaca tacccttcgc caagcgctac 360 ctcatagaca agggattagt gccaatggaa ggcgacgagg agctgaaaat gcttgccttt 420 gatatcgaga cgctctacca tgagggcgag gagttcgccg aggggccaat ccttatgata 480 agctacgccg acgaggaagg ggccagggtg ataacgtgga agaacgcgga tctgccctac 540 gttgacgtcg tctcgacgga gagggagatg ataaagcgct tcctaaaggt ggtcaaagag 600 aaagatcctg acgtcctaat aacctacaac ggcgacaact tcgacttcgc ctacctaaaa 660 aaacgctgtg aaaagcttgg aataaacttc acgctcggaa gggacggaag cgagccgaag 720 attcagagga tgggcgacag gtttgccgtc gaagtgaagg gacggataca cttcgatctc 780 tatcctgtga taagacggac gataaacctg cccacataca cgcttgaggc cgtttatgaa 840 gccgtcttcg gtcagccgaa ggagaaggtc tacgctgagg agatagctac agcttgggag 900 agcggtgaag gccttgagag agtagccaga tactcgatgg aagatgcgaa ggtcacatac 960 gagcttggga aggagttttt ccctatggag gcccagcttt ctcgcttaat cggccagtcc 1020 ctctgggacg tctcccgctc cagcactggc aacctcgttg agtggttcct cctcaggaag 1080 gcctacgaga ggaatgagct ggccccgaac aagcccgatg aaaaggagct ggccagaaga 1140 cgacagagct atgaaggagg ctatgtaaaa gagcccgaga gagggttgtg ggagaacata 1200 gtgtacctag attttagatc tctgtacccc tcaatcatca tcacccacaa cgtctcgccg 1260 gatactctca acagggaagg atgcaaggaa tatgacgttg ccccccaggt cggtcaccgc 1320 ttctgcaagg acttcccagg atttatcccg agcctgcttg gagacctcct agaggagagg 1380 cagaagataa agaagaagat gaaggccacg attgacccga tcgagaggaa gctcctcgat 1440 tacaggcaga gggccatcaa gatcctggcc aacagctact acggttacta cggctatgca 1500 agggcgcgct ggtactgcaa ggagtgtgca gagagcgtaa cggcctgggg aagggagtac 1560 ataacgatga ccatcagaga gatagaggaa aagtacggct ttaagtaat ctacagcgac 1620 accgacggat tttgccac atacctgga gccgatgctg aaccgtcaa aaagaaggcg 1680 atggagttcc tcaagtatat caacgccaaa ctcccgggcg cgcttgagct cgagtacgag 1740 ggctctaca aacgcggctt cttcgtcacg aagagaagt acgcggtgat agacgaggaa 1800 ggcaagataa caacgcgcgg acttgagat gtgaggcgcg actggagcga gatgcgaaa 1860 gagacgcagg cgaggttct tgaagctttg ctaaggacg gtgacgtcga gaaggccgtg 1920 aggatagtca aagaagttac cgaaagctg agcaagtacg aggttccccc ggagaagctg 1980 gtgatccacg agcagataac gagggattta aaggactaca aggcaccgg tccccacgtt 2040 gccgttgcca agaggttggc cgcgagagga gtcaaatac gccctggaac ggtgataagc 2100 tacatcgtgc tcaagggctc tgggaggata ggcgacaggg cgataccgtt cgacgagttc 2160 gacccgacga agcacaagta cgacgccgag tactacattg agaaccaggt tctcccagcc 2220 gttgagagaa ttctgagagc ctcggttac cgcaaggaag acctgcgcta ccagaagacg 2280 agacaggttg gtctgggagc ctggctgaag ccgaagggaa cttga 2325 <210> 2 <211> 774 <212> PRT <213> Thermococcus sp. KS-1 <400> 2 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asn Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Thr 50 55 60 Val Lys Arg Ala Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Glu His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Asp Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Glu Gly Ala Arg Val Ile Thr Trp Lys Asn Ala 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Arg Glu Met Ile Lys 180 185 190 Arg Phe Leu Lys Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Thr Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Val Phe Gly Gln Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Ala Thr Ala Trp Glu Ser Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Tyr 305 310 315 320 Glu Leu Gly Lys Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Ser Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Asp Glu Lys Glu Leu Ala Arg Arg Arg Gln Ser Tyr 370 375 380 Glu Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Pro Serr Ile Leu Gly Asp Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Lys Lys Met Lys Ala Thr Ile Asp Pro Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Only With Lys And Leu Only Asn Ser Tyr Tyr Gly Tyr 485,490,495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Only Trp Gly Arg Glu Tyr Ile Met Thr Ile Arg Glu Ile 515,520,525 Glu Glu Lys Tyr Gly Phe Lys Val Ile Tyr Ser Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Ala Asp Ala Glu Thr Val Lys Lys Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565,570,575 Leu Glu Tyr Glu Gly Phe Tyr Lys Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Lys Ala Val 625 630 635 640 Arg Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690,695,700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Arg Ala Phe Gly Tyr Arg Lys 740,745,750 Glu Asp Leu Arg Tyr Gln Lys Thr Gln Arg Val Gly Leu Gly Ala Trp 755,760,765 Leu Lys Pro Lys Gly Thr 770 <210> 3 <211> 2325 <212> DNA <213> Thermococcus celer <400> 3 atgatccctcg acgctgacta catcaccgaa gatgggaagc ccgtcgtgag gatattcagg 60 aaggaagg gcgagttcag atcgactac gatgact tcgagcccta catctacgcc 120 ctcctgaagg acgattcggc catcgaggag gtgaagagga taaccgttga gcgccacggg 180 aaggccgtca gggttaagcg ggtggagaag gtcgaaaaga agttcctcaa caggccgata 240 gaggtctgga agctctactt caatcacccg caggacgttc cggcgataag ggacgagata 300 aggaagcatc cggccgtcgt tgatatctac gagtacgaca tccccttcgc caagcgctac 360 ctcatcgata aggggctcgt cccgatggag ggggaggagg agctcaaact gatggccttc 420 gacatcgaga ccctctacca cgagggagac gagttcgggg aggggccgat cctgatgata 480 agctacgccg acggggacgg ggcgagggtc ataacctgga agaagatcga cctcccctac 540 gtcgacgtcg tctcgaccga gaaggagatg ataaagcgct tcctccaggt ggtgaaggag 600 aaggacccgg acgtgctcgt aacttacaac ggcgacaact tcgacttcgc ctacctgaag 660 agacgctccg aggagcttgg attgaagttc atcctcggga gggacgggag cgagcccaag 720 atccagcgca tgggcgaccg cttcgccgtc gaggtgaagg ggaggataca cttcgacctc 780 tacccggtga taaggcgcac cgtgaacctg ccgacctaca cgctcgaggc ggtctacgag 840 gccatcttcg ggaggccaaa ggagaaggtc tacgccgggg agatagtgga ggcctgggaa 900 accggcgagg gtcttgagag ggttgcccgc tactccatgg aggacgcaaa ggttaccttc 960 gagctcggga gggagttctt cccgatggag gcccagctct cgaggctcat cggccagggt 1020 ctctgggacg tctcccgctc gagcaccggc aacctggtcg agtggttcct cctgaggaag 1080 gcctacgaga ggaacgaact ggccccgaac aagccgagcg gccgggaagt ggagatcagg 1140 aggcgtggct acgccggtgg ttacgttaag gagccggaga ggggtttatg ggagaacatc 1200 gtgtacctcg actttcgctc tctttacccc tccatcatca taacccacaa cgtctcgccc 1260 gataccctaa acagggaggg ctgtgagaac tacgacgtcg ccccccaggt ggggcataag 1320 ttctgcaaag attttccggg cttcatcccg agcctgctcg gaggcctgct tgaggagagg 1380 cagaagataa agcggaggat gaaggcctct gtggatcccg ttgagcggaa gctcctcgat 1440 tacaggcaga gggccatcaa gatactggcc aacagcttct acggatacta cggctacgcg 1500 agggcgaggt ggtactgcag ggagtgcgcg gagagcgtta ccgcctgggg cagggagtac 1560 atcgataggg tcatcaggga gctcgaggag aagttcggct tcaaggtgct ctacgcggac 1620 acggacggac tgcacgccac gatccccggg gcggacgccg ggaccgtcaa ggagagggcg 1680 agggggttcc tgagatacat caaccccaag ctccccggcc tcctggagct cgagtacgag 1740 gggttctacc tgaggggttt cttcgtgacg aagaagaagt acgcggtcat agacgaggag 1800 ggcaagataa ccacgcgcgg cctcgagata gtcaggcggg actggagcga ggtggccaag 1860 gagacgcagg cgagggtcct ggaggcgata ctgaggcacg gtgacgtcga ggaggccgtt 1920 agaatcgtca gggaggtac cgaaaagctg agcaagtacg aggttccgcc ggagaaactg gtgatccacg agcagataac gagggatttg aggactaca aagccacggg accgcacgtg gcggtggcga agcgcctggc cgggaggggg gtaaggatac gccccgggac ggtgataagc 2100. tacatcgtcc tcaagggctc cggaggata ggggacaggg cgattccctt cgacgagttc gacccgacta agcacaggta cgacgccgac tactacatcg agaaccaggt tctgccagcc gtcgaga tcctgaaggc cttcggctac cgcaaggagg acctgaata ccagaagacg aggcaggtgg gcctgggtgc gtggctcaac gcggggaagg ggtga 2325. <210> 4 <211> 774 <212> PRT <213> Thermococcus cell <400> 4 Met Ile Leu Asp Ala Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Val 1 5 10 15 Arg Ile Phe Arg Lys Glu Lys Gly Glu Phe Arg Ile Asp Tyr Asp Arg 20 25 30 Asp Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Arg Ile Thr Val Glu Arg His Gly Lys Ala Val Arg 50 55 60 Val Lys Arg Val Glu Lys Val Glu Lys Phe Leu Asn Arg Pro Ile 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Asn His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Glue Ile Arg Lys Pro Ala Val Val Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Glu Glu Glu Leu Lys Leu Met Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Asp Glu Phe Gly Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Gly Asp Gly Ala Arg Val Ile Thr Trp Lys Lys Ile 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Lys Glu Met Ile Lys 180 185 190 Arg Phe Leu Gln Val Val Lys Glu Lys Asp Pro Asp Val Leu Val Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Arg Arg Ser Glu 210 215 220 Glu Leu Gly Leu Lys Phe Ile Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Val Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Ile Phe Gly Arg Pro Lys Glu 275 280 285 Lys Val Tyr Ala Gly Glu Ile Val Glu Ala Trp Glu Thr Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Phe 305 310 315 320 Glu Leu Gly Arg Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Gly Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Ser Gly Arg Glu Val Glu Ile Arg Arg Arg Gly Tyr 370 375 380 Ala Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Thr His 405 410 415 Asn Will Be Pro Asp Thr Asn Arg Glu Gly Cys Glu Asn Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Lys Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Pro Ile Ser Leu Gly Gly Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Arg Arg Met Lys Ala Ser Val Asp Pro Val Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Only With Lys And Leu With Asn Ser Phe Tyr Gly Tyr 485,490,495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Arg Glu Cys Ala Glu Ser 500 505 510 Val Thr Only Trp Gly Arg Glu Tyr Ile Asp Arg Val Ile Arg Glu Leu 515,520,525 Glu Glu Lys Phe Gly Phe Lys Val Leu Tyr Ala Asp Thr Asp Gly Leu 530 535 540 His Ala Thr Ile Pro Gly Ala Asp Ala Gly Thr Val Lys Glu Arg Ala 545 550 555 560 Arg Gly Phe Leu Arg Tyr Ile Asn Pro Lys Leu Pro Gly Leu Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Leu Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Val Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Ile Leu Arg His Gly Asp Val Glu Glu Ala Val 625 630 635 640 Arg Ile Val Arg Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Arg Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Gly 675 680 685 Arg Gly Val Arg Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Arg Tyr Asp Ala Asp Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Lys Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Lys Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Asn Ala Gly Lys Gly 770 <210> 5 <211> 2328 <212> DNA <213> Thermococcus siculi <400> 5 atgatcctcg acacggacta catcacggaa gatgggaaac ccgtcataag gatattcaag 60 aaagagaacg gcgagttcaa gatcgagtac gacaggactt ttgaacccta catctacgcc 120 ctcctgaagg acgactccgc gattgaggat gttaaaaaga taaccgccga gaggcacgga 180 acggtggtga aggtcaagcg cgccgaaaag gtgcagaaga agttcctagg caggccggtt 240 gaagtctgga agctctactt cacccacccc caagatgtcc cggcgataag ggacaagatt 300 aggaagcatc cagctgtaat tgacatctac gagtacgaca taccattcgc caagcgctac 360 ctcatcgaca agggcctgat tccgatggag ggtgaagaag agcttaagat gctcgcttc 420 gacattgaga cgctctacca tgagggtgag gagttcgccg aggggcctat tctgatgata 480 agctacgccg acgagagcga ggcacgcgtc atcacctgga agaaaatcga cctcccctac 540 gttgacgtcg tctcaacgga gaaggagatg ataaagcgct tcctccgcgt tgtgaaggag 600 aaagatcccg atgtcctcat aacctacaac ggcgacaact tcgacttcgc ctacctgaag 660 aagcgctgtg aaaagcttgg aataaacttc ctccttggaa gggacgggag cgagccgaag 720 atccagagaa tgggtgaccg cttcgccgtt gaggtgaagg ggaggataca cttcgacctc 780 tatcctgtaa taaggcgcac gataaacctg ccgacctaca tgcttgaggc agtctacgag 840 gccatctttg ggaagccaaa ggagaaggtt tacgccgagg agatagccac cgcttgggaa 900 accggagagg gccttgagag ggtggctcgc tactctatgg aggacgcgaa ggtcacgttt 960 gagcttggaa aggagttctt cccgatggag gcccaacttt cgaggttggt cggccagagc 1020 ttctgggatg tcgcgcgctc aagcacgggc aatctggtcg agtggttcct cctcaggaag 1080 gcctacgaga ggaacgagct ggctccaaac aagccctctg gaagggaata tgacgagagg 1140 cgcggtggat acgccggcgg ctacgtcaag gaaccggaaa agggcctgtg ggagaacata 1200 gtctacctcg actataaatc tctctacccc tcaatcatca tcacccacaa cgtctcgccc 1260 gataccctca accgcgaggg ctgtaaggag tatgacgtag ctccacaggt cggccaccgc 1320 ttctgcaagg actttccagg cttcatcccg agcctgctcg gggatctcct ggaggagagg 1380 cagaagataa agaggaagat gaaggcaaca attgacccga tcgagagaaa gctccttgat 1440 tacaggcaac gggccatcaa gatccttcta aatagttttt acggctacta cggctacgca 1500 agggctcgct ggtactgcaa ggagtgtgcc gagagcgtta cggcatgggg aagggaatat 1560 atcaccatga caatcaggga aatagaagag aagtatggct ttaaagtact ttatgcggac 1620 actgacggct tcttcgcgac gattcccggg gaagatgccg agaccatcaa aaagagggcg 1680 atggagttcc tcaagtacat aaacgccaaa ctccccggtg cgctcgaact tgagtacgag 1740 gacttctaca ggcgcggctt cttcgtcacc aagagaat acgcggttat cgacgaggag 1800 ggcaagataa caacgcgcgg gctggagatc gtcaggcgcg actggagcga gatagccaag 1860 gagacgcagg cgcgggttct ggaggccctt ctgaggacg gtgacgtcga agaggccgtg 1920 agcatagtca aagaagtgac cgagaagctg agcaagtacg aggttccgcc ggagaagctc 1980 gttatccacg agcagataac gcgcgagctg aaggactaca agcaacggg accacacgtg 2040 gcgatagcga agaggttagc cgcgagaggc gtcaaatcc gccccgggac agtcatcagc 2100 tacatcgtgc tcaagggctc cgggaggata ggcgacaggg cgattcctt cgacgagttc 2160 gaccccacga agcacaagta cgatgcagag tactacatcg agaaccaggt tctacctgcc 2220 gtcgagagga ttctgaggc cttcggctat cgcggtgagg agctcagata ccagaagacg 2280 aggcaggttg gacttggggc gtggctgaag ccgagggga aggggtga 2328 <210> 6 <211> 775 <212> PRT <213> Thermococcus siculi <400> 6 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Asp Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Lys 50 55 60 Val Lys Arg Ala Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Lys His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Ile Pro 115 120 125 Met Glu Gly Glu Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Ser Glu Ala Arg Val Ile Thr Trp Lys Lys Ile 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Lys Glu Met Ile Lys 180 185 190 Arg Phe Leu Arg Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Leu Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Met Leu Glu Ala Val Tyr Glu Ala Ile Phe Gly Lys Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Ala Thr Ala Trp Glu Thr Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Phe 305 310 315 320 Glu Leu Gly Lys Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Val Gly Gln Ser Phe Trp Asp Val Ala Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Ser Gly Arg Glu Tyr Asp Glu Arg Arg Gly Gly Tyr 370 375 380 Ala Gly Gly Tyr Val Lys Glu Pro Glu Lys Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Tyr Lys Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Arg Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Leu Asn Ser Phe Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Thr Met Thr Ile Arg Glu Ile 515 520 525 Glu Glu Lys Tyr Gly Phe Lys Val Leu Tyr Ala Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Glu Asp Ala Glu Thr Ile Lys Lys Arg Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Asp Phe Tyr Arg Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Glu Ala Val 625 630 635 640 Ser Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Glu Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Ile Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Lys Ala Phe Gly Tyr Arg Gly 740 745 750 Glu Glu Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Lys Pro Lys Gly Lys Gly 770 775 <210> 7 <211> 774 <212> PRT <213> Thermococcus kodakarensis <400> 7 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Phe Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Thr 50 55 60 Val Lys Arg Val Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Glu His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Asp Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr Glu Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Glu Gly Ala Arg Val Ile Thr Trp Lys Asn Val 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Arg Glu Met Ile Lys 180 185 190 Arg Phe Leu Arg Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Ala Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Val Phe Gly Gln Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Thr Thr Ala Trp Glu Thr Gly Glu Asn 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Tyr 305 310 315 320 Glu Leu Gly Lys Glu Phe Leu Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Ser Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Asp Glu Lys Glu Leu Ala Arg Arg Arg Gln Ser Tyr 370 375 380 Glu Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Pro Serr Ile Leu Gly Asp Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Lys Lys Met Lys Ala Thr Ile Asp Pro Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Only With Lys And Leu Only Asn Ser Tyr Tyr Gly Tyr 485,490,495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Met Thr Ile Lys Glu Ile 515,520,525 Glu Glu Lys Tyr Gly Phe Lys Val Ile Tyr Ser Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Ala Asp Ala Glu Thr Val Lys Lys Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Glu Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Lys Ala Val 625 630 635 640 Arg Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Arg Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Ser Ala Trp 755 760 765 Leu Lys Pro Lys Gly Thr 770 <210> 8 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 8 tagaattgaa gaa 13 <210> 9 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 9 tggccatagc tac 13 <210> 10 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 10 gtcatctgcg acc 13 <210> 11 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 11 ttcgcgcttg gac 13 <210> 12 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 12 cgcgaaccgt tag 13 <210> 13 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 13 ttgcagcctc taa 13 <210> 14 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 14 tctactagta cga 13 <210> 15 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 15 gtaggttcta ctg 13 <210> 16 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 16 gccaatatca agt 13 <210> 17 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 17 ctatcttgct ggt 13 <210> 18 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 18 gttctcatag gta 13 <210> 19 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 19 gtctatgaac caa 13 <210> 20 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 20 cggagcgctt att 13 <210> 21 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 21 tatgccatga gga 13 <210> 22 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 22 atacgactcg gag 13 <210> 23 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 23 gatggaactc agc 13 <210> 24 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 24 ggacctgcat gaa 13 <210> 25 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 25 tagactggaa ctt 13 <210> 26 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 26 gaattacctc gtt 13 <210> 27 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 27 aggatcaggc tac 13 <210> 28 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 28 acgcgtagaa gag 13 <210> 29 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 29 cttcgagact tac 13 <210> 30 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 30 gacggctaac tcc 13 <210> 31 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 31 ttagcattct ctt 13 <210> 32 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 32 gcaaggcata gta 13 <210> 33 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 33 acctagatat gga 13 <210> 34 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 34 acgccaaggc gta 13 <210> 35 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 35 tatgacggat ccg 13 <210> 36 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 36 cctccattag aga 13 <210> 37 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 37 attgaatact ctg 13 <210> 38 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 38 gagatgagaa gaa 13 <210> 39 <211> 13 <212> DNA <213> Artificial Sequence ...
Claims
1. (a) (i) non-mutated sequence reads generated from a first copy of at least one target template nucleic acid molecule; and (ii) mutant sequence reads generated from a second copy of the at least one target template nucleic acid molecule, wherein the second copy is mutated; obtaining data including (b) analyzing the mutant sequence reads, wherein analyzing the mutant sequence reads includes: (i) a pre-clustering step of identifying mutant sequence reads that contain one or more common signature mutations that occur more frequently in mutated sequence reads compared to non-mutated sequence reads by assigning the mutant sequence reads to groups to obtain pre-clustered groups, wherein each member of the same group has a reasonable likelihood of originating from the same mutant target template nucleic acid molecule; and (ii) identifying mutated sequence reads in the pre-clustered groups that are likely to originate from the same at least one mutated target template nucleic acid molecule by identifying mutated sequence reads that contain one or more common signature mutations; and, (c) constructing a sequence of at least a portion of at least one target template nucleic acid molecule from non-mutated sequence reads using information obtained from analysis of the mutant sequence reads; (i) creating an assembly graph from unmutated reads, and (ii) mapping mutated reads to the assembly graph; The process comprising: a computer-implemented method for determining the sequence of said at least one target template nucleic acid molecule, said method comprising:
2. The method of claim 1 , wherein the pre-clustering step comprises Markov clustering or Louvain clustering.
3. The method of claim 1 or 2, wherein members of the same group are mapped to a common position on the assembly graph and / or share a common mutation pattern.
4. The method of claim 3, wherein the mutant sequence reads sharing a common mutation pattern are mutant sequence reads that contain at least 1, at least 2, at least 3, at least 4, at least 5, or at least k common signature k-mers and / or common signature mutations.
5. 5. The method of claim 4, wherein the signature k-mer is a k-mer that does not occur in non-mutated sequence reads but occurs at least 2 times, at least 3 times, at least 4 times, at least 5 times, or at least 10 times in mutant sequence reads.
6. 6. The method of claim 4 or 5, wherein the signature mutation is a nucleotide that appears at least 2, at least 3, at least 4, at least 5, or at least 10 times in the mutant sequence reads but does not appear at the corresponding position in the non-mutated sequence reads.
7. The method of claim 6 , wherein the signature mutation is a co-occurring mutation.
8. 8. The method of claim 6 or 7, wherein a signature mutation is ignored if at least one, at least two, at least three, or at least five nucleotides at corresponding positions in mutant sequence reads that share the signature mutation differ from each other.
9. The method of any one of claims 6 to 8, wherein a signature mutation is ignored if the signature mutation is an unexpected mutation.
10. The method of any one of claims 1 to 9, wherein the assembly graph comprises nodes computed from non-mutated sequence reads.
11. 11. The method of claim 10, wherein each valid path through the assembly graph comprising nodes represents at least a portion of the sequence of at least one target template nucleic acid molecule.
12. A method described in any one of claims 1 to 11, wherein step (b)(ii) further comprises determining mutated sequence reads that map to a common region of the assembly graph.
13. The first mutant sequence read, (i) the probability that a given nucleotide in a mutant sequence read is mutated; (ii) the probability that a given nucleotide in a mutant sequence read is misread, given that the nucleotide is misread; and / or (iii) The probability that a nucleotide in the mutant sequence read was misread.
13. The method of any one of claims 1 to 12, wherein the first mutant sequence read is identified as likely to originate from the same at least one target template nucleic acid molecule as the second mutant sequence read based on:
14. A computer program adapted to carry out the method according to any one of claims 1 to 13.
15. 15. A computer readable medium containing the computer program of claim 14.
16. A computer comprising a processor adapted to carry out the method of any one of claims 1 to 13.
Citation Information
Patent Citations
Label-free on-target pharmacological methods
JP2013529891A
Sequencing process
JP2017517282A