Sequencing algorithm
By sequencing non-mutation and mutation sequence reads of paired samples and combining them with computer program analysis, the difficulty of assembling nucleic acid molecule sequences in repetitive regions in the existing technology was solved, accurate and rapid sequence assembly was achieved, and the processing process was simplified.
Patent Information
- Application Number
- CN201980067627.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-20
- Filing Date
- 2019-08-12
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2039-11-14
AI Technical Summary
Existing technologies have difficulty accurately, quickly and efficiently assembling sequences of nucleic acid molecules containing repetitive regions. In particular, when using next-generation sequencing technology, it is difficult to distinguish between repetitive sequences or sequences of different nucleic acid molecules corresponding to similar sequence reads.
By providing paired samples, sequencing of non-mutation and mutation sequence reads is performed separately, and non-mutation sequences are assembled using mutation sequence read analysis information. In combination with computer programs and methods, the number of target template nucleic acid molecules in the sample is controlled or standardized to improve assembly accuracy.
Accurate, rapid and efficient sequence assembly of target template nucleic acid molecules is achieved, artifacts in the consensus sequence are reduced and the calculation process is simplified.
Smart Images

Figure GDA0004930012920000271 
Figure GDA0004930012920000281 
Figure GDA0004930012920000301
Abstract
Description
Technical Field
[0001] The present invention relates to methods for determining the sequence of at least one target template nucleic acid molecule using non-mutated sequence reads and mutant sequence reads. The present invention also relates to methods for determining the sequence of at least one target template nucleic acid molecule in a sample, the methods involving controlling or normalizing the number of target template nucleic acid molecules in a sample. The present invention also relates to computer programs suitable for performing the methods, computer-readable media comprising the computer programs, and computer-implemented methods. Background Art
[0002] The ability to sequence nucleic acid molecules is a valuable tool in numerous applications. However, determining the exact sequence of nucleic acid molecules containing undefined structures, such as those containing repetitive regions, can be difficult. Resolving structural variation, such as the haploid structure of diploid and polyploid organisms, can also be challenging.
[0003] Many of the more modern technologies (so-called next generation sequencing technologies) can only accurately sequence short nucleic acid molecules. Longer nucleic acid sequences can be sequenced using next generation sequencing technologies, but this is often difficult. Next generation sequencing technologies can be used to generate short sequence reads corresponding to the sequence of a portion of a nucleic acid molecule, and a complete sequence can be assembled from the short sequence reads. When a nucleic acid molecule contains repetitive regions, it may not be clear to the user whether two sequence reads with similar sequences correspond to two repeats in a longer sequence or two repeats of the same sequence. Similarly, a user may want to sequence two similar nucleic acid molecules at the same time, and it may be difficult to determine whether two sequence reads with similar sequences correspond to the sequence of the same original nucleic acid molecule or the sequence of two different original nucleic acid molecules.
[0004] Sequencing assisted by mutagenesis (SAM) technology can be used to assist in assembling sequences from short sequence reads. Typically, SAM involves introducing mutations into a target template nucleic acid sequence. The introduced mutation pattern can help users of this method assemble the sequence of a nucleic acid molecule from short sequence reads.
[0005] For example, where the template nucleic acid molecule comprises repetitive regions, the repetitive regions can be distinguished from each other by different mutation patterns, thereby enabling the repetitive regions to be correctly resolved and assembled.
[0006] Typically, SAM technology involves mutating copies of a target template nucleic acid molecule and then assembling the sequences of the mutated copies based on their mutation patterns. Users can then create a consensus sequence from the sequences of the mutated copies. Because different mutated copies will contain mutations at different positions, the consensus sequence can be representative of the original template nucleic acid molecule. However, the consensus sequence may contain artifacts from the mutation process. Furthermore, creating the consensus sequence involves the use of complex and processing-intensive computer programs.
[0007] Therefore, there remains a need for methods for determining the sequence of at least one target template nucleic acid molecule, wherein sequence reads can be assembled accurately, quickly, and efficiently. Summary of the Invention
[0008] The present inventors have developed a new and improved method for determining the sequence of at least one target template nucleic acid molecule. Therefore, in a first aspect of the present invention, a method for determining the sequence of at least one target template nucleic acid molecule is provided, the method comprising:
[0009] (a) providing paired samples, each sample comprising at least one target template nucleic acid molecule;
[0010] (b) sequencing a region of at least one target template nucleic acid molecule in a first sample of the paired samples to provide a non-mutation sequence read;
[0011] (c) introducing a mutation into at least one target template nucleic acid molecule in the second sample of the paired samples to provide at least one mutated target template nucleic acid molecule;
[0012] (d) sequencing a region of at least one mutated target template nucleic acid molecule to provide a mutation sequence read;
[0013] (e) analyzing the mutant sequence reads and using information obtained from analyzing the mutant sequence reads to assemble a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads.
[0014] Thus, in a second aspect of the present invention, there is provided a method for generating a sequence of at least one target template nucleic acid molecule, the method comprising:
[0015] (a) obtaining data, which includes:
[0016] (i) non-mutated sequence reads; and
[0017] (ii) mutant sequence reads;
[0018] (b) analyzing the mutant sequence reads and using information obtained from analyzing the mutant sequence reads to assemble a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads.
[0019] In a third aspect of the invention there is provided a computer program adapted to carry out the method of the invention.
[0020] In a fourth aspect of the invention there is provided a computer readable medium embodying the computer program of the invention.
[0021] In a fifth aspect of the invention, there is provided a computer-implemented method comprising the method of the invention.
[0022] In a sixth aspect of the present invention, there is provided a method for determining the sequence of at least one target template nucleic acid molecule, the method comprising:
[0023] (a) providing at least one sample, the at least one sample comprising at least one target template nucleic acid molecule;
[0024] (b) sequencing a region of at least one target template nucleic acid molecule; and
[0025] (c) assembling the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule,
[0026] in:
[0027] (i) providing at least one sample comprising at least one target template nucleic acid molecule comprises: controlling the number of target template nucleic acid molecules in the at least one sample; and / or
[0028] (ii) providing at least one sample by combining two or more subsamples, and normalizing the number of target template nucleic acid molecules in each subsample.
[0029] In a sixth aspect of the present invention, there is provided a method for determining the sequence of at least one target template nucleic acid molecule, the method comprising:
[0030] (a) providing at least one sample, the at least one sample comprising at least one target template nucleic acid molecule;
[0031] (b) sequencing a region of at least one target template nucleic acid molecule; and
[0032] (c) assembling the sequence of at least a portion of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule,
[0033] in:
[0034] (i) providing at least one sample comprising at least one target template nucleic acid molecule comprises: controlling the number of target template nucleic acid molecules in the at least one sample; and / or
[0035] (ii) providing at least one sample by combining two or more subsamples, and normalizing the number of target template nucleic acid molecules in each subsample. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1Shown are the mutation levels achieved using three different polymerases in the presence or absence of dPTP. Panel A shows data obtained using Taq (Jena Biosciences), Panel B shows data obtained using LongAmp (New England Biolabs), and Panel C shows data obtained using Primestar GXL (Takara). The dark grey bars show the results obtained in the absence of dPTP, and the light grey bars show the results obtained in the presence of 0.5 mM dPTP.
[0037] Figure 2 Describes the mutation rates achieved by dPTP mutagenesis of templates with varying G+C content using Thermococcus polymerase (Primestar GXL; Takara). For a low-GC template (33% GC) from Staphylococcus aureus (S. aureus), a median mutation rate of ~7% was observed, while for other templates the median was approximately 8%;
[0038] Figure 3 is a sequence listing;
[0039] Figure 4 describes the lengths of the fragments obtained using the method described in Example 5;
[0040] Figure 5 Depicts the distribution of values using variational inference on simulated data. Panel A shows the values of M inferred using variational inference on the simulated data. The true value of identity ([1,1], [2,2], [3,3], [4,4]) is 0.895, the true value of transition ([1,3], [2,4], [3,1], [4,2]) is 0.1, and the true value of transversion (all other entries) is 0.005. Panel B shows the values of z inferred using variational inference on the simulated data. For same[1:5], the true value of z is 1, and for same[91:95], the true value of z is 0;
[0041] Figure 6 is the precision-recall plot of simulated data using a cutoff value ranging from 100 to 10,000 with a unit of 100. 2,000 tests were performed for each threshold, including 1,000 read pairs that were indeed derived from the same template and 1,000 read pairs that were not derived from the same template;
[0042] Figure 7 is a flow chart illustrating a method for determining the sequence of at least one target template nucleic acid molecule of the present invention;
[0043] Figure 8is a flow chart illustrating a method for generating a sequence of at least one target template nucleic acid molecule of the present invention;
[0044] Figure 9 The assembly map is depicted in panel A, and the mutant sequence reads are mapped to the assembly map in panel B;
[0045] Figure 10 Depicted are the sizes of target nucleic acid molecules amplified using adapters annealed to each other (right line) or using standard adapters (left line);
[0046] Figure 11 is a graph depicting the linear relationship between the sample dilution factor and the number of unique templates observed. A starting sample of target template nucleic acid molecules is serially diluted and ultimately sequenced to identify and quantify the number of unique templates in each dilution;
[0047] Figure 12 Figures 2 and 3 show the normalization of template counts between samples in a pool. (A) Shows the unique template counts for 66 barcoded bacterial genomes, determined from pooled samples before normalization. (B) Shows the template counts for the same samples after normalization (expressed per megabase (Mb) of genome content), demonstrating much less variability.
[0048] Figure 13 A workflow for assembling bacterial genomes according to the present invention is shown;
[0049] Figure 14 Comparison of assembly statistics from 65 bacterial genomes is shown for standard read assembly versus the assembly of the present invention (Morphoseq assembly);
[0050] Figure 15 Exemplary assembly metrics for the assembly of bacterial genomes are shown for short read assemblies compared to the present assembly;
[0051] Figure 16An exemplary workflow of the present invention for generating synthetic long reads is shown. (a) Preparation of long mutant templates. The target genomic DNA is first tagged and fragmented to generate long templates containing terminal adapters. The template is then amplified in the presence of a mutagenic nucleotide analog dPTP, which is randomly incorporated opposite the A and G residues of the two product chains (mutagenetic PCR). This step also introduces: (i) a sample tag, and; (ii) an additional adapter sequence at the end of the template to facilitate downstream amplification of products containing P bases. Further amplification is performed in the absence of dPTP (recovery PCR), during which the template P residues are replaced by natural nucleotides to generate conversion mutations (shown in red). The sample is then size-selected (8kb-10kb) to limit it to a fixed number of unique templates and then selectively enriched to create many copies of each unique molecule. (b) Preparation, sequencing and analysis of short read libraries. The long mutant templates are subjected to short read sequencing through further tag fragmentation and library amplification. In this step, the very end fragments derived from the full-length template and the random "internal" fragments are amplified and barcoded separately using different primers targeting the original template end adapters (dark gray) and the internal tag fragmentation adapters (light gray). These two libraries are sequenced, a non-mutation reference library is generated in parallel, and a custom algorithm is used to reconstruct synthetic long reads. This involves creating an assembly graph from the reference data, on which the mutant reads are mapped and connected together by different overlapping mutation patterns. The final synthetic long read corresponds to the identification path through the non-mutation assembly graph. DETAILED DESCRIPTION
[0052] General Definition
[0053] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0054] Generally, the term "comprising" is intended to mean including, but not limited to. For example, the phrase "a method for determining the sequence of at least one target template nucleic acid molecule comprises [certain steps]" should be interpreted to mean that the method includes the recited steps, but additional steps may be performed.
[0055] In some embodiments of the present invention, the word "comprising" is replaced with "consisting of." The term "consisting of" is intended to be limiting. For example, the phrase "a method for determining the sequence of at least one target template nucleic acid molecule consists of [certain steps]" should be understood to mean that the method includes the recited steps and does not perform other steps.
[0056] Method for determining the sequence of at least one target template nucleic acid molecule
[0057] In some aspects, the present invention provides methods for determining the sequence of at least one target template nucleic acid molecule or methods for generating the sequence of at least one target template nucleic acid molecule.
[0058] For the purposes of the present invention, the terms "determine" and "generate" can be used interchangeably. However, a method of "determining" a sequence generally includes steps, such as a sequencing step, while a method of "generating" a sequence may be limited to steps that can be implemented by a computer.
[0059] The method can be used to determine or generate the complete sequence of at least one target template nucleic acid molecule. Alternatively, the method can be used to determine or generate a partial sequence, i.e., the sequence of a portion of at least one target template nucleic acid molecule. For example, if it is not possible or not straightforward to determine the complete sequence, the user can determine that the sequence of a portion of at least one target template nucleic acid molecule is useful or even sufficient for their purpose.
[0060] For the purposes of the present invention, a "nucleic acid molecule" refers to a polymeric form of nucleotides of any length. Nucleotides can be deoxyribonucleotides, ribonucleotides, or analogs thereof. Preferably, at least one target template nucleic acid molecule consists of deoxyribonucleotides or ribonucleotides. Even more preferably, at least one target template nucleic acid molecule consists of deoxyribonucleotides, i.e., at least one target template nucleic acid molecule is a DNA molecule.
[0061] At least one "target template nucleic acid molecule" can be any nucleic acid molecule that the user wants to sequence. The "at least one target template nucleic acid molecule" can be single-stranded, or can be part of a double-stranded complex. If the at least one target template nucleic acid molecule is composed of deoxyribonucleotides, it can form part of a double-stranded DNA complex. In this case, one chain (e.g., a coding chain) will be considered to be at least one target template nucleic acid molecule, and the other chain is a nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule. At least one target template nucleic acid molecule can be a DNA molecule corresponding to a gene, can contain introns, can be an intergenic region, can be an intragenic region, can be a genomic region spanning multiple genes, or can actually be the entire genome of an organism.
[0062] The terms "at least one target template nucleic acid molecule" and "at least one target template nucleic acid molecules" are considered synonymous and can be used interchangeably herein.
[0063] In the method of the present invention, any number of at least one target template nucleic acid molecules can be sequenced simultaneously. Therefore, in one embodiment of the present invention, at least one target template nucleic acid molecule comprises a plurality of target template nucleic acid molecules. Alternatively, at least one target template nucleic acid molecule comprises at least 10, at least 20, at least 50, at least 100 or at least 250 target template nucleic acid molecules. Alternatively, at least one target template nucleic acid molecule comprises 10 to 1000, 20 to 500 or 50 to 100 target template nucleic acid molecules.
[0064] Methods for determining the sequence of at least one target template nucleic acid molecule may comprise:
[0065] (a) providing paired samples, each sample comprising at least one target template nucleic acid molecule;
[0066] (b) sequencing a region of at least one target template nucleic acid molecule in a first sample of the paired samples to provide a non-mutation sequence read;
[0067] (c) introducing a mutation into at least one target template nucleic acid molecule in the second sample of the paired samples to provide at least one mutated target template nucleic acid molecule;
[0068] (d) sequencing a region of at least one mutated target template nucleic acid molecule to provide a mutation sequence read;
[0069] (e) analyzing the mutant sequence reads and using information obtained from analyzing the mutant sequence reads to assemble a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads.
[0070] A method for generating a sequence of at least one target template nucleic acid molecule may comprise:
[0071] (a) obtaining data, which includes:
[0072] (i) non-mutated sequence reads; and
[0073] (ii) mutant sequence reads;
[0074] (b) analyzing the mutant sequence reads and using information obtained from analyzing the mutant sequence reads to assemble a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads.
[0075] Providing paired samples, each sample containing at least one target template nucleic acid molecule
[0076] A method for determining the sequence of at least one target template nucleic acid molecule may include the steps of providing a pair of samples, each sample comprising at least one target template nucleic acid molecule.
[0077] The methods of the present invention use information obtained from analyzing mutant sequence reads to assemble the sequence of at least a portion of at least one target template nucleic acid molecule from non-mutated sequence reads. The methods of the present invention may include introducing a mutation into at least one target template nucleic acid molecule in the second sample of the paired sample. Thus, sequencing a region of at least one mutated target template nucleic acid molecule in the second sample of the paired sample can be used to provide a mutant sequence read, while sequencing a region of at least one mutated non-mutated target template nucleic acid molecule in the first sample of the paired sample can be used to provide a non-mutated sequence read.
[0078] In order for a user to use information obtained from analyzing mutant sequence reads from a second sample to assemble a sequence comprising primarily non-mutated sequence from a first sample, some mutant sequence reads and some non-mutated sequence reads will correspond to the same original target template nucleic acid molecule.
[0079] For example, if a user wishes to determine the sequence of target template nucleic acid molecules A and B, a first sample will contain template nucleic acid molecules A and B, and a second sample will contain template nucleic acid molecules A and B. A and B in the first sample can be sequenced to provide non-mutated sequence reads of A and B, and A and B in the second sample can be mutated and sequenced to provide mutated sequence reads of A and B.
[0080] Since the first sample in the pair of samples and the second sample in the pair of samples both contain at least one target template nucleic acid molecule, the pair of samples can be derived from the same target organism or taken from the same original sample.
[0081] For example, if the user intends to sequence at least one target template nucleic acid molecule in a sample, the user can obtain paired samples from the same original sample. Alternatively, the user can replicate at least one target template nucleic acid molecule in the original sample before obtaining the paired sample from the original sample. The user can intend to sequence various nucleic acid molecules from a specific organism, such as Escherichia coli (E. coli). If this is the case, the first sample of the paired sample can be an E. coli sample from one source, and the second sample of the paired sample can be an E. coli sample from a second source.
[0082] The paired sample can be derived from any source that contains or is suspected of containing the at least one target template nucleic acid molecule. The paired sample can comprise a nucleic acid molecule sample derived from a human, such as a sample extracted from a human patient's skin swab. Alternatively, the paired sample can be derived from other sources, such as a water source. Such a sample may contain billions of template nucleic acid molecules. Using the method of the present invention, each of the target nucleic acid molecules in these billions of target template nucleic acid molecules can be sequenced simultaneously, so there is no upper limit to the number of target template nucleic acid molecules that can be used for the method of the present invention.
[0083] In one embodiment, multiple pairs of samples can be provided. For example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 15, 20, 25, 50, 75 or 100 pairs of samples can be provided. Alternatively, less than 100, less than 75, less than 50, less than 25, less than 20, less than 15, less than 11, less than 10, less than 9, less than 8, less than 7, less than 6, less than 5 or less than 4 samples are provided. Alternatively, 2 pairs to 100 pairs, 2 pairs to 75 pairs, 2 pairs to 50 pairs, 2 pairs to 25 pairs, 5 pairs to 15 pairs or 7 pairs to 15 pairs of samples are provided.
[0084] In the case of providing multiple pairs of samples, different sample labels can be used to label at least one target template nucleic acid molecule in different pairs of samples. For example, if the user intends to provide two pairs of samples, all or substantially all of at least one target template nucleic acid molecule in the first pair of samples can be labeled with sample label A, and all or substantially all of at least one target template nucleic acid molecule in the second pair of samples can be labeled with sample label B. Sample labels are discussed in more detail under the heading "Sample Labels and Barcodes."
[0085] Control the number of target template nucleic acid molecules in the sample
[0086] As described above, the sequencing method of the present invention includes using the information obtained by analyzing the corresponding mutant sequence reads to assemble the sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutant sequence reads. Generally, the target template nucleic acid molecules in the sample can be assembled to generate the sequence of one or more larger nucleic acid molecules present in the sample. In a representative embodiment, the target template nucleic acid molecules can be assembled to generate the sequence of a genome. A certain limited amount of data is generated in the form of sequencing reads obtained by performing a sequencing run. In order to assemble the sequence of the target template nucleic acid molecule from the sequencing reads obtained from the target template nucleic acid molecule (thereby assembling the target template nucleic acid molecule to generate the sequence of one or more larger target template nucleic acid molecules), it is preferred to ensure that the coverage of the target template nucleic acid molecule in the sequencing reads is sufficient (i.e., sufficient to assemble the sequence) without generating excessive redundant (i.e., repeated) sequencing reads for each target template nucleic acid molecule. For example, if the sample contains too many target template nucleic acid molecules, so that a sufficient number of sequencing reads cannot be generated from each target template nucleic acid molecule, the sequence of each target template nucleic acid molecule may not be assembled (i.e., each template may not have enough data). On the other hand, if the sample contains too few target template nucleic acid molecules, although it may be possible to assemble each target template nucleic acid molecule, it may not be possible to assemble the target template nucleic acid molecules to produce the sequence of a larger nucleic acid molecule, for example, it may not be possible to produce a genome sequence (i.e., there may be too much data for each template, and therefore insufficient data for the entire sample).
[0087] In view of these, it is advantageous for the user to be able to control the number of the unique target template nucleic acid molecules present in the first sample of this paired sample and / or the second sample of this paired sample.Then the user can select the optimal number of the unique target template nucleic acid molecules present in the first sample of this paired sample and / or the second sample of this paired sample.The optimal number of unique target template nucleic acid molecules can depend on the many different factors that the user will pay attention to.For example, if the target template nucleic acid molecules are longer, they will be more difficult to order-checking, and the user may wish to select a smaller number of unique target template nucleic acid molecules.
[0088] Thus, the method of the present invention may comprise the step of providing a pair of samples, each sample comprising at least one target template nucleic acid molecule, the step comprising controlling the number of target template nucleic acid molecules in the first sample and / or the second sample of the pair of samples.
[0089] The number of the target template nucleic acid molecules in the first sample of the paired sample may be useful. However, it is particularly preferred that, for the second sample of the paired sample (that is, a sample comprising at least one target template nucleic acid molecule to which a mutation will be introduced), the number of the target template nucleic acid molecules in the second sample of the paired sample is controlled. In the method of the present invention, at least one target template nucleic acid molecule in the second sample of the paired sample is mutated and used to reconstruct the sequence of the target template nucleic acid molecule. In this case, the number of the target template nucleic acid molecules in the second sample of the paired sample may be crucial. Therefore, it may be particularly advantageous to control the number of the target template nucleic acid molecules in the second sample of the paired sample.
[0090] Similarly, in one aspect of the present invention, there is provided a method for determining the sequence of at least one target template nucleic acid molecule, the method comprising:
[0091] (a) providing at least one sample, the at least one sample comprising at least one target template nucleic acid molecule;
[0092] (b) sequencing a region of at least one target template nucleic acid molecule; and
[0093] (c) assembling the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule,
[0094] Wherein the step of providing at least one sample comprising at least one target template nucleic acid molecule comprises: controlling the number of target template nucleic acid molecules in the at least one sample.
[0095] Similarly, in one aspect of the present invention, there is provided a method for determining the sequence of at least one target template nucleic acid molecule, the method comprising:
[0096] (a) providing at least one sample, the at least one sample comprising at least one target template nucleic acid molecule;
[0097] (b) sequencing at least a portion of a region of at least one target template nucleic acid molecule; and
[0098] (c) assembling the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule,
[0099] Wherein the step of providing at least one sample comprising at least one target template nucleic acid molecule comprises: controlling the number of target template nucleic acid molecules in the at least one sample.
[0100] For the purposes of this application, the phrase "controlling" the number of "target template nucleic acid molecules" in a sample refers to providing a desired number of target template nucleic acid molecules in the sample. According to certain specific embodiments, this may include manipulating or adjusting the sample so that it contains a desired number of target template nucleic acid molecules (e.g., by diluting the sample or combining the sample with another sample that also contains target template nucleic acid molecules).
[0101] It should be understood that "controlling the number of target template nucleic acid molecules" may not be completely precise because, for example, it is difficult to obtain an exact number of template nucleic acid molecules by diluting a sample using conventional techniques. However, if the user finds that the sample contains approximately twice the desired target template nucleic acid molecules, the user can dilute the sample and obtain a diluted sample containing approximately half the number of target template nucleic acid molecules present in the original sample (e.g., 45% to 55% of the number of target template nucleic acid molecules present in the original sample).
[0102] Controlling the number of target template nucleic acid molecules can include measuring the number of target template nucleic acid molecules in a sample (e.g., a user can measure the number of target template nucleic acid molecules in a first sample of a paired sample, a second sample in a paired sample, or at least one sample). The term "measuring" can be replaced by the term "estimating" in this article. Typically, measuring the number of target template nucleic acid molecules in a sample is used as a part of the step of controlling the number of target template nucleic acid molecules in a sample, and the step of controlling the number of target template nucleic acid molecules in a sample can be used to help the user ensure that the sample comprises several target template nucleic acid molecules that are suitable (i.e., within the required range) for a specific sequencing method. However, this step of controlling the number of target template nucleic acid molecules does not require complete accuracy. The method for approximating the number of target template nucleic acid molecules in a control sample will contribute to improving the method for sequencing the target template nucleic acid molecules. In one embodiment, "measuring the number of target template nucleic acid molecules" refers to determining the number of target template nucleic acid molecules in a sample to be at least within the correct order of magnitude, i.e., within 10 times, or more preferably within 5 times, 4 times, 3 times, or 2 times, compared to the true number. More preferably, the number of target template nucleic acid molecules in a sample can be determined to be within at least 50%, or at least 40%, or at least 30%, or at least 25%, or at least 20%, or at least 15%, or at least 10% of the actual number. Any method can be used to measure the number of target template nucleic acid molecules in a sample.
[0103] The sample (e.g., the first sample of a paired sample, the second sample of a paired sample, or at least one sample) can be diluted before or during the measurement of the number of target template nucleic acid molecules in the sample. For example, if the user believes that the sample contains a large number of target template nucleic acid molecules, he may wish to dilute the sample to obtain a sample with an appropriate number of target template nucleic acid molecules, thereby accurately measuring by, for example, sequencing. Therefore, a diluted sample can be provided. Therefore, the number of target template nucleic acid molecules can be measured in the diluted sample to determine the number of target template nucleic acid molecules in the sample.
[0104] According to certain embodiments, it may be advantageous to prepare more than one diluted sample, each sample having a different dilution factor. For example, if the user does not know how many target template nucleic acid molecules are present in the sample, he may wish to prepare a dilution series and measure the number of target template nucleic acid molecules in each dilution (i.e., each diluted sample). Therefore, measuring the number of target template nucleic acid molecules can include preparing a dilution series of a first sample of a paired sample, a second sample of the paired sample, or at least one sample to provide a dilution series comprising a diluted sample. The dilution series can include 1 to 50, 1 to 25, 1 to 20, 1 to 15, 1 to 10, 1 to 5 diluted samples, 5 to 25, 5 to 20, 5 to 15, or 5 to 10 diluted samples.
[0105] Such a dilution series can be prepared by performing a serial dilution. Alternatively, the sample can be diluted 2-fold to 20-fold, 5-fold to 15-fold, or approximately 10-fold. For example, to obtain a dilution series of 10 samples, each diluted 10-fold, the user would prepare a 10-fold dilution of the sample, then separate a portion of the diluted sample, then dilute it 10-fold again, and so on, until 10 diluted samples are obtained.
[0106] A user can prepare 10 dilution samples but only determine the number of target template nucleic acid molecules in less than 10 dilution samples. For example, if a user determines the number of target template nucleic acid molecules in 5 dilution samples and accurately determines the number of target template nucleic acid molecules in the fifth dilution sample, there is no need to further determine the number of target template nucleic acid molecules in any other dilution sample. In other embodiments, the user can correlate the results from multiple dilution samples to be more confident in the results. Advantageously, this can also provide the user with information about the dynamic range over which the number of target template nucleic acid molecules in a sample can be accurately determined under a given set of conditions. However, the user can only perform a single dilution to accurately determine the number of target template nucleic acid molecules in a sample.
[0107] According to some specific embodiments, the number of target template nucleic acid molecules in the sample (or diluted sample) can be measured by determining the molar concentration of the target template nucleic acid molecules in the sample. This can, for example, be accomplished by electrophoresis. According to a specific embodiment, the number of target template nucleic acid molecules in the sample can be determined by high-resolution microfluidic electrophoresis (Highresolution microfluidic electrophoresis), thus the sample can be loaded into a microchannel, and the target template nucleic acid molecules can be separated by electrophoresis, and detected by its fluorescence. The suitable system for measuring the number of target template nucleic acid molecules in this way includes Agilent 2100Bioanalyzer and Agilent 4200Tapestation.
[0108] In alternative embodiments, the number of target template nucleic acid molecules can be measured by sequencing the target template nucleic acid molecules in a first sample of a paired sample, a second sample of the paired sample, at least one sample, or one or more diluted samples.
[0109] According to a specific embodiment, the method can include measuring the number of target template nucleic acid molecules by sequencing the target template nucleic acid molecules in one or more diluted samples.
[0110] The target template nucleic acid can be sequenced using any sequencing method. Examples of possible sequencing methods include: Maxam Gilbert Sequencing, Sanger Sequencing, sequencing comprising bridge amplification (e.g., bridge PCR), or any high throughput sequencing (HTS) method, such as described in Maxam AM, Gilbert W (February 1977), "A new method for sequencing DNA", Proc. Natl. Acad. Sci. USA 74(2):560-4; Sanger F, Coulson AR (May 1975), "A rapid method for determining sequences in DNA by primed synthesis with DNA polymerase", J. Mol. Biol. 94(3):441-8; and Bentley DR, Balasubramanian S, et al. (2008), "Accurate whole human genome sequencing using reversible terminator chemistry", Nature, 456(7218):53-59.
[0111] Measuring the number of target template nucleic acid molecules can include: amplifying the target template nucleic acid molecules in the first sample of the paired samples, the second sample of the paired samples, at least one sample, or one or more diluted samples, and then sequencing them (or from another perspective, the amplified target template nucleic acid molecules). Amplifying the target template nucleic acid molecules provides the user with multiple copies of the target template nucleic acid molecules, allowing the user to sequence the target template nucleic acid molecules more accurately (because sequencing technology is not completely accurate, sequencing multiple copies of the target template nucleic acid sequence and then calculating the consensus sequence from the sequences of these copies improves accuracy). Preparing multiple copies of a fixed number of unique target template nucleic acid molecules in the sample and sequencing a portion of the entire (amplified) sample can obtain sequence information from all target template nucleic acid molecules.
[0112] Suitable methods for amplifying at least one target template nucleic acid molecule are known in the art. For example, PCR is typically used. PCR is described in more detail below under the heading "Introducing mutations into at least one target template nucleic acid molecule."
[0113] In a typical embodiment, the sequencing step can involve bridge amplification. Optionally, the bridge PCR step is performed using an extension time of greater than 5 seconds, greater than 10 seconds, greater than 15 seconds, or greater than 20 seconds. An example of using bridge PCR is in an Illumina genome analysis sequencer. Preferably, paired-end sequencing is used.
[0114] Measuring the number of target template nucleic acid molecules can include fragmenting the target template nucleic acid molecules in the first sample of the paired sample, the second sample of the paired sample, at least one sample, or one or more diluted samples. For example, this may be particularly advantageous in cases where the sequencing platform excludes the use of long nucleic acid molecules as templates. Fragmentation can be performed using any suitable technique. Fragmentation can be performed using restriction digestion or using PCR with primers complementary to at least one internal region of at least one mutated target nucleic acid molecule. Preferably, fragmentation is performed using a technique that produces arbitrary fragments. The term "arbitrary fragment" refers to randomly generated fragments, such as fragments generated by tag fragmentation. Fragments generated using restriction endonucleases are not "arbitrary" because restriction digestion occurs at a specific DNA sequence defined by the restriction endonuclease used. Even more preferably, fragmentation is performed by tag fragmentation. If fragmentation is performed by tag fragmentation, the tag fragmentation reaction optionally introduces an adapter region into the at least one mutated target nucleic acid molecule. The adapter region is a short DNA sequence that can encode, for example, an adapter to allow sequencing of at least one target nucleic acid molecule using Illumina technology.
[0115] In a specific embodiment, measuring the number of target template nucleic acid molecules includes: amplifying and fragmenting the target template nucleic acid molecules in the first sample of a paired sample, the second sample of the paired sample, at least one sample, or one or more diluted samples, and then sequencing the target template nucleic acid molecules (or from another perspective, the amplified and fragmented target template nucleic acid molecules). Amplification and fragmentation can be performed in any order before sequencing. In one embodiment, measuring the number of target template nucleic acid molecules includes: amplifying the target template nucleic acid molecules in the first sample of a paired sample, the second sample of the paired sample, at least one sample, or one or more diluted samples, and then fragmenting and then sequencing. Alternatively, measuring the number of target template nucleic acid molecules includes: fragmenting the target template nucleic acid molecules in the first sample of a paired sample, the second sample of the paired sample, at least one sample, or one or more diluted samples, and then amplifying and then sequencing. Alternatively, amplification and fragmentation can be performed simultaneously, i.e., in a single step. When the target template nucleic acid molecules are very long (e.g., too long to be sequenced using conventional techniques), it is useful for the method to fragment the target template nucleic acid molecules and then amplify them.
[0116] Measuring the number of target template nucleic acid molecules can include identifying the total number of target template nucleic acid molecules in a sample. However, preferably, measuring the number of target template nucleic acid molecules includes identifying the number of unique target template nucleic acid molecule sequences in a first sample of a paired sample, a second sample of the paired sample, at least one sample, or one or more diluted samples. As described above, when the at least one target template nucleic acid sequence is part of a sample comprising many different target template nucleic acid sequences, determining the sequence of the at least one target template nucleic acid sequence is more difficult. Therefore, reducing the number of unique target template nucleic acid molecules makes the method of determining the sequence of the at least one target template nucleic acid molecule simpler.
[0117] As discussed elsewhere herein, introducing mutations into a target template nucleic acid sequence can facilitate assembly of at least a portion of the target template nucleic acid sequence. For example, mutating a target template nucleic acid molecule can be particularly beneficial when identifying whether sequence reads are likely to originate from the same target template nucleic acid molecule, or whether sequence reads are likely to originate from different target template nucleic acid molecules. Thus, according to certain embodiments of this aspect of the invention, where the number of target template nucleic acid molecules is measured by sequencing, introducing mutations into the target template nucleic acid molecule can be beneficial. Thus, in certain such embodiments, measuring the number of target template nucleic acid molecules can comprise mutating the target template nucleic acid molecule.
[0118] The target template nucleic acid molecule can be mutated by any convenient means. In particular, the target template nucleic acid molecule can be mutated as described elsewhere herein. According to a particularly preferred embodiment, mutations can be introduced by using a low-bias DNA polymerase. In other or alternative embodiments, the target template nucleic acid molecule can be mutated in the presence of nucleotide analogs, such as amplification of the target template nucleic acid molecule in the presence of dPTP.
[0119] According to a preferred embodiment, measuring the number of target template nucleic acid molecules may include:
[0120] (ii) mutating the target template nucleic acid molecule to provide a mutated target template nucleic acid molecule;
[0121] (ii) sequencing a region of the mutated target template nucleic acid molecule; and
[0122] (iii) identifying the number of unique mutated target template nucleic acid molecules based on the number of unique mutated target template nucleic acid molecule sequences.
[0123] In order to quantify the number of target template nucleic acid molecules in a sample, the user does not need the complete sequence of each target template nucleic acid molecule. Specifically, what is needed is enough information about the sequences of the different target template nucleic acid molecules (or, where applicable, the amplified and fragmented target template nucleic acid molecules) in the sample to allow the user to estimate the total number of target template nucleic acid molecules and / or the number of unique target template nucleic acid molecules. Therefore, the user can choose to sequence only the region of each target template nucleic acid molecule. For example, in certain embodiments, the user can choose to sequence the terminal region of each unique target template nucleic acid molecule or fragmented target template nucleic acid molecule as part of the step of measuring the number of unique target template nucleic acid molecules. Therefore, the user can sequence the 3' terminal region and / or 5' terminal region of the target template nucleic acid molecule or fragmented target template nucleic acid molecule as part of the step of measuring the number of target template nucleic acid molecules. The terminal region of the target template nucleic acid molecule comprises a continuous segment of nucleotides of the desired length and the terminal (e.g., 5' terminal or 3' terminal) nucleotide in the target template nucleic acid molecule and the nucleotide at the most 5' end or the most 3' end in the target template nucleic acid molecule.
[0124] According to certain representative embodiments, measuring the number of target template nucleic acid molecules can include introducing a barcode (also referred to herein as a unique molecular tag or unique molecular identifier, as described below) or a pair of barcodes into the target template nucleic acid molecule (or in other words, labeling the target template nucleic acid molecule with a barcode or a pair of barcodes) to provide barcoded target template nucleic acid molecules. As described elsewhere herein, the barcodes are suitably degenerate and substantially each target template nucleic acid molecule can comprise a unique or substantially unique sequence, such that each (or substantially each) target template nucleic acid molecule is labeled with a different barcode sequence. The barcodes can be introduced into the target template nucleic acid molecules as described elsewhere herein. In specific embodiments, the barcode sequence can be introduced into the end of the target template nucleic acid molecule, i.e., as an additional 5' end introduced into the 5' end (or 5'-most) nucleotide or as an additional 3' end introduced into the 3' end (or 3'-most) nucleotide in the target template nucleic acid molecule.
[0125] In a preferred embodiment, target template nucleic acid molecules labeled with a barcode sequence can be sequenced to measure the number of target template nucleic acid molecules in a sample. More particularly, regions of target template nucleic acid molecules comprising a barcode sequence can be sequenced to measure the number of target template nucleic acid molecules in a sample. The barcode sequence is substantially unique, and the target template nucleic acid molecules are labeled with the barcode sequence, thereby introducing a substantially unique (and therefore countable) sequence into the target template nucleic acid molecules. Therefore, the number of unique barcodes identified by sequencing according to this embodiment can allow the number of unique target template nucleic acid molecules in a sample to be determined.
[0126] Thus, according to certain embodiments, measuring the number of target template nucleic acid molecules may include:
[0127] (i) sequencing a region of a barcoded target template nucleic acid molecule, the barcoded target template nucleic acid molecule comprising a barcode or a pair of barcodes; and
[0128] (ii) identifying the number of unique barcoded target template nucleic acid molecules based on the number of unique barcodes or pairs of barcodes.
[0129] In another embodiment, one or more bar codes may not be used to determine the number of the target template nucleic acid molecules present in the sample. In a specific representative embodiment, the number of target template nucleic acid molecules can be determined by sequencing the terminal region of the target template nucleic acid molecules. Alternatively, the number of the unique terminal sequences that the user identifies exists is then determined, and / or the user is for example drawing a spectrum with reference to genome with respect to the sequence of the terminal region. In the case where theory is not desired to be bound, it is believed that this method can allow determining the number of target template nucleic acid molecules, because the sequence of each target template nucleic acid molecule can originate from the different sites in the reference sequence.
[0130] Furthermore, the sequencing step according to this aspect of the invention can be a "coarse" sequencing step, in that the user may not require precise sequence information in order to be able to measure the number of target template nucleic acid molecules in a sample. As a representative example, the sequencing step can be performed on a poorly amplified set of molecules, which can allow the step to be performed more quickly and / or at a lower cost.
[0131] Alternatively, measuring the number of unique target template nucleic acid molecules in a sample can include sequencing terminal regions of a barcoded target template nucleic acid molecule comprising a barcode or a pair of barcodes. Thus, reference to sequencing terminal regions of a target template nucleic acid molecule can include sequencing terminal regions of a barcoded target template nucleic acid molecule that can comprise a barcode or a pair of barcodes.
[0132] Once the number of unique target template nucleic acid molecules in the sample is determined, the sample can be adjusted to control the number of target template nucleic acid molecules in the sample so that the sample contains a desired number of unique target template nucleic acid molecules. According to certain embodiments, this can include a step of diluting the sample. Thus, controlling the number of target template nucleic acid molecules in the sample can include measuring the number of target template nucleic acid molecules in the sample and diluting the sample so that the sample contains a desired number of target template nucleic acid molecules.
[0133] As described above, the sample according to this aspect of the invention may be any sample, and in particular may be the first sample or the second sample according to the method of the invention. Thus, according to a particular embodiment, controlling the number of target template nucleic acid molecules in the first sample of the paired sample and / or the second sample of the paired sample comprises: measuring the number of target template nucleic acid molecules and diluting the first sample of the paired sample and / or the second sample of the paired sample such that the first sample of the paired sample and / or the second sample of the paired sample contain the desired number of target template nucleic acid molecules.
[0134] Combine subsamples to provide a sample
[0135] A sample can be provided by combining several subsamples. This can allow target template nucleic acid molecules from multiple samples (e.g., from multiple sources) to be sequenced simultaneously, which in turn can achieve greater sample throughput, thereby reducing the cost and time required to determine the sequence of the target template nucleic acid molecule.
[0136] Therefore, the method of the present invention can be performed on a sample provided by merging two or more subsamples. According to certain embodiments, a first sample in a paired sample can be provided by merging two or more subsamples. In another embodiment, a second sample in the paired sample can be provided by merging two or more subsamples. Therefore, a first sample and / or a second sample can be provided by merging two or more subsamples. Alternatively, a first sample and a second sample can be obtained from the merged sample and subjected to the method of the present invention.
[0137] Thus, this aspect of the invention allows the sequence of at least one target template nucleic acid molecule from each of two or more smaller samples that are combined to provide the sample to be assayed.
[0138] One problem associated with pooled samples for sequencing is that each sample may contain a different number of target nucleic acid molecules. Therefore, it may be beneficial for the pooled sample to contain target template nucleic acid molecules from each of its constituent subsamples in a desired amount, more specifically in a desired ratio. In other words, it may be beneficial for the pooled sample to contain an appropriate number (i.e., within a desired range) of unique target template nucleic acid molecules from each of its subsamples, so that a specific sequencing method can be used to sequence the target template nucleic acid molecules from each subsample in the pooled sample.
[0139] As a representative example, two separate subsamples may be provided: sample Y and sample Z. If the total number of target template nucleic acid molecules in sample Y is 100 times the total number of target template nucleic acid molecules in sample Z, then combining equal amounts of sample Y and sample Z and performing a sequencing method on the combined sample is expected to result in the number of sequencing reads generated for the target template nucleic acid molecules in sample Y being 100 times the number of sequencing reads generated for the target template nucleic acid molecules in sample Z. Therefore, combining samples in this manner may not only result in insufficient sequencing reads being generated for sample Z to enable the sequence assembly step to be performed using the sequence reads obtained from sample Z, but it may also complicate the sequence assembly step performed on the sequence reads obtained from sample Y.
[0140] Thus, the methods of the present invention may include the step of normalizing the number of target template nucleic acid molecules in each pooled subsample to provide the first sample of the paired samples and / or the second sample of the paired samples.
[0141] More generally, however, the present invention provides a method for determining the sequence of at least one target template nucleic acid molecule, the method comprising:
[0142] (a) providing at least one sample, the at least one sample comprising at least one target template nucleic acid molecule;
[0143] (b) sequencing a region of at least one target template nucleic acid molecule; and
[0144] (c) assembling the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule,
[0145] Wherein at least one sample is provided by combining two or more subsamples, and the number of target template nucleic acid molecules in each subsample is normalized.
[0146] For the purposes of this application, the phrases "normalizing the number of target template nucleic acid molecules in each subsample" and "normalizing the number of target template nucleic acid molecules in each combined subsample" refer to combining the subsamples in such a manner as to provide the total number of target template nucleic acid molecules in the combined sample derived from each subsample in the desired amount. In some embodiments, the number of unique target template nucleic acid molecules is normalized. "Unique target template nucleic acid molecules" are target template nucleic acid molecules comprising different nucleic acid sequences. Optionally, each target template nucleic acid molecule in at least one target template nucleic acid molecule is a unique target template nucleic acid molecule. Unique target template nucleic acid molecules can differ in sequence by only a single nucleotide, or can be substantially different from each other.
[0147] The normalization step can advantageously allow the number of target template nucleic acid molecules from each subsample to be provided in a desired ratio. According to certain embodiments, this can include manipulating or adjusting each subsample so that when combined, the combined sample contains the desired number of target template nucleic acid molecules from each subsample. From another perspective, it can be seen that this step allows the number of target template nucleic acid molecules in the combined sample from each of two or more subsamples to be controlled, or the number of target template nucleic acid molecules in at least one sample from two or more subsamples to be controlled.
[0148] In another aspect, the present invention therefore provides a method for determining the sequence of at least one target template nucleic acid molecule, the method comprising:
[0149] (a) providing at least one sample, the at least one sample comprising at least one target template nucleic acid molecule;
[0150] (b) sequencing a region of at least one target template nucleic acid molecule; and
[0151] (c) assembling the sequence of at least one target template nucleic acid molecule from the sequence of a region of the at least one target template nucleic acid molecule, wherein the step of providing at least one sample comprising the at least one target template nucleic acid molecule comprises combining two or more subsamples and controlling the number of target template nucleic acid molecules in at least one sample from the two or more subsamples.
[0152] According to certain embodiments, standardizing the number of target template nucleic acid molecules in each subsample can include providing a similar number of target template nucleic acid molecules in the combined sample from each subsample (i.e., in a ratio of about 1: 1). This embodiment may be particularly useful, for example, in cases where each subsample is derived from a sample comprising a genome of similar size. However, in alternative embodiments, the number of target template nucleic acid molecules can be provided in different amounts, i.e., the number of target template nucleic acid molecules from the first subsample can be provided at a higher abundance than the number of target template nucleic acid molecules from the second subsample. For example, if the first subsample is derived from a larger genome and the second subsample is derived from a sample comprising a smaller genome, this embodiment may be desirable.
[0153] It will be understood that "normalizing the number of target template nucleic acid molecules in each pooled subsample" may not be completely accurate because, for example, it may be difficult to measure the number of target template nucleic acid molecules in each subsample. However, if the user finds that the subsamples contain approximately twice the desired target template nucleic acid molecules, the user can normalize the number of target template nucleic acid molecules in the subsamples so that the number of target template nucleic acid molecules in the pooled sample is approximately half the number of target template nucleic acid molecules present in the subsamples (e.g., 45% to 55% of the number of target template nucleic acid molecules present in the subsamples).
[0154] In a broad sense, normalizing the number of target template nucleic acid molecules in each subsample can be considered equivalent to controlling the number of target template nucleic acid molecules in each subsample provided in the combined sample. Thus, normalizing the number of target template nucleic acid molecules can include measuring the number of target template nucleic acid molecules in each subsample.
[0155] According to certain embodiments, as described elsewhere herein, particularly in the context of methods for controlling the number of target template nucleic acid molecules in a sample, the number of target template nucleic acid molecules in a subsample can be measured.
[0156] In a preferred embodiment, standardizing the number of target template nucleic acid molecules in each subsample can include labeling target template nucleic acid molecules from different subsamples with different sample tags. A sample tag is a tag used to label most or all of the target template nucleic acid molecules of at least one target template nucleic acid molecule in a sample. Labeling target template nucleic acid molecules in different subsamples with different sample tags can allow for distinguishing between template target nucleic acid molecules originating from different subsamples. Sample tags can therefore be particularly useful in this aspect of the present invention because their use can allow for simultaneous measurement of the number of target template nucleic acid molecules in each of two or more subsamples. In particular, sample tags can allow for measurement of the number of target template nucleic acid molecules in each of two or more subsamples in a single sample. Preferably, the target template nucleic acid molecules can be labeled with the sample tags before the subsamples are combined. Therefore, in a specific embodiment, the present aspect of the present invention can include: preparing preliminary pools of subsamples, each preliminary pool of subsamples comprising target template nucleic acid molecules labeled with a sample tag; and measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pools.
[0157] Viewed from another perspective, the present invention provides a method for measuring the number of target template nucleic acid molecules in two or more subsamples, the method comprising:
[0158] (a) labeling target template nucleic acid molecules from two or more different subsamples with different sample tags;
[0159] (b) combining two or more subsamples to provide a pre-pool of subsamples; and
[0160] (c) Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool.
[0161] Alternatively, two or more preliminary pools may be prepared, eg, each preliminary pool comprising subsamples provided in different amounts or ratios, and / or consisting of different subsamples (eg, different combinations of subsamples).
[0162] According to certain embodiments, the techniques described elsewhere herein for measuring the number of target template nucleic acid molecules in a sample can be used to measure the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool (particularly when the number of target template nucleic acid molecules in the sample is controlled). In this regard, the skilled artisan will understand that the target template nucleic acid molecules from each sample are distinguishable based on the sample tags contained in each sample, and therefore measuring the number of target template nucleic acid molecules labeled with any given sample tag in the preparatory pool can be performed by adapting the method to measure the total number of target template nucleic acid molecules present in a particular sample.
[0163] In this regard, according to certain embodiments, the preparatory pool can be diluted before or during the measurement of the number of target template nucleic acid molecules labeled with each sample tag. The dilution can be performed as described elsewhere herein. For example, in certain embodiments, the preparatory pool can be serially diluted to provide a series of dilutions comprising the diluted preparatory pool.
[0164] As mentioned elsewhere, two or more different pre-pools can be prepared. Each pre-pool can be diluted to a different extent, for example according to a different serial dilution.
[0165] According to a particularly preferred embodiment, the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool can be measured by sequencing the (sample-tagged) target template nucleic acid molecules labeled in the preparatory pool or the diluted preparatory pool. Sequencing can be performed according to any convenient sequencing method, such as the method described elsewhere herein. Preferably, sequencing the labeled target template nucleic acid molecules can include sequencing the sample tags of the labeled target template nucleic acid molecules.
[0166] In a specific embodiment, measuring the number of target template nucleic acid molecules labeled with each sample label in the preliminary pool may include an amplification step. Suitable methods for amplifying labeled target template nucleic acid molecules are known in the art and, for example, can be amplified as described elsewhere herein. In certain embodiments, measuring the number of target template nucleic acid molecules labeled with each sample label in the preliminary pool may include amplifying the target template nucleic acid molecules and then sequencing them.
[0167] In certain embodiments, the target template nucleic acid molecules in the subsamples can be amplified, i.e., before combining two or more subsamples to provide a combined preliminary sample. Amplification can be performed before labeling the target template nucleic acid molecules in the subsamples with a sample label, or in certain preferred embodiments, can be performed simultaneously with labeling the target template nucleic acid molecules in the subsamples with a sample label (e.g., using PCR primers comprising a sample barcode). In other embodiments, the target template nucleic acid molecules labeled with the sample label can be amplified before providing the combined preliminary sample.
[0168] According to another embodiment, measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool may include: amplifying the target template nucleic acid molecules labeled with the sample tags in the preparatory pool, ie, after combining two or more subsamples.
[0169] Optionally, two or more amplification steps can be performed, such as a first amplification performed before or simultaneously with labeling the target template nucleic acid molecules in the subsample with the sample label, and a second amplification to amplify the target template nucleic acid molecules labeled with the sample label (as described above, the second amplification can be performed on the subsample or the pooled preliminary sample).
[0170] After amplification, measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool can include sequencing the target template nucleic acid molecules labeled with each sample tag (i.e., target template nucleic acid molecules labeled with the sample tags) in the preparatory pool or the diluted preparatory pool. In a preferred embodiment, measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool can therefore include amplifying and then sequencing the target template nucleic acid molecules labeled with each sample tag in the preparatory pool or the diluted preparatory pool.
[0171] Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool can include a fragmentation step. Preferably, after preparing the combined sample, the target template nucleic acid molecules in the combined sample are fragmented. Fragmentation can be performed using any suitable technique, including any technique described elsewhere herein.
[0172] In a specific embodiment, measuring the number of target template nucleic acid molecules labeled with each sample tag can include: performing an amplification and fragmentation step before sequencing the target template nucleic acid molecules in the preliminary pool or the diluted preliminary pool. According to a preferred embodiment, therefore, before merging two or more sub-samples to provide a combined preliminary sample and sequencing the target template nucleic acid, the target nucleic acid molecules in the sub-samples can be amplified, fragmented and labeled with the sample tags. Amplification and fragmentation can be performed in any order. In one embodiment, the target template nucleic acid molecules in the sub-samples can be amplified and then fragmented, or first fragmented and then amplified and then labeled with the sample tags. In a further embodiment, the target template nucleic acid molecules can be amplified, fragmented and labeled simultaneously, i.e., in a single step. A particularly preferred method for amplifying, fragmenting and labeling the target template nucleic acid molecules in a single step can be performed using tag fragmentation and PCR, in particular using PCR primers comprising sample tags. Therefore, the amplified and fragmented target nucleic acid molecules after this step will be labeled with the sample tags and, once incorporated into the combined preliminary sample, can be identified as originating from a specific sub-sample, for example, when sequencing.
[0173] Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool may include identifying the number of target template nucleic acid molecules (optionally unique target template nucleic acid molecules) with each sample tag (i.e., labeled with each sample tag) in the preparatory pool (or diluted preparatory pool). However, preferably, measuring the number of target template nucleic acid molecules with each sample tag includes identifying the number of unique target template nucleic acid sequences with each sample tag in the preparatory pool (or diluted preparatory pool).
[0174] As discussed elsewhere, mutating target template nucleic acid molecules can be particularly beneficial, for example, in identifying whether sequence reads are likely to originate from the same target template nucleic acid molecule or different target template nucleic acid molecules. Thus, this can be beneficial in determining the number of target template nucleic acid molecules in a preparative pool that originate from a particular subsample.
[0175] Therefore, according to certain embodiments, measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool (or diluted preliminary pool) can include mutating the target template nucleic acid molecules. In certain embodiments, the target template nucleic acid molecules in the preliminary sample merged can be mutated. However, mutating the target template nucleic acid molecules can preferably occur in a subsample, i.e., before merging two or more samples to provide the merged sample. In a particularly preferred embodiment, the target template nucleic acid molecules can be mutated before or simultaneously with the target template nucleic acid molecules labeled with the sample tag. It may be preferred not to mutate the sample tag sequence used to label the target template nucleic acid molecules. Target template nucleic acid molecules can be mutated in any convenient manner, including any manner described elsewhere herein. Therefore, in one embodiment, mutations can be introduced using a low-bias DNA polymerase. In other embodiments, mutating the target template nucleic acid molecules can be included in the presence of nucleotide analogs, such as amplifying the target template nucleic acid molecules in the presence of dPTP.
[0176] According to a specific embodiment, measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparation pool may include:
[0177] (i) mutating the target template nucleic acid molecule to provide a mutated target template nucleic acid molecule;
[0178] (ii) sequencing a region of the mutated target template nucleic acid molecule; and
[0179] (iii) identifying the number of target template nucleic acid molecules having a unique mutation for each sample tag based on the number of target template nucleic acid molecules having a unique mutation labeled with each sample tag.
[0180] As outlined in more detail above, in order to quantify target template nucleic acid molecules, it is not necessary to obtain the complete sequence of each target template nucleic acid molecule. Simply sequencing the terminal region of each labeled target template nucleic acid molecule as part of the step of measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool may be sufficient. Therefore, the user can choose to sequence only the terminal region of each target template nucleic acid molecule. As described above, the sample tags will preferably be sequenced.
[0181] According to certain representative embodiments, measuring the number of target template nucleic acid molecules can include introducing a barcode or a pair of barcodes into the target template nucleic acid molecules to provide barcoded, sample-tagged target template nucleic acid molecules. Barcodes suitable for this step and methods for introducing barcodes into target template nucleic acid molecules are described in more detail elsewhere herein.
[0182] Preferably, the barcode can be introduced into the target template nucleic acid molecule before merging the subsamples, that is, before merging the subsamples to provide a temporary merged sample. The barcode and the sample label can be introduced into the target template nucleic acid molecule in any order. For example, in one embodiment, the barcode can be introduced into the target template nucleic acid molecule and the sample label is subsequently introduced. In another embodiment, the sample label can be introduced into the target template nucleic acid molecule and the barcode is subsequently introduced. In other embodiments, the sample label and the barcode label can be introduced simultaneously. In any case, in certain embodiments, the target template nucleic acid molecule from the subsample can be marked simultaneously with the sample label and the barcode. In this respect, it should be noted that the sample label is particularly advantageous for identifying the specific target template nucleic acid molecule in the preparatory sample as originating from a specific subsample, while the barcode may be particularly advantageous for allowing measurement of the number of unique target template nucleic acid molecules from each subsample.
[0183] Therefore, according to a particularly preferred embodiment, measuring the number of target template nucleic acid molecules labeled with each sample tag may include:
[0184] (i) sequencing a region of the barcoded, sample-tagged target template nucleic acid molecule; and
[0185] (ii) identifying the number of unique barcoded target template nucleic acid molecules having each sample tag based on the number of unique barcodes or paired barcode sequences associated with each sample tag.
[0186] As described elsewhere herein, the sequencing step that measures the number of target template nucleic acid molecules can be a "coarse" sequencing step, as the user may not need precise sequence information to be able to measure the number of target template nucleic acid molecules in a sample. Instead, it is sufficient for sequencing to identify sample tags, barcodes, and / or target template nucleic acid molecules.
[0187] In certain representative embodiments, once the number of target template nucleic acid molecules comprising different sample tags has been measured, the ratio of the number of target template nucleic acid molecules comprising different sample tags can be calculated. In further representative embodiments, once the number of target template nucleic acid molecules comprising different sample tags has been measured, it is possible to determine the number of target template nucleic acid molecules (in the combined preliminary sample) generated by each subsample, thereby calculating the number of target template nucleic acid molecules present in each subsample.
[0188] Information about the ratio of target template nucleic acid molecules containing different sample labels and / or information about the number of target template nucleic acid molecules produced by each subsample can be used to prepare a combined sample for use in the methods of the present invention. In particular, such information can be used in a normalization step to normalize the number of target template nucleic acid molecules provided in each of two or more subsamples in the combined sample, thereby providing the target template nucleic acid molecules from each subsample in the combined sample at a desired ratio.
[0189] Thus, it will be seen that the present invention provides a method for determining the sequence of at least one target template nucleic acid molecule, the method comprising:
[0190] (a) providing at least one sample, the at least one sample comprising at least one target template nucleic acid molecule;
[0191] (b) sequencing a region of at least one target template nucleic acid molecule; and
[0192] (c) assembling the sequence of at least one target template nucleic acid molecule from the sequence of a region of at least one target template nucleic acid molecule, wherein at least one sample is provided by:
[0193] (i) providing a combined preliminary sample by combining two or more subsamples;
[0194] (ii) measuring the number of target template nucleic acid molecules in a combined preliminary sample generated from each of the two or more subsamples; and
[0195] (iii) combining two or more subsamples;
[0196] Wherein the number of target template nucleic acid molecules in the sample from each subsample is normalized.
[0197] As described above, standardizing the number of target template nucleic acid molecules in a sample by merging two or more subsamples can include providing the target template nucleic acid molecules from each subsample in a desired ratio. According to certain embodiments, the sample formed by merging two or more subsamples can be considered to be a re-merged sample, wherein the target template nucleic acid molecules in each subsample are provided in the re-merged sample in a desired ratio (i.e., after providing a preliminary pool and measuring the number of target template nucleic acid molecules in the preliminary pool generated by each of the two or more subsamples). Therefore, measuring the number of target template nucleic acid molecules in the subsamples standardizes the number of target template nucleic acid molecules in the sample from each subsample when the subsamples are re-merged.
[0198] According to current aspect of the present invention, sample can be provided by merging two or more subsamples.Therefore, in order to provide the sample (i.e. the sample merged) for the method of the present invention, 2 or more, preferably 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000 or more subsamples can be merged.According to some embodiments, 2 to 5000, 10 to 1000 or 25 to 150 subsamples can be merged.
[0199] The term "combining two or more subsamples" does not require combining the entire subsample with another subsample to provide a sample, but rather preferably refers to obtaining an aliquot of each subsample and combining the aliquots to provide a sample. Similarly, references to introducing barcodes or tags into target template nucleic acid molecules in a subsample or mutating target template nucleic acid molecules in a subsample are understood to mean performing such steps on an aliquot or portion of a subsample.
[0200] According to certain specific embodiments, "merging two or more subsamples " can include diluting the subsample and combining the diluted subsamples to provide a sample. In a further embodiment, the term can include obtaining aliquots of a sample and diluting the aliquots, and merging the diluted aliquots of the subsamples to provide a sample. Diluting the subsamples (or aliquots) can include a separate dilution step performed before merging the subsamples (or aliquots) to provide a sample. However, it will be appreciated that merging two or more subsamples (or aliquots) to provide a sample may actually reduce the concentration of the target template nucleic acid molecules from each subsample provided in the sample, and may therefore, represent a dilution step. The technician will be able to determine the extent to which each subsample may need to be diluted, including any dilution that may occur due to merging two or more subsamples (or aliquots).
[0201] Sequencing a region of at least one target template nucleic acid molecule or at least one mutated target template nucleic acid molecule
[0202] The method for determining the sequence of at least one target template nucleic acid molecule can include the steps of sequencing a region of at least one target template nucleic acid molecule in a first sample of a paired sample to provide a non-mutated sequence read and / or sequencing a region of at least one mutated target template nucleic acid molecule to provide a mutated sequence read.
[0203] The sequencing step can be performed using any sequencing method. Examples of possible sequencing methods include: Maxam-Gilbert sequencing, Sanger sequencing, sequencing comprising bridge amplification (e.g., bridge PCR), or any high-throughput sequencing (HTS) method, such as described in Maxam AM, Gilbert W (February 1977), "A new method for sequencing DNA", Proc. Natl. Acad. Sci. USA 74(2):560-4; Sanger F, Coulson AR (May 1975), "A rapid method for determining sequences in DNA by primed synthesis with DNA polymerase", J. Mol. Biol. 94(3):441-8; and Bentley DR, Balasubramanian S, et al. (2008), "Accurate whole human genome sequencing using reversible terminator chemistry", Nature, 456(7218):53-59.
[0204] In a typical embodiment, at least one, or preferably both, sequencing steps involve bridge amplification. Optionally, a bridge PCR step is performed using an extension time of greater than 5 seconds, greater than 10 seconds, greater than 15 seconds, or greater than 20 seconds. An example of using bridge PCR is in an Illumina genome analysis sequencer.
[0205] Optionally, step (i) of sequencing a region of at least one target template nucleic acid molecule in the first sample of the paired samples to provide a non-mutated sequence read and step (ii) of sequencing a region of at least one mutated target template nucleic acid molecule to provide a mutant sequence read are performed using the same sequencing method. Alternatively, step (i) of sequencing a region of at least one target template nucleic acid molecule in the first sample of the paired samples to provide a non-mutated sequence read and step (ii) of sequencing a region of at least one mutated target template nucleic acid molecule to provide a mutant sequence read are performed using different sequencing methods.
[0206] Alternatively, the step (i) of sequencing a region of at least one target template nucleic acid molecule in the first sample of the paired sample to provide a non-mutated sequence reading and the step (ii) of sequencing a region of at least one mutated target template nucleic acid molecule to provide a mutated sequence reading can be performed using more than one sequencing method. For example, a first sequencing method can be used to sequence a portion of at least one target template nucleic acid molecule in the first sample of the paired sample, and a second sequencing method can be used to sequence a portion of at least one target template nucleic acid molecule in the first sample of the paired sample. Similarly, a first sequencing method can be used to sequence a portion of at least one mutated target template nucleic acid molecule, and a second sequencing method can be used to sequence a portion of at least one mutated target template nucleic acid molecule.
[0207] Optionally, the step (i) of sequencing the region of at least one target template nucleic acid molecule in the first sample of the paired sample to provide a non-mutated sequence reading and the step (ii) of sequencing the region of at least one mutated target template nucleic acid molecule to provide a mutant sequence reading are performed at different times. Alternatively, steps (i) and (ii) can be performed almost simultaneously, for example within one year relative to each other. The first sample of the paired sample and the second sample of the paired sample do not have to be performed simultaneously with each other. In the case where both samples are from the same organism, they can be provided at substantially different times (even several years apart), so the two sequencing steps may also be separated by several years. In addition, even if the first sample in the paired sample and the second sample in the paired sample are both from the same original sample, the biological samples can be stored for a period of time, and therefore, there is no need to perform the sequencing steps simultaneously.
[0208] The mutant sequence reads and / or non-mutated sequence reads can be single end or paired end sequence reads.
[0209] Alternatively, the mutant sequence reads and / or the non-mutant sequence reads are greater than 50 bp, greater than 100 bp, greater than 500 bp, less than 200,000 bp, less than 15,000 bp, less than 1,000 bp, between 50 bp and 200,000 bp, between 50 bp and 15,000 bp, or between 50 bp and 1,000 bp. The longer the read length, the easier it is to use the information obtained from analyzing the mutant sequence reads to assemble the sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutant sequence reads. For example, if an assembly graph is used, using longer sequence reads will make it easier to identify valid paths that align with the assembly graph. For example, as described in more detail below, identifying valid paths that align with the assembly graph can include identifying characteristic k-mers, and larger read lengths can allow for longer k-mers.
[0210] Optionally, the sequencing step is performed using a sequencing depth of 0.1 to 500 reads, 0.2 to 300 reads, or 0.5 to 150 reads per nucleotide of at least one target template nucleic acid molecule. The greater the sequencing depth, the higher the accuracy of the determined / generated sequence, but assembly may be more difficult.
[0211] Introducing a mutation into at least one target template nucleic acid molecule
[0212] The method may include the step of introducing a mutation into at least one target template nucleic acid molecule in a second sample of the paired samples to provide at least one mutated target template nucleic acid molecule.
[0213] The mutation may be a substitution mutation, an insertion mutation, or a deletion mutation. For the purposes of the present invention, the term "substitution mutation" shall be interpreted as meaning that a nucleotide is replaced by a different nucleotide. For example, the conversion of the sequence ATCC to the sequence AGCC introduces a single substitution mutation. For the purposes of the present invention, the term "insertion mutation" shall be interpreted as meaning that at least one nucleotide is added to the sequence. For example, the conversion of the sequence ATCC to the sequence ATTCC is an example of an insertion mutation (an additional T nucleotide is inserted). For the purposes of the present invention, the term "deletion mutation" shall be interpreted as meaning that at least one nucleotide is removed from the sequence. For example, the conversion of the sequence ATTCC to ATCC is an example of a deletion mutation (T nucleotide is removed). Preferably, the mutation is a substitution mutation.
[0214] The phrase "introducing mutations into at least one target template nucleic acid molecule" refers to exposing at least one target template nucleic acid molecule in the second sample of the paired sample to conditions that cause the at least one target template nucleic acid molecule to mutate. This can be achieved using any suitable method. For example, mutations can be introduced by chemical mutagenesis and / or enzymatic mutagenesis.
[0215] Alternatively, the step of introducing mutations into at least one target template nucleic acid molecule mutates 1% to 50%, 3% to 25%, 5% to 20%, or about 8% of the nucleotides of at least one target template nucleic acid molecule. Alternatively, at least one mutated target template nucleic acid molecule comprises 1% to 50%, 3% to 25%, 5% to 20%, or about 8% mutations.
[0216] The user can determine how many mutations are included in the at least one mutated target template nucleic acid molecule and / or the extent to which the step of introducing the mutation into the at least one target template nucleic acid molecule mutates the at least one target template nucleic acid molecule by performing the following steps: introducing the mutation into a nucleic acid molecule of known sequence, sequencing the resulting nucleic acid molecule, and determining the percentage of the total number of nucleotides that have changed compared to the original sequence.
[0217] Optionally, the step of introducing mutations into at least one target template nucleic acid molecule mutates at least one target template nucleic acid molecule in a substantially random manner. Optionally, at least one mutated target template nucleic acid molecule comprises a substantially random pattern of mutations.
[0218] If at least one mutated target template nucleic acid molecule comprises a substantially random mutation pattern, the at least one mutated target template nucleic acid molecule comprises a substantially random mutation pattern over its entire length. For example, a user can determine whether at least one mutated target template nucleic acid molecule comprises a substantially random mutation pattern by mutating a test nucleic acid molecule of known sequence to provide a mutated test nucleic acid molecule. The sequence of the mutated test nucleic acid molecule can be compared to the test nucleic acid molecule to determine the position of each mutation. The user can then determine whether the mutation occurs at a substantially similar level over the entire length of the mutated test nucleic acid molecule by:
[0219] (i) Calculate the distance between each mutation;
[0220] (ii) calculating the average value of the distances;
[0221] (iii) subsampling the distances without replacement to a smaller number, such as 500 or 1000;
[0222] (iv) constructing simulation sets of 500 or 1000 distances from the geometric distribution whose means, given by the method of moments, match the means previously calculated from the observed distances; and
[0223] (v) Compute the Kolmolgorov-Smirnov from two distributions.
[0224] If D < 0.15, D < 0.2, D < 0.25 or D < 0.3, depending on the length of the non-mutated reads, the at least one mutated target template nucleic acid molecule can be considered to contain a substantially random mutation pattern.
[0225] Similarly, if the resulting at least one mutated target template nucleic acid molecule comprises a substantially random mutation pattern, then the step of introducing mutations into the at least one target template nucleic acid molecule mutates the at least one target template nucleic acid molecule in a substantially random manner. Whether the step of introducing mutations into the at least one target template nucleic acid molecule does in fact mutate the at least one target template nucleic acid molecule in a substantially random manner can be determined by performing the step of introducing mutations into the at least one target template nucleic acid molecule on a test nucleic acid molecule of known sequence to provide a mutated test nucleic acid molecule. The user can then sequence the mutated test nucleic acid molecule to identify which mutations have been introduced and determine whether the mutated test nucleic acid molecule comprises a substantially random mutation pattern.
[0226] Alternatively, the target template nucleic acid molecule of at least one mutation comprises an unbiased mutation pattern. Alternatively, the step of introducing mutation into at least one target template nucleic acid molecule introduces mutation in an unbiased manner. If the mutation type introduced is random, the target template nucleic acid molecule of at least one mutation comprises an unbiased mutation pattern. If the mutation introduced is a substitution mutation, then if similar ratios of A (adenosine) nucleotides, T (thymine) nucleotides, C (cytosine) nucleotides and G (guanine) nucleotides are introduced, the mutation introduced is random. By the phrase "similar ratios of A (adenosine) nucleotides, T (thymine) nucleotides, C (cytosine) nucleotides and G (guanine) nucleotides are introduced", we mean that the number of the introduced adenosine nucleotides, the number of the thymine nucleotides, the number of the cytosine nucleotides and the number of the guanine nucleotides are within 20% of each other (for example, 20 A nucleotides, 18 T nucleotides, 24 C nucleotides and 22 G nucleotides can be introduced).
[0227] Whether the step of introducing mutations into at least one target template nucleic acid molecule actually mutates the at least one target template nucleic acid molecule in an unbiased manner can be determined by performing the step of introducing mutations into at least one target template nucleic acid molecule on a test nucleic acid molecule of known sequence to provide a mutated test nucleic acid molecule. The user can then sequence the mutated test nucleic acid molecule to identify which mutations have been introduced and determine whether the mutated test nucleic acid molecule contains an unbiased mutation pattern.
[0228] It is useful that even if the step of introducing mutation into at least one target template nucleic acid molecule introduces unevenly distributed mutations, a method for producing the sequence of at least one target template nucleic acid molecule can be used. Therefore, in one embodiment, the target template nucleic acid molecule of at least one mutation comprises unevenly distributed mutations. Alternatively, the step of introducing mutation into at least one mutation target template nucleic acid molecule introduces unevenly distributed mutations. If mutations are introduced in a biased manner, it is considered that mutations are "unevenly distributed", that is, the number of adenosine nucleotides, the number of thymine nucleotides, the number of cytosine nucleotides, and the number of guanine nucleotides introduced are not within 20% of each other. Whether the target template nucleic acid molecule of at least one mutation comprises unevenly distributed mutations, or whether the step of introducing mutations into at least one target template nucleic acid molecule introduces unevenly distributed mutations, can be determined in a similar manner to the above-mentioned step of determining whether mutations are introduced into at least one target template nucleic acid molecule in an unbiased manner.
[0229] Similarly, the method of generating a sequence of at least one target template nucleic acid molecule can be used even when the mutant sequence reads and / or the non-mutant sequence reads include an uneven distribution of sequencing errors. Thus, in one embodiment, the mutant sequence reads and / or the non-mutant sequence reads include an uneven distribution of sequencing errors. Similarly, in one embodiment, the steps of sequencing a region of at least one target template nucleic acid molecule and / or sequencing a region of at least one mutant target template nucleic acid molecule introduce an uneven distribution of sequencing errors.
[0230] Whether a particular step of sequencing a region of at least one target template nucleic acid molecule and / or sequencing a region of at least one mutant target template nucleic acid molecule introduces an uneven distribution of sequence errors will likely depend on the accuracy of the sequencing instrument and may be known to the user. However, the user can investigate whether the step of sequencing a region of at least one target template nucleic acid molecule and / or sequencing a region of at least one mutant target template nucleic acid molecule introduces an uneven distribution of sequence errors by performing the sequencing method on a nucleic acid molecule of known sequence and comparing the resulting sequence reads with the sequence reads of the original nucleic acid molecule of known sequence. The user can then apply the probability function discussed in Example 6 to determine the values of M and E. If the value of E and the value of the matrix model are not equal or substantially not equal (within 10% of each other), then the step of sequencing a region of at least one target template nucleic acid molecule introduces an uneven distribution of sequence errors.
[0231] Introducing mutations into at least one target template nucleic acid molecule by chemical mutagenesis can be achieved by exposing at least one target template nucleic acid to a chemical mutagen. Suitable chemical mutagens include mitomycin C (MMC), N-methyl-N-nitrosourea (MNU), nitrous acid (NA), diepoxybutane (DEB), 1,2,7,8-diepoxyoctane (DEO), ethylmethanesulfonic acid (EMS), methylmethanesulfonic acid (MMS), N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), 4-nitroquinoline 1-oxide (4-NQO), 2-methoxy-6-chloro-9(3-[ethyl-2-chloroethyl]-aminopropylamino)-acridine dihydrochloride (ICR-170), 2-aminopurine (2A), bisulfite, and hydroxylamine (HA). For example, when a nucleic acid molecule is exposed to bisulfite, the bisulfite deaminates cytosine to form uracil, effectively introducing a CT substitution mutation.
[0232] As mentioned above, the step of introducing mutation into at least one target template nucleic acid molecule can be carried out by enzymatic mutagenesis.Alternatively, enzymatic mutagenesis is carried out using DNA polymerase.For example, some DNA polymerases are prone to error (being low-fidelity polymerases), and mutation is introduced using an error-prone DNA polymerase to replicate at least one target template nucleic acid molecule.Taq polymerase is an example of low-fidelity polymerase, and can be by using Taq polymerase, for example, by PCR to replicate at least one target template nucleic acid molecule to carry out the step of introducing mutation into at least one target template nucleic acid molecule.
[0233] The DNA polymerase can be a low bias DNA polymerase, which will be discussed in more detail below.
[0234] If a DNA polymerase is used to perform the step of introducing a mutation into at least one target template nucleic acid molecule, the at least one target template nucleic acid molecule can be incubated with the DNA polymerase and appropriate primers under conditions suitable for the DNA polymerase to catalyze the production of at least one mutated target template nucleic acid molecule.
[0235] Suitable primers include short nucleic acid molecules that are complementary to regions flanking the at least one target nucleic acid molecule or to regions flanking a nucleic acid molecule that is complementary to the at least one target nucleic acid molecule. For example, if the at least one target template nucleic acid molecule is part of a chromosome, the primers will be complementary to the following regions: a region of the chromosome immediately adjacent to the 3' end of the at least one target template nucleic acid molecule to the 3' end and a region of the chromosome immediately adjacent to the 5' end of the at least one target template nucleic acid molecule to the 5' end, or the primers will be complementary to the following regions: a region of the chromosome immediately adjacent to the 3' end of the nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule to the 3' end and a region of the chromosome immediately adjacent to the 5' end of the nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule to the 5' end.
[0236] Suitable conditions include temperatures at which the DNA polymerase can replicate at least one target template nucleic acid molecule, for example, temperatures between 40°C and 90°C, 50°C and 80°C, 60°C and 70°C, or approximately 68°C.
[0237] The step of introducing a mutation into at least one target template nucleic acid molecule may comprise multiple rounds of replication. For example, the step of introducing a mutation into at least one target template nucleic acid molecule preferably comprises:
[0238] i) performing rounds of replication on at least one target template nucleic acid molecule to provide at least one nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule; and
[0239] ii) performing rounds of replication on the at least one target template nucleic acid molecule to provide copies of the at least one target template nucleic acid molecule.
[0240] Optionally, the step of introducing mutations into at least one target template nucleic acid molecule comprises replicating at least one target template nucleic acid molecule for at least 2 rounds, at least 4 rounds, at least 6 rounds, at least 8 rounds, at least 10 rounds, less than 10 rounds, less than 8 rounds, about 6 rounds, 2 to 8 rounds, or 1 to 7 rounds. The user may choose to use a small number of rounds of replication to reduce the possibility of introducing amplification bias.
[0241] Optionally, the step of introducing mutations into at least one target template nucleic acid molecule comprises replication at a temperature of 60°C to 80°C for at least 2 rounds, at least 4 rounds, at least 6 rounds, at least 8 rounds, at least 10 rounds, less than 10 rounds, less than 8 rounds, about 6 rounds, 2 to 8 rounds, or 1 to 7 rounds.
[0242] Alternatively, the step of introducing mutations into at least one target template nucleic acid molecule is performed using polymerase chain reaction (PCR). PCR is a process that involves performing multiple rounds of the following steps to replicate a nucleic acid molecule:
[0243] a) unzipping;
[0244] b) annealing; and
[0245] c) Extension and elongation.
[0246] Nucleic acid molecules (e.g., at least one target template nucleic acid molecule) are mixed with appropriate primers and a polymerase. In the denaturation step, the nucleic acid molecules are heated to a temperature above 90°C to denature the double-stranded nucleic acid molecules (separate into two strands). In the annealing step, the nucleic acid molecules are cooled to a temperature below 75°C, such as 55°C to 70°C, about 55°C, or about 68°C, to allow the primers to anneal to the nucleic acid molecules. In the extension and elongation step, the nucleic acid molecules are heated to a temperature above 60°C to allow the DNA polymerase to catalyze the primer extension and add nucleotides complementary to the template strand.
[0247] Alternatively, the step of introducing a mutation into at least one target template nucleic acid molecule comprises: replicating at least one target template nucleic acid molecule using Taq polymerase under error-prone reaction conditions. For example, the step of introducing a mutation into at least one target template nucleic acid molecule may comprise: 2+ Mg 2+ Alternatively, PCR can be performed using Taq polymerase in the presence of unequal dNTP concentrations (e.g., excess cytosine, guanine, adenine, or thymine).
[0248] Obtain data containing non-mutation sequence reads and mutation sequence reads
[0249] The method of the present invention may include the step of obtaining data comprising non-mutation sequence reads and mutation sequence reads.Non-mutation sequence reads and mutation sequence reads may be obtained from any source.
[0250] Alternatively, non-mutated sequence reads are obtained by sequencing a region of at least one target template nucleic acid molecule in a first sample of the paired samples. Alternatively, mutated sequence reads are obtained by introducing a mutation into at least one target template nucleic acid molecule in a second sample of the paired samples to provide at least one mutated target template nucleic acid molecule, and sequencing a region of the at least one mutated target template nucleic acid molecule.
[0251] Optionally, the non-mutated sequence reads include the sequence of a region of at least one target template nucleic acid molecule in a first sample of a paired sample, the mutant sequence reads include the sequence of a region of at least one mutated target template nucleic acid molecule in a second sample of the paired sample, and the paired samples are taken from the same original sample or derived from the same organism.
[0252] Analyze mutant sequence reads and use the information obtained from analyzing mutant sequence reads to assemble sequences
[0253] As described above, the first sample and the second sample comprise at least one target template nucleic acid molecule. Therefore, the mutation pattern present in the mutant sequence reads can help the user assemble the sequence of at least a portion of the at least one target template nucleic acid molecule.
[0254] As described above, assembling a sequence may be difficult if, for example, regions of the sequence are similar to one another or the sequence contains repetitive portions. However, a user may be able to more efficiently assemble a sequence from non-mutant sequence reads using information obtained from mutant sequence reads corresponding to the non-mutant sequence reads. For example, mutant sequence reads may be used to identify nodes computed from non-mutant sequence reads that form part of a valid path through the sequence assembly graph.
[0255] According to certain embodiments, the information from multiple mutation reads can be used to assemble a sequence. As described in more detail below, it is possible to identify mutation sequence reads of target template nucleic acid molecules that may be derived from the same mutation. According to certain embodiments, mutation sequence reads can be assembled, and / or a consensus sequence can be generated by multiple mutation sequence reads. In a specific embodiment, a long mutation read (i.e., a synthetic long mutation read) can be reconstructed from multiple partially overlapping mutation reads of target template nucleic acid molecules derived from the same mutation to provide information to assemble the sequence. This synthetic long read can correspond to the identification path of the non-mutation assembly graph in series, as discussed elsewhere herein.
[0256] Prepare assembly diagram
[0257] The steps of analyzing the mutant sequence reads and using information obtained from analyzing the mutant sequence reads to assemble a sequence of at least a portion of at least one target template nucleic acid molecule from non-mutated sequence reads include preparing an assembly graph.
[0258] For the purposes of the present invention, an "assembly graph" is a graph comprising nodes and paths calculated from non-mutated sequence reads, which paths (in the case of valid paths) may correspond to portions of at least one target template nucleic acid molecule. For example, a node may represent a consensus sequence calculated from assembled non-mutated sequence reads.
[0259] Nodes can be calculated from non-mutated sequence reads. However, if some of the at least one target template nucleic acid molecules are not sequenced correctly, there may not be enough non-mutated sequence reads to assemble the complete sequence of the at least one target template nucleic acid molecule. If this is the case, nodes can be calculated from a combination of non-mutated sequence reads and mutant sequence reads, wherein the mutant sequence reads are used to supplement the region of the assembly graph that represents the missing non-mutated sequence reads. Alternatively, nodes are calculated from non-mutated sequence reads and mutant sequence reads. Using nodes calculated from non-mutated sequence reads alone is beneficial because non-mutated sequence reads accurately correspond to the original target template nucleic acid molecules. Therefore, an assembly graph composed of nodes calculated from non-mutated sequence reads can avoid artifacts introduced by the mutation step.
[0260] exist Figure 9 An illustration of a suitable assembly diagram is provided in Panel A of .
[0261] Optionally, the nodes of the assembly graph are unitigs. For the purposes of the present invention, the term "unitig" means a portion of at least one target template nucleic acid molecule whose sequence can be defined with high confidence. For example, the nodes of the assembly graph may include unitigs corresponding to the consensus sequence of all or part of one or more non-mutated sequence reads and / or all or part of one or more mutant sequence reads. Preferably, the nodes of the assembly graph include unitigs corresponding to the consensus sequence of all or part of one or more non-mutated sequence reads.
[0262] The assembly graph can be an overlapping graph, a unitig graph, or a weighted graph. For example, the assembly graph can be a de Bruijn graph.
[0263] Identify nodes that form part of a valid path in the collusion assembly graph
[0264] Using information obtained from analyzing mutant sequence reads to assemble the sequence of at least a portion of at least one target template nucleic acid molecule from non-mutant sequence reads can include: using information obtained from analyzing mutant sequence reads to identify nodes calculated from non-mutant sequence reads, each node forming part of a valid path through the assembly graph. Each valid path through the assembly graph can represent the sequence of a portion of at least one target template nucleic acid molecule. If the assembly graph includes multiple putative paths from node to node, the order of the nodes can be determined using information obtained from analyzing mutant sequence reads. In further embodiments, the information obtained from analyzing mutant sequence reads can be used to determine the copy number of a given sequence in a genome.
[0265] Optionally, analyzing the mutant sequence reads includes identifying mutant sequence reads of a target template nucleic acid molecule that may be derived from the same at least one mutation. The method of the present invention can result in providing a plurality of mutant sequence reads, wherein the plurality of mutant sequence reads include mutant sequences corresponding to the same region, i.e., a plurality of groups of mutant sequence reads corresponding to the same region. Some of the mutant sequence reads in the group may overlap, and some of the mutant sequence reads in the group may be duplicates. When the group of mutant sequence reads is mapped to the assembly graph, as shown in FIG. Figure 9 As shown in B, they can be used to identify valid paths through the assembly graph because they can link nodes computed from non-mutated sequence reads.
[0266] Thus, optionally, analyzing the mutant sequence reads comprises: identifying mutant sequence reads of target template nucleic acid molecules that are likely derived from the same at least one mutation. Alternatively, using information obtained from analyzing the mutant sequence reads to identify nodes that form part of a valid path of the collusion assembly graph may comprise:
[0267] (i) Nodes are calculated from non-mutation sequence reads;
[0268] (ii) mapping mutant sequence reads to the assembly graph;
[0269] (iii) identifying mutant sequence reads of a target template nucleic acid molecule that may be derived from the same at least one mutation; and
[0270] (iv) identifying nodes connected by mutant sequence reads that are likely derived from mutant sequence reads of the same at least one mutated target template nucleic acid molecule,
[0271] Nodes connected by mutation sequence reads are likely to originate from the same target template nucleic acid molecule with at least one mutation and constitute part of an effective path connecting the assembly graph.
[0272] Optionally, mutant sequence reads of target template nucleic acid molecules that are likely derived from the same mutation are grouped.
[0273] Identify mutant sequence reads of target template nucleic acid molecules that may be derived from the same mutation
[0274] As discussed, analyzing mutant sequence reads can include identifying mutant sequence reads that are likely derived from target template nucleic acid molecules that share at least one mutation.
[0275] Alternatively, if the mutant sequence reads have a common mutation pattern, the mutant sequence reads may be derived from the same at least one mutated target template nucleic acid molecule. Alternatively, the mutant sequence reads having a common mutation pattern include a common characteristic k-mer or a common characteristic mutation. Preferably, the mutant sequence reads having a common mutation pattern include at least 1, at least 2, at least 3, at least 4, at least 5, or at least k common characteristic k-mers and / or common characteristic mutations.
[0276] When a sample is provided by combining two or more subsamples, identifying mutant sequence reads that may be derived from the same at least one mutant target template nucleic acid molecule may have particular use. In certain embodiments, such a step may be used when determining the sequence of at least one target template nucleic acid molecule in a sample provided by combining two or more subsamples. More specifically, such a step may be used when determining the sequence of at least one target template nucleic acid molecule from each of two or more subsamples that are combined to provide a sample. Such a step may also have particular use when measuring the number of target template nucleic acid molecules in a sample from each of two or more subsamples where a target template nucleic acid molecule in the subsample has mutated.
[0277] Characteristic k-mer or characteristic mutation
[0278] The mutant sequence reads having a common mutation pattern may include common characteristic k-mers and / or common characteristic mutations. Preferably, the mutant sequence reads having a common mutation pattern include at least 1, at least 2, at least 3, at least 4, at least 5, or at least k common characteristic k-mers and / or common characteristic mutations.
[0279] In the context of the present invention, a "k-mer" refers to a nucleic acid sequence of length k that is included in a sequence read. A "signature k-mer" can be a k-mer that does not appear in non-mutation sequence reads but appears at least twice in mutant sequence reads. In one embodiment, a signature k-mer is a k-mer that appears at least n times more frequently in mutant sequence reads than in non-mutation sequence reads, where n is any integer, such as 2, 3, 4, or 5. Alternatively, a signature k-mer is a k-mer that appears at least twice, at least three times, at least four times, at least five times, or at least ten times in mutant sequence reads. Thus, a user can determine whether mutant sequence reads contain a common signature k-mer by dividing mutant sequence reads into k-mers and non-mutation sequence reads into k-mers. The user can then compare the mutant sequence read k-mers with the non-mutation sequence read k-mers and determine which k-mers appear in the mutant sequence read k-mers but not in the non-mutation sequence read k-mers (or which k-mers appear more frequently in the mutant sequence read k-mers than in the non-mutation read k-mers). The user can then evaluate the k-mers that appear in the mutant sequence read k-mer but do not appear (or appear less frequently) in the non-mutation sequence read k-mer and count them. Any k-mer that appears at least two, at least three, at least four, at least five, or at least ten times in the mutant sequence read k-mer but does not appear in the non-mutation sequence read k-mer is a feature k-mer. Any k-mer that appears less than k times, less than 5 times, less than 4 times, less than 3 times, or once in the mutant sequence read k-mer but does not appear (or appears less frequently) in the non-mutation sequence read k-mer can be the result of a sequencing error and should be ignored.
[0280] The k value can be selected by the user and can be any value. Optionally, the k value is at least 5, at least 10, at least 15, less than 100, less than 50, less than 25, 5 to 100, 10 to 50, or 15 to 25. Typically, the user will select a k value that is as long as possible while ensuring that one or more low sequencing errors are included in the portion of the k-mer in the read. Preferably, in the read that includes sequencing errors, the proportion of k-mer is less than 50%, less than 40%, less than 30%, between 0% and 50%, between 0% and 40%, or between 0% and 30.
[0281] A "signature mutation" can be a nucleotide that appears at least twice in a mutant sequence read but not at the corresponding position in a non-mutated sequence read. In one embodiment, a signature mutation is a mutation that appears at least n times more frequently in a mutant sequence read than in a non-mutated sequence read, where n is any integer, such as 2, 3, 4, or 5. Alternatively, a signature mutation is a mutation that appears at least two, at least three, at least four, at least five, or at least ten times in a mutant read but does not appear (or appears less frequently) at the corresponding position in a non-mutated read.
[0282] Alternatively, the signature mutation is a co-occurring mutation. "Co-occurring mutations" are two or more signature mutations that occur in the same mutant sequence read. For example, if the mutant sequence read contains three signature mutations, it contains three co-occurring mutation pairs or one co-occurring mutation 3-tuple. If the mutant sequence read contains four signature mutations, it contains six co-occurring mutation pairs, four co-occurring mutation 3-tuples, and one co-occurring mutation 4-tuple.
[0283] Alternatively, a signature mutation can be ignored if it does not meet certain criteria, indicating that the identified signature mutation is spurious or does not contribute to assembling the sequence of at least a portion of at least one target template nucleic acid molecule.
[0284] Alternatively, if at least one, at least two, at least three, or at least five nucleotides at corresponding positions in the mutant sequence reads having the signature mutation differ from each other, the signature mutation is ignored. For example, if two mutant sequence reads overlap and have a common signature mutation in the overlap, the overlapping nucleotides should be identical. If their level of identity is low, an error may have occurred, and the mutant sequence read should be ignored. For example, a one nucleotide difference may be tolerated because it may be a simple sequencing error.
[0285] Alternatively, if the signature mutation is an unexpected mutation, the signature mutation is ignored. The phrase "unexpected mutation" refers to a mutation that is unlikely to occur using a particular procedure for introducing a mutation into at least one target template nucleic acid molecule. For example, if the procedure for introducing a mutation into at least one target template nucleic acid molecule is performed using a chemical mutagen that only replaces guanine with adenine, then any replacement of cytosine is unexpected, and mutant sequence reads containing such a mutation should be ignored.
[0286] Optionally, identifying mutant sequence reads of the target template nucleic acid molecule that may be derived from the same at least one mutation includes identifying mutant sequence reads corresponding to a specific region of the at least one target template nucleic acid molecule. For example, a user may be interested only in identifying mutant sequence reads that contain a signature mutation in a region that overlaps with other mutant sequence reads, and may ignore signature mutations occurring in other regions.
[0287] In general, mutant sequence reads of characteristic mutation groups with larger intersections and smaller symmetric differences are more likely to originate from the same target template nucleic acid molecule with at least one mutation. For two mutant sequence reads A and B with characteristic mutations SM(A) and SM(B), it can be assumed that A and B originate from the same target template nucleic acid molecule with at least one mutation if the following conditions are met:
[0288] Intersection(SM(A),SM(B))>=C
[0289] as well as
[0290] Symmetric difference (SM(A), SM(B)) < intersection (SM(A), SM(B))
[0291] Wherein, C is greater than 4, greater than 5, less than 20 or less than 10, and SM(X) is a set of characteristic mutations of mutant sequence read X, which may be a subset of characteristic mutations of X.
[0292] Alternatively, a collection of co-occurring mutations can be used instead of signature mutations in the equations below.
[0293] Intersection(SM(A),SM(B))>=C
[0294] as well as
[0295] Symmetric difference (SM(A), SM(B)) < C2*intersection (SM(A), SM(B))
[0296] Wherein C2 is less than 3, less than 2, or less than or equal to 1.5, SM(X) is the set of co-occurring mutations of mutant sequence reads X, which may be a subset of the characteristic mutations of X.
[0297] Mutant sequence reads having a common characteristic k-mer or a common characteristic mutation can be grouped together. Preferably, mutant sequence reads are grouped together if they have at least 1, at least 2, at least 3, at least 4, at least 5, or at least k common characteristic k-mers and / or common characteristic mutations. In this embodiment, "k" is the length of the k-mer used.
[0298] Determine the likelihood that two mutant sequence reads originate from the same mutant target template nucleic acid molecule
[0299] Mutated sequence reads of target template nucleic acid molecules that are likely to be derived from the same mutation can be identified by calculating the following odds ratio:
[0300] The probability that a mutant sequence read is derived from the same mutated target template nucleic acid molecule: the probability that a mutant sequence read is not derived from the same mutated target template nucleic acid molecule.
[0301] If the odds ratio exceeds a threshold, the mutant sequence reads are likely derived from the same at least one mutated target template nucleic acid molecule. Similarly, if the odds ratio of a first mutant sequence read and a second mutant sequence read is higher than that of the first mutant sequence read and other mutant sequence reads mapped to the same region of the assembly graph, the first mutant sequence read is likely derived from the same at least one target template nucleic acid molecule as the second mutant sequence read.
[0302] The threshold applied can be at any level. In practice, the user will determine the threshold for any given sequencing method based on their requirements.
[0303] For example, the user can determine the required degree of stringency. If the user is using the method to determine or generate the sequence of at least one target template nucleic acid molecule, for which accuracy is not important, the selected threshold value can be much lower than the threshold value in the following case: the user is using the method to generate or determine the sequence of at least one target template nucleic acid molecule for which accuracy is crucial. If the user is using the method to determine or generate the sequence of the target template nucleic acid in the sample, for example, to determine whether the sample contains multiple bacterial strains or only one bacterial strain, a lower degree of accuracy may be required compared to the following case: the user is using the method to determine or generate the sequence of a specific variant gene to determine how it differs from the natural gene. Therefore, the threshold value can be changed (determined) based on the required stringency.
[0304] Similarly, the user can vary the threshold value based on the mutation rate used in the step of introducing mutations into at least one target template nucleic acid molecule. If the mutation rate is higher, it is easier to determine whether two mutant sequence reads originate from the same mutant target template nucleic acid molecule, and thus a higher probability threshold value can be used.
[0305] Similarly, the user can change the threshold value based on the size of the at least one target template nucleic acid molecule. The larger the size of the at least one target template nucleic acid molecule, the more difficult it is to sequence the entire length without any sequencing errors, so the user may want to use a higher threshold value for longer at least one target template nucleic acid molecule.
[0306] Similarly, users can change the threshold value based on time constraints and resource constraints. If these constraints are high, users may be satisfied with a lower threshold value, thereby providing a less accurate sequence.
[0307] Additionally, the user can vary the threshold value based on the error rate of the step of sequencing at least one region of the mutated target template to provide mutant sequence reads. If the error rate is high, the user can set a higher threshold value than when the error rate is low. This is because, if the error rate is high, there may be less information about whether two mutant sequence reads are derived from the same mutated target template nucleic acid molecule, especially if the errors are biased in a manner similar to the introduced mutations.
[0308] Optionally, identifying mutant sequence reads of the target template nucleic acid molecule that are likely derived from the same mutation comprises using a probability function based on the following parameters:
[0309] a. Matrix of nucleotides at each position of the mutant sequence reads (N) and assembly graph;
[0310] b. The probability (M) of mutating a given nucleotide (i) to read nucleotide (j);
[0311] c. the probability (E) of incorrectly reading a given nucleotide (i) so that the nucleotide (j) is read conditional on the nucleotide being incorrectly read; and
[0312] d. The probability of incorrectly reading the nucleotide at position Y (Q).
[0313] The probability function can be used to determine the odds ratio:
[0314] The probability that mutant sequence reads are derived from the same mutant target template nucleic acid molecule: the probability that mutant sequence reads are not derived from the same mutant target template nucleic acid molecule.
[0315] Alternatively, the Q value is obtained by performing a statistical analysis on the mutant sequence reads and the non-mutant sequence reads, or based on existing knowledge of the accuracy of the sequencing method. For example, Q depends on the accuracy of the sequencing method used. Therefore, the user can determine the Q value by sequencing a nucleic acid molecule of a known sequence and determining the average number of nucleotides that are incorrectly read. Alternatively, the user can select a subset of mutant sequence reads and non-mutant sequence reads and compare them. The difference between the mutant sequence reads and the non-mutant sequence reads may be due to sequencing errors or the introduction of mutations. The user can use statistical analysis to approximately estimate the number of differences due to sequencing errors.
[0316] Alternatively, the values of M and E are estimated based on a statistical analysis performed on a subset of mutant sequence reads and non-mutant sequence reads, wherein the subset includes mutant sequence reads and non-mutant sequence reads that are selected because they map to the same region of the reference assembly map. An example of how to determine M and E is provided in Example 6. In short, the user can perform a statistical analysis on a subset of mutant sequence reads and non-mutant sequence reads to obtain the best-fit values of M and E (via unsupervised learning). Since unsupervised learning can be a computationally expensive process, it is advantageous to perform this step on a subset of mutant sequence reads and non-mutant sequence reads and then apply the values of M and E to the complete set of mutant sequence reads and non-mutant sequence reads.
[0317] Optionally, the statistical analysis is performed using Bayesian inference, Monte Carlo methods such as Hamiltonian Monte Carlo, variational inference, or maximum likelihood simulation of Bayesian inference.
[0318] Optionally, identifying mutant sequence reads of target template nucleic acid molecules that are likely derived from the same mutation comprises using machine learning or neural networks; for example, as described in detail in Russell & Norvig "Artificial Intelligence, a modern approach".
[0319] Pre-clustering
[0320] Optionally, the method includes a pre-clustering step. For example, the user can perform an initial calculation to assign mutant sequence reads into groups, wherein each member of the same group has a reasonable likelihood of being derived from the same at least one mutated target template nucleic acid molecule. The mutant sequence reads in each group can be mapped to a common position on the assembly map and / or have a common mutation pattern. If two mutant sequence reads in a group are mapped to the same region, or they overlap in the assembly map, they are mapped to a common position on the assembly map. The likelihood threshold applied in the pre-clustering step can be lower than the likelihood threshold applied in the step of identifying mutant sequence reads that may be derived from the same at least one mutated target template nucleic acid molecule, that is, the pre-clustering step can be a lower stringency step compared to the step of identifying mutant sequence reads that may be derived from the same at least one mutated target template nucleic acid molecule.
[0321] Optionally, identifying mutant sequence reads that are likely to be derived from the same mutation target template nucleic acid molecule is constrained by the results of the pre-clustering step. For example, a user can apply a lower stringency pre-clustering step to group mutant sequence reads that map to a common region of the assembly graph and have a reasonable likelihood of being derived from the same at least one mutation target template nucleic acid molecule. The user can then apply a higher stringency step of identifying mutant sequence reads that are likely to be derived from the same at least one mutation target template nucleic acid molecule to each member of a group of members to see which of them are indeed likely to be derived from the same at least one mutation target template nucleic acid molecule. The advantage of using a pre-clustering step is that a higher stringency step will use more processing power than a lower stringency step, and in this example, a higher stringency step is only applied to mutant sequence reads that are assigned to the same group by a lower stringency step, thereby reducing the required overall processing power.
[0322] Optionally, the pre-clustering step includes Markov clustering or Louvain clustering ( https: / / micans.org / mcl / as well as https: / / arxiv.org / abs / 0803.0476 ).
[0323] Optionally, the pre-clustering step is performed by assigning the mutant sequence reads to the same group having at least 1, at least 2, at least 3, at least 5, or at least k characteristic k-mers or at least 1, at least 2, at least 3, or at least 5 characteristic mutations as described above. Optionally, if the mutant sequence reads have a common mutation pattern, they are reasonably likely to originate from the same at least one mutated target template nucleic acid molecule, and the mutant sequence reads having a common mutation pattern are mutant sequence reads that contain at least 1, at least 2, at least 3, at least 5, or at least k common characteristic k-mers or common characteristic mutations.
[0324] Alternatively, as described under the heading "Signature k-mer or signature mutation", a signature k-mer is a k-mer that does not appear (or appears less frequently) in non-mutation sequence reads but appears at least twice (optionally at least three times, at least four times, at least five times, or at least ten times) in mutation sequence reads. Alternatively, a signature mutation is a nucleotide that appears at least twice (optionally at least three times, at least four times, at least five times, or at least ten times) in mutation sequence reads but does not appear (or appears less frequently) at the corresponding position in non-mutation sequence reads.
[0325] Ignore the inferred paths of the collusion assembly graph
[0326] In some embodiments of the invention, identifying nodes that form part of a valid path of the collusive assembly graph comprises ignoring putative paths of the collusive assembly graph.
[0327] For example, the putative paths of the collusion assembly graph can be ignored in the following cases:
[0328] (i) their ends do not match those present in the end sequence library;
[0329] (ii) they are the result of template collisions;
[0330] (iii) they are longer or shorter than expected; and / or
[0331] (iv) They have atypical coverage depths.
[0332] The term "template collision" refers to the situation where two putative pathways of the concatenated assembly graph are identified that correspond to one or more identical mutated sequence reads or mutated sequence reads containing the same mutation pattern (the two putative pathways collide).
[0333] Ignore putative paths that do not match the ends of the collusion assembly graph
[0334] The method can include preparing a library of sequences of the end pairs of at least one mutated target template nucleic acid molecule. For example, the library can specify that the first at least one target template nucleic acid molecule has an A-terminal sequence and a B-terminal sequence, while the second at least one target template nucleic acid molecule has a C-terminal sequence and a D-terminal sequence. The library can be prepared by performing paired-end sequencing on at least one target template nucleic acid molecule. Alternatively, the method includes sequencing the ends of at least one target template nucleic acid molecule using paired-end sequencing.
[0335] In this embodiment, identifying nodes that constitute part of a valid path of the concatenated assembly graph includes ignoring putative paths that have mismatched termini, i.e., where the sequences of the termini of the putative path do not correspond to one of the paired termini in the library. For example, if the library specifies that a first at least one target template nucleic acid molecule has an A-terminal sequence and a B-terminal sequence, and a second at least one target template nucleic acid molecule has a C-terminal sequence and a D-terminal sequence, then a putative path that pairs the A terminus with the D terminus would be an incorrect path and should be ignored.
[0336] To disregard putative pathways with mismatched ends, the user may map the sequence of the ends of the at least one target template nucleic acid molecule onto the assembly map. Alternatively, to assist the user in assembling the sequence of at least a portion of the at least one target template nucleic acid molecule from non-mutated sequence reads, the user may also wish to map the sequence of the ends of the at least one target template nucleic acid molecule onto the assembly map to identify the starting and ending positions of each at least one target template nucleic acid molecule on the assembly map.
[0337] Alternatively, at least one target template nucleic acid molecule comprises at least one barcode. Alternatively, at least one target template nucleic acid molecule comprises a barcode at each end. The term "at each end" means that the barcode is substantially present at both ends close to the at least one target template nucleic acid molecule, for example, within 50 base pairs, within 25 base pairs, or within 10 base pairs of the end of the at least one target template nucleic acid molecule. If at least one target template nucleic acid molecule comprises at least one barcode, it is easier for the user to determine whether the putative path has a mismatched end. This is because the end sequences are more distinctive, and it is easier to determine whether the sequences of the two ends that appear to be mismatched are indeed mismatched, or whether sequencing errors have been introduced into the sequence of one of the ends.
[0338] Barcodes and sample labels
[0339] For the purposes of the present invention, a barcode (also referred to herein as a "unique molecular tag" or "unique molecular identifier") is a degenerate or randomly generated sequence of nucleotides. A target template nucleic acid molecule may comprise 1, 2, or 3 barcodes. According to certain embodiments, each barcode may have a sequence that is different from every other barcode generated. However, in other embodiments, two or more barcode sequences may be identical, i.e., the barcode sequence may appear more than once. For example, at least 90% of the barcode sequences may be different from the sequence of every other barcode sequence. It is only required that the barcodes are appropriately degenerate so that each target template nucleic acid molecule comprises a barcode with a unique or substantially unique sequence compared to every other target template nucleic acid molecule in the paired sample. Thus, marking (or labeling) the target template nucleic acid molecules with barcodes allows the target template nucleic acid molecules to be distinguished from each other, thereby facilitating the methods discussed elsewhere herein. Therefore, the barcodes may be considered unique molecular tags (UMT). Barcodes can be 5, 6, 7, 8, 5 to 25, 6 to 20, or more nucleotides in length.
[0340] Alternatively, as described above, at least one target template nucleic acid molecule in different paired samples can be labeled with different sample tags.
[0341] For the purposes of the present invention, a sample label is a label used to label a majority of at least one target template nucleic acid molecule in a sample. Different sample labels can be used in other samples to distinguish which at least one target template nucleic acid molecule is derived from which sample. The sample label is a known nucleotide sequence. The sample label can be 5, 6, 7, 8, 5 to 25, 6 to 20, or more nucleotides in length.
[0342] Alternatively, the method of the present invention includes a step of introducing at least one bar code or sample tag into at least one target template nucleic acid molecule. Any suitable method can be used to introduce at least one bar code or sample tag, including PCR, tag fragmentation and physical shearing or restriction digestion of the target nucleic acid, followed by connection with adapters (optionally, sticky end connection). For example, a first set of primers capable of hybridizing with at least one target nucleic acid molecule can be used to perform PCR on at least one target template nucleic acid molecule. Primers can be used to introduce at least one bar code or sample tag into each of at least one target template nucleic acid molecule by PCR, the primers comprising a portion (5' end portion) comprising a bar code, a sample tag and / or an adapter, and a portion (3' end portion) having a sequence that can hybridize (optionally complementary) with at least one target nucleic acid molecule. Such primers will hybridize with at least one target template nucleic acid molecule, and then PCR primer extension will provide at least one target template nucleic acid molecule comprising a bar code and / or a sample tag. Further cycles of PCR performed with these primers can be used to optionally add other bar codes or sample tags to the other end of at least one target template nucleic acid molecule. Primers may be degenerate, ie, the 3' terminal portions of the primers may be similar but not identical to each other.
[0343] At least one barcode or sample tag can be introduced using tag fragmentation. At least one barcode or sample tag can be introduced by using direct tag fragmentation; or by introducing a defined sequence by tag fragmentation followed by two cycles of PCR using primers comprising a portion capable of hybridizing to a defined sequence and a portion containing a barcode, sample tag and / or adapter. At least one barcode or sample tag can be introduced by restricting the original at least one target nucleic acid molecule and then connecting a nucleic acid comprising a barcode and / or sample tag. Restriction digestion of the original at least one target nucleic acid molecule should be performed so that the digestion produces a nucleic acid molecule comprising a region to be sequenced (at least one target template nucleic acid molecule). At least one barcode or sample tag can be introduced by shearing at least one target nucleic acid molecule, followed by end repair, A-tail ligation, and then connecting a nucleic acid comprising a barcode and / or sample tag.
[0344] Ignore inferred paths due to template collisions
[0345] The method can include ignoring putative pathways resulting from template collisions. As described above, the term "template collision" refers to a situation where two putative pathways in the colluding assembly graph are identified that correspond to one or more identical mutant sequence reads or mutant sequence reads containing the same mutation pattern (the two putative pathways collide). Because each valid pathway should contain a unique set of mutant sequence reads, it is likely that at least one of the two colliding putative pathways is incorrect. For these reasons, ignoring putative pathways resulting from template collisions may reduce the number of incorrect pathways identified.
[0346] Similarly, it is possible that two different at least one mutant target template nucleic acid molecules may have similar or identical mutation patterns because they did not receive many mutations during the step of introducing the mutations into the at least one target template nucleic acid molecule, or because the mutations they received by chance were the same. If this is the case, template collisions will again be observed. In this case, it is not practical to use the information obtained by analyzing these at least one mutant target template nucleic acid molecules with poor mutations to assemble the sequence of at least a portion of the at least one target template nucleic acid molecule from non-mutated sequences, and putative paths corresponding to nodes calculated from non-mutated sequence reads derived from such at least one mutant target template nucleic acid molecules with poor mutations should be ignored.
[0347] Ignore inferred paths that are longer or shorter than expected
[0348] The length of at least one target template nucleic acid molecule can be known or predictable.
[0349] This length can be defined by analyzing the length of at least one target template nucleic acid molecule in a laboratory environment. For example, the user can use gel electrophoresis to separate a sample of at least one target template nucleic acid molecule, and use this sample for the method of the present invention. In this case, all of at least one target template nucleic acid molecules to be determined or to produce a sequence will be within a known size range. For example, the user can extract a band from the gel exposed to gel electrophoresis, and this band corresponds to at least one target template nucleic acid molecule with a length of 6000bp-14000bp or 18000bp-12000bp. Alternatively or additionally, a variety of methods for determining the size of nucleic acid molecules can be used, including gel electrophoresis, to quantify the size of at least one target template nucleic acid molecule. For example, the user can use an instrument, such as an Agilent Bioanalyzer or a FemtoPulse machine.
[0350] When the size of at least one target template nucleic acid molecule is known or predictable, putative pathways longer and shorter than the defined length are likely incorrect and should be ignored.
[0351] Ignore putative paths with atypical coverage depth
[0352] The method of the present invention may include the step of amplifying at least one mutated target template nucleic acid molecule, i.e., duplicating at least one mutated target nucleic acid molecule to provide a copy of at least one mutated target template nucleic acid molecule. For example, the method may include amplifying at least one mutated target template nucleic acid molecule using PCR. Amplification may result in some of the at least mutated target template nucleic acid molecule being replicated more times than others. If some of the at least one mutated target template nucleic acid molecule are amplified to a greater extent (have a higher coverage depth) than the at least one mutated target template nucleic acid molecule, a greater number of mutation sequence reads will be associated with the inferred path corresponding to those at least one mutated target template nucleic acid molecules compared to the at least one mutated target template nucleic acid molecule. Similarly, it is expected that the coverage depth will be consistent over the length of the at least one template nucleic acid molecule. Therefore, it is expected that different parts of the effective path will have a similar number of mutation sequence reads associated therewith (similar coverage depth). If the inferred path includes a portion with a low coverage depth and a portion with a high coverage depth, the two portions may not correspond to the same effective path, and the inferred path is wrong and should be ignored.
[0353] Assembling a sequence of at least a portion of at least one target template nucleic acid molecule
[0354] Optionally, the sequence of at least a portion of at least one target template nucleic acid molecule is assembled from non-mutated sequence reads that form part of a valid path of the concatenated assembly graph.
[0355] Optionally, the method does not comprise generating a consensus sequence from the mutant sequence reads. Optionally, the method does not comprise the step of assembling the sequence of at least one mutant target template nucleic acid molecule or a substantial portion of at least one mutant target template nucleic acid molecule.
[0356] By "consensus sequence" is meant a sequence comprising at each position a possible nucleotide defined by analysis of a set of aligned sequence reads, eg, the most frequently occurring nucleotide at each position in a set of aligned sequence reads.
[0357] The method comprises the steps of assembling the sequence of at least a portion of at least one target template nucleic acid molecule from nodes, the nodes constituting a valid path through the assembly graph. Alternatively, the step of assembling the sequence of at least a portion of at least one target template nucleic acid molecule comprises assembling the sequence of at least a portion of at least one target template nucleic acid molecule from nodes constituting a portion of a valid path through the assembly graph.
[0358] Optionally, assembling the sequence of at least a portion of at least one target template nucleic acid molecule includes identifying an "endwall". The endwall is a position on the assembly map corresponding to a plurality of "end reads + int reads" (the end read corresponds to one of the ends of at least one target template nucleic acid molecule, and the internal read corresponds to an internal sequence (i.e., a sequence that is not at the end of at least one target template nucleic acid molecule)). End reads can be generated using, for example, paired end sequencing methods. Optionally, the endwall is identified as a position on the assembly map where at least 5 end reads are mapped. Optionally, the endwall is identified as a position on the assembly map where 2 to 4 end reads are mapped and at least 5 end reads or internal reads are mapped. Optionally, the step of assembling the sequence of at least a portion of the at least one target template nucleic acid molecule includes: assembling the sequence of at least a portion of the at least one target template nucleic acid molecule from a node that constitutes part of a valid path through the assembly map, and the assembling step starts at the endwall.
[0359] As described above, a valid path of a concatenated assembly graph may comprise connected nodes. When a series of connected nodes form (consisting of one or more nodes) a single path of a concatenated assembly graph (e.g., wherein the nodes of the graph may be unitig), then the sequence covered by the connected nodes represents at least a portion of at least one target template nucleic acid molecule. The sequence can then be determined by using standard techniques, such as canu( https: / / github.com / marbl / canu ) or miniasm (https: / / github.com / lh3 / miniasm) to assemble these parts by connecting the nodes. For example, users can prepare a consensus sequence from the nodes that form a valid path.
[0360] Alternatively, the assembled sequence includes nodes calculated primarily from non-mutation sequence reads. If the sequence is assembled from nodes calculated from more than 50% non-mutation sequence reads, the assembled sequence will contain nodes calculated primarily from non-mutation sequence reads. It is advantageous to assemble sequences from nodes calculated primarily from non-mutation sequence reads because the assembled sequence is more likely to accurately correspond to the original at least one target template nucleic acid molecule sequence. However, if it is not possible to map the non-mutation sequence reads to a portion of the putative path of the assembly graph, the sequence of the missing portion can be assembled from the nodes calculated from the mutation sequence reads. Preferably, the assembled sequence includes nodes calculated from more than 50%, more than 60%, more than 70%, more than 80%, more than 90%, more than 98%, 50% to 100%, 60% to 100%, 70% to 100% or 80% to 100% non-mutation sequence reads.
[0361] Amplifying at least one target template nucleic acid molecule
[0362] The method may include: before the step of sequencing the region of at least one target template nucleic acid molecule in the first sample of the paired samples, a step of amplifying the at least one target template nucleic acid molecule. The method may include: before the step of sequencing the region of at least one target template nucleic acid molecule in the second sample of the paired samples, a step of amplifying the at least one target template nucleic acid molecule.
[0363] Suitable methods for amplifying at least one target template nucleic acid molecule are known in the art. For example, PCR is typically used. PCR is described in more detail above under the heading "Introducing mutations into at least one target template nucleic acid molecule."
[0364] Fragmenting at least one target template nucleic acid molecule
[0365] The method may include: before the step of sequencing the region of at least one target template nucleic acid molecule in the first sample of the paired samples, the step of fragmenting the at least one target template nucleic acid molecule. Alternatively, the method may include: before the step of sequencing the region of at least one mutated target template nucleic acid molecule in the second sample of the paired samples, the step of fragmenting the at least one mutated target template nucleic acid molecule.
[0366] Any suitable technique can be used to fragment at least one target template nucleic acid molecule. Fragmentation can be performed using restriction digestion or using PCR with primers complementary to at least one internal region of at least one mutated target nucleic acid molecule. Preferably, fragmentation is performed using a technique that produces arbitrary fragments. The term "arbitrary fragment" refers to randomly generated fragments, such as fragments generated by tag fragmentation. The fragments generated using restriction endonucleases are not "arbitrary" because restriction digestion occurs at a specific DNA sequence defined by the restriction endonuclease used. Even more preferably, fragmentation is performed by tag fragmentation. If fragmentation is performed by tag fragmentation, the tag fragmentation reaction optionally introduces an adapter region into the at least one mutated target nucleic acid molecule. The adapter region is a short DNA sequence that can encode, for example, an adapter to allow the at least one mutated target nucleic acid molecule to be sequenced using Illumina technology.
[0367] Low-bias DNA polymerase
[0368] As described above, mutations can be introduced using low-bias DNA polymerases. Low-bias DNA polymerases can introduce mutations uniformly and randomly, which can be beneficial in the methods of the present invention because if mutations are introduced in a uniformly random manner, the likelihood that any given portion of the template nucleic acid molecule will have a unique mutation pattern is higher. As described above, unique mutation patterns can be used to identify efficient pathways for assembling a graph.
[0369] Additionally, methods using DNA polymerases with high template amplification bias may be limited. DNA polymerases with high template amplification bias replicate and / or mutate some target template nucleic acid molecules better than others, and therefore sequencing methods using such highly biased DNA polymerases may not sequence some target template nucleic acid molecules well.
[0370] A low bias DNA polymerase can have low template amplification bias and / or low mutation bias.
[0371] Low mutation bias
[0372] A low-bias DNA polymerase exhibiting low mutational bias is a DNA polymerase that can mutate adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine at similar rates. In one embodiment, the low-bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at similar rates.
[0373] Alternatively, the low bias DNA polymerase can mutate adenine and thymine, adenine and guanine, adenine and cytosine, thymine and guanine, thymine and cytosine, or guanine and cytosine, respectively, at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1: 1. Preferably, the low bias DNA polymerase can mutate guanine or adenine, respectively, at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1: 1. Preferably, the low bias DNA polymerase is capable of mutating thymine and cytosine at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively.
[0374] In this embodiment, in the step of introducing mutations into a plurality of target template nucleic acid molecules, the low bias DNA polymerase mutates adenine nucleotides and thymine nucleotides, adenine nucleotides and guanine nucleotides, adenine nucleotides and cytosine nucleotides, thymine nucleotides and guanine nucleotides, thymine nucleotides and cytosine nucleotides, or guanine nucleotides and cytosine nucleotides in at least one target template nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1:1, respectively. Preferably, the low bias DNA polymerase mutates guanine nucleotides and adenine nucleotides in at least one target template nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1: 1. Preferably, the low bias DNA polymerase mutates thymine nucleotides and cytosine nucleotides in at least one target template nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2, or about 1: 1.
[0375] Alternatively, the low bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at a ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1. Preferably, the low bias DNA polymerase can mutate adenine, thymine, guanine, and cytosine at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3.
[0376] In this embodiment, in the step of introducing mutations into at least one target template nucleic acid molecule in the second sample of the paired samples, the low bias DNA polymerase can mutate adenine nucleotides, thymine nucleotides, guanine nucleotides and cytosine nucleotides in at least one target template nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1, respectively. Preferably, the low bias DNA polymerase mutates adenine nucleotides, thymine nucleotides, guanine nucleotides and cytosine nucleotides in at least one target template nucleic acid molecule at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3.
[0377] Adenine, thymine, cytosine and / or guanine can be replaced by another nucleotide. For example, if a low-bias DNA polymerase is capable of mutating adenine, enzymatic mutagenesis using a low-bias DNA polymerase can replace at least one adenine nucleotide in a nucleic acid molecule with thymine, guanine, or cytosine. Similarly, if a low-bias DNA polymerase is capable of mutating thymine, enzymatic mutagenesis using a low-bias DNA polymerase can replace at least one thymine nucleotide with adenine, guanine, or cytosine. If a low-bias DNA polymerase is capable of mutating guanine, enzymatic mutagenesis using a low-bias DNA polymerase can replace at least one adenine nucleotide with thymine, guanine, or cytosine. If a low-bias DNA polymerase is capable of mutating cytosine, enzymatic mutagenesis using a low-bias DNA polymerase can replace at least one cytosine nucleotide with thymine, guanine, or adenine.
[0378] Low bias DNA polymerases may not be able to directly replace nucleotides, but they may still be able to mutate nucleotides by replacing the corresponding nucleotides on the complementary strand. For example, if the target template nucleic acid molecule comprises thymine, an adenine nucleotide will be present in the corresponding position of the at least one nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule. The low bias DNA polymerase may be able to replace the adenine nucleotide of the at least one nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule with guanine, so that when the at least one nucleic acid molecule that is complementary to the at least one target template nucleic acid molecule is replicated, this will result in a cytosine being present in the corresponding replicated at least one target template nucleic acid molecule where there was originally a thymine (thymine to cytosine substitution).
[0379] In one embodiment, the low-bias DNA polymerase mutates 1% to 15%, 2% to 10%, or approximately 8% of the nucleotides in at least one target template nucleic acid. In this embodiment, enzymatic mutagenesis using the low-bias DNA polymerase is performed in such a manner that 1% to 15%, 2% to 10%, or approximately 8% of the nucleotides in at least one target template nucleic acid are mutated. For example, if the user desires to mutate approximately 8% of the nucleotides in the target template nucleic acid molecules, and the low-bias DNA polymerase mutates approximately 1% of the nucleotides per round of replication, then the step of introducing mutations into the plurality of target template nucleic acid molecules by enzymatic mutagenesis can include performing 8 rounds of replication in the presence of the low-bias DNA polymerase.
[0380] In one embodiment, the low bias DNA polymerase is capable of mutating 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5% of the nucleotides in at least one target template nucleic acid molecule per round of replication. In one embodiment, the low bias DNA polymerase is capable of mutating 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5% of the nucleotides in at least one target template nucleic acid molecule per round of replication. The actual amount of mutations that occur per round may vary, but may average 0% to 3%, 0% to 2%, 0.1% to 5%, 0.2% to 3%, or about 1.5%.
[0381] Can DNA polymerase mutate nucleotides, and if so, at what rate?
[0382] Whether a low-bias DNA polymerase is capable of mutating a certain percentage of nucleotides in at least one target template nucleic acid molecule per round of replication can be determined by amplifying a nucleic acid molecule of known sequence in the presence of a low-bias DNA polymerase for a certain number of rounds of replication. The resulting amplified nucleic acid molecule can then be sequenced, and the percentage of nucleotides mutated during each round of replication can be calculated. For example, a nucleic acid molecule of known sequence can be amplified using 10 rounds of PCR in the presence of a low-bias DNA polymerase. The resulting nucleic acid molecule can then be sequenced. If the resulting nucleic acid molecule contains 10% nucleotides that differ from the corresponding nucleotides in the original known sequence, the user will understand that the low-bias DNA polymerase is capable of mutating, on average, 1% of the nucleotides in at least one target template nucleic acid molecule per round of replication. Similarly, to determine whether a low-bias DNA polymerase mutates a certain percentage of nucleotides in at least one target template nucleic acid molecule in a given method, the user can perform the method on a nucleic acid molecule of known sequence and use sequencing to determine the percentage of nucleotides mutated after the method is completed.
[0383] If a low bias DNA polymerase is used to amplify a nucleic acid molecule, it provides a nucleic acid molecule in which a nucleotide (such as adenine) is substituted or deleted in some cases, then the low bias DNA polymerase is capable of mutating the nucleotide. Preferably, the term "mutation" refers to the introduction of a substitution mutation, and in some embodiments, the term "mutation" can be replaced with "introduction of a substitution in..."
[0384] If, during the step of introducing a mutation into a plurality of target template nucleic acid molecules using a low-bias DNA polymerase, the low-bias DNA polymerase mutates a nucleotide, such as an adenine, in at least one target template nucleic acid molecule, then this step results in at least one target template nucleic acid molecule having a mutation in which the nucleotide is mutated. For example, if, during the step of introducing a mutation into a plurality of target template nucleic acid molecules using a low-bias DNA polymerase, the low-bias DNA polymerase mutates an adenine in at least one target template nucleic acid molecule, then this step results in at least one target template nucleic acid molecule having a mutation in which at least one adenine is substituted or deleted.
[0385] In order to determine whether DNA polymerase can introduce certain mutations, the technician only needs to use nucleic acid molecules of known sequence to test DNA polymerase. Suitable nucleic acid molecules of known sequence are fragments of bacterial genomes such as Escherichia coli MG1655 from known sequences. The technician can use PCR to amplify nucleic acid molecules of known sequence in the presence of a low-bias DNA polymerase. The technician can then order-check the amplified nucleic acid molecules and determine whether their sequence is identical to the original known sequence. If not, the technician can determine the nature of the mutation. For example, if the technician wishes to determine whether DNA polymerase can use nucleotide analogs to mutate adenine, the technician can use PCR to amplify nucleic acid molecules of known sequence in the presence of nucleotide analogs, and order-check the amplified nucleic acid molecules obtained. If the amplified DNA has a mutation at a position corresponding to an adenine nucleotide in the known sequence, the technician will know that DNA polymerase can use nucleotide analogs to mutate adenine.
[0386] Rate ratio can be calculated in a similar manner. For example, if the technician wishes to determine the rate ratio of guanine nucleotide and cytosine nucleotide mutations, the technician can use PCR to amplify nucleic acid molecules with known sequences in the presence of a low-bias DNA polymerase. The technician can then sequence the resulting amplified nucleic acid molecules and identify how many guanine nucleotides have been replaced or deleted, and how many cytosine nucleotides have been replaced or deleted. The rate ratio is the ratio of the number of guanine nucleotides that have been replaced or deleted to the number of cytosine nucleotides that have been replaced or deleted. For example, if 16 guanine nucleotides have been replaced or deleted, and 8 cytosine nucleotides have been replaced or deleted, then the guanine nucleotide and cytosine nucleotide have mutated at a rate ratio of 16:8 or 2:1, respectively.
[0387] Use of nucleotide analogs
[0388] Low bias DNA polymerases may not be able to directly (at least not at high frequency) substitute nucleotides for other nucleotides, but low bias DNA polymerases may still be able to mutate nucleic acid molecules using nucleotide analogs. Low bias DNA polymerases may be able to substitute nucleotides for other natural nucleotides (i.e., cytosine, guanine, adenine, or thymine) or for nucleotide analogs.
[0389] For example, a low-bias DNA polymerase can be a high-fidelity DNA polymerase. Typically, high-fidelity DNA polymerases tend to introduce very few mutations because they are highly accurate. However, the inventors have discovered that some high-fidelity DNA polymerases may still be able to mutate the target template nucleic acid molecule because they may be able to introduce nucleotide analogs into the target template nucleic acid molecule.
[0390] In one embodiment, the high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, between 0% and 0.0015%, or between 0% and 0.001% mutations per round of replication in the absence of nucleotide analogs.
[0391] In one embodiment, the low bias DNA polymerase is capable of incorporating nucleotide analogs into at least one target template nucleic acid molecule. In one embodiment, the low bias DNA polymerase is capable of incorporating nucleotide analogs into at least one target template nucleic acid molecule. In one embodiment, the low bias DNA polymerase is capable of using nucleotide analogs to mutate adenine, thymine, guanine, and cytosine. In one embodiment, the low bias DNA polymerase is capable of using nucleotide analogs to mutate adenine, thymine, guanine, and / or cytosine in at least one target template nucleic acid molecule. In one embodiment, the DNA polymerase replaces guanine, cytosine, adenine, and / or thymine with nucleotide analogs. In one embodiment, the DNA polymerase replaces guanine, cytosine, adenine, and / or thymine with nucleotide analogs.
[0392] Nucleotide analogs are incorporated into at least one target template nucleic acid molecule and can be used to mutate nucleotides, because nucleotide analogs can replace existing nucleotides and nucleotide analogs can be paired with the nucleotides in the opposite chain. For example, dPTP can be incorporated into nucleic acid molecules (can replace thymine or cytosine) instead of pyrimidine nucleotides. Once in a nucleic acid chain, when being an imino tautomeric form, dPTP can be paired with adenine. Therefore, when forming a complementary chain, this complementary chain can have adenine at a position complementary to dPTP. Similarly, once in a nucleic acid chain, when being an amino tautomeric form, dPTP can be paired with guanine. Therefore, when forming a complementary chain, this complementary chain can have guanine at a position complementary to dPTP.
[0393] For example, if a dPTP is introduced into at least one target template nucleic acid molecule of the present invention, when forming at least one nucleic acid molecule complementary to the at least one target template nucleic acid molecule, the at least one nucleic acid molecule complementary to the at least one target template nucleic acid molecule will contain an adenine or guanine (depending on whether the dPTP is in its amino form or imino form) at a position complementary to the dPTP in the at least one target template nucleic acid molecule. When replicating the at least one nucleic acid molecule complementary to the at least one target template nucleic acid molecule, the resulting replica of the at least one target template nucleic acid molecule will contain a thymine or cytosine at a position complementary to the dPTP in the at least one target template nucleic acid molecule. Thus, a mutation to thymine or cytosine can be introduced into the mutated at least one target template nucleic acid molecule.
[0394] Alternatively, if a dPTP is introduced into at least one nucleic acid molecule complementary to at least one target template nucleic acid molecule when forming a replica of at least one target template nucleic acid molecule, the replica of the at least one target template nucleic acid molecule will contain an adenine or guanine (depending on the tautomeric form of the dPTP) at a position complementary to the dPTP in the at least one nucleic acid molecule complementary to the at least one target template nucleic acid molecule. Thus, a mutation to an adenine or guanine can be introduced into the mutated at least one target template nucleic acid molecule.
[0395] In one embodiment, low biased DNA polymerase can replace cytosine or thymine with nucleotide analogs. In another embodiment, low biased DNA polymerase uses nucleotide analogs to introduce guanine nucleotide or adenine nucleotide with a rate ratio of 0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2 or about 1:1. Guanine nucleotide or adenine nucleotide can be introduced by the low biased DNA polymerase that makes it and nucleotide analogs (such as dPTP) relatively pairing. In another embodiment, low biased DNA polymerase uses nucleotide analogs to introduce guanine or adenine nucleotide with a rate ratio of 0.7-1.3:0.7-1.3.
[0396] A skilled artisan can use routine methods to determine whether a low-bias DNA polymerase is capable of incorporating a nucleotide analog into at least one target template nucleic acid molecule, or use routine methods to mutate adenine, thymine, guanine and / or cytosine in at least one target template nucleic acid molecule with a nucleotide analog.
[0397] For example, to determine whether a low-bias DNA polymerase is capable of incorporating nucleotide analogs into at least one target template nucleic acid molecule, a technician can use a low-bias DNA polymerase to amplify a nucleic acid molecule for two rounds of replication. The first round of replication should be performed in the presence of nucleotide analogs, and the second round of replication should be performed in the absence of nucleotide analogs. The resulting amplified nucleic acid molecules can be sequenced to see whether mutations have been introduced, and if so, how many mutations have been introduced. The user should repeat the experiment without nucleotide analogs and compare the number of mutations introduced with and without nucleotide analogs. If the number of mutations introduced with nucleotide analogs is significantly higher than the number of mutations introduced without nucleotide analogs, the user can conclude that the low-bias DNA polymerase is capable of incorporating nucleotide analogs. Similarly, a technician can use nucleotide analogs to determine whether a DNA polymerase incorporates nucleotide analogs or mutates adenine, thymine, guanine, and / or cytosine. The skilled person only needs to carry out the method in the presence of nucleotide analogues and see whether the method results in mutations at positions originally occupied by adenine, thymine, guanine and / or cytosine.
[0398] If the user wishes to mutate at least one target template nucleic acid molecule using nucleotide analogs, the method can include the step of amplifying at least one target template nucleic acid molecule using a low bias DNA polymerase, wherein the step of amplifying at least one target template nucleic acid molecule using a low bias DNA polymerase is performed in the presence of nucleotide analogs, and the step of amplifying at least one target template nucleic acid molecule provides at least one target template nucleic acid molecule comprising nucleotide analogs.
[0399] Suitable nucleotide analogs include dPTP (2'deoxy-P-nucleoside-5'-triphosphate), 8-oxo-dGTP (7,8-dihydro-8-oxoguanine), 5Br-dUTP (5-bromo-2'-deoxy-uridine-5'-triphosphate), 2OH-dATP (2-hydroxy-2'-deoxyadenosine-5'-triphosphate), dKTP (9-(2-deoxy-β-D-ribofuranosyl)-N6-methoxy-2,6,-diaminopurine-5'-triphosphate) and dITP (2'-deoxyinosine 5'-triphosphate). The nucleotide analog can be dPTP. Nucleotide analogs can be used to introduce the substitution mutations described in Table 1.
[0400] Table 1
[0401] Nucleotides Replacement 8-Oxo-dGTP A:T to C:G and T:A to G:C dPTP A:T to G:C and G:C to A:T 5Br-dUTP A:T to G:C and T:A to C:G 2OH-dATP A:T to C:G, G:C to T:A, and A:T to G:C dITP A:T to G:C and G:C to A:T dKTP A:T to G:C and G:C to A:T
[0402] Different nucleotide analogs can be used alone or in combination to introduce different mutations into at least one target template nucleic acid molecule. Thus, a low-bias DNA polymerase can use nucleotide analogs to introduce guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations. A low-bias DNA polymerase may optionally use nucleotide analogs to introduce guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations.
[0403] Low bias DNA polymerases may be capable of introducing guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at rates of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2:0.8-1.2, or approximately 1:1:1:1, respectively. Preferably, the low bias DNA polymerase is capable of introducing guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at rate ratios of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively. Suitable methods for determining whether a low bias DNA polymerase is capable of introducing substitution mutations and at what rate ratios are described under the heading "Whether a DNA polymerase is capable of mutating nucleotides, and if so, at what rate."
[0404] In some methods, the low bias DNA polymerase introduces guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at rates of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, 0.6-1.4:0.6-1.4:0.6-1.4:0.6-1.4, 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, 0.8-1.2:0.8-1.2:0.8-1.2, or about 1:1:1:1, respectively. Preferably, the low bias DNA polymerase introduces guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations at a rate ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively. Suitable methods for determining whether substitution mutations have been introduced and at what rate ratios are described under the heading "Whether a DNA polymerase can mutate nucleotides, and if so, at what rate."
[0405] Typically, when low-bias DNA polymerase uses nucleotide analogs to introduce mutations, this requires more than one round of replication. In the first round of replication, low-bias DNA polymerase introduces nucleotide analogs to replace nucleotides, and in the second round of replication, the nucleotide analogs are paired with natural nucleotides to introduce substitution mutations in the complementary chain. The second round of replication can be carried out in the presence of nucleotide analogs. However, the method may further include a step of amplifying at least one target template nucleic acid molecule comprising nucleotide analogs in the second sample of the paired sample in the absence of nucleotide analogs. The step of amplifying at least one target template nucleic acid molecule comprising nucleotide analogs in the absence of nucleotide analogs can be carried out using low-bias DNA polymerase.
[0406] Low template amplification bias
[0407] A low-bias DNA polymerase can have a low template amplification bias. A low-bias DNA polymerase has a low template amplification bias if it can amplify different target template nucleic acid molecules with a similar degree of success per cycle. A high-bias DNA polymerase may have difficulty amplifying template nucleic acid molecules that contain a high G:C content or a significant degree of secondary structure. In one embodiment, the low-bias DNA polymerase has a low template amplification bias for template nucleic acid molecules that are less than 25,000 nucleotides in length, less than 10,000 nucleotides in length, 1 to 15,000 nucleotides in length, or 1 to 10,000 nucleotides in length.
[0408] In one embodiment, in order to determine whether DNA polymerase has low template amplification bias, technicians can use DNA polymerase to amplify a series of different sequences, and check whether different sequences are amplified at different levels by sequencing the amplified DNA obtained. For example, technicians can select a series of short (possibly 50 nucleotides) nucleic acid molecules with different characteristics, including nucleic acid molecules with high GC content, nucleic acid molecules with low GC content, nucleic acid molecules with a large degree of secondary structure, and nucleic acid molecules with a low degree of secondary structure. The user can then use DNA polymerase to amplify those sequences and quantify the amplification level of each nucleic acid molecule in the nucleic acid molecule. In one embodiment, if the levels are within 25%, 20%, 10%, or 5% of each other, the DNA polymerase has a low template amplification bias.
[0409] Alternatively, in one embodiment, a DNA polymerase has low template amplification bias if it is able to amplify a fragment of 7 kbp-10 kbp with a Kolmolgorov-Smirnov D of less than 0.1, less than 0.09, or less than 0.08. The Kolmolgorov-Smirnov D with which a particular low-bias DNA polymerase is able to amplify a fragment of 7 kbp-10 kbp can be determined using the assay provided in Example 4.
[0410] Low biased DNA polymerase can be a high-fidelity DNA polymerase.High-fidelity DNA polymerase is a kind of DNA polymerase that is not very prone to error, therefore when high-fidelity DNA polymerase is used for amplification of target template nucleic acid molecules in the absence of nucleotide analogs, a large amount of sudden changes can not be introduced usually.High-fidelity DNA polymerase is not usually used in the method for introducing sudden changes, because it is generally believed that mistake-prone DNA polymerase is more effective.However, the application has demonstrated that some high-fidelity polymerase can use nucleotide analogs to introduce sudden changes, and compared with mistake-prone DNA polymerase (such as Taq polymerase), those sudden changes can be introduced with lower bias.
[0411] High-fidelity DNA polymerases offer other advantages. When used with nucleotide analogs, high-fidelity DNA polymerases can be used to introduce mutations, but in the absence of nucleotide analogs, high-fidelity DNA polymerases can replicate the target template nucleic acid molecule with high accuracy. This means that users can use the same DNA polymerase to efficiently mutate at least one target template nucleic acid molecule and amplify the mutated at least one target template nucleic acid molecule with high accuracy. If a low-fidelity DNA polymerase is used to mutate the target template nucleic acid molecule, it may be necessary to remove the low-fidelity DNA polymerase from the reaction mixture before amplifying the target template nucleic acid molecule.
[0412] High-fidelity DNA polymerases can have proofreading and reading activity. Proofreading and reading activity can help DNA polymerases amplify target template nucleic acid sequences with high precision. For example, low-bias DNA polymerases can include a proofreading and reading domain. The proofreading and reading domain can confirm whether the nucleotides added by the polymerase are correct (checking that they are correctly paired with the corresponding nucleic acid of the complementary chain). If incorrect, the nucleotides are excised from the nucleic acid molecule. The inventors have surprisingly found that in some DNA polymerases, the proofreading and reading domain will accept the pairing of natural nucleotides and nucleotide analogs. The structure and sequence of suitable proofreading and reading domains are known to the technician. DNA polymerases comprising a proofreading and reading domain include members of DNA polymerase families I, II, and III, such as Pfu polymerase (derived from Pyrococcus furiosus), T4 polymerase (derived from bacteriophage T4), and the Thermococcal polymerase described in detail below.
[0413] In one embodiment, the high-fidelity DNA polymerase introduces less than 0.01%, less than 0.0015%, less than 0.001%, between 0% and 0.0015%, or between 0% and 0.001% mutations per round of replication in the absence of nucleotide analogs.
[0414] In addition, the low-bias DNA polymerase may comprise a processivity enhancing domain. The processivity enhancing domain allows the DNA polymerase to amplify the target template nucleic acid molecule more quickly. This is advantageous because it allows the method of the present invention to be performed more quickly.
[0415] Thermococcal polymerase
[0416] In one embodiment, the low-bias DNA polymerase is a fragment or variant of a polypeptide comprising SEQ ID NO. 2, SEQ ID NO. 4, SEQ ID NO. 6, or SEQ ID NO. 7. The polypeptides of SEQ ID NOs. 2, 4, 6, and 7 are Pyrococcus polymerases. The polymerases of SEQ ID NOs. 2, 4, 6, or 7 are high-fidelity, low-bias DNA polymerases that can mutate target template nucleic acid molecules by incorporating nucleotide analogs (e.g., dPTP). The polymerases of SEQ ID NOs. 2, 4, 6, or 7 are particularly advantageous because they have low mutation bias and low template amplification bias. The polymerases of SEQ ID NOs. 2, 4, 6, or 7 are also highly processive and contain a proofreading domain, meaning they can rapidly and accurately amplify mutated target template nucleic acid molecules in the absence of nucleotide analogs.
[0417] The low bias DNA polymerase can comprise a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 contiguous amino acids of the following sequence:
[0418] a. Sequence of SEQ ID NO. 2;
[0419] b. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.2;
[0420] c. Sequence of SEQ ID NO.4;
[0421] d. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.4;
[0422] e. Sequence of SEQ ID NO.6;
[0423] f. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.6;
[0424] g. the sequence of SEQ ID NO.7; or
[0425] h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO. 7.
[0426] Preferably, the low bias DNA polymerase comprises a fragment of at least 700 consecutive amino acids of the following sequence:
[0427] a. Sequence of SEQ ID NO. 2;
[0428] b. a sequence that is at least 98% or at least 99% identical to SEQ ID NO.2;
[0429] c. Sequence of SEQ ID NO.4;
[0430] d. a sequence that is at least 98% or at least 99% identical to SEQ ID NO.4;
[0431] e. Sequence of SEQ ID NO.6;
[0432] f. a sequence that is at least 98% or at least 99% identical to SEQ ID NO.6;
[0433] g. the sequence of SEQ ID NO.7; or
[0434] h. A sequence that is at least 98%, or at least 99%, identical to SEQ ID NO. 7.
[0435] Low-bias DNA polymerases may include:
[0436] a. Sequence of SEQ ID NO. 2;
[0437] b. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.2;
[0438] c. Sequence of SEQ ID NO.4;
[0439] d. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.4;
[0440] e. Sequence of SEQ ID NO.6;
[0441] f. a sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO.6;
[0442] g. the sequence of SEQ ID NO.7; or
[0443] h. A sequence that is at least 95%, at least 98%, or at least 99% identical to SEQ ID NO. 7.
[0444] Preferably, the low bias DNA polymerase comprises:
[0445] a. Sequence of SEQ ID NO. 2;
[0446] b. a sequence that is at least 98% or at least 99% identical to SEQ ID NO.2;
[0447] c. Sequence of SEQ ID NO.4;
[0448] d. a sequence that is at least 98% or at least 99% identical to SEQ ID NO.4;
[0449] e. Sequence of SEQ ID NO.6;
[0450] f. a sequence that is at least 98% or at least 99% identical to SEQ ID NO.6;
[0451] g. the sequence of SEQ ID NO.7; or
[0452] h. A sequence that is at least 98%, or at least 99%, identical to SEQ ID NO. 7.
[0453] The low bias DNA polymerase can be Pyrococcus polymerase or a derivative thereof. The DNA polymerases of SEQ ID NOs. 2, 4, 6, and 7 are Pyrococcus polymerases. Pyrococcus polymerase is advantageous because it is generally a high-fidelity polymerase that can be used to introduce mutations using nucleotide analogs with low mutation and template amplification bias.
[0454] Thermococcus polymerase is a polymerase having the polypeptide sequence of a polymerase isolated from a strain of the genus Pyrococcus. A derivative of the Pyrococcus polymerase can be a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of the Pyrococcus polymerase, or a fragment of at least 400, at least 500, at least 600, at least 700, or at least 750 consecutive amino acids of the Pyrococcus polymerase that is at least 95%, at least 98%, at least 99%, or 100% identical to the Pyrococcus polymerase. A derivative of the Pyrococcus polymerase can be at least 95%, at least 98%, at least 99%, or 100% identical to the Pyrococcus polymerase. A derivative of the Pyrococcus polymerase can be at least 98% identical to the Pyrococcus polymerase.
[0455] In the context of the present invention, the Pyrococcus polymerase from any strain can be effective. In one embodiment, the Pyrococcus polymerase is derived from a Pyrococcus strain selected from the group consisting of T. kodakarensis, T. celer, T. siculi, and T. sp KS-1. The Pyrococcus polymerases from these strains are described in SEQ ID NO. 2, SEQ ID NO. 4, SEQ ID NO. 6, and SEQ ID NO. 7.
[0456] Alternatively, the low bias DNA polymerase is a polymerase that has high catalytic activity at temperatures between 50°C and 90°C, between 60°C and 80°C, or about 68°C.
[0457] Example
[0458] Example 1 - Mutagenesis of Nucleic Acid Molecules Using PrimeStar GXL or Other Polymerases
[0459] DNA molecules are fragmented to an appropriate size (eg, 10 kb) using tag fragmentation, and defined sequence primer sites (adapters) are ligated to each end.
[0460] The first step is a tag fragmentation reaction to fragment the DNA. Under the following conditions, 50 ng of high molecular weight genomic DNA from one or more bacterial strains in a volume of 4 μl or less is subjected to tag fragmentation. 50 ng of DNA is combined with 4 μl of Nextera transposase (diluted to 1:50) and 8 μl of 2× tag fragmentation buffer (20 mM Tris [pH 7.6], 20 mM MgCl, 20% (v / v) dimethylformamide) in a total volume of 16 μl. The reaction is incubated at 55°C for 5 minutes, 4 μl of NT buffer (or 0.2% SDS) is added to the reaction, and the reaction is incubated at room temperature for 5 minutes.
[0461] Tag fragmentation reactions were cleaned using SPRIselect beads (Beckman Coulter) according to the manufacturer's instructions, with left-side size selection performed using 0.6 volumes of beads and DNA eluted in molecular-grade water.
[0462] PCR was then performed with standard dNTPs and dPTPs in a limited 6-cycle reaction using Primestar GXL, with 12.5 ng of tagged, fragmented, and purified DNA added in a total reaction volume of 25 μl containing: 1× GXL buffer; 200 μM each of dATP, dTTP, dGTP, and dCTP; 0.5 mM dPTP, and 0.4 μM custom primers (Table 2).
[0463] Table 2:
[0464]
[0465] Table 2. Custom primers used for mutagenic PCR on a 10 kbp template. XXXXXX is a defined, sample-specific 6-8 nt barcode (sample tag) sequence. NNNNNN is a 6 nt random nucleotide region.
[0466] The reaction was thermally cycled in the presence of Primestar GXL as follows: initial gap extension at 68°C for 3 minutes, followed by 6 cycles of 98°C for 10 seconds, 55°C for 15 seconds, and 68°C for 10 minutes.
[0467] The next stage was a dPTP-free PCR to remove the dPTP from the template and replace it with a conversion mutation ("recovery PCR"). The PCR reaction was cleaned with SPRIselect beads to remove excess dPTP and primers, and then amplified for an additional 10 cycles (minimum 1, maximum 20) using primers that annealed to the ends of the fragments introduced during the dPTP incorporation cycles (Table 3).
[0468] Table 3
[0469] i7 Flow Cell Primers CAAGCAGAAGACGGCATACGA i5 flow cell primers AATGATACGGCGACCACCGA
[0470] Then carry out gel extraction step, with the fragment of amplification and mutation selected in the required size range by size, for example 7kb-10kb. Gel extraction can be completed manually or by automation system (for example BluePippin). Then carry out another round of PCR, 16-20 cycles ("enrichment PCR").
[0471] After amplifying a defined number of long mutant templates, the templates are randomly fragmented to generate a set of overlapping shorter fragments for sequencing. Fragmentation is performed by tag fragmentation.
[0472] The long DNA fragments from the previous step are subjected to a standard tag fragmentation reaction (e.g., Nextera XT or Nextera Flex), except that the reaction is divided into three pools for PCR amplification. This enables selective amplification of fragments derived from each end of the original template (including the sample tag), as well as internal fragments from the long template that have been fragmented with new tags at both ends. This effectively creates three pools for sequencing on an Illumina instrument (e.g., MiSeq or HiSeq).
[0473] The method was repeated using standard Taq (Jena Biosciences) and a mixture of Taq and a proofreading polymerase (DeepVent) called LongAmp (New England Biolabs).
[0474] The data obtained from this experiment are Figure 1 Depicted. No dPTP was used as a control. Reads were mapped against the E. coli genome, achieving a median mutation rate of ~8%.
[0475] Example 2 - Comparison of mutation frequencies of different DNA polymerases
[0476] Mutagenesis was performed with a range of different DNA polymerases (Table 4). Genomic DNA from E. coli strain MG1655 was tagged to generate long fragments and beads were cleaned as described in Example 1. Six cycles of "mutagenesis PCR" were then performed in the presence of 0.5 mM dPTP, followed by SPRIselect bead purification and an additional 14-16 cycles of "recovery PCR" in the absence of dPTP. The resulting long mutant templates were then subjected to a standard tag fragmentation reaction (see Example 1), and the "internal" fragments were amplified and sequenced on an Illumina MiSeq instrument.
[0477] The mutation rates are described in Table 4, where the frequencies of base substitutions by dPTP mutagenesis reactions are normalized as measured using Illumina sequencing of DNA from a known reference genome. For Taq polymerase, only ~12% of mutations occurred at the template G+C site, even when used in a buffer optimized for Pyrococcus polymerase. Thermococcus-like polymerases produced 58%-69% mutations at the template G+C site, while the polymerase derived from Pyrococcus produced 88% mutations at the template G+C site.
[0478] Enzymes were obtained from Jena Biosciences (Taq), Takara (Primestar variants), Merck Millipore (KOD DNA polymerase), and New England Biolabs (Phusion).
[0479] Taq was tested using the supplied buffer, and Primestar GXL buffer (Takara) was also used for this experiment. All other reactions were performed using the standard buffer provided for each polymerase.
[0480] Table 4
[0481]
[0482] Example 3 - Determination of dPTP mutagenesis rate
[0483] We performed dPTP mutagenesis on a series of genomic DNA samples with varying levels of G+C content (33%-66%) using Thermococcus polymerase (Primestar GXL; Takara) under a single set of reaction conditions. Mutagenesis and sequencing were performed as described in Example 1, except that 10 cycles of "recovery PCR" were performed. As expected, despite the diversity in G+C content, the mutation rates were roughly similar between samples (median rate of 7%-8%) ( Figure 2 ).
[0484] Example 4 - Measuring Template Amplification Bias
[0485] The template amplification bias of two polymerases was measured: Kapa HiFi, a proofreading polymerase commonly used in Illumina sequencing protocols, and PrimeStar GXL, a KOD family polymerase known for its ability to amplify long fragments. In the first experiment, Kapa HiFi was used to amplify a limited number of E. coli genomic DNA templates, approximately 2 kbp in size. The ends of these amplified fragments were then sequenced. A similar experiment was performed using PrimeStar GXL on fragments of approximately 7 kbp to 10 kbp from E. coli. The position of each end sequence read was determined by mapping relative to the E. coli reference genome. The distances between adjacent fragment ends were measured. These distances were compared to a set of distances randomly sampled from a uniform distribution. D was compared using the nonparametric Kolmolgorov-Smirnov test. When two samples come from the same distribution, the value of D is close to zero. For the low-bias PrimeStar polymerase, D = 0.07 was observed when measuring 50,000 fragment ends, compared to a uniform random sample of 50,000 genomic positions. For Kapa HiFi polymerase, we observed D = 0.14 at the ends of 50,000 fragments.
[0486] Example 5 - Measuring the size range of reconstruction
[0487] Mutant sequence reads and non-mutant sequence reads are generated, and the sequence of the non-mutant sequence reads is determined using computer-implemented method steps.
[0488] To generate mutant sequence reads, mutant target template nucleic acid molecule fragments were generated using the method described in Fragment Example 1, except that the fragment size range was limited to 1 kb-2 kb. The mutant target template nucleic acid molecule fragments were sequenced using an Illumnia MiSeq with a V2500 circulating flow cell.
[0489] To generate non-mutation sequence reads, the following steps were performed. The first step was a tag fragmentation reaction to fragment the DNA. 50 ng of high molecular weight genomic DNA from one or more bacterial strains in a volume of 4 μl or less was subjected to tag fragmentation under the following conditions. 50 ng of DNA was mixed with 4 μl of Nextera transposase (diluted to 1:50) and 8 μl of 2× tag fragmentation buffer (20 mM Tris [pH 7.6], 20 mM MgCl, 20% (v / v) dimethylformamide) in a total volume of 16 μl. The reaction was incubated at 55°C for 5 minutes, 4 μl of NT buffer (or 0.2% SDS) was added to the reaction, and the reaction was incubated at room temperature for 5 minutes.
[0490] The labeling reaction was cleaned using SPRIselect beads (Beckman Coulter) according to the manufacturer's instructions, with 0.6 volumes of beads used for left size selection and DNA elution in molecular grade water. The long DNA fragments from the previous step were subjected to a standard label fragmentation reaction (e.g., NexteraXT or NexteraFlex), except that the reaction was divided into three pools for PCR amplification. This enables selective amplification of the fragments derived from each end of the original template (including the sample label) and the internal fragments of the long template that had been newly labeled fragmented at both ends. This effectively creates three pools for sequencing on an Illumina instrument (e.g., MiSeq or HiSeq).
[0491] The sequence of the target template nucleic acid molecule was determined by pre-clustering the mutant sequence reads into read groups and then de novo assembling each group of mutant reads using steps 1 and 2 of the A5-miseq assembly pipeline (Coil et al 2015 Bioinformatics). This analysis generated 53,053 virtual fragments with a length distribution of Figure 4 shown.
[0492] Example 6 - Test Probability Algorithm
[0493] A probabilistic algorithm is used to determine whether two mutant sequence reads originate from the same original at least one template nucleic acid molecule. The details of the probabilistic algorithm are as follows.
[0494] Given two non-mutated sequence reads S1 and S2 that have been aligned to a non-mutated reference sequence R, the model described here attempts to determine whether S1 and S2 have been sequenced from the same template nucleic acid molecule with at least one mutation or from different templates. The alignment of these three sequences can be represented as a 3×N matrix N of alignment positions, such as a single nucleotide s 1,i :s 2,j :rk N 3-tuples, the aligned nucleotides appear in the same column y of N, for example n .,y For convenience, we define the nucleotides A, C, G, and T to be mapped to integers 1, 2, 3, and 4, so that A is mapped to 1, C is mapped to 2, and so on. This mapping is implicit in the rest of the description below. Then, we define two 4×4 probability matrices: M and E. Each entry m i,j records the probability that nucleotide i mutates to nucleotide j through the mutagenesis process, i, j ∈ {A, C, G, T}. Similarly, the entry e i,j Record the conditional probability of incorrectly reading nucleotide i as nucleotide j, i, j ∈ {A, C, G, T}, conditional on the incorrect reading of the nucleotide. Further, define a 2×N matrix Q, where the entry q 1,y and q 2,y represents the probability that the nucleotide at position y of sequences S1 and S2 reported by the sequencing instrument is incorrectly read. Finally, z∈{0,1} is used as an indicator value for whether two sequence reads originate from the same mutant template, where z=1 indicates that S1 and S2 have been sequenced from the same template fragment, and z=0 indicates that S1 and S2 have been sequenced from different template fragments.
[0495] The values of Q and N are provided / determined by the sequencing and subsequent reading profiling process, but the values of M, E, and z are generally unknown. Fortunately, these values (and any other unknown parameters) can be estimated from the data using any of a variety of techniques. A prior distribution can be imposed on the values of the unknown parameters based on knowledge of the mutational process. A Dirichlet distribution is imposed on the rows of M such that: 1,· ~Dirichlet(α+β, 1-β, 1-α, 1-β), where the entries correspond to the events A→A (no mutation), A→C (transversion), A→G (transition), A→T (transversion). Here, α is an unknown hyperparameter for the transition rate, and β is an unknown hyperparameter for the transversion rate. The complete prior for M is specified as:
[0496] m1,·~Dirichlet(α+β,1-β,1-α,1-β)
[0497] m2,·~Dirichlet(1-β,α+β,1-β,1-α)
[0498] m3,·~Dirichlet(1-α,1-β,α+β,1-β)
[0499] m4,·~Dirichlet(1-β, 1-α, 1-β, α+β)
[0500] Experimenters often have access to prior knowledge of the mutation process (e.g., knowledge of the properties of the polymerase or other mutagen), and this can allow for the application of hyperpriors on the α and β terms. Priors on M may have more general structures. Uniform priors are applied to the matrices E as well as z.
[0501] Given the above notation, the likelihood of the data given the model can be expressed as:
[0502] P(N,Q|M,E,z)=π =1 (z)f(N,Q|M,E,i)+(1-z)g(N,Q|M,E,i)
[0503] in:
[0504]
[0505]
[0506] Here, the center point of the matrix subscript represents all members of the row or column, and vector multiplication represents the dot product. 1_{} is the indicator function, which is 1 if the expression in the subscript is true, and 0 otherwise.
[0507] Combining the likelihood with the aforementioned prior yields the elements necessary for Bayesian inference about unknown values. There are many methods for implementing Bayesian inference, including exact methods for analytically tractable posterior probability distributions and a range of Monte Carlo and related methods for approximating posterior distributions. In the present case, the model is implemented in the Stan modeling language (see Code Listing X1), which facilitates inference using Hamiltonian Monte Carlo as well as variational inference using mean-field and full-order approximations. The variational inference approximation used depends on stochastic gradient descent to maximize the evidence lower bound (ELBO) (Kucukelbir et al 2015 https: / / arxiv.org / abs / 1506.03431), which requires the probability model to be continuous and differentiable. To meet this requirement, z is implemented as a continuous parameter on the support [0, 1], and a β(0.1, 0.1) distribution is applied for sparsification before concentrating the posterior mass of z around 0 and 1. This approach of using a continuous relaxation of a discrete random variable is called a “concrete distribution” and is used in [1]. https: / / arxiv.org / abs / 161.00712 Fitting the model using variational inference on a collection of about 100 simulated sequence alignments of at least 100 bases in length takes only a few minutes of CPU time on a laptop computer and yields the posterior for the unknown parameters. Figure 5 The posterior distribution of the model parameters is shown.
[0508] Although variational inference is faster than many Monte Carlo methods, it is not yet sufficient to analyze the millions of sequence reads generated in a typical sequencing run. Therefore, a faster method was developed to calculate the probability that two reads r0 and r1 are derived from or not from the same target template nucleic acid molecule with at least one mutation. Given the mutagenesis process and sequencing errors, these probabilities can be expressed as:
[0509] P 相同_模板 (r0, r1)=P(N, Q|M, E, z=1)=Π =1 f(N, Q|M, E, i) (Equation 1)
[0510] P 不同_模板 (r0, r1)=P(N, Q|M, E, z=0)=Π =1 g(N, QIM, E, i) (Equation 2)
[0511] Here, the values of M and E have been fixed to the maximum a posteriori probability or similar values with high a posteriori probability, which are determined by Bayesian (or maximum likelihood) inference using a small subset of the total data set. The values of N and Q correspond to the alignment of r0 and r1 with the reference sequence. The log-odds score of two reads derived from a common template can then be simply calculated as:
[0512] Score = log(P 相同_模板 )-log(P 不同_模板 ) (Formula 3)
[0513] If the pairwise score of mutant sequence reads is above a certain predetermined threshold, they are considered to originate from the same at least one target template nucleic acid molecule. In the current case, this value is set to 1000. Tests on simulated data show that this log-odds score can distinguish whether two mutant reads originate from the same at least one target template nucleic acid molecule with high precision and recall ( Figure 6 ).
[0514] Example 7 - Preferential amplification of longer templates using two identical primer binding sites and a single primer sequence
[0515] As mentioned above, tag fragmentation can be used to fragment DNA molecules and simultaneously introduce primer binding sites (adapters) into the ends of the fragments. The Nextera tag fragmentation system (Illumina) utilizes a transposase loaded with one of two unique adaptors (referred to herein as x and Y). This generates a random product mixture, some of which have identical end sequences (XX, YY), while others have unique ends (XY). The standard Nextera protocol uses two different primer sequences to selectively amplify "XY" products that contain different adaptors (necessary for sequencing using Illumina technology) on each end. However, a single primer sequence can also be used to amplify "XX" or "YY" fragments with identical end adaptors.
[0516] To generate long mutant templates containing identical end adapters, 50 ng of high molecular weight genomic DNA (E. coli strain MG1655) was first tag-fragmented and cleaned up using SPRIselect beads as described in Example 1. Five cycles of "mutagenesis PCR" were then performed with standard dNTPs and dPTPs as described in detail in Example 1, except that a single primer sequence was used (Table 5).
[0517] The PCR reaction was cleaned up with SPRIselect beads to remove excess dPTP and primers, and then an additional 10 cycles of "recovery PCR" were performed in the absence of dPTP to replace the dPTP in the template with the conversion mutation. The recovery PCR was performed with a single primer that annealed to the end of the fragment introduced during the dPTP incorporation cycle, thereby enabling selective amplification of the mutant template generated in the previous PCR step.
[0518] Table 5:
[0519]
[0520] Table 5. Primers used to generate mutant templates with the same basic adapter structure at both ends. Primer "single_mut" was used for mutagenesis PCR on DNA fragments generated by Nextera tag fragmentation. This primer contains a 5' portion that introduces an additional primer binding site at the end of the fragment. Primer "single_rec" anneals to this site and is used during recovery PCR to selectively amplify mutant templates generated with the single_mut primer. XXXXXXXXXXXXX is a defined, sample-specific 13 nt tag sequence. NNN is a 3 nt random nucleotide region.
[0521] As a control, a mutant template with different adapters at each end was generated using the same protocol as above, except that two different primer sequences were used during both the mutagenesis PCR (see Table 2) and the recovery PCR (see Table 3). The final PCR product was cleaned up with SPRIselect beads and analyzed on a high-sensitivity DNA chip using a 2100 Bioanalyzer system (Agilent). Figure 10 As shown, templates generated using the same end adapters were significantly longer on average than the control samples containing dual adapters. The smallest detectable control template size was ~800 bp, while no templates under 2000 bp were observed in the single adapter samples.
[0522] Mutant templates with identical end adapters (blue) and a control template with dual adapters were run on an Agilent 2100 Bioanalyzer (High Sensitivity DNA Kit) to compare size characteristics. The use of identical end adapters inhibited amplification of templates <2 kbp. Data are presented in Figure 10 middle.
[0523] Example 8 - Sample dilution and end sequencing to quantify DNA template
[0524] Initial samples of long mutant templates for analysis are diluted to a defined number of unique template molecules in preparation for downstream processing, sequencing, and analysis, ensuring that each template generates sufficient sequence data for efficient template assembly.
[0525] First, long mutation templates were prepared from human genomic DNA (genome NA12878) using the methods outlined in Example 7. Five cycles of mutagenic PCR and six cycles of recovery were performed, followed by gel extraction to select templates in the 8-10 kb size range. Templates flanked by identical adapter sequences were generated using the primers shown in Table 5.
[0526] The template sample of the selected size is then serially diluted in 10-fold steps, and DNA sequencing is then used to determine the number of unique templates present in each dilution. This involves first amplifying the diluted sample to produce many copies of each unique template. PCR is performed using a single primer (5'-CAAGCAGAAGACGGCATACGA-3') that anneals to the ends of the fragments introduced in the previous recovery PCR step, thereby selectively amplifying templates that have completed the dPTP incorporation and substitution process to produce conversion mutations. A total of 16-30 PCR cycles (depending on the sample dilution factor) are required to generate enough material for downstream processing.
[0527] Each PCR product is then fragmented using a standard tag fragmentation reaction (see Example 1), and fragments from the ends of the template (including the sample tag and the unique molecular tag) are selectively amplified in preparation for Illumina sequencing. This is achieved using a pair of primers, one primer that specifically anneals to the end of the original template (5'-CAAGCAGAAGACGGCATACGA-3') and the other primer that anneals to the adapter introduced during the tag fragmentation process (i5 custom index primer; Table 2). After the samples are sequenced on the Illumina MiSeq instrument, the unique template is identified based on the sequence information corresponding to the very end of the original template molecule. To this end, a clustering algorithm (such as vsearch) is used to group together reads with the same sequence that may come from the same original unique template. Other types of sequence information, such as unique molecular tags, can also be used for this purpose. As Figure 11 As shown, a clear linear relationship is observed between the sample dilution factor and the number of unique templates observed. Using this information, the precise dilution factor required to control the number of mutant target template nucleic acid molecules in the second sample to the desired number of unique templates can be determined in preparation for subsequent sequencing and template assembly.
[0528] Example 9 - Dilution and end sequencing to normalize pooled template samples
[0529] The sample dilution and end sequencing method described above was used to quantify multiple template pools in the pooled preparative samples. This information was then used to normalize the number of templates between individual samples in the pooled sample.
[0530] First, as described in Example 5, the genomic DNA samples from 96 different bacterial strains were subjected to tag fragmentation and 5 rounds of mutagenesis PCR. For each reaction, a single primer (single_mut design; Table 5) with a unique sample tag was used. The mutagenesis products of each sample tag were then merged, and the merged samples were cleaned with SPRIselect beads to remove excessive dPTP and primers. Subsequently, the single_rec primer (Table 5) was used to perform 6 rounds of recovery PCR and gel extraction to select the template of the 8kb-10kb size range. The template sample merged was then diluted at 1:1000, and terminal sequencing was performed to determine the number of the unique template present in each bacterial strain in the dilution pool. This was achieved using the method outlined in Example 7.
[0531] It was found that the template counts varied widely between strains in the dilution pool, ranging from undetectable templates in several strains to more than 1000 unique templates in other strains. 66 strains with non-zero template counts were selected for normalization.
[0532] Based on the observed number of templates and the known genome size of each strain, a normalization pool was prepared by combining different volumes of sample-tagged mutagenic PCR products, with the goal of obtaining a constant number of unique templates per unit genome content (e.g., per Mb) for each strain. The normalization pool was then subjected to end sequencing as described above, and the number of unique templates for each strain was determined. As expected, the variation in template counts between strains after normalization was minimal ( Figure 12 ).
[0533] Example 10 - Assembling bacterial genome sequences using assembly algorithms
[0534] Bacterial strains and DNA preparation
[0535] DNA from 62 bacterial strains was obtained from the BEI resource. These strains were isolates sequenced as part of the Human Microbiome Project. They represent a range of GC content (25% to 69%), with more details provided in Table 6.
[0536] Table 6
[0537]
[0538]
[0539]
[0540]
[0541] As controls, three additional strains with well-characterized genomes (E. coli K12 MG1655, S. aureus ATCC 25923, and H. volcanii DS2) were included, which also cover a wide range of GC content. DNA was prepared from these strains using the Qiagen DNeasy UltraClean Microbial Kit according to the manufacturer's instructions with the following modifications. Overnight cultures (20 mL per strain) were centrifuged at 3200 g for 5 minutes to obtain cell pellets, and each pellet was washed with 5 mL of sterile 0.9% sodium chloride solution. Each pellet was resuspended in 300 ul of PowerBead solution before continuing with the manufacturer's protocol. For E. coli and S. aureus, DNA was eluted with 50 uL of elution buffer preheated to 42°C, while H. volcanii DNA was eluted in 35 uL of elution buffer.
[0542] DNA concentrations were measured for all samples using the Quant-iT PicoGreen dsDNA kit (Thermo Scientific). For a subset of species, DNA purity and molecular weight were also assessed by Nanodrop (Thermo Scientific) spectrophotometry and agarose gel electrophoresis.
[0543] Morphoseq library preparation
[0544] Tag fragmentation to generate long fragments
[0545] DNA from each bacterial genome was arrayed into a 96-well plate and the concentration was standardized to 10 ng / ul. Escherichia coli MG1655 DNA was included in two independent wells to provide an internal control for sample processing and downstream data analysis. Tag fragmentation was performed using Nextera DNA tag fragmentation enzyme (Tagment Enzyme; Illumina) diluted 1 to 50 in storage buffer (5 mM Tris-HCl [pH 8.0], 0.5 mM EDTA, 50% (v / v) glycerol). For each sample, a 16 μl tag fragmentation reaction containing 50 ng DNA and 4 μl diluted TDE1 was prepared in 1 × tag fragmentation buffer (10 mM Tris-HCl [pH 7.6], 10 mM MgCl, 10% (v / v) dimethylformamide). Each reaction was incubated at 55°C for 5 minutes and then cooled to 10°C. SDS was added to a final concentration of 0.04% and the reaction was incubated for an additional 15 minutes at 25° C. The reaction was left cleaned using SPRIselect magnetic beads (Beckman Coulter) (0.6× volume of magnetic beads) and eluted in 20 μl of molecular grade water according to the manufacturer's instructions.
[0546] Mutagenesis of long DNA fragments
[0547] PCR incorporating the mutagenic nucleotide analog dPTP was performed as follows. 5 μl of each of the cleaned tag fragmentation reactions described above was used as a template in a 25 μl PCR reaction containing 0.625U PrimeStar GXL polymerase, 1× Primestar GXL buffer, and 0.2mM dNTPs (all purchased from Takara), as well as 0.5mM dPTP (TriLink Biotechnologies) and 0.4mM Morphoseq index primers (see Table 7; unique index for each sample). A single primer was used during mutagenic PCR to amplify a template containing the same Nextera tag fragmentation adapter sequence at both ends. The reaction was performed under the following cycling conditions: 3 minutes at 68°C, followed by 10 seconds at 98°C, 15 seconds at 55°C, and 10 minutes at 68°C for 5 cycles.
[0548] At this point, equal volumes of each reaction (4 μl) were combined into a single pool, which was then subjected to another SPRIselect Left Bead Cleanup using a 0.6× volume of beads. The purified pool was eluted in 45 μl of molecular grade water and quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific).
[0549] Then, in the absence of dPTP, the template sample comprising dPTP that is merged is further amplified, thereby replacing nucleotide analogs with natural dNTPs, and generating conversion mutations by the contradictory base pairing properties of dPTP. This " recovery " PCR comprises 1.25U PrimeStar GXL polymerase, 1× Primestar GXL buffer and 0.2mM dNTP (Takara), as well as 0.4 μM recovery primers (see Table 7) and 10ng of the template sample merged in a total volume of 50 μl. The reaction is carried out at 98°C for 10 seconds, 55°C for 15 seconds and 68°C for 10 minutes for 6 cycles.
[0550] Long template size selection
[0551] The PCR products were recovered according to size selection to remove unwanted short fragments using DNA gel electrophoresis. 25 μl of the recovered PCR reaction and DNA size standards were loaded onto a 0.9% agarose gel and run overnight (900 min) at 18 V in 1× TBE buffer. Gel slices corresponding to the 8 kb-10 kb size region were cut and DNA was extracted using Wizard SV Gel and PCR Clean-Up Kit (Promega) according to the manufacturer's instructions. Size-selected DNA was quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific) and the size range was confirmed using a Bioanalyzer High Sensitivity DNA Chip (Agilent).
[0552] Template normalization and quantification
[0553] The following method was used to assess the abundance of templates between samples tagged with a single sample within the pooled and size-selected products. First, the size-selected DNA was diluted to 0.1 pg / μl, and 2 μl of the dilution (0.2 pg) was used as input for the enrichment PCR to make many copies of each unique template. Preliminary experiments showed that this dilution level limited the diversity of unique templates sufficiently to allow accurate template quantification from the sequence output of a single Illumina MiSeq run. The 50 μl enrichment PCR also contained 1.25 U PrimeStar GXL polymerase, 1× Primestar GXL buffer and 0.2 mM dNTPs (Takara), as well as 0.4 μM enrichment primers (see Table 7). The enrichment primers were designed to anneal to the fragment end adapters introduced in the previous recovery PCR step, thereby selectively amplifying templates that had completed the dPTP incorporation and substitution process to produce conversion mutations. The reaction was performed for 22 cycles at 98°C for 10 seconds, 55°C for 15 seconds, and 68°C for 10 minutes, followed by purification by SPRIselect left bead cleanup using a 0.6× volume of beads and elution into 20 μl of molecular grade water. The samples were then quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific) and the size range was confirmed using a Bioanalyzer High Sensitivity DNA Chip (Agilent).
[0554] Next, the full-length enrichment product was fragmented by a second tag fragmentation reaction, and fragments derived from the ends of the original template (including the sample barcode) were amplified for Illumina sequencing. As described above, tag fragmentation was performed for the generation of long templates, except that 2 ng of starting DNA was used instead of 50 ng. After SDS treatment, the end library PCR reaction was prepared by adding KAPA HiFiHotStart ReadyMix (Kapa Biosystems) to a final concentration of 1×, together with 0.23 μM enrichment primers (annealed to the Illumina p7 flow cell adapter at the end of the full-length template) and 0.23 μM custom i5 index primers (annealed to the internal adapter introduced during the second round of tag fragmentation; see Table 7). The reaction was cycled as follows: 72°C for 3 minutes; 98°C for 30 seconds; 98°C for 15 seconds, 55°C for 30 seconds and 72°C for 30 seconds for 12 cycles; followed by a final extension at 72°C for 5 minutes. The terminal library was purified and quantified as above to obtain the full-length enriched product.
[0555] Illumina sequencing was performed on a MiSeq using V3 chemistry and generated 2×75nt paired-end reads. Unique template counts were determined for each individual bacterial genome sample in the dilution pool by first demultiplexing the end read data based on index 1 (i7) read sequence and then mapping the two read sequences (corresponding to the extreme ends of the original genomic insert) to the publicly available reference genome for each strain. The number of unique templates was calculated by counting the number of unique mapping start sites (corresponding to the start or end of the template), noting that two sites were expected for each template.
[0556] The observed template counts for each genome in the dilution pool varied, ranging from undetectable templates for some samples to over 1,000 unique templates for other samples. For simplicity, 66 samples with non-zero template counts were selected for further processing, sequencing, and assembly. Based on the observed template counts and the known genome size of each of these samples, normalization pools were prepared by combining different volumes of original barcoded mutagenic PCR products with the goal of obtaining a constant number of unique templates per unit genome content (e.g., per Mb) for each strain. To verify that normalization had been successful, the normalized pools were further processed for template quantification by repeating all subsequent steps of library preparation and sequencing described above (recovery PCR, size selection, template dilution and enrichment, end library preparation, Illumina sequencing, and analysis). As expected, the variation in template counts between strains after normalization was small ( Figure 11 ).
[0557] Template bottlenecking, enrichment, and short-read library processing
[0558] Based on the template quantification data in the standardized sample library and the known long fragment sizes, we chose to process a total of 1.5 million unique templates for Morphoseq sequencing and assembly. This will ensure at least 20 times the theoretical long template coverage of each individual genome (up to 90 times). To this end, the final long template sample was prepared by diluting the recovered PCR product selected by size in the previous step to 750,000 templates / μl and using 2 μl of the dilution as input for enrichment PCR to prepare many copies of each unique template. Enrichment PCR was performed as described above, except that 16 amplification cycles were performed instead of 22 amplification cycles.
[0559] To process the final long template sample for short read (Illumina) sequencing, the barcoded end library was first prepared, purified and quantified according to the methods outlined in the previous section. A second library containing randomly generated internal fragments from the long template was prepared using the Nextera DNA Flex Library Prep Kit (Illumina) with some modifications to the manufacturer's protocol. Specifically, the BLT (Bead-Linked Transposome) reagent was diluted in molecular grade water at a ratio of 1 to 50, and 10 μl of this dilution was used in a tag fragmentation reaction with 10 ng of long template DNA. Twelve library amplification cycles were performed using custom i5 index primers and custom i7 index primers (Table 7) instead of standard Illumina adapters.
[0560] Preparation of non-mutated reference library
[0561] Reference libraries were generated for all 66 genomes included in the final Morphoseq pool. Library preparation was performed according to the steps described above for the in-house Morphoseq library using 10 ng of genomic DNA as input, but with further modifications to the NexteraDNA Flex method. Specifically, the Illumina TB1 buffer was replaced with a custom tag fragmentation buffer (see above), the kit polymerase was replaced with KAPA HiFi HotStart ReadyMix (1× final concentration; KapaBiosystems), and the Illumina sample purification beads (SPB) were replaced with SPRIselect magnetic beads (Beckman Coulter). The thermal cycling conditions for reference library amplification were as follows: 72°C for 3 minutes; 98°C for 30 seconds; 98°C for 15 seconds, 55°C for 30 seconds, and 72°C for 30 seconds for 12 cycles; followed by a final extension at 72°C for 5 minutes.
[0562] To normalize the reference library, equal volumes of each sample were first pooled, and the pooled libraries were sequenced using the MiSeq Reagent Nano Kit (Illumina), generating 2 × 150 nt paired-end reads using the MiSeq V2 chemistry. The resulting sequence data were demultiplexed to determine the read counts for each individual genome. These counts were then used to prepare a normalization pool by combining different volumes of each original reference library, aiming to achieve equal coverage of each genome.
[0563] Illumina sequencing
[0564] Final samples for Illumina sequencing were prepared by combining a normalized reference pool, a morphoseq end library, and a morphoseq internal library at a molar ratio of 1:1:20, respectively. Sequencing was performed at the Ramaciotti Centre for Genomics at the University of New South Wales (Sydney, Australia) using a NovaSeq 6000 instrument and S1 flow cells to generate 2×150 nt paired-end reads.
[0565] Assembly of bacterial genomes
[0566] An overview of the bacterial genome assembly workflow is shown in Figure 13 shown.
[0567] Non-mutation reference assembly
[0568] The genome of each bacterial strain is assembled from 150 base pair reads of non-mutated end pairing. Initial quality filtering is performed using bbduk v36.99 to remove low-quality sequences and excise library adapters. Reads are demultiplexed using a custom python script and assembled using MEGAHIT v1.1.3 in conjunction with the following custom parameters: trimming level = 3, low local ratio = 0.1 and max-tip-len = 280, which are selected to reduce the complexity of the resulting genome map and to facilitate better mapping of mutant sequences in the next stage (as described below). The resulting graphical fragment assembly (gfa file) is used as an input to VG (index) v1.14.0 to create an index suitable for mapping. The resulting map is referred to as an "indexed non-mutated reference assembly map" or simply an "index map."
[0569] Generation of synthetic long reads (morphoreads)
[0570] Using VG(map) v1.14.0 with default parameters, the mutant reads from each end library (end reads) and the pooled internal library (internal reads) were mapped to their corresponding indexed VG bacterial genome assembly to generate a pair of graphical alignment map (GAM) files for each sample. The data from each GAM pair were combined with information from the corresponding non-mutated reference assembly, processed using custom tools, and stored in an HDF5-formatted database, which facilitated parallel processing of many of the remaining steps in reconstructing the sequence of the original template. The morphoread generation process consists of three main stages: "end wall recognition," "priming," and "extension."
[0571] The nature of the process used to fragment the target DNA into long fragments and generate the final short read library creates a situation where the sequence at the very end of any original template will only be found in the second read of the paired Illumina library. When these reads are mapped to the reference genome, they appear to suddenly accumulate at positions corresponding to the ends of the original long DNA template. These positions are called "end walls" and are identified by finding the end groups and internal read groups that map to the same positions in the reference assembly. Any site that maps to at least five end reads in the above pattern is marked as an end wall. Internal reads are used to increase the mapping counts at sites with two to four mapped end reads, and these sites are also marked as end walls if the total increased count is at least five.
[0572] The endwalls indicate where in the reference assembly the algorithm will begin building synthetic long reads, but whenever two or more templates have the same starting or ending position, it is possible to have a single endwall corresponding to more than one original DNA template. Each DNA template will have a unique mutation pattern, so the reads derived from a given template will contain a subset of its pattern, which will appear as transition mismatches in the VG map. The "priming" stage analyzes these mutation patterns in the ends and the internal reads at each endwall, clustering reads with similar patterns together and creating a single short (400bp-600bp) morphoread instance for each cluster. Each morphoread instance contains a directed acyclic graph-based representation of the mapped mutation reads, which contains what is called a "consensus graph". The structure of the consensus graph is roughly equivalent to a subgraph of the index graph, and the position of the reads in the consensus graph is equivalent to the mapped position of the reads relative to the index graph. The main difference between the consensus graph and its corresponding subgraph of the index graph is that the edges between nodes in the consensus graph represent the paths of mapped reads through the index graph, and whenever such a path follows a loop in the index graph, the nodes in that loop are duplicated, effectively unrolling the loop in the index graph, eliminating any cycles. Thus, each node in the index graph is equivalent to multiple nodes in the consensus graph, and edges in the consensus graph often (but not always) are equivalent to edges in the index graph. The consensus graph stores information about the mutant reads that index assembly and mapping, and can therefore be used to create a "consensus sequence," which corresponds to a path through the index graph (i.e., does not contain any mutations), as well as a "mutation set," which contains the common mutation patterns found in all included internal and end reads.
[0573] During the "extend" phase, the algorithm moves along the consensus graph, starting from the end wall, and iteratively adds end reads and internal reads to the morphoread if they match the consensus sequence (>90% identity, >=100 bp overlap) and their mutational pattern shares at least 3 mutations with the mutation set and contains no more than 5 mutations that differ from the mutation set. A large number of different mutations is required to reduce the impact of errors from single reads masquerading as mutations, and also because reads tested for inclusion in the morphoread may map to nodes that are beyond the end of the current consensus graph and may contain mutations not yet included in the morphoread's mutation set. Each time a new read is included in the morphoread, a new node can be added to the consensus graph, so the consensus segment may become longer. The algorithm continues to advance along the extended consensus graph until an end read is incorporated into the morphoread, indicating that the far end of the original long DNA template has been reached, or no reads can be found to continue extension. The final consensus segment for each morphoread is written to a FASTA file, and all morphoreads shorter than 500 bp are discarded. The algorithm also generates a BAM file containing the positions of the included end and internal reads written to the consensus sequence and some summary statistics for each morphoread.
[0574] Hybrid genome assembly
[0575] High-quality morphoreads as well as non-mutated reference reads were combined into hybrid genome assemblies using Unicycler v0.4.6 with default parameters.
[0576] result
[0577] The Morphoseq approach consistently produced assemblies with significantly fewer but larger branches compared to short-read-only assemblies (Kruskal Wallis, p < 0.001) ( Figure 14 ). For Morphoseq and short-read-only assemblies, the median maximum branch length as a percentage of genome size was 55.84% and 10.15%, respectively, while the median number of branches was 17 and 192, respectively. Exemplary assembly metrics for bacterial genomes can be found at Figure 15 Found in.
[0578] Table 7
[0579]
[0580]
[0581]
[0582]
[0583] Table 7: Primers used in this study
[0584] a. Sample tag sequences are shown in bold.
[0585] b. During mutagenesis PCR, unique Morphoseq index primers were used for each sample.
[0586] c. Unique combinations of custom i7 index and custom i5 index primers were used for each non-mutated reference library. Sequence Listing <110> Illumina Singapore Pte Ltd <120> Sequencing algorithm <130> N411618WO <150> GB1907101.8 <151> 2018-08-13 <150> GB1813171.4 <151> 2018-08-13 <160> 291 <170> PatentIn version 3.5 <210> 1 <211> 2325 <212> DNA <213> Thermococcus sp. KS-1 <400> 1 atgatcctcg acactgacta cataactgag aatggaaaac ccgtcataag gattttcaag 60 aaggagaacg gcgagtttaa gattgagtac gataggactt ttgaacccta catttacgcc 120 ctcctgaagg acgattctgc cattgaggag gtcaagaaga taaccgccga gaggcacgga 180 acggttgtaa cggttaagcg ggctgaaaag gttcagaaga agttcctcgg gagaccagtt 240 gaggtctgga aactctactt tactcaccct caggacgtcc cagcgataag ggacaagata 300 cgagagcatc cagcagttat tgacatctac gagtacgaca tacccttcgc caagcgctac 360 ctcatagaca agggattagt gccaatggaa ggcgacgagg agctgaaaat gcttgccttt 420 gatatcgaga cgctctacca tgagggcgag gagttcgccg aggggccaat ccttatgata 480 agctacgccg acgaggaagg ggccagggtg ataacgtgga agaacgcgga tctgccctac 540 gttgacgtcg tctcgacgga gagggagatg ataaagcgct tcctaaaggt ggtcaaagag 600 aaagatcctg acgtcctaat aacctacaac ggcgacaact tcgacttcgc ctacctaaaa 660 aaacgctgtg aaaagcttgg aataaacttc acgctcggaa gggacggaag cgagccgaag 720 attcagagga tgggcgacag gtttgccgtc gaagtgaagg gacggataca cttcgatctc 780 tatcctgtga taagacggac gataaacctg cccacataca cgcttgaggc cgtttatgaa 840 gccgtcttcg gtcagccgaa ggagaaggtc tacgctgagg agatagctac agcttgggag 900 agcggtgaag gccttgagag agtagccaga tactcgatgg aagatgcgaa ggtcacatac 960 gagcttggga aggagtttt ccctatggag gcccagcttt ctcgcttaat cggccagtcc 1020 ctctgggacg tctcccgctc cagcactggc aacctcgttg agtggttcct cctcaggaag 1080 gcctacgaga ggaatgagct ggccccgaac aagcccgatg aaaaggagct ggccagaga 1140 cgacagagct atgaggagg ctatgtaaaa gagcccgaga gaggttgtg ggagaacata 1200 gtgtacctag attttagatc tctgtacccc tcaatcatca tcaccacaa cgtctcgccg 1260 gatactctca acagggagg atgcaggaa tatgacgttg cggcaggt cggtcaccgc 1320 ttctgcaagg acttcccagg atttacccg agcctgcttg gagacctcct agaggagagg 1380 cagagata agagagat gaggccacg attgacccg tcgagagaa gctcctcgat 1440 tacaggcaga gggccatca gatcctggcc aacagctact acggttacta cggctatgca 1500 agggcgcgct ggtactgcaa ggagtgtgca gagagcgtaa cggcctgggg aagggagtac 1560 ataacgatga ccatcagaga gatagaggaa aagtacggct ttaagtaat ctacagcgac 1620 accgacggat tttgccac atacctgga gccgatgctg aaccgtcaa aaagaaggcg 1680 atggagttcc tcaagtatat caacgccaaa ctcccgggcg cgcttgagct cgagtacgag 1740. ggcttctaca aacgcggctt cttcgtcacg aagaagagt acgcggtgat agcgagga ggcaagataa caacgcgcgg acttgagatt gtgaggcgcg actggagcga gatagcgaaa gagacgcagg cgagggttct tgaagctttg ctaaaggacg gtgacgtcga gaaggccgtg aggatagtca aagaagttac cgaaaagctg agcaagtacg aggttccgcc ggagaagctg gtgatccacg agcagatac gagggattta aaggactaca aggcaaccgg tccccacgtt gccgttgcca agaggttggc cgcgagagg gtcaaaatac gccctggac ggtgataagc 2160. tcatcgtgc tcaagggctc tgggaggata ggcgacaggg cgataccgtt cgacgagttc gacccgacga agcacaagta cgacgccgag tactacattg agaaccaggt tctcccagcc gttgagaga ttctgagagc cttcggttac cgcaaggaag acctgcgcta ccagaagacg agacaggttg gtctgggagc ctggctgaag ccgaagga cttga <210> 2 <211> 774 <212> PRT <213> Thermococcus sp. KS-1 <400> 2 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asn Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Thr 50 55 60 Val Lys Arg Ala Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Glu His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Asp Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Glu Gly Ala Arg Val Ile Thr Trp Lys Asn Ala 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Arg Glu Met Ile Lys 180 185 190 Arg Phe Leu Lys Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Thr Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Val Phe Gly Gln Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Ala Thr Ala Trp Glu Ser Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Tyr 305 310 315 320 Glu Leu Gly Lys Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Ser Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Asp Glu Lys Glu Leu Ala Arg Arg Arg Gln Ser Tyr 370 375 380 Glu Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Lys Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Only With Lys And Leu Only Asn Ser Tyr Gly Tyr 485,490,495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Only Trp Gly Arg Glu Tyr Ile Met Thr Ile Arg Glu Ile 515,520,525 Glu Glu Lys Tyr Gly Phe Lys Val Ile Tyr Ser Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Ala Asp Ala Glu Thr Val Lys Lys Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565,570,575 Leu Glu Tyr Glu Gly Phe Tyr Lys Arg Gly Phe Phe Val Thr Lys Lys 580,585,590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595,600,605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Lys Ala Val 625 630 635 640 Arg Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Arg Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Lys Pro Lys Gly Thr 770 <210> 3 <211> 2325 <212> DNA <213> Thermococcus celer <400> 3 atgatccctcg acgctgacta catcaccgaa gatgggaagc ccgtcgtgag gatattcagg 60 aaggagaagg gcgagttcag atcgactac gatgact tcgagcccta catctacgcc 120 ctcctgaagg acgattcggc catcgaggag gtgaagagga taaccgttga gcgccacggg 180 aaggccgtca gggttaagcg ggtggagaag gtcgaaaga agttcctca caggccgata 240 gaggtctgga agctctactt caatcacccg caggacgttc cggcgataag ggacgagata 300 aggaagcatc cggccgtcgt tgatatcc gagtacgaca tccccttcgc caagcgctac 360 ctcatcgata aggggctcgt cccgatggag ggggaggagg agctcaact gatggccttc 420 gatacgaga ccctctacca cgagggac gagttcgggg aggggccgat cctgatgata 480 agctacgccg acgggacgg ggcgagggtc ataacctgga agagatcga cctccctac 540 gtcgacgtcg tctcgaccga gaggagatg ataaagcgct tcctccaggt ggtgaggag 600 aaggacccgg acgtgctcgt aacttacaac ggcgacaact tcgactcgc ctacctgaag 660 agacgctccg aggagcttgg attgaagttc atcctcggga gggacgggag cgagcccaag 720 atccagcgca tgggcgaccg cttcgccgtc gaggtgaagg ggaggataca cttcgacctc 780 tacccggtga taaggcgcac cgtgaacctg ccgacctaca cgctcgaggc ggtctacgag 840 gccatcttcg ggaggccaaa ggagaaggtc tacgccgggg agatagtgga ggcctgggaa 900 accggcgagg gtcttgagag ggttgcccgc tactccatgg aggacgcaaa ggttaccttc 960 gagctcggga gggagttctt cccgatggag gcccagctct cgaggctcat cggccagggt 1020 ctctgggacg tctcccgctc gagcaccggc aacctggtcg agtggttcct cctgaggaag 1080 gcctacgaga ggaacgaact ggccccgaac aagccgagcg gccgggaagt ggagatcagg 1140 aggcgtggct acgccggtgg ttacgttaag gagccggaga ggggtttatg ggagaacatc 1200 gtgtacctcg actttcgctc tctttacccc tccatcatca taacccacaa cgtctcgccc 1260 gataccctaa acagggaggg ctgtgagaac tacgacgtcg ccccccaggt ggggcataag 1320 ttctgcaaag attttccggg cttcatcccg agcctgctcg gaggcctgct tgaggagagg 1380 cagaagataa agcggaggat gaaggcctct gtggatcccg ttgagcggaa gctcctcgat 1440 tacaggcaga gggccatcaa gatactggcc aacagcttct acggatacta cggctacgcg 1500 agggcgaggt ggtactgcag ggagtgcgcg gagagcgtta ccgcctgggg cagggagtac 1560 atcgataggg tcatcaggga gctcgaggag aagttcggct tcaaggtgct ctacgcggac 1620 acggacggac tgcacgccac gatccccggg gcggacgccg ggaccgtcaa ggagagggcg 1680 agggggttcc tgagatacat caaccccaag ctccccggcc tcctggagct cgagtacgag 1740 gggttctacc tgaggggttt cttcgtgacg aagaagaagt acgcggtcat agacgaggag 1800 ggcaagataa ccacgcgcgg cctcgagata gtcaggcggg actggagcga ggtggccaag 1860 gagacgcagg cgagggtcct ggaggcgata ctgaggcacg gtgacgtcga ggaggccgtt 1920 agaatcgtca gggaggtaac cgaaaagctg agcaagtacg aggttccgcc ggagaaactg 1980 gtgatccacg agcagataac gagggatttg agggactaca aagccacggg accgcacgtg 2040 gcggtggcga agcgcctggc cgggaggggg gtaaggatac gccccgggac ggtgataagc 2100 tacatcgtcc tcaagggctc cggaaggata gggacaggg cgattcctt cgacgagttc 2160 gacccgacta agcacaggta cgacgccgac tactacatcg agaaccaggt tctgccagcc 2220 gtcgagagga tcctgaaggc cttcggctac cgcaggagg acctgaaata ccagaagacg 2280 aggcaggtgg gcctgggtgc gtggctcac gcggggaagg ggtga 2325 <210> 4 <211> 774 <212> PRT <213> Thermococcus celer <400> 4 Met Ileu Asp Ala Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Val 1 5 10 15 Arg Ile Phe Arg Lys Glu Lys Gly Glu Phe Arg Ile Asp Tyr Asp Arg 20 25 30 Asp Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Arg Ile Thr Val Glu Arg His Gly Lys Ala Val Arg 50 55 60 Val Lys Arg Val Glu Lys Val Glu Lys Phe Leu Asn Arg Pro Ile 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Asn His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Glu Ile Arg Lys His Pro Ala Val Val Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Glu Glu Glu Leu Lys Leu Met Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Asp Glu Phe Gly Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Gly Asp Gly Ala Arg Val Ile Thr Trp Lys Lys Ile 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Lys Glu Met Ile Lys 180 185 190 Arg Phe Leu Gln Val Val Lys Glu Lys Asp Pro Asp Val Leu Val Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Arg Arg Ser Glu 210 215 220 Glu Leu Gly Leu Lys Phe Ile Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Val Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Ile Phe Gly Arg Pro Lys Glu 275 280 285 Lys Val Tyr Ala Gly Glu Ile Val Glu Ala Trp Glu Thr Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Phe 305 310 315 320 Glu Leu Gly Arg Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Gly Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Ser Gly Arg Glu Val Glu Ile Arg Arg Arg Gly Tyr 370 375 380 Ala Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Glu Asn Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Lys Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Gly Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Arg Arg Met Lys Ala Ser Val Asp Pro Val Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Ala Asn Ser Phe Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Arg Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Asp Arg Val Ile Arg Glu Leu 515 520 525 Glu Glu Lys Phe Gly Phe Lys Val Leu Tyr Ala Asp Thr Asp Gly Leu 530 535 540 His Ala Thr Ile Pro Gly Ala Asp Ala Gly Thr Val Lys Glu Arg Ala 545 550 555 560 Arg Gly Phe Leu Arg Tyr Ile Asn Pro Lys Leu Pro Gly Leu Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Leu Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Val Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Ile Leu Arg His Gly Asp Val Glu Glu Ala Val 625 630 635 640 Arg Ile Val Arg Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Arg Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Gly 675 680 685 Arg Gly Val Arg Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Arg Tyr Asp Ala Asp Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Lys Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Lys Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Asn Ala Gly Lys Gly 770 <210> 5 <211> 2328 <212> DNA <213> Thermococcus siculi <400> 5 atgatcctcg acacggacta catcacggaa gatgggaaac ccgtcataag gatattcaag 60 aaagagaacg gcgagttcaa gatcgagtac gacaggactt ttgaacccta catctacgcc 120 ctcctgaagg acgactccgc gattgaggat gttaaaaaga taaccgccga gaggcacgga 180 acggtggtga aggtcaagcg cgccgaaaag gtgcagaaga agttcctagg caggccggtt 240 gaagtctgga agctctactt cacccacccc caagatgtcc cggcgataag ggacaagatt 300 aggaagcatc cagctgtaat tgacatctac gagtacgaca taccattcgc caagcgctac 360 ctcatcgaca agggcctgat tccgatggag ggtgaagaag agcttaagat gctcgccttc 420 gacattgaga cgctctacca tgagggtgag gagttcgccg aggggcctat tctgatgata 480 agctacgccg acgagagcga ggcacgcgtc atcacctgga agaaaatcga cctcccctac 540 gttgacgtcg tctcaacgga gaaggagatg ataaagcgct tcctccgcgt tgtgaaggag 600 aaagatcccg atgtcctcat aacctacaac ggcgacaact tcgacttcgc ctacctgaag 660 aagcgctgtg aaaagcttgg aataaacttc ctccttggaa gggacgggag cgagccgaag 720 atccagagaa tgggtgaccg cttcgccgtt gaggtgaagg ggaggataca cttcgacctc 780 tatcctgtaa taaggcgcac gataaacctg ccgacctaca tgcttgaggc agtctacgag 840 gccatctttg ggaagccaaa ggagaaggtt tacgccgagg agatagccac cgcttgggaa 900 accggagagg gccttgagag ggtggctcgc tactctatgg aggacgcgaa ggtcacgttt 960 gagcttggaa aggagttctt cccgatggag gcccaacttt cgaggttggt cggccagagc 1020 ttctgggatg tcgcgcgctc aagcacgggc aatctggtcg agtggttcct cctcaggaag 1080 gcctacgaga ggaacgagct ggctccaac aagccctg gaagggaata tgacgagagg 1140 cgcggtggat acgccggcgg ctacgtcaag gaaccggaaa agggctgtg ggagaacata 1200 gtctaccctcg actataaatc tctctacccc tcaatcatca tcaccacaa cgtctcgccc 1260 gataccctca accgcgagggg ctgtaggag tatgacgtag ctccacaggt cggccaccgc 1320 ttctgcagg actttccagg cttcatcccg agcctgctcg gggatctcct ggaggagagg 1380 spagagata agagagat gaggxaca attgacccga tcgagagaa gctccttgat 1440 tacaggcaac gggccatca gatccttcta atagttttt acggctacta cggctacgca 1500 agggctcgct ggtactgca ggagtgtgcc gagagcgtta cggcatgggg aagggaatat 1560 atcaccatga caatcaggga atagagag aagtatggct ttaaagtact ttatgcggac 1620 actgacggct tcttcgcgac gattcccggg gagatgccg agaccatcha aaagaggggcg 1680 atggagttcc tcaagtacat aaacgccaaa ctcccggtg cgctcgaact tgagtacgag 1740 gacttctaca ggcgcggctt cttcgtcacc aagagaat acgcggttat cgacgaggag 1800 ggcaagataa caacgcgcgg gctggagatc gtcaggcgcg actggagcga gatagccaag gagacgcagg cgcgggttct ggaggccctt ctgaaggacg gtgacgtcga agaggccgtg agcatagtca aagaagtgac cgagaagctg agcatagtc aggttccgcc ggagaagctc gttatccacg agcagatac gcgcgagctg aaggactaca aggcaacggg accacacgtg gcgatagcga agaggttagc cgcgagaggc gtcaaaatcc gccccgggac agtcatcagc 2160. tacatcgtgc tcaagggctc cgggaggata ggcgacaggg cgattccctt cgacgagttc gaccccacga agcacaagta cgatgcagag tactacatcg agcaccaggt tctacctgcc gtcgaga ttctgaaggc cttcggctat cgcggtgagg agctcagata ccagaagacg aggcaggttg gacttggggc gtggctgaag ccgaagggga aggggtga 2328. <210> 6 <211> 775 <212> PRT <213> Thermococcus siculi <400> 6 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Ile Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Asp Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Lys 50 55 60 Val Lys Arg Ala Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Lys His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Ile Pro 115 120 125 Met Glu Gly Glu Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr His Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Ser Glu Ala Arg Val Ile Thr Trp Lys Lys Ile 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Lys Glu Met Ile Lys 180 185 190 Arg Phe Leu Arg Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Leu Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Met Leu Glu Ala Val Tyr Glu Ala Ile Phe Gly Lys Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Ala Thr Ala Trp Glu Thr Gly Glu Gly 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Phe 305 310 315 320 Glu Leu Gly Lys Glu Phe Phe Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Val Gly Gln Ser Phe Trp Asp Val Ala Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Ser Gly Arg Glu Tyr Asp Glu Arg Arg Gly Gly Tyr 370 375 380 Ala Gly Gly Tyr Val Lys Glu Pro Glu Lys Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Tyr Lys Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Arg Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Leu Asn Ser Phe Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Thr Met Thr Ile Arg Glu Ile 515 520 525 Glu Glu Lys Tyr Gly Phe Lys Val Leu Tyr Ala Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Glu Asp Ala Glu Thr Ile Lys Lys Arg Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Asp Phe Tyr Arg Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Glu Ala Val 625 630 635 640 Ser Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Glu Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Ile Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Lys Ala Phe Gly Tyr Arg Gly 740 745 750 Glu Glu Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Gly Ala Trp 755 760 765 Leu Lys Pro Lys Gly Lys Gly 770 775 <210> 7 <211> 774 <212> PRT <213> Thermococcus kodakarensis <400> 7 Met Ile Leu Asp Thr Asp Tyr Ile Thr Glu Asp Gly Lys Pro Val Ile 1 5 10 15 Arg Ile Phe Lys Lys Glu Asn Gly Glu Phe Lys Ile Glu Tyr Asp Arg 20 25 30 Thr Phe Glu Pro Tyr Phe Tyr Ala Leu Leu Lys Asp Asp Ser Ala Ile 35 40 45 Glu Glu Val Lys Lys Ile Thr Ala Glu Arg His Gly Thr Val Val Thr 50 55 60 Val Lys Arg Val Glu Lys Val Gln Lys Lys Phe Leu Gly Arg Pro Val 65 70 75 80 Glu Val Trp Lys Leu Tyr Phe Thr His Pro Gln Asp Val Pro Ala Ile 85 90 95 Arg Asp Lys Ile Arg Glu His Pro Ala Val Ile Asp Ile Tyr Glu Tyr 100 105 110 Asp Ile Pro Phe Ala Lys Arg Tyr Leu Ile Asp Lys Gly Leu Val Pro 115 120 125 Met Glu Gly Asp Glu Glu Leu Lys Met Leu Ala Phe Asp Ile Glu Thr 130 135 140 Leu Tyr Glu Glu Gly Glu Glu Phe Ala Glu Gly Pro Ile Leu Met Ile 145 150 155 160 Ser Tyr Ala Asp Glu Glu Gly Ala Arg Val Ile Thr Trp Lys Asn Val 165 170 175 Asp Leu Pro Tyr Val Asp Val Val Ser Thr Glu Arg Glu Met Ile Lys 180 185 190 Arg Phe Leu Arg Val Val Lys Glu Lys Asp Pro Asp Val Leu Ile Thr 195 200 205 Tyr Asn Gly Asp Asn Phe Asp Phe Ala Tyr Leu Lys Lys Arg Cys Glu 210 215 220 Lys Leu Gly Ile Asn Phe Ala Leu Gly Arg Asp Gly Ser Glu Pro Lys 225 230 235 240 Ile Gln Arg Met Gly Asp Arg Phe Ala Val Glu Val Lys Gly Arg Ile 245 250 255 His Phe Asp Leu Tyr Pro Val Ile Arg Arg Thr Ile Asn Leu Pro Thr 260 265 270 Tyr Thr Leu Glu Ala Val Tyr Glu Ala Val Phe Gly Gln Pro Lys Glu 275 280 285 Lys Val Tyr Ala Glu Glu Ile Thr Thr Ala Trp Glu Thr Gly Glu Asn 290 295 300 Leu Glu Arg Val Ala Arg Tyr Ser Met Glu Asp Ala Lys Val Thr Tyr 305 310 315 320 Glu Leu Gly Lys Glu Phe Leu Pro Met Glu Ala Gln Leu Ser Arg Leu 325 330 335 Ile Gly Gln Ser Leu Trp Asp Val Ser Arg Ser Ser Thr Gly Asn Leu 340 345 350 Val Glu Trp Phe Leu Leu Arg Lys Ala Tyr Glu Arg Asn Glu Leu Ala 355 360 365 Pro Asn Lys Pro Asp Glu Lys Glu Leu Ala Arg Arg Arg Gln Ser Tyr 370 375 380 Glu Gly Gly Tyr Val Lys Glu Pro Glu Arg Gly Leu Trp Glu Asn Ile 385 390 395 400 Val Tyr Leu Asp Phe Arg Ser Leu Tyr Pro Ser Ile Ile Ile Thr His 405 410 415 Asn Val Ser Pro Asp Thr Leu Asn Arg Glu Gly Cys Lys Glu Tyr Asp 420 425 430 Val Ala Pro Gln Val Gly His Arg Phe Cys Lys Asp Phe Pro Gly Phe 435 440 445 Ile Pro Ser Leu Leu Gly Asp Leu Leu Glu Glu Arg Gln Lys Ile Lys 450 455 460 Lys Lys Met Lys Ala Thr Ile Asp Pro Ile Glu Arg Lys Leu Leu Asp 465 470 475 480 Tyr Arg Gln Arg Ala Ile Lys Ile Leu Ala Asn Ser Tyr Tyr Gly Tyr 485 490 495 Tyr Gly Tyr Ala Arg Ala Arg Trp Tyr Cys Lys Glu Cys Ala Glu Ser 500 505 510 Val Thr Ala Trp Gly Arg Glu Tyr Ile Thr Met Thr Ile Lys Glu Ile 515 520 525 Glu Glu Lys Tyr Gly Phe Lys Val Ile Tyr Ser Asp Thr Asp Gly Phe 530 535 540 Phe Ala Thr Ile Pro Gly Ala Asp Ala Glu Thr Val Lys Lys Lys Ala 545 550 555 560 Met Glu Phe Leu Lys Tyr Ile Asn Ala Lys Leu Pro Gly Ala Leu Glu 565 570 575 Leu Glu Tyr Glu Gly Phe Tyr Glu Arg Gly Phe Phe Val Thr Lys Lys 580 585 590 Lys Tyr Ala Val Ile Asp Glu Glu Gly Lys Ile Thr Thr Arg Gly Leu 595 600 605 Glu Ile Val Arg Arg Asp Trp Ser Glu Ile Ala Lys Glu Thr Gln Ala 610 615 620 Arg Val Leu Glu Ala Leu Leu Lys Asp Gly Asp Val Glu Lys Ala Val 625 630 635 640 Arg Ile Val Lys Glu Val Thr Glu Lys Leu Ser Lys Tyr Glu Val Pro 645 650 655 Pro Glu Lys Leu Val Ile His Glu Gln Ile Thr Arg Asp Leu Lys Asp 660 665 670 Tyr Lys Ala Thr Gly Pro His Val Ala Val Ala Lys Arg Leu Ala Ala 675 680 685 Arg Gly Val Lys Ile Arg Pro Gly Thr Val Ile Ser Tyr Ile Val Leu 690 695 700 Lys Gly Ser Gly Arg Ile Gly Asp Arg Ala Ile Pro Phe Asp Glu Phe 705 710 715 720 Asp Pro Thr Lys His Lys Tyr Asp Ala Glu Tyr Tyr Ile Glu Asn Gln 725 730 735 Val Leu Pro Ala Val Glu Arg Ile Leu Arg Ala Phe Gly Tyr Arg Lys 740 745 750 Glu Asp Leu Arg Tyr Gln Lys Thr Arg Gln Val Gly Leu Ser Ala Trp 755 760 765 Leu Lys Pro Lys Gly Thr 770 <210> 8 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 8 tagaattgaa gaa 13 <210> 9 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 9 tggccatagc tac 13 <210> 10 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 10 gtcatctgcg acc 13 <210> 11 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 11 ttcgcgcttg gac 13 <210> 12 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 12 cgcgaaccgt tag 13 <210> 13 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 13 ttgcagcctc taa 13 <210> 14 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 14 tctactagta cga 13 <210> 15 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 15 gtaggttcta ctg 13 <210> 16 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 16 gccaatatca agt 13 <210> 17 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 17 ctatcttgct ggt 13 <210> 18 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 18 gttctcatag gta 13 <210> 19 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 19 gtctatgaac caa 13 <210> 20 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 20 cggagcgctt att 13 <210> 21 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 21 tatgccatga gga 13 <210> 22 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 22 atacgactcg gag 13 <210> 23 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 23 gatggaactc agc 13 <210> 24 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 24 ggacctgcat gaa 13 <210> 25 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 25 tagactggaa ctt 13 <210> 26 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 26 gaattacctc gtt 13 <210> 27 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 27 aggatcaggc tac 13 <210> 28 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 28 acgcgtagaa gag 13 <210> 29 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 29 cttcgagact tac 13 <210> 30 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 30 gacggctaac tcc 13 <210> 31 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 31 ttagcattct ctt 13 <210> 32 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 32 gcaaggcata gta 13 <210> 33 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 33 acctagatat gga 13 <210> 34 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 34 acgccaaggc gta 13 <210> 35 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 35 tatgacggat ccg 13 <210> 36 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 36 cctccattag aga 13 <210> 37 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 37 attgaatact ctg 13 <210> 38 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 38 gagatgagaa gaa 13 <210> 39 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 39 tctgagtagc cgg 13 <210> 40 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 40 aataggtagt acg 13 <210> 41 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 41 gtcgaagaag tcc 13 <210> 42 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 42 tactgcatct cgt 13 <210> 43 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 43 gacgtattag agc 13 <210> 44 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 44 cctgcattat tcg 13 <210> 45 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 45 acgaatgatg ctc 13 <210> 46 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 46 tactagcaga gat 13 <210> 47 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 47 ctcctcatct tcc 13 <210> 48 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 48 tcctctgcgc tgc 13 <210> 49 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 49 ccttctcagt ccg 13 <210> 50 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 50 cagcttcata gcg 13 <210> 51 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 51 ttgactctcg cgc 13 <210> 52 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 52 tatcctgagc gat 13 <210> 53 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 53 aacgcctagc cga 13 <210> 54 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 54 ccgaagacgt cat 13 <210> 55 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 55 gagttctcca gat 13 <210> 56 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 56 tgcatccgcg ctt 13 <210> 57 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 57 cctgaactca agt 13 <210> 58 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 58 ggtcgtatgc gta 13 <210> 59 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 59 aggcctctct acc 13 <210> 60 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 60 gtactccatc caa 13 <210> 61 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 61 cagcggacgc gct 13 <210> 62 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 62 atctctctta gca 13 <210> 63 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 63 aagcaataat aat 13 <210> 64 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 64 aaggcgactc cga 13 <210> 65 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 65 acgtctctag gag 13 <210> 66 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 66 ccatcagacc tct 13 <210> 67 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 67 acttaatcgt act 13 <210> 68 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 68 tggaattctc caa 13 <210> 69 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 69 ccatacgatc agg 13 <210> 70 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 70 ttatggagca ata 13 <210> 71 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 71 gctcggcgtt cga 13 <210> 72 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 72 ttggccagtc gct 13 <210> 73 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 73 cagatacgta gag 13 <210> 74 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 74 aatgctatta tcc 13 <210> 75 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 75 gcagcatgcc gat 13 <210> 76 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 76 ggagagttac ctc 13 <210> 77 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 77 gagagtccat gat 13 <210> 78 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 78 caatctattc tga 13 <210> 79 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 79 gctcttagta tcc 13 <210> 80 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 80 ccatagttat ggt 13 <210> 81 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 81 tgcgagatcg aag 13 <210> 82 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 82 agagaagtcg agt 13 <210> 83 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 83 ggtaactcca tat 13 <210> 84 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 84 tgctattcca ggc 13 <210> 85 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 85 aaccgcgagg ctc 13 <210> 86 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 86 ttctagagat acc 13 <210> 87 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 87 ttcgctcaag tat 13 <210> 88 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 88 cagagaaggc gca 13 <210> 89 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 89 tagaattggc ctc 13 <210> 90 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 90 ggccattctc cag 13 <210> 91 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 91 tccaacgcgc gtt 13 <210> 92 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 92 gccgcagatt acg 13 <210> 93 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 93 gcagttcgaa cgc 13 <210> 94 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 94 ttctctctgc agg 13 <210> 95 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 95 taagctacca gcg 13 <210> 96 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 96 ctgcatgagg ttg 13 <210> 97 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 97 ttgcctagcg agg 13 <210> 98 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 98 caactgaatt agg 13 <210> 99 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 99 aagcggtcct ctt 13 <210> 100 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 100 aatggaagga ccg 13 <210> 101 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 101 gagttagtaa gtt 13 <210> 102 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 102 ttcctaattc caa 13 <210> 103 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 103 gttctggttc gct 13 <210> 104 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 104 gttcatctct tcc 13 <210> 105 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 105 attccgagga aga 13 <210> 106 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 106 cttagccgag aga 13 <210> 107 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 107 gtctgctacg ctt 13 <210> 108 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 108 atggcgccgc gca 13 <210> 109 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 109 taattggtta tct 13 <210> 110 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 110 tcggttataa gtc 13 <210> 111 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 111 tgcctgagaa cgt 13 <210> 112 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 112 agatgcggtt aac 13 <210> 113 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 113 atggaatagg cga 13 <210> 114 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 114 agagatgcga tcg 13 <210> 115 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 115 ctccaactaa cgt 13 <210> 116 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 116 gccttgctac tgg 13 <210> 117 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 117 cttcgtctct acg 13 <210> 118 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 118 acgctcatag cct 13 <210> 119 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 119 gtcgaagata agg 13 <210> 120 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 120 gccggagtcc tcg 13 <210> 121 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 121 tatacggcga cct 13 <210> 122 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 122 aggtagatat tcg 13 <210> 123 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 123 ttaaggtact gct 13 <210> 124 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 124 cggatctggt ata 13 <210> 125 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 125 gaggtctcgg agg 13 <210> 126 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 126 ggcatcgatg gac 13 <210> 127 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 127 gatctccgat ata 13 <210> 128 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 128 gattcggaat act 13 <210> 129 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 129 ctgcgatccg gcc 13 <210> 130 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 130 gatccggttg caa 13 <210> 131 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 131 cgtcaggctt gac 13 <210> 132 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 132 tcggcaaggc gag 13 <210> 133 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 133 gaacggcgaa cgc 13 <210> 134 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 134 cctcaagcgg act 13 <210> 135 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 135 gaagccagat ggt 13 <210> 136 <211> 13 <212> DNA <213> Artificial Sequence <220> <223> Sample tag sequence <400> 136 tgctcatacc aat 13 <210> 137 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> i7 custom index primer (Table 2) <220> <221> misc_feature <222> (25)..(36) <223> n is a, c, g or t <400> 137 caagcagaag acggcatacg agatnnnnnn nnnnnngtct cgtgggctcg g 51 <210> 138 <211> 55 <212> DNA <213> Artificial Sequence <220> <223> i5 custom index primer (Table 2) <220> <221> misc_feature <222> (30)..(41) <223> n is a, c, g, or t <400> 138 aatgatacgg cgaccaccga gatctacacn nnnnnnnnnn ntcgtcggca gcgtc 55 <210> 139 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> i7 flow cell primer (Table 3) <400> 139 caagcagaag acggcatacg a 21 <210> 140 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> i5 flow cell primer (Table 3) <400> 140 aatgatacgg cgaccaccga 20 <210> 141 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> single_mut primer for mutagenesis (Table 5) <220> <221> misc_feature <222> (19)..(34) <223> n is a, c, g, or t <400> 141 tcggtctgcg cctctagcnn nnnnnnnnnn nnnngtctcg tgggctcgga g 51 <210> 142 <211> 42 <212> DNA <213> Artificial Sequence <220> <223> single_rec primer for recovery (Table 5) <400> 142 caagcagaag acggcatacg agattcggtc tgcgcctcta gc 42 <210> 143 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A1 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 143 tcggtctgcg cctctagcnn nctctatcga cgtagtctcg tgggctcgga g 51 <210> 144 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A2 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 144 tcggtctgcg cctctagcnn ntaagtctgg tctagtctcg tgggctcgga g 51 <210> 145 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A3 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 145 tcggtctgcg cctctagcnn nacctgcgta acctgtctcg tgggctcgga g 51 <210> 146 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A4 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 146 tcggtctgcg cctctagcnn ncgtctctag gatggtctcg tgggctcgga g 51 <210> 147 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A5 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 147 tcggtctgcg cctctagcnn ntcattaggt atatgtctcg tgggctcgga g 51 <210> 148 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A6 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 148 tcggtctgcg cctctagcnn naagtattcc atgagtctcg tgggctcgga g 51 <210> 149 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A7 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 149 tcggtctgcg cctctagcnn nttctggtac ttcagtctcg tgggctcgga g 51 <210> 150 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A8 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 150 tcggtctgcg cctctagcnn natgcctcct gcttgtctcg tgggctcgga g 51 <210> 151 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A9 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 151 tcggtctgcg cctctagcnn ntggtaatac gcctgtctcg tgggctcgga g 51 <210> 152 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A10 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 152 tcggtctgcg cctctagcnn nactgacgat tggtgtctcg tgggctcgga g 51 <210> 153 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A11 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 153 tcggtctgcg cctctagcnn nttagagtag ttgcgtctcg tgggctcgga g 51 <210> 154 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_A12 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 154 tcggtctgcg cctctagcnn naagccgttg aatagtctcg tgggctcgga g 51 <210> 155 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B1 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 155 tcggtctgcg cctctagcnn ntagcctcgc tctcgtctcg tgggctcgga g 51 <210> 156 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B2 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 156 tcggtctgcg cctctagcnn ncttggcctt gcaagtctcg tgggctcgga g 51 <210> 157 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B3 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 157 tcggtctgcg cctctagcnn nctatcttca actggtctcg tgggctcgga g 51 <210> 158 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B4 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 158 tcggtctgcg cctctagcnn natccatacg gactgtctcg tgggctcgga g 51 <210> 159 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B5 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 159 tcggtctgcg cctctagcnn ncgctcgctc atatgtctcg tgggctcgga g 51 <210> 160 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B6 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 160 tcggtctgcg cctctagcnn ncgtatcgaa ttcagtctcg tgggctcgga g 51 <210> 161 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B7 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 161 tcggtctgcg cctctagcnn nattcttctc ggtagtctcg tgggctcgga g 51 <210> 162 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B8 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 162 tcggtctgcg cctctagcnn ncaagttgca gcaggtctcg tgggctcgga g 51 <210> 163 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B9 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 163 tcggtctgcg cctctagcnn nactaatctg gtacgtctcg tgggctcgga g 51 <210> 164 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B10 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 164 tcggtctgcg cctctagcnn ncaggaagat tagtgtctcg tgggctcgga g 51 <210> 165 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B11 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 165 tcggtctgcg cctctagcnn naataactag cttggtctcg tgggctcgga g 51 <210> 166 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_B12 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 166 tcggtctgcg cctctagcnn ntacgactta ctaagtctcg tgggctcgga g 51 <210> 167 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_C1 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 167 tcggtctgcg cctctagcnn nctcggcttc tcctgtctcg tgggctcgga g 51 <210> 168 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_C2 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 168 tcggtctgcg cctctagcnn nttcctctct atcagtctcg tgggctcgga g 51 <210> 169 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_C3 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 169 tcggtctgcg cctctagcnn natggattcc tagagtctcg tgggctcgga g 51 <210> 170 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_C4 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 170 tcggtctgcg cctctagcnn nttcttgagt aagggtctcg tgggctcgga g 51 <210> 171 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_C5 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 171 tcggtctgcg cctctagcnn nactactacg aagggtctcg tgggctcgga g 51 <210> 172 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_C6 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 172 tcggtctgcg cctctagcnn ncatcgctat cgttgtctcg tgggctcgga g 51 <210> 173 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_C7 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 173 tcggtctgcg cctctagcnn naagttccgc attagtctcg tgggctcgga g 51 <210> 174 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_C8 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 174 tcggtctgcg cctctagcnn nacttaagtt gaaggtctcg tgggctcgga g 51 <210> 175 <211> 51 <212> DNA <213> Artificial Sequence <220> <223> Morphoseq_index_C9 <220> <221> misc_feature <222> (19)..(21) <223> n is a, c, g, or t <400> 175 tcggtctgcg cctctagcnn ntgagtaatt cgacgtctcg tgggctcgga g 51 <210> 176 <211> 51 <212> DNA <213> Artificial Sequence <220...
Claims
1. A method for determining the sequence of at least one target template nucleic acid molecule, the method comprising: (a) providing paired samples, each sample comprising at least one target template nucleic acid molecule; (b) sequencing a region of the at least one target template nucleic acid molecule in a first sample of the paired samples to provide a non-mutation sequence read; (c) introducing a mutation into the at least one target template nucleic acid molecule in the second sample of the paired samples to provide at least one mutated target template nucleic acid molecule; (d) sequencing a region of the at least one mutated target template nucleic acid molecule to provide a mutation sequence read; and (e) analyzing the mutant sequence reads and using information obtained from analyzing the mutant sequence reads to assemble a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads, The analyzing the mutant sequence reads comprises: identifying mutant sequence reads that are likely derived from the same at least one mutation in the target template nucleic acid molecule based on a common mutation pattern or by determining whether an odds ratio exceeds a predetermined threshold, the odds ratio comprising a ratio of a probability that the mutant sequence reads are derived from the same mutation in the target template nucleic acid molecule compared to a probability that the mutant sequence reads are not derived from the same mutation in the target template nucleic acid molecule; and grouping mutant sequence reads of target template nucleic acid molecules that are likely derived from the same at least one mutation, said assembling a sequence of at least a portion of at least one target template nucleic acid molecule from said non-mutated sequence reads comprises mapping the mutant sequence reads in each group onto an assembly graph, The first sample of the paired samples and / or the second sample of the paired samples include two or more synthesized sub-samples.
2. A method for generating a sequence of at least one target template nucleic acid molecule, the method comprising: (a) obtaining data from paired samples, wherein the data comprises: (i) non-mutation sequence reads; and (ii) mutation sequence reads; and (b) analyzing the mutant sequence reads and using information obtained from analyzing the mutant sequence reads to assemble a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads, The analyzing the mutant sequence reads comprises: identifying mutant sequence reads that are likely derived from the same at least one mutation in the target template nucleic acid molecule based on a common mutation pattern or by determining whether an odds ratio exceeds a predetermined threshold, the odds ratio comprising a ratio of a probability that the mutant sequence reads are derived from the same mutation in the target template nucleic acid molecule compared to a probability that the mutant sequence reads are not derived from the same mutation in the target template nucleic acid molecule; and grouping mutant sequence reads of target template nucleic acid molecules that are likely derived from the same at least one mutation, said assembling a sequence of at least a portion of at least one target template nucleic acid molecule from said non-mutated sequence reads comprises mapping the mutant sequence reads in each group onto an assembly graph, The first sample of the paired samples and / or the second sample of the paired samples include two or more synthesized sub-samples.
3. The method according to claim 1 or 2, characterized in that The step of assembling a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads comprises preparing an assembly graph.
4. The method according to claim 3, characterized in that The assembly graph includes nodes corresponding to consensus sequences calculated from non-mutated sequence reads, and each valid path through the assembly graph including the nodes represents a sequence of at least a portion of at least one target template nucleic acid molecule.
5. The method according to claim 4, characterized in that The node is unitig.
6. The method according to claim 3, characterized in that Assembling a sequence of at least a portion of at least one target template nucleic acid molecule from the non-mutated sequence reads includes using information obtained from analyzing the mutated sequence reads to identify nodes that form part of a valid path through the assembly graph.
7. The method according to claim 4, characterized in that The sequence of at least a portion of at least one target template nucleic acid molecule is assembled from nodes that form part of a valid path stringing through the assembly graph.
8. The method according to claim 1, characterized in that The paired samples are taken from the same original sample or derived from the same organism.
9. The method according to claim 2, characterized in that The non-mutated sequence reads include the sequence of a region of at least one target template nucleic acid molecule in a first sample of the paired samples, and the mutant sequence reads include the sequence of a region of at least one mutated target template nucleic acid molecule in a second sample of the paired samples, and the paired samples are taken from the same original sample or derived from the same organism.
10. The method according to claim 6, characterized in that Using information obtained from analyzing the mutant sequence reads to identify nodes that form part of a valid path through the assembly graph includes: (i) Nodes are calculated from non-mutation sequence reads; (ii) mapping the mutant sequence reads to the assembly map; (iii) identifying mutant sequence reads of a target template nucleic acid molecule that may be derived from the same at least one mutation; and (iv) identifying nodes connected by mutant sequence reads of target template nucleic acid molecules that are likely derived from the same at least one mutation, Nodes connected by mutant sequence reads are likely to originate from the same target template nucleic acid molecule with at least one mutation and constitute part of an effective path connecting the assembly graph.
11. The method according to claim 1 or 2, characterized in that If the mutant sequence reads have a common mutation pattern, the mutant sequence reads are likely derived from the same mutated target template nucleic acid molecule, or wherein analyzing the mutant sequence reads includes identifying mutant sequence reads having a common mutation pattern.
12. The method according to claim 11, characterized in that Mutated sequence reads with a common mutation pattern include at least one common signature k-mer and / or a common signature mutation.
13. The method according to claim 12, characterized in that A characteristic k-mer is a k-mer that does not appear in the non-mutation sequence reads but appears at least twice in the mutation sequence reads.
14. The method according to claim 12, characterized in that A signature mutation is a nucleotide that appears at least twice in the mutant sequence reads and does not appear at the corresponding position in the non-mutated sequence reads.
15. The method according to claim 14, characterized in that The characteristic mutation is a commensal mutation.
16. The method according to claim 14, characterized in that If at least two nucleotides at corresponding positions in the mutant sequence reads having the signature mutation are different from each other, the signature mutation is ignored.
17. The method according to claim 14, characterized in that If a signature mutation is an unexpected mutation, the signature mutation is ignored.
18. The method according to claim 14, characterized in that The step of identifying mutant sequence reads of the target template nucleic acid molecule that may be derived from the same at least one mutation comprises identifying mutant sequence reads corresponding to a specific region of the at least one target template nucleic acid molecule.
19. The method according to claim 1 or 2, characterized in that The method includes determining a probability that two mutant sequence reads are derived from the same mutant target template nucleic acid molecule by determining whether an odds ratio exceeds a predetermined threshold.
20. The method according to claim 19, wherein If the odds ratio of a first mutant sequence read to a second mutant sequence read is higher than the odds ratio of the first mutant sequence read to other mutant sequence reads mapping to the same region of the assembly graph, the mutant sequence reads are likely derived from the same mutant target template nucleic acid molecule.
21. The method according to claim 19, wherein The predetermined threshold is determined based on one or more of the following factors: (i) the required degree of stringency; and / or (ii) an error rate in the step of sequencing a region of the at least one mutated target template nucleic acid molecule to provide a mutation sequence read; and / or (iii) the mutation rate used in the step of introducing mutations into at least one target template nucleic acid molecule; and / or (iv) the size of the at least one target template nucleic acid molecule; and / or (v) time limits; and / or (vi) Resource limitations.
22. The method according to claim 1 or 2, characterized in that Identifying mutant sequence reads of a target template nucleic acid molecule that are likely derived from the same mutation includes using a probability function based on the following parameters: e. a matrix (N) of the nucleotides at each position of the mutant sequence reads and the assembly map; f. The probability (M) of mutating a given nucleotide (i) to read nucleotide (j); g. the probability (E) that a given nucleotide (i) is read incorrectly, and thus that nucleotide (j) is read conditional on the incorrect reading of that nucleotide; and h. The probability of incorrectly reading the nucleotide at position Y (Q).
23. The method according to claim 22, characterized in that The Q value is obtained by performing a statistical analysis on the mutant sequence reads and the non-mutated sequence reads, or is obtained based on prior knowledge of the accuracy of the sequencing method.
24. The method according to claim 22, characterized in that The values of M and E are estimated based on statistical analysis of a subset of the mutant sequence reads and non-mutated sequence reads, wherein the subset includes mutant sequence reads and non-mutated sequence reads that were selected because they mapped to the same region of the assembly map.
25. The method according to claim 24, characterized in that The statistical analysis is performed using Bayesian inference, Monte Carlo methods, or variational inference.
26. The method according to claim 25, characterized in that The statistical analysis was performed using Hamiltonian Monte Carlo.
27. The method according to claim 25, characterized in that The statistical analysis was performed using maximum likelihood simulation of Bayesian inference.
28. The method according to claim 1 or 2, characterized in that Identifying mutant sequence reads of a target template nucleic acid molecule that are likely derived from the same mutation includes using machine learning or a neural network.
29. The method according to claim 1 or 2, characterized in that The method includes a pre-clustering step.
30. The method according to claim 29, wherein Identification of mutated sequence reads of target template nucleic acid molecules that are likely to be derived from the same mutation is constrained by the results of the pre-clustering step.
31. The method according to claim 29, wherein The pre-clustering step includes assigning mutant sequence reads to groups, wherein each member of the same group has a reasonable likelihood of being derived from the same mutant target template nucleic acid molecule.
32. The method according to claim 29, wherein The pre-clustering step includes Markov clustering or Leuven clustering.
33. The method according to claim 31, characterized in that Each member of the same group maps to a common position on the assembly map and / or has a common mutation pattern.
34. The method according to claim 33, wherein Mutated sequence reads with a common mutation pattern are mutant sequence reads that include at least one common signature k-mer and / or a common signature mutation.
35. The method according to claim 34, wherein A characteristic k-mer is a k-mer that does not appear in the non-mutation sequence reads but appears at least twice in the mutation sequence reads.
36. The method according to claim 34, wherein A signature mutation is a nucleotide that appears at least twice in the mutant sequence reads and does not appear at the corresponding position in the non-mutated sequence reads.
37. The method according to claim 36, wherein The characteristic mutation is a commensal mutation.
38. The method according to claim 36, characterized in that If at least two nucleotides at corresponding positions in the mutant sequence reads having the signature mutation are different from each other, the signature mutation is ignored.
39. The method according to claim 36, wherein If a signature mutation is an unexpected mutation, the signature mutation is ignored.
40. The method according to claim 36, wherein The step of identifying mutant sequence reads of the target template nucleic acid molecule that may be derived from the same at least one mutation comprises identifying mutant sequence reads corresponding to a specific region of the at least one target template nucleic acid molecule.
41. The method according to claim 1 or 2, characterized in that The method includes sequencing the ends of the at least one target template nucleic acid molecule using paired-end sequencing.
42. The method according to claim 41, wherein The method includes mapping the sequence of the end of the at least one target template nucleic acid molecule onto an assembly map.
43. The method according to claim 1 or 2, characterized in that The at least one target template nucleic acid molecule comprises a barcode at each end.
44. The method according to claim 43, wherein The method comprises mapping the sequence of the ends of the at least one target template nucleic acid molecule onto an assembly map and comprising a barcode at substantially each end.
45. The method according to claim 6, wherein Identifying nodes that form part of a valid path through the assembly graph includes ignoring putative paths with mismatched ends.
46. The method according to claim 6, wherein Identifying nodes that form part of a valid path through the assembly graph includes ignoring putative paths due to template collisions.
47. The method according to claim 6, wherein Identifying nodes that form part of a valid path through the assembly graph includes ignoring putative paths that are longer or shorter than expected.
48. The method according to claim 6, wherein Identifying nodes that form part of a valid path through the assembly graph includes ignoring putative paths with atypical coverage depths.
49. The method according to claim 1 or 2, characterized in that The at least one mutated target template nucleic acid molecule comprises 1% to 50% mutations.
50. The method according to claim 1 or 2, characterized in that The at least one mutated target template nucleic acid molecule comprises a non-uniform distribution of mutations.
51. The method according to claim 1 or 2, characterized in that The mutant sequence reads and / or the non-mutated sequence reads comprise unevenly distributed sequencing errors.
52. The method according to claim 1 or 2, characterized in that The step of introducing mutations into at least one mutated target template nucleic acid molecule introduces a non-uniform distribution of mutations.
53. The method according to claim 1 or 2, characterized in that The steps of sequencing a region of the at least one target template nucleic acid molecule and / or sequencing a region of the at least one mutated target template nucleic acid molecule introduce unevenly distributed sequencing errors.
54. The method according to claim 1 or 2, characterized in that The at least one mutated target template nucleic acid molecule comprises a random pattern of mutations.
55. The method according to claim 1 or 2, characterized in that Several pairs of samples are available.
56. The method according to claim 55, characterized in that The at least one target template nucleic acid molecule in different paired samples is labeled with different sample tags.
57. The method according to claim 1, wherein The method further comprises a step of amplifying the at least one target template nucleic acid molecule before the step of sequencing the region of the at least one target template nucleic acid molecule in the first sample of the paired samples.
58. The method according to claim 1, wherein The method further comprises a step of amplifying the at least one target template nucleic acid molecule before the step of sequencing the region of the at least one mutated target template nucleic acid molecule in the second sample of the paired samples.
59. The method according to claim 1, wherein The method further comprises the step of fragmenting the at least one target template nucleic acid molecule before the step of sequencing a region of the at least one target template nucleic acid molecule in the first sample of the paired samples.
60. The method according to claim 1, wherein The method further comprises the step of fragmenting the at least one target template nucleic acid molecule before the step of sequencing the region of the at least one mutated target template nucleic acid molecule in the second sample of the paired samples.
61. The method according to claim 1, wherein The method further comprises the step of fragmenting the at least one mutated target template nucleic acid molecule before the step of sequencing the region of the at least one mutated target template nucleic acid molecule in the second sample of the paired samples.
62. The method according to claim 1 or 2, characterized in that The at least one target template nucleic acid molecule is larger than 2 kbp.
63. The method according to claim 1, wherein The step of introducing a mutation into the at least one target template nucleic acid molecule in the second sample of the paired samples is performed by chemical mutagenesis or enzymatic mutagenesis.
64. The method according to claim 63, wherein The enzymatic mutagenesis is performed using DNA polymerase.
65. The method according to claim 64, characterized in that The DNA polymerase is a low bias DNA polymerase.
66. The method according to claim 65, characterized in that The low bias DNA polymerase introduces substitution mutations.
67. The method according to claim 65, characterized in that The low bias DNA polymerase mutates adenine nucleotides, thymine nucleotides, guanine nucleotides and cytosine nucleotides in the at least one target template nucleic acid molecule at a ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, respectively.
68. The method according to claim 65, characterized in that The low bias DNA polymerase mutates adenine nucleotides, thymine nucleotides, guanine nucleotides and cytosine nucleotides in the at least one target template nucleic acid molecule at a ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively.
69. The method according to claim 65, characterized in that The low bias DNA polymerase mutates between 1% and 15% of the nucleotides in the at least one target template nucleic acid molecule.
70. The method according to claim 65, wherein The low bias DNA polymerase mutates 0% to 3% of the nucleotides in the at least one target template nucleic acid molecule per round of replication.
71. The method according to claim 65, characterized in that The low bias DNA polymerase incorporates nucleotide analogs into the at least one target template nucleic acid molecule.
72. The method according to claim 65, wherein The low bias DNA polymerase uses nucleotide analogs to mutate adenine, thymine, guanine and / or cytosine in the at least one target template nucleic acid molecule.
73. The method according to claim 65, characterized in that The low bias DNA polymerase replaces guanine, cytosine, adenine and / or thymine with nucleotide analogs.
74. The method according to claim 65, characterized in that The low bias DNA polymerase uses nucleotide analogs to incorporate guanine nucleotides or adenine nucleotides at a ratio of 0.5-1.5:0.5-1.5, respectively.
75. The method according to claim 65, wherein The low bias DNA polymerase uses nucleotide analogs to incorporate guanine nucleotides or adenine nucleotides at a ratio of 0.7-1.3:0.7-1.3, respectively.
76. The method according to any one of claims 65, wherein The method includes a step of amplifying the at least one target template nucleic acid molecule in the second sample of the paired samples using a low-bias DNA polymerase, the step of amplifying the at least one target template nucleic acid molecule using a low-bias DNA polymerase is performed in the presence of the nucleotide analog, and the step of amplifying the at least one target template nucleic acid molecule provides the at least one target template nucleic acid molecule in the second sample of the paired samples containing the nucleotide analog.
77. The method according to claim 65, characterized in that The nucleotide analog is dPTP.
78. The method according to claim 77, characterized in that The low-bias DNA polymerase introduces guanine to adenine substitution mutations, cytosine to thymine substitution mutations, adenine to guanine substitution mutations, and thymine to cytosine substitution mutations.
79. The method according to claim 78, characterized in that The low-bias DNA polymerase introduces guanine to adenine substitution mutation, cytosine to thymine substitution mutation, adenine to guanine substitution mutation and thymine to cytosine substitution mutation at a rate ratio of 0.5-1.5:0.5-1.5:0.5-1.5:0.5-1.5, respectively.
80. The method of claim 78, wherein: The low-bias DNA polymerase introduces guanine to adenine substitution mutation, cytosine to thymine substitution mutation, adenine to guanine substitution mutation, and thymine to cytosine substitution mutation at a rate ratio of 0.7-1.3:0.7-1.3:0.7-1.3:0.7-1.3, respectively.
81. The method according to claim 65, characterized in that The low-bias DNA polymerase is a high-fidelity DNA polymerase.
82. The method according to claim 81, characterized in that In the absence of nucleotide analogs, the high-fidelity DNA polymerase introduces less than 0.01% mutations per round of replication.
83. The method according to claim 81, characterized in that The method comprises the further step of amplifying the at least one target template nucleic acid molecule comprising nucleotide analogs in the absence of nucleotide analogs.
84. The method according to claim 83, characterized in that The step of amplifying the at least one target template nucleic acid molecule comprising nucleotide analogs in the absence of nucleotide analogs is performed using a low bias DNA polymerase.
85. The method according to claim 65, wherein The method provides at least one mutated target template nucleic acid molecule, and the method further comprises the step of amplifying the mutated at least one mutated target template nucleic acid molecule using the low bias DNA polymerase.
86. The method according to claim 65, characterized in that The low bias DNA polymerase has low template amplification bias.
87. The method according to claim 65, wherein The low-bias DNA polymerase comprises a proofreading domain and / or a processivity-enhancing domain.
88. The method according to claim 65, wherein The low bias DNA polymerase comprises a fragment of at least 400 consecutive amino acids of the following sequence: Sequence of SEQ ID NO. 2; Sequence of SEQ ID NO. 4; The sequence of SEQ ID NO. 6; or The sequence of SEQ ID NO.
7.
89. The method according to claim 65, characterized in that The low bias DNA polymerases include: The sequence of SEQ ID NO. 2; Sequence of SEQ ID NO. 4; The sequence of SEQ ID NO. 6; or Sequence of SEQ ID NO.
7.
90. The method according to claim 89, wherein The low-bias DNA polymerase includes the sequence of SEQ ID NO.
2.
91. The method according to claim 89, wherein The low-bias DNA polymerase includes the sequence of SEQ ID NO.
4.
92. The method according to claim 89, wherein The low-bias DNA polymerase includes the sequence of SEQ ID NO.
6.
93. The method according to claim 89, wherein The low-bias DNA polymerase includes the sequence of SEQ ID NO.
7.
94. The method according to claim 65, characterized in that The low-bias DNA polymerase is Thermococcus polymerase or a derivative thereof.
95. The method according to claim 94, characterized in that The low bias DNA polymerase is Thermococcus polymerase.
96. The method according to claim 94, wherein The Thermococcus polymerase is derived from a Thermococcus strain selected from the group consisting of Thermococcus kodalurica, Thermococcus tachyzoites, Thermococcus thermophilus, and Thermococcus sp. KS-1.
97. The method according to claim 1, wherein The step of providing paired samples includes controlling the number of target template nucleic acid molecules in a first sample of the paired samples, each sample including at least one target template nucleic acid molecule.
98. The method according to claim 1, wherein The step of providing paired samples includes controlling the number of target template nucleic acid molecules in a second sample of the paired samples, each sample including at least one target template nucleic acid molecule.
99. The method according to claim 1, characterized in that The method further comprises the step of normalizing the number of target template nucleic acid molecules in each of the combined subsamples to provide a first sample of the paired samples and / or a second sample of the paired samples.
100. The method according to claim 1, wherein Controlling the number of target template nucleic acid molecules in the first sample of the paired samples and / or the second sample of the paired samples includes: measuring the number of target template nucleic acid molecules and diluting the first sample of the paired samples and / or the second sample of the paired samples so that the first sample of the paired samples and / or the second sample of the paired samples contain the desired number of target template nucleic acid molecules.
101. The method according to claim 99, wherein Normalizing the number of target template nucleic acid molecules in each of the subsamples includes labeling target template nucleic acid molecules from different subsamples with different sample labels.
102. The method according to claim 101, characterized in that The method includes: preparing a preliminary pool of the subsamples, which will form the first sample of the paired samples and / or the second sample of the paired samples; and measuring the number of target template nucleic acid molecules labeled with each sample tag in the preliminary pool.
103. The method according to claim 102, characterized in that Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool includes serially diluting the preparatory pool to provide a serial dilution solution comprising the diluted preparatory pool.
104. The method according to claim 102, characterized in that Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool includes sequencing the target template nucleic acid molecules in the preparatory pool or the diluted preparatory pool.
105. The method according to claim 104, characterized in that Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparation pool includes: amplifying the target template nucleic acid molecules and then sequencing them.
106. The method according to claim 104, characterized in that Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool includes amplifying, fragmenting, and then sequencing the target template nucleic acid molecules.
107. The method according to claim 102, characterized in that Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool includes: identifying the number of unique target template nucleic acid molecule sequences with each sample tag.
108. The method according to claim 102, characterized in that Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool includes mutating the target template nucleic acid molecules.
109. The method according to claim 108, characterized in that Mutating the target template nucleic acid molecule tag comprises: amplifying the target template nucleic acid molecule in the presence of nucleotide analogs.
110. The method according to claim 109, characterized in that The nucleotide analogue is dPTP.
111. The method according to claim 102, characterized in that Measuring the number of target template nucleic acid molecules labeled with each sample tag in the preparatory pool comprises: (i) mutating the target template nucleic acid molecule to provide a mutated target template nucleic acid molecule; (ii) sequencing the region of the mutated target template nucleic acid molecule; and (iii) identifying the number of unique mutated target template nucleic acid molecules having each sample tag based on the number of unique mutated target template nucleic acid molecules.
112. The method according to claim 102, characterized in that Measuring the number of target template nucleic acid molecules includes introducing a barcode into the target template nucleic acid molecules to provide barcoded, sample-tagged target template nucleic acid molecules.
113. The method according to claim 112, characterized in that Measuring the number of target template nucleic acid molecules includes introducing pairs of barcodes into the target template nucleic acid molecules to provide barcoded, sample-tagged target template nucleic acid molecules.
114. The method according to claim 112, characterized in that Measuring the number of target template nucleic acid molecules labeled with each sample tag includes: (i) sequencing a region of the barcoded, sample-tagged target template nucleic acid molecule; and (ii) identifying the number of unique barcoded target template nucleic acid molecules having each sample tag based on the number of unique barcodes or paired barcode sequences associated with each sample tag.
115. The method according to claim 101, wherein The method includes calculating a ratio of the number of target template nucleic acid molecules comprising different sample tags.
116. The method of claim 1, wherein The first sample and / or the second sample of the paired samples are provided by recombining the subsamples so that the number of target template nucleic acid molecules in each of the subsamples is in a desired ratio.
117. The method according to claim 101, characterized in that The method includes labeling target template nucleic acid molecules from different samples before combining the subsamples.
118. A computer implemented method comprising the method of claim 2.
Citation Information
Patent Citations
A method for nucleic acid sequence analysis
WO2002079502A1