System and method for determining nucleic acid sequences based on mutant sequence reads

The method aligns non-variant sequence reads with variant reads to determine the most likely sequence of a nucleic acid template, addressing errors in repetitive regions by iteratively updating probabilities, thus improving sequencing accuracy and precision.

JP2026509691APending Publication Date: 2026-03-25ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-05
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Conventional nucleotide sequencing techniques fail to accurately represent repetitive regions due to introducing errors, often producing a 'mean sequence' instead of an accurate representation of the nucleic acid template, particularly when using error correction or polishing methods.

Method used

A method for determining the sequence of a nucleic acid template by aligning non-variant sequence reads to variant sequence reads, incorporating per-base accuracy scores, and applying an optimization process to iteratively update the probability that non-mutant sequence reads originate from the same copy as mutant sequence reads, thereby determining the most likely sequence.

Benefits of technology

Accurately predicts the sequence of the nucleic acid template by statistically removing mutations, improving alignment accuracy and confidence in base calls, especially in regions with multiple copies or haplotypes, and enhancing the precision of sequencing repetitive regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509691000001_ABST
    Figure 2026509691000001_ABST
Patent Text Reader

Abstract

Methods and systems for determining the sequence of a nucleic acid template by removing mutations found in mutant sequence reads are described herein. In some embodiments, the methods and systems include the steps of: aligning a non-mutant sequence read of the nucleic acid template with a mutant sequence read; determining the most likely sequence of the nucleic acid template and a base-by-base precision score correlated with the probability that a base matches the nucleic acid template, thereby determining the sequence of the nucleic acid template.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Citation by reference of any priority application Any application in which a foreign or domestic priority claim is identified in the application data sheet filed with this application is incorporated herein by reference under 37 CFR 1.57.

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 489,550, filed on 10 March 2023, which is incorporated herein by reference in its entirety.

[0003] This disclosure relates in general to the field of nucleic acid sequencing. More specifically, this disclosure relates to the field of sequencing nucleic acid templates by removing mutations found in mutant sequence reads. [Background technology]

[0004] explanation Conventional nucleotide sequencing techniques, known as "error correction" or "polishing," are used to correct errors unintentionally introduced into the sequencing process by noisy or error-prone measurements of the primary sequence signal. However, such techniques have been criticized for being ineffective, particularly in repetitive regions. In some cases, using such methods produces a "mean sequence" of copies of the repetitive region rather than an accurate representation of any particular copy of the nucleic acid template containing the repetitive region. [Overview of the project]

[0005] The methods disclosed herein each have several aspects, none of which alone carry their desirable attributes. Without limiting the claims, some notable features are briefly considered here. Numerous other embodiments are contemplated, including embodiments having fewer, additional, and / or different components, steps, features, objectives, benefits, and advantages. Components, aspects, and steps can be differently arranged and ordered. After considering this discussion, particularly after reading the section entitled "Best Mode for Carrying Out the Invention," one will understand what advantages the features of the systems and methods disclosed herein provide over other known systems and methods.

[0006] A computer-implemented method for determining the sequence of a nucleic acid template by removing mutations found in variant sequence reads is described herein. In some embodiments, the method comprises aligning non-variant sequence reads of the nucleic acid template to variant sequence reads, wherein the variant sequence reads contain one or more mutations, and determining the most likely sequence of the nucleic acid template and, for each base of the most likely sequence, determining a per-base accuracy score that correlates with the probability that the base matches the nucleic acid template, thereby determining the sequence of the nucleic acid template.

[0007] In some embodiments, the variant sequence reads are from a mutant nucleic acid template. In some embodiments, the one or more mutations include mutations introduced during sample preparation. In some embodiments, the one or more mutations include mutations randomly introduced during sample preparation by mutagenesis. In some embodiments, the one or more mutations include mutations introduced inadvertently. In some embodiments, the nucleic acid template is longer than the variant and non-variant sequence reads.

[0008] In some embodiments, determining the most likely sequence of a nucleic acid template involves determining the most likely base for each site associated with the alignment of non-variant sequence reads to variant sequence reads. In some embodiments, determining the most likely sequence of a nucleic acid template involves determining the most likely sequence for a subset of sites associated with the alignment of non-variant sequence reads to variant sequence reads.

[0009] In some embodiments, determining the most likely sequence of a nucleic acid template and determining a per-base accuracy score for each base of the most likely sequence involves, for each site, summarizing the probabilities of each base at each site associated with non-variant sequence reads.

[0010] In some embodiments, determining the most likely sequence of a nucleic acid template and determining a per-base accuracy score for each base of the most likely sequence involves incorporating information related to the probability of a mutation at each site of the variant sequence reads.

[0011] In some embodiments, determining the most likely sequence of a nucleic acid template and determining a per-base accuracy score for each base of the most likely sequence involves incorporating information that describes the probability of errors associated with the sample preparation or sequencing process.

[0012] In some embodiments, determining the most likely sequence of a nucleic acid template and determining a per-base accuracy score for each base of the most likely sequence involves determining the most likely base for sites of variant sequence reads that do not align to non-variant sequence reads, based on the statistically inferred probability of a mutation at the site.

[0013] In some embodiments, determining the most likely sequence of a nucleic acid template and determining a per-base accuracy score for each base of the most likely sequence involves incorporating information that describes the variation in mutation rates between variant sequence reads or the unequal mutation rates of adenine, cytosine, guanine, and thymine.

[0014] In some embodiments, the nucleic acid template is derived from a nucleic acid sample containing two or more copies of the repeating region of the nucleic acid template.

[0015] In some embodiments, determining the most likely sequence of a nucleic acid template and determining a base-by-base precision score for each base of the most likely sequence includes estimating the probability that a non-mutant sequence read originated from the same copy of the repeating region as the mutant sequence read to which the non-mutant sequence read is aligned.

[0016] In some embodiments, determining the most likely sequence of a nucleic acid template and determining a base-by-base precision score for each base in the most likely sequence involves applying an optimization process, which iteratively updates the estimate of the probability that a non-mutant sequence read originated from the same copy of the mutant sequence read to which the non-mutant sequence read is aligned, while updating the inference of the most likely base. In some embodiments, the optimization process involves applying an expectation maximization method. In some embodiments, the optimization process incorporates information describing variations in read length.

[0017] In some embodiments, the contribution of each non-mutant sequence read to the probability calculation of the most likely base is weighted by a weighting scheme. In some embodiments, the weighting scheme relates to the sequence identity of the non-mutant sequence read and the current iterative estimation of the nucleic acid template sequence.

[0018] In some embodiments, determining the most likely sequence of a nucleic acid template and determining a base-by-base precision score for each base of the most likely sequence includes selecting a subset of non-mutant sequence reads from non-mutant sequence reads derived from different copies of the repeat region. In some embodiments, selecting a subset of non-mutant sequence reads includes scoring the alignment of the non-mutant sequence reads to the mutant sequence reads and selecting the top-scoring alignment.

[0019] In some embodiments, the method includes determining the sequence of copies of a repeating region of a nucleic acid template. In some embodiments, two or more copies of the repeating region include two or more haplotypes, and the method includes determining the sequence of the haplotypes of a nucleic acid sample.

[0020] In some embodiments, the accuracy score per base includes a Quality Value (QV) per base. In some embodiments, the method further includes determining an overall accuracy score associated with the most likely sequence.

[0021] In some embodiments, the method further includes assembling the non-mutant sequence reads of the nucleic acid template before aligning them to the mutant sequence reads. In some embodiments, aligning the non-mutant sequence reads of the nucleic acid template to the mutant sequence reads includes aligning the mutant sequence reads to the assembly of non-mutant sequence reads.

[0022] In some embodiments, the method further includes assembling the mutant sequence reads before aligning the non-mutant sequence reads of the nucleic acid template to the mutant sequence reads. In some embodiments, aligning the non-mutant sequence reads of the nucleic acid template to the mutant sequence reads includes aligning the non-mutant sequence reads to the assembly of mutant sequence reads.

[0023] In some embodiments, the method further includes aligning the non-mutant sequence reads of the nucleic acid template to the mutant sequence reads before aligning the non-mutant sequence reads to the mutant sequence reads.

[0024] In some embodiments, the method further includes assembling the mutant sequence reads before aligning the non-mutant sequence reads of the nucleic acid template to the mutant sequence reads, generating a consensus of the assembled mutant sequence reads, and aligning the non-mutant sequence reads to the consensus. In some embodiments, aligning the non-mutant sequence reads of the nucleic acid template to the mutant sequence reads includes replacing the consensus with the mutant sequence reads in the alignment of the non-mutant sequence reads to the consensus. In some embodiments, assembling the mutant sequence reads includes aligning the mutant sequence reads to a reference sequence.

[0025] In some embodiments, the method further includes sequencing a non-mutant nucleic acid template molecule and a mutant nucleic acid template molecule to obtain a non-mutant sequence read and a mutant sequence read.

[0026] In some embodiments, the method includes aligning non-mutant and mutant sequence reads to a reference sequence; determining the most likely base for each site related to the alignment of the non-mutant sequence read to the mutant sequence read, and incorporating information describing the variation in mutation rates between mutant sequence reads or the unequal mutation rates of adenine, cytosine, guanosine, and thymine; and determining further rendered sequences by comparing k-mers from the initial rendered sequence with k-mers from non-mutant sequence reads.

[0027] In some embodiments, the method includes aligning non-mutant and mutant sequence reads to a reference sequence; realigning non-mutant sequence reads having mapping positions that directly overlap with mutant sequence reads to mutant sequence reads; scoring the alignment of non-mutant sequence reads to mutant sequence reads and selecting the top-scoring alignment; and determining the most likely base for each site associated with the alignment of non-mutant sequence reads to mutant sequence reads, thereby determining the rendered sequence.

[0028] In some embodiments, the method includes determining the most likely base for each site related to the alignment of a non-mutant sequence read to a mutant sequence read, and determining the rendered sequence by applying an expectation maximization (EM) optimization process, the EM optimization process iteratively updates the estimate of the probability that the non-mutant sequence read originated from the same copy of the mutant sequence read to which the non-mutant sequence read is aligned, while updating the inference of the most likely base.

[0029] In some embodiments, the method includes generating a consensus of assembled mutant sequence reads, aligning a non-mutant sequence read to the consensus, replacing the consensus with the mutant sequence read in the alignment of the non-mutant sequence read to the consensus, determining the most likely base for each site related to the alignment of the non-mutant sequence read to the mutant sequence read, and determining the rendered sequence by applying an expectation maximization (EM) optimization process, wherein the EM optimization process iteratively updates the estimate of the probability that the non-mutant sequence read originated from the same copy as the mutant sequence read to which the non-mutant sequence read is aligned, while updating the inference of the most likely base.

[0030] In another embodiment, a system comprising a processor configured to perform the method is disclosed herein. In some embodiments, the method includes aligning a non-mutant sequence read of a nucleic acid template to a mutant sequence read of a mutant nucleic acid template, wherein the mutant sequence read contains one or more mutations; determining the most likely sequence of the nucleic acid template; determining a base-by-base precision score for each base of the most likely sequence, which correlates with the probability that the base matches the nucleic acid template; and thereby determining the sequence of the nucleic acid template.

[0031] In further embodiments, computer-readable media are disclosed herein. In some embodiments, the computer-readable media includes instructions that, when executed by a processor, perform a method comprising: aligning a non-mutant sequence read of a nucleic acid template to a mutant sequence read of a mutant nucleic acid template, wherein the mutant sequence read contains one or more mutations; determining the most likely sequence of the nucleic acid template; determining a base-by-base precision score for each base of the most likely sequence, which correlates with the probability that the base matches the nucleic acid template; and thereby determining the sequence of the nucleic acid template. [Brief explanation of the drawing]

[0032] [Figure 1] This document describes an embodiment of a method for determining the sequence of a nucleic acid template by removing mutations found in a mutant sequence read. [Figure 2] The present invention illustrates an embodiment of a system comprising a processor configured to perform a method for determining the sequence of a nucleic acid template by removing mutations found in a mutant sequence read. [Figure 3] This example demonstrates how to determine the sequence of a nucleic acid template using mutant and non-mutant sequence reads. [Figure 4] This example shows how to determine the sequence at a site containing a heterozygous transition SNP. [Figure 5] This example demonstrates how to determine the sequence of an actual human genome dataset. [Modes for carrying out the invention]

[0033] All patents, applications, published applications, and other publications referenced herein are incorporated herein by reference in their entirety. Where any term or phrase is used herein in a manner that contradicts or otherwise contradicts the definitions contained in the patents, applications, published applications, and other publications incorporated herein by reference, the use herein shall prevail over the definitions incorporated herein by reference.

[0034] Some embodiments relate to systems and methods for determining the nucleotide sequence of a nucleic acid template. In some cases, sequencing a nucleic acid template may involve introducing random mutations into the template so that the template can be fragmented into smaller sequences that are easier for next-generation sequencing processes to perform. Incorporating random mutations may help to reconstruct the sequence of the mutant nucleic acid template by aligning the mutant fragments with each other. However, in some cases, the sequence of the original non-mutant template molecule is desirable. Embodiments relate to methods for determining the non-mutant sequence of the original nucleic acid template.

[0035] One embodiment of determining the non-mutant sequence of the original nucleic acid template may include fragmenting the non-mutant nucleic acid template and determining the sequence of the non-mutant fragments. The system and method can then align the non-mutant sequence reads from the nucleic acid template with the mutant sequence reads of the nucleic acid template to determine the most probable sequence of the non-mutant nucleic acid template. This process may include determining a base-by-base precision score for each base of the most probable non-mutant nucleic acid template, which correlates with the probability that the base matches the nucleic acid template. Using this base-by-base precision score, the most probable sequence of the non-mutant nucleic acid template can be determined. This is described in more detail in the following sections.

[0036] definition As used herein, “align,” “align,” or “to align” refers to arranging two or more nucleic acid sequences to identify regions of similarity.

[0037] A "nucleic acid template" can be any nucleic acid molecule that the user intends to sequence. A "nucleic acid template" can be single-stranded or part of a double-stranded complex. If the nucleic acid template consists of deoxyribonucleotides, it can form part of a double-stranded DNA complex. In this case, one strand (e.g., the coding strand) can be considered the nucleic acid template, and the other strand can be a nucleic acid molecule complementary to the nucleic acid template. A nucleic acid template can be a DNA molecule corresponding to a gene, may contain introns, may be an intergeneric region, an intrageneric region, a genomic region spanning multiple genes, or in practice, the entire genome of an organism. A nucleic acid template may contain repetitive sequences, such as variable-number tandem repeats (VNTRs) or short tandem repeats (STRs). A nucleic acid template can originate from a nucleic acid sample.

[0038] A "nucleic acid sample" includes any sample of nucleic acid. A nucleic acid sample may be a sample of nucleic acid derived from humans, for example, a sample extracted from a skin swab of a human patient. A nucleic acid sample may be a sample of nucleic acid derived from non-human sources. A nucleic acid sample may originate from a source such as a sample from a water source. Such a sample may contain hundreds, thousands, millions, or billions of nucleic acid template molecules. A nucleic acid sample may contain a single genome, or it may contain multiple non-identical genomes (e.g., from different organisms or different species or strains). A nucleic acid sample may contain multiple haplotypes from a single organism.

[0039] As used herein, “repeating region” refers to a region within a nucleic acid template whose nucleotide sequence is copied and, in at least one other instance, present in a nucleic acid sample. In other words, a nucleic acid sample may have two or more copies of a repeating region of a nucleic acid template. In some embodiments, the two copies of the repeating region do not have to be identical and may, for example, contain a single base difference. In some embodiments, the repeating region contains a repeating sequence (e.g., VNTR or STR), where, for example, the repeats are adjacent to each other or arranged in tandem. For example, a nucleic acid sample may contain a nucleic acid template having a repeating region containing a repeating sequence, and the nucleic acid sample may contain another copy of the repeating region having the repeating sequence in another location (e.g., on another chromosome or at a different location on the same chromosome). As another example, in some embodiments, the repeating region may contain the entire length of a chromosome. In such embodiments, a “copy” of the repeating region is another copy of the chromosome found in the nucleic acid sample (e.g., a second haplotype).

[0040] A "mutant nucleic acid template" refers to a nucleic acid template that contains one or more mutations, such as mutated bases, compared to the original sequence of the nucleic acid template. In other words, a mutant nucleic acid template is a mutated copy of the nucleic acid template.

[0041] As used herein, “mutation” means a change in a nucleic acid sequence. For example, a mutation includes the substitution of one nucleotide base with another. A mutation may be a transition mutation (A to G, G to A, C to T, or T to C). In another embodiment, a mutation is a transversion mutation. In some embodiments, mutations are intentionally introduced into nucleic acids through a sample preparation step that randomly introduces mutations into the nucleic acid molecule. In some embodiments, mutations are introduced unintentionally, such as by errors or unintended results during sample preparation or library preparation. In some embodiments, mutations are caused by unintended errors during sequencing of the nucleic acid template, such as base calling errors.

[0042] As used herein, “mutant sequence read” is also referred to herein as “Dream” or “Dream Sequence” and refers to a sequence read that contains mutations compared to a nucleic acid template. In some embodiments, the mutant sequence read is a sequence read of a mutant nucleic acid template. In some embodiments, the mutant sequence read is obtained by sequencing a region of the mutant nucleic acid template. In some embodiments, the mutant sequence read is obtained by sequencing a region of the nucleic acid template and contains mutations resulting from sequencing errors. The mutant sequence read may have a length shorter than the length of the nucleic acid template. In some embodiments, the mutant sequence read has a length of about 100 to 600 nt, for example, about 150 nt or about 300 nt. In some embodiments, the mutant sequence read may be assembled by aligning the mutant sequence reads. Any method known in the art (for example, the method described in U.S. Patent Application Publication 2021 / 0174905(A1)) may be used for such alignment.

[0043] As used herein, “non-mutant sequence read” refers to a sequence read of a nucleic acid template. In some embodiments, non-mutant sequence reads are obtained by sequencing a region of the nucleic acid template. Non-mutant sequence reads may have a length shorter than the length of the nucleic acid template, for example, about 100 to 600 nt, for example, about 150 nt or about 300 nt.

[0044] "Rendering" refers to the process of determining the most likely sequence of a nucleic acid template using alignment of non-mutant and mutant sequence reads. In some embodiments, rendering includes determining the most likely sequence of a nucleic acid template by replacing mutations in mutant sequence reads with the most likely bases.

[0045] The “most likely base” refers to the nucleic acid base that is calculated to have the highest probability of matching the nucleic acid template from which the mutant sequence read originates (in some cases, a specific copy of the nucleic acid template). For example, the most likely base may be the maximum a posteriori base as described herein.

[0046] The “most likely sequence” refers to the nucleic acid sequence (e.g., the non-mutant nucleic acid template that existed before the mutation was introduced) that is calculated to have the highest probability of matching the nucleic acid template from which the mutant sequence read ultimately originates (in some cases, a specific copy of the nucleic acid template). Therefore, in some embodiments, the most likely sequence does not contain the mutation (it is statistically predicted that it will not). Therefore, in some embodiments, the most likely sequence can be generated by statistically removing the mutation from the mutant sequence read and estimating the sequence of the (non-mutant) nucleic acid template.

[0047] Method for determining the sequence of a nucleic acid template using mutant and non-mutant sequence reads. Embodiments relate to a method for sequencing a nucleic acid template molecule using a sample preparation step that intentionally introduces random mutations into a copy of a sample nucleic acid. For example, a method for introducing random mutations into a target nucleic acid molecule is described in U.S. Patent Application Publication 2021 / 0010008(A1). In some embodiments, the mutations are introduced by Illumina® Complete Long Read sequencing sample preparation chemistry. The introduction of such random mutations may be useful in sequence assembly, for example, as described in U.S. Patent Application Publication 2021 / 0174905(A1). For example, the creation of a mutated copy of a nucleic acid template may be particularly useful for performing nucleic acid sequencing when the nucleic acid template contains repeating regions, for example, when the nucleic acid sample contains two or more non-identical copies (e.g., haplotypes) of repeating regions.

[0048] Accordingly, some embodiments of the systems and methods disclosed herein include the step of obtaining one or more nucleic acid samples (including obtaining one nucleic acid sample that is divided into two nucleic acid samples, including two nucleic acid samples from the same subject). Some embodiments include the step of randomly introducing mutations into nucleic acid molecules in one or more nucleic acid samples to generate mutant nucleic acid samples. Some embodiments include keeping the nucleic acid samples non-mutated by not randomly introducing mutations into nucleic acid molecules in the nucleic acid samples. Some embodiments herein also include sequencing the nucleic acid molecules (e.g., nucleic acid templates) of the (non-mutated) nucleic acid samples to generate non-mutated sequence reads. Some embodiments disclosed herein also include sequencing the nucleic acid molecules (e.g., mutant nucleic acid templates) of mutant nucleic acid samples to generate mutant sequence reads. Sequencing of mutant and non-mutated nucleic acid samples can be performed simultaneously or sequentially. Sequencing can be achieved by any method known to those skilled in the art, including next-generation sequencing (NGS).

[0049] While mutant sequence reads are useful for determining the sequence of a nucleic acid template, it is sometimes desirable to determine the sequence of the original nucleic acid template molecule that existed before the introduction of random mutations. To achieve this, mutant sequence reads can be aligned with non-mutant sequence reads in a process referred to herein as "rendering" to determine the sequence of the nucleic acid template.

[0050] However, determining the alignment of mutant sequence reads to non-mutant sequence reads can be ambiguous in some situations. For example, if the nucleic acid template has multiple copies of a repeating region (e.g., a repeating region containing a repeating sequence), and / or if the nucleic acid sample contains two or more haplotypes of the nucleic acid template, the alignment of non-mutant sequence reads to mutant sequence reads can be ambiguous because mutant sequence reads can align at many different positions on the nucleic acid template.

[0051] This disclosure relates to a system and method for accurately predicting the sequence of a nucleic acid template given one or more mutant sequence reads and one or more non-mutant sequence reads from the same nucleic acid sample. In some embodiments, the method includes aligning the mutant sequence reads and non-mutant sequence reads, and statistically determining the most likely sequence of the nucleic acid template and a base-by-base accuracy score corresponding to each base of the most likely sequence.

[0052] Various computer implementation methods for determining the correct sequence of a nucleic acid template by removing mutations found in a mutant sequence read are described herein. In some embodiments, the mutant sequence read contains one or more nucleic acid mutations. In some embodiments, the mutant sequence read is a mutant nucleic acid template. In some embodiments, one or more mutations include mutations introduced during sample preparation. In some embodiments, one or more mutations are randomly introduced during sample preparation by mutagenesis. In some embodiments, the mutations include nucleic acid mutations that are intentionally introduced. The method for introducing such mutations may be any method known in the art. In some embodiments, the mutations include substitution mutations. In some embodiments, the mutations include transition mutations.

[0053] In some embodiments, the nucleic acid template is longer than the mutant sequence read and / or non-mutant sequence read. For example, in some embodiments, the non-mutant sequence read and mutant sequence read are constructed from a range of more than 10 bp, more than 100 bp, less than 200 bp, less than 300 bp, less than 400 bp, less than 500 bp, less than 600 bp, less than 700 bp, less than 1,000 bp, less than 3,000 bp, or any of the aforementioned values. In some embodiments, the non-mutant sequence read and mutant sequence read are approximately 300 bp long. In some embodiments, the non-mutant sequence read and mutant sequence read are approximately 150 bp long.

[0054] In some embodiments, the methods described herein are used, for example, to remove mutations unintentionally introduced into the sequencing process by noisy or error-prone measurements of the primary sequence signal. In such embodiments, the methods described herein may be considered an improved method of error correction or polishing.

[0055] Alignment to the reference sequence and / or consensus sequence Some embodiments include a system or method for determining the sequence of a nucleic acid template by removing mutations found in a mutant sequence read by the method shown in Figure 1. As illustrated, the method begins in step S100, where sequence reads (e.g., non-mutant sequence reads and / or mutant sequence reads) are aligned to a reference sequence.

[0056] Alignment can be performed by methods known to those skilled in the art. The reference sequence may be derived from a different nucleic acid sample than the nucleic acid template, or from the same nucleic acid sample. In some embodiments, the reference sequence is derived from the same biological species as the sample nucleic acid template. In some embodiments, the reference sequence is a publicly available reference sequence.

[0057] In some embodiments, the system and method further include step S110 of assembling mutant sequence reads, generating a consensus of the assembled mutant sequence reads, and aligning non-mutant sequence reads to the consensus.

[0058] Previous studies on sequence analysis by mutagenesis, such as the one disclosed in Keith JM, Adams P, Bryant D, Cochran DAE, Lala GH, Mitchelson KR. Algorithms for sequence analysis via mutagenesis. Bioinformatics. 2004 05;20(15):2401-2410. doi:10.1093 / bioinformatics / bth258, have developed techniques for accurately calling a consensus sequence from a set of mutant sequences. To determine the most likely sequence of a nucleic acid template, previous approaches involved constructing a multiple alignment of mutant sequence reads via mapping to a reference sequence or via de novo multiple alignment, calling a consensus, copying the consensus sequence to each corresponding site in the aligned mutant sequence reads, and thereby determining the most likely sequence of the nucleic acid template. When a reference sequence is used, it can be constructed independently, from non-mutant sequence reads, or from mutant sequence reads.

[0059] In some embodiments, instead of directly using consensus to determine the most likely base or sequence, a consensus of mutant sequence reads is used instead to improve the accuracy of the alignment of non-mutant sequence reads to mutant sequence reads. Briefly, in step S110, the consensus of mutant sequence reads is invoked using, for example, the method disclosed in Keith et al. (2004) to generate the consensus of mutant sequence reads. Further in step S110, the consensus is then used as a mapping target for non-mutant sequence reads. Further in step S110, the alignment to the consensus is translated back to the original (non-consensus) mutant sequence read, and the alignment of non-mutant sequence reads to the original mutant sequence read is used in the sequencing method described herein.

[0060] One advantage of this approach over directly performing alignment of non-mutant sequence reads to the original mutant sequence reads is that consensus can have much higher sequence similarity to non-mutant sequence reads derived from the same haplotype. As further described herein, nucleic acid samples from which nucleic acid templates originate may contain multiple copies of repetitive regions, such as multiple haplotypes, which can make the alignment of non-mutant sequence reads to mutant sequence reads ambiguous or inaccurate. Higher sequence identity means that mapping and alignment are more accurate, computations are faster, and fewer non-mutant sequence reads may be included that end up being aligned to mutant sequence reads from other copies of the repetitive region (e.g., other haplotypes). An advantage of this approach over using consensus methods to determine the most likely sequence is that the use of non-mutant sequence reads can provide much higher confidence and accuracy in the resulting base call.

[0061] Alignment of mutant and non-mutant sequence reads In some embodiments, the methods and systems described herein further include step S120 of aligning a non-mutant sequence read of a nucleic acid template with a mutant sequence read. The mechanism for generating such alignment may be any method known to those skilled in the art.

[0062] In some embodiments, after step S110 of generating a consensus sequence, step S120 of aligning non-mutant sequence reads to mutant sequence reads proceeds by replacing the consensus with mutant sequence reads in the alignment of non-mutant sequence reads to the consensus. In some embodiments, no consensus sequence is generated, and the non-mutant sequence reads and mutant sequence reads are aligned by mapping them to each other by methods known to those skilled in the art.

[0063] In some embodiments, the alignment of mutant sequence reads to non-mutant sequence reads may include coverage of about 30 times that of the non-mutant sequence reads. In some embodiments, the alignment includes coverage of about 10 times, about 15 times, about 20 times, about 30 times, about 40 times, about 50 times, about 60 times, or more, or a range constructed from any of the aforementioned values. In some embodiments, the alignment includes the above-mentioned coverage of individual haplotypes of the nucleic acid template sample through the non-mutant sequence reads, for example, 15 times coverage.

[0064] In some embodiments, the method further includes assembling the non-mutant sequence reads to generate an assembly of non-mutant sequence reads prior to the alignment in step S120. In some embodiments, the non-mutant sequence reads are assembled using a reference sequence. In some embodiments, step S120, which aligns the non-mutant sequence reads to mutant sequence reads, may include aligning the mutant sequence reads to the assembly of non-mutant sequence reads.

[0065] In some embodiments, prior to the alignment of non-mutant sequence reads to mutant sequence reads in step S120, the method includes assembling the mutant sequence reads to generate an assembly of mutant sequence reads. In some embodiments, the mutant sequence reads are assembled by using a reference sequence. In some embodiments, the alignment of non-mutant sequence reads to mutant sequence reads in step S120 may include aligning the non-mutant sequence reads to the assembly of mutant sequence reads.

[0066] Approach to multiple copies of iterative regions In some embodiments, the nucleic acid template is derived from a nucleic acid sample containing two or more copies of the repeating region of the nucleic acid template. For example, the nucleic acid sample may contain two or more haplotypes. As another example, the nucleic acid sample may contain a repeating region containing a sequence (e.g., a repeating sequence) that has a copy at a different location on the nucleic acid molecule in the nucleic acid sample (e.g., another chromosome or a different region of the same chromosome, in other words, a different location of the genome found in the nucleic acid sample). In some embodiments, each copy of the repeating region has slight variations from one another so that the nucleic acid sample contains two or more non-identical copies of the repeating region.

[0067] A challenge faced by methods for sequencing target nucleic acids using mutant and non-mutant sequence reads is the uncertainty of whether any given non-mutant sequence read covering a repetitive region originates from the same "copy" of the repetitive region as the mutant sequence read covering the same region. "Same copy" can be conceptualized as meaning that an error-free non-mutant sequence read is 100% identical to the homologous portion of an aligned (unknown) error-free (mutant removed) mutant read sequence. This challenge can be particularly difficult when sequence reads cover long repetitive regions of the nucleic acid template.

[0068] Furthermore, a problem arising from mutagenesis (and error-prone sequencing) is that a specific change in a repeating region can alter the sequence of one repeating region to be more similar to another copy of the repeating region that may exist elsewhere in the nucleic acid sample (e.g., on a different chromosome, a different region on the same chromosome, or a different haplotype). This also contributes to ambiguity in the alignment of non-mutant sequence reads with mutant sequence reads.

[0069] When the above methods are applied in their simplest forms, a non-mutant sequence read covering one copy of a repeating region may align with a mutant sequence read covering another copy of the repeating region, introducing errors. Alignment of a non-mutant sequence read to an inaccurate copy of a repeating region in a mutant sequence read can hinder the accurate identification of the mutant site within the mutant sequence read, potentially resulting in numerous low-confidence (low QV) base calls (e.g., bases with low base-by-base precision scores) in the most likely sequence.

[0070] When the nucleic acid sample from which the nucleic acid template originates contains two or more copies of a repeating region, various exemplary approaches to determining the most likely sequence of the nucleic acid template are described in further detail below. For example, an approach to this problem may include step S130 of applying an optimization process, or step S130 of selecting a subset of non-mutant sequence reads as further described herein.

[0071] In some embodiments, the method includes determining the sequence of copies of a repeating region in a nucleic acid template. In non-limiting examples, the systems and methods described herein may be used to determine the sequence of copies of a repeating region in a nucleic acid template, where the repeating region (e.g., a repeating region containing a repeating sequence) has another non-identical copy at a different location in the nucleic acid sample (e.g., another location in the genome). In some embodiments, two or more copies of the repeating region comprise two or more haplotypes, and the method includes determining the sequence of haplotypes in the nucleic acid sample. In non-limiting examples, a human-derived nucleic acid sample may comprise two haplotypes. Embodiments of the systems and methods disclosed herein can advantageously determine the sequence of one or more haplotypes.

[0072] Application of the optimization process After the process has aligned the non-mutant sequence read and the mutant sequence read in step S120, the system and method may proceed to step S130, which applies an optimization process (including iterative steps S140, S150, and S160). In some embodiments, in step S130, the optimization process iteratively updates the estimate of the probability that the non-mutant sequence read originated from the same copy as the mutant sequence read to which the non-mutant sequence read is aligned, while updating the most likely base inference.

[0073] As previously stated herein, a challenge faced by polishing or error correction techniques related to rendering, alignment, and alignment is the handling of repetitive regions (similar but not identical copies of a repetitive region present in different abundances in the nucleic acid sample). For this reason, it may be advantageous to determine (or estimate the probability that) a non-mutant sequence read originates from the same copy of the repetitive region as the aligned mutant sequence read, and to incorporate this estimation information into the rendering process so that sequence reads that may originate from a different copy of the repetitive region do not contribute much to the final result. Various exemplary techniques, such as step S130 in which an optimization model is applied, are described below. In some embodiments, any combination of these methods is used to estimate whether a non-mutant sequence read originates from the same copy of the repetitive region as the mutant sequence read.

[0074] This can be illustrated by referring to the following non-limiting, exemplary examples. In one example, there are two similar but non-identical repeat regions, r1 and r2, where the r1 sequence is present in only one location, but the r2 sequence is present in 10 locations in the genome of the nucleic acid sample. Alternatively, in another example in a metagenomic setting, r2 is present in the nucleic acid sample in a species or strain at a 10-fold higher abundance than the species or strain containing r1. In both such examples, the alignment of non-mutant sequence reads to mutant sequence reads typically produces alignments of non-mutant sequence reads from both r1 and r2 to mutant sequence reads from both r1 and r2. The problem arises when determining the sequence of a nucleic acid template containing r1 because 10-fold more non-mutant sequence reads from r2 align (inaccurately) to the mutant sequence reads of r1 than non-mutant sequence reads from r1. If there is no technique to select the correct non-mutant sequence read that originates from the same copy of the repeating region r1, the non-mutant sequence read from r2 will dominate the signal when rendering r1, and the r2 sequence will be received as a byproduct of rendering to the most likely predicted sequence of r1. This phenomenon occurs when using standard error correction or assembly polishing techniques and has been reported by the academic community to degrade the accuracy of repeating regions. However, techniques that can weight the contribution of non-mutant sequence reads to rendering, such as those described herein, have the ability to select the correct read for use in rendering in these situations.

[0075] As a non-limiting example, the method may include step S140 for each non-mutant sequence read aligned to one or more mutant sequence reads, estimating the probability that the non-mutant sequence read was generated from the same copy of the repeat region (e.g., the same haplotype) as the one or more mutant sequence reads to which the non-mutant sequence read is aligned. In some embodiments, step S140 includes estimating the probability that the non-mutant sequence read aligned to one or more mutant sequence reads at a given site was generated by sequencing a nucleic acid molecule containing the same copy of the repeat region as the one or more mutant sequence reads to which the non-mutant sequence read is aligned. In some embodiments, the non-mutant sequence read covers the repeat region. For example, step S140 may estimate the probability that the non-mutant sequence read originates from the same haplotype as the mutant sequence read to which the non-mutant sequence read is aligned. In step S140, various statistical models can be used to estimate the probability that the non-mutant sequence read originates from the same copy of the repeat region as the mutant sequence read to which the non-mutant sequence read is aligned. Therefore, the systems and methods described herein may include step S140 of estimating the probability that the non-mutant sequence read originated from a copy of the same repeat region as the mutant sequence read to which the non-mutant sequence read is aligned.

[0076] In some embodiments, the method further includes step S150, in which the contribution of each non-mutant sequence read to the probability calculation of the most likely base is weighted by a weighting scheme. For example, in step S150, non-mutant sequence reads aligned to mutant sequence reads may be given unequal weights when statistically inferring the most likely base. For example, if, for a particular site in the alignment of non-mutant sequence reads to mutant sequence reads, some non-mutant sequence reads have the base "G" and some have the base "A", then the non-mutant sequence reads having the base "G" may be given more weight when statistically inferring the most likely base for that site.

[0077] In some embodiments, in step S150, the weighting scheme relates to the sequence identity of the non-mutant sequence read and the current iterative estimation of the nucleic acid template sequence. For example, in step S150, the weighting scheme may relate to the estimated probability that the non-mutant sequence read originates from the same copy of the iterative region (e.g., the same haplotype and / or location in the genome) as the mutant sequence read to which the non-mutant sequence read is aligned. For example, in some embodiments, each pairwise alignment of a non-mutant sequence read to a mutant sequence read has one weight based on the estimated probability that the non-mutant sequence read and the mutant sequence read originate from the same copy of the iterative region. Such estimated probabilities may be iteratively updated as described above. In step S150, the weighting scheme may also relate to the current iterative estimation of the most likely sequence of the nucleic acid template, the most likely sequence which is iteratively updated as described herein. Therefore, in some embodiments, in step S150, the weighting scheme can iteratively update the weighted contribution of each non-mutant sequence read aligned to one or more mutant sequence reads at a given site, based on the current estimated probability of sequence identity of the non-mutant sequence read and the currently predicted most likely sequence of the nucleic acid template.

[0078] In step S150, the contribution of non-mutant sequence reads may be weighted based on current estimated probabilities, such as the estimated probabilities based on the current iteration of the optimization process. Thus, when determining the most likely sequence, the estimated probabilities for non-mutant sequence reads may then be taken into consideration. In step S160, the process proceeds to update the inference for the most likely base for the site. In some embodiments, the inference for the most likely base involves calculating a probability distribution over bases (e.g., over A, C, G, T) for the site. In some embodiments, the most likely base is the base with the highest statistical likelihood of matching the nucleic acid template at the site of the base. In some embodiments, in step S160, the method may incorporate the weighted contribution of non-mutant sequence reads when updating the inference for the most likely base for the site. In some embodiments, in step S160, the inference for the most likely base is updated for each site in a subset of sites.

[0079] In some embodiments, in step S140, the updated probabilities may be re-estimated using the updated inference for the most likely base in step S160. In step S150, the updated estimated probabilities may be used to weight the contribution of non-mutant sequence reads, and in step S160, the inference for the most likely base may be updated, for example, for a second site or subset of sites. In some embodiments, the process is iterated until the resulting change in the probability distribution across the most likely sequence or base falls below a convergence threshold.

[0080] The advantage of using the optimization process described in step S130 is that the optimization process proceeds iteratively toward a specific copy of the repeating region of the nucleic acid template in the nucleic acid template sample, rather than determining an average sequence across multiple copies. By proceeding toward a specific copy of the repeating region, the method allows for the determination of the most likely sequence of a specific copy of the repeating region in the nucleic acid template, rather than generating an "average" of the repeating region across different copies of the repeating region in the nucleic acid sample. Thus, in some embodiments, the method and system described herein include determining the sequence of a copy of the repeating region (e.g., one copy of two or more non-identical copies). In some embodiments, if two or more copies of the repeating region in the nucleic acid sample contain two or more haplotypes, the method includes determining the sequence of a haplotype in the nucleic acid sample (e.g., one haplotype of two or more non-identical haplotypes).

[0081] As an exemplary example, in some embodiments, the optimization process in step S130 includes applying an expectation maximization method. Expectation maximization (EM) is a statistical inference method commonly used when there are dependencies between parameters in a probabilistic model that prevent the direct use of analytical techniques to calculate the expected value, mean, variance, etc. As a non-restrictive example, under the EM iterative approach, in step S130, the alignment of mutant sequence reads to non-mutant sequence reads is taken as input. The sites are evaluated site by site to find sites with high confidence regarding the most likely base for the site. In step S160, the inference of the most likely base can be updated by calculating a probability distribution over the bases (e.g., over A, C, G, T) for the site. Then, in step S140, the probability of alignment of the non-mutant sequence read can be re-estimated using the probability distribution over the bases (or the set of such distributions among all aligned sites). The updated alignment can be re-evaluated for reliable regions. Then, the most likely base for the site can be determined using sites with sufficiently high confidence. This process can be repeated until a convergence threshold is reached.

[0082] In some embodiments, in step S130, the optimization process incorporates information describing variations in read length. Problems can arise in real (unsimulated) datasets, where sequence reads may be of unequal length due to the properties of sequencing chemistry and preliminary data analysis steps (e.g., read trimming), and information about read length may generally be taken into account by the optimization process. Therefore, information about such variations in sequence read length can be incorporated into the optimization process. For example, in some embodiments, the EM method uses a pseudo-likelihood defined on a fixed data dimension (read length).

[0083] Incorporate summaries and information from non-mutant sequence reads. In some embodiments, the process of calculating the most likely base inference includes summarizing the probabilities of each base at sites associated with the non-mutant sequence read. For example, in some embodiments, step S160, which updates the most likely base inference, includes step S161, which summarizes the probabilities of each base at each site associated with the non-mutant sequence read, site by site. For example, in some embodiments, for each site associated with the non-mutant sequence read, the probabilities of nucleotide A, C, G, or T being at a particular position in the non-mutant sequence read are summarized. In some embodiments, the summarization may be expressed as formula (3) in Example 1.

[0084] In some embodiments, the process of calculating the most likely base prediction includes steps of incorporating information related to the probability of mutation at a site, information related to the probability of errors (e.g., unintentionally introduced mutations or sequencing errors), and / or information related to mutation rates and variations in local sequence composition. For example, following the site-by-site analysis in step S161, the process may include step S162 of incorporating information related to the probability of mutation at each site in the mutant sequence read. For example, in the case of intentionally introduced mutations, the probability of mutation at each site in the mutant sequence read can be statistically estimated based on knowledge of the mutation rate of the mutagenesis process used in the sample preparation formulation.

[0085] In some embodiments, step S162 includes incorporating information describing the probability of errors related to the sample preparation or sequencing process. For example, in some embodiments, step S162 includes the method considering changes other than intentionally introduced mutations (e.g., unintentionally introduced mutations). In some embodiments, step S162 includes the method considering sequence changes, such as those introduced by the use of PCR during sample preparation and sequencing, or other causes of sequencing errors. In some embodiments, in step S162, such other changes may be considered in combination with the probability of intentionally introduced mutations at certain sites.

[0086] Depending on the mutagenesis process, sample preparation, library preparation, and sequencing method used, the probability of intentionally introduced mutations may differ from the probability of unintended changes to the sequence in the mutant sequence read. For example, in step S162, the method can account for at least four different possible scenarios occurring at each site in the mutant sequence read: 1) intentional mutations (e.g., transition mutations) and unintended mutations (e.g., polymerase errors); 2) the absence of unintended and intentional mutations; 3) both intentional and unintended mutations; and 4) a case where there are neither intentional nor unintended mutations.

[0087] In some embodiments, this process involves determining the most likely base for a site in a mutant sequence read that does not align to a non-mutant sequence read, based on a statistically estimated probability of mutation at that site. For example, in some embodiments, some sites in a mutant sequence read may not align to any non-mutant sequence read (may not be mapped). In some embodiments, if no information about the site is available from the non-mutant sequence read, the most likely base for the site may be determined based only on the mutant sequence read covering the site, in combination with a statistically estimated probability that mutation occurred at the site. In some embodiments, a base-by-base precision score for a site is also based on the same process.

[0088] In some embodiments, step S162 may include incorporating information describing the variation in mutation rates between mutant sequence reads, or the unequal mutation rates of adenine, cytosine, guanosine, and thymine. In some embodiments, step S162 may include incorporating information describing the variation in mutation rates between mutant sequence reads when statistically predicting the most likely sequence. For example, a mutagenesis process may have an average mutation rate of, for example, 5%, but some mutant sequence reads may randomly have more or less mutations than average, such as 3% or 8%. Similarly, step S162 may include incorporating information describing the unequal mutation rates of adenine, cytosine, guanosine, and thymine when statistically predicting the most likely sequence. For example, a mutagenesis process may generate a variety of mutation rates between A, C, G, and T nucleotide sites, and this information can be incorporated into the method described herein.

[0089] Determining the most likely sequence and accuracy score for each base. In some embodiments, after the optimization process in step S130, the method further includes step S170, which determines i) the most likely sequence of the nucleic acid template, and ii) for each base of the most likely sequence, a base-by-base precision score that correlates with the probability that the base matches the nucleic acid template.

[0090] In some embodiments, the base-by-base precision score includes a base-by-base quality value (QV), also known as the "PHRED" score. In some embodiments, the base-by-base precision score estimates the statistical confidence in the prediction for each base. In some embodiments, the method further includes determining an overall precision score associated with the most likely sequence. In some embodiments, the overall precision score is determined based on the base-by-base precision score of each base in the most likely sequence. For example, the overall precision score may be the mean or median of the base-by-base precision scores, or another composite value of the base-by-base precision scores.

[0091] In some embodiments, step S170, which determines the most likely sequence of the nucleic acid template, includes determining the most likely base for each site related to the alignment of the non-mutant sequence read to the mutant sequence read, as described herein. In some embodiments, the method includes determining the most likely sequence progression for each site. For example, the method may include statistically inferring the most likely base for a site (location in the nucleic acid sequence) by analyzing the alignment of the mutant sequence read to the non-mutant sequence read at that site, and then moving on to another site. For example, in some embodiments, the method includes determining the maximum posterior probability base for each site, and the most likely sequence includes one or more maximum posterior probability bases. In some embodiments, determining the most likely base for each site includes calculating an inference of the most likely base for the site, such as by calculating a probability distribution over bases (e.g., A, C, G, T) at the site. In some embodiments, the most likely sequence includes one or more most likely bases determined in step S130, for example, by applying an optimization process.

[0092] In some embodiments, step S170, which determines the most likely sequence of the nucleic acid template, includes determining the most likely sequence of a subset of sites related to the alignment of the non-mutant sequence read to the mutant sequence read. Thus, in some embodiments, the method includes determining that the most likely sequence proceeds subset by subset. As used herein, a subset of sites refers to a group of sites contained within the most likely sequence. In some embodiments, the subset of sites is smaller than the total number of sites in the most likely sequence.

[0093] In some embodiments, the most likely sequence and base-by-base precision scores are stored in a file on a computer storage medium (e.g., a computer hard drive, e.g., a rotating magnetic disk drive or a solid-state drive). In some embodiments, the file is stored in the format of a BAM, SAM, CRAM, or VCF file.

[0094] Selection of a subset of non-mutant sequence reads In some embodiments, the systems and methods of the present invention include a method for selecting a subset of non-mutant sequence reads from non-mutant sequence reads derived from different copies of a repeating region. In some embodiments, selecting a subset of non-mutant sequence reads includes selecting a subset of non-mutant sequence reads that are aligned to one or more mutant sequence reads at a site in order to contribute to the statistical inference of the most likely base of the nucleic acid template at that site. In some embodiments, other unselected non-mutant sequence reads aligned to mutant sequence reads at a site are excluded from or contribute little to the statistical inference of the most likely base of the nucleic acid template at that site. In some embodiments, the process of selecting a subset of non-mutant sequencing reads is repeated at a second site after the most likely base for the site has been statistically inferred as described herein. In some embodiments, the subset of non-mutant sequences is selected from non-mutant sequence reads derived from different copies of a repeating region in a nucleic acid sample.

[0095] In some embodiments, selecting a subset of non-mutant sequence reads includes scoring the alignment of non-mutant sequence reads to mutant sequence reads and selecting the top-scoring alignment. In some embodiments, the scoring scheme is specific to the alignment of non-mutant sequence reads to mutant sequence reads at a particular site (e.g., a site where the most likely base is predicted). For example, in some embodiments, the alignment of non-mutant sequence reads to mutant sequence reads is scored using an alignment scoring scheme, and the best "n" alignments are selected and scored under the alignment scoring scheme.

[0096] In embodiments of the present invention, alignments can be scored using various scoring schemes. As a non-limiting example, in some embodiments, alignments are scored using a standard scoring scheme. In some embodiments, the scoring scheme may be a 4x4 permutation matrix and an affine gap penalty score. In some embodiments, the scoring scheme is a probability generated by integrating the alignments using a method such as a paired hidden Markov model (paired HMM). In some embodiments, the alignment of non-mutant sequence reads to mutant sequence reads is scored using another method known in the art.

[0097] As a non-limiting example, after scoring the alignments, an iterative process is applied to select the top-scoring alignments for each site. For example, for a particular site, the top-scoring "n" alignments for that particular site are selected, where "n" refers to a determined number of alignments. For example, "n" may be 20, for example, so that the top 20 alignments are selected. "n" may be a range consisting of 2, 5, 10, 15, 20, 25, 30, 40, 50, or any of the aforementioned values. In some embodiments, "n" is related to the coverage depth of the non-mutant sequence reads used for the alignment. For example, in some embodiments, if "n" corresponds to a 50% coverage depth and the coverage is 40x, then "n" is 20.

[0098] In some embodiments, after the top-score alignment is selected for a site, the most likely base (e.g., the highest posterior probability base) is determined as described herein. This most likely base may be incorporated into the most likely sequence. In some embodiments, this process is repeated for further sites using the current estimate of the most likely sequence. For example, this process may be repeated until the number of bases changing as a result of the iteration falls below a certain percentage. A base-by-base precision score may be determined for each base and may be further described above. In some embodiments, the most likely sequence and the base-by-base precision score are determined in this manner.

[0099] system In other embodiments, systems are disclosed herein. For example, a system may comprise one or more processors configured to perform any of the methods described herein, which are computer implementable. In some embodiments, a system is disclosed herein comprising a processor configured to perform a method comprising: a) aligning a non-mutant sequence read of a nucleic acid template to a mutant sequence read such that the mutant sequence read contains one or more mutations; and b) i) determining the most likely sequence of the nucleic acid template, and ii) for each base of the most likely sequence, a base-by-base precision score that correlates with the probability that the base matches the nucleic acid template. thereby, instructions for performing a method comprising determining the sequence of the nucleic acid template are disclosed herein.

[0100] Figure 2 is a block diagram of an exemplary computing device 200 that may be used in connection with an exemplary sequencing system. The computing device 200 may be configured to sequence a nucleic acid template by removing mutations found in a mutant sequence read. The general architecture of the computing device 200 shown in Figure 2 includes the configuration of computer hardware and software components. The computing device 200 may include more (or fewer) elements than those shown in Figure 2. However, not all of these general conventional elements need to be shown to provide a valid disclosure.

[0101] As illustrated, the computing device 200 includes a processing unit 210, a network interface 220, a computer-readable media drive 230, an input / output device interface 240, a display 250, and an input device 260, all of which can communicate with each other via a communication bus. The network interface 220 may provide connectivity to one or more networks or computing systems. The processing unit 210 may therefore receive information and commands from other computing systems or services via the network. The processing unit 210 may also communicate with memory 270 and further provide output information from an optional display 250 via the input / output device interface 240. The input / output device interface 240 may also accept input from an optional input device 260, such as a keyboard, mouse, digital pen, microphone, touchscreen, gesture recognition system, speech recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0102] Memory 270 may include computer program instructions (grouped as modules or components in some embodiments) that the processing unit 210 executes to implement one or more embodiments. Memory 270 generally includes RAM, ROM, and / or other persistent, auxiliary, or non-transient computer-readable media. Memory 270 may store an operating system 272 that provides computer program instructions for use by the processing unit 210 in the overall management and operation of the computing device 200. Memory 270 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0103] For example, in one embodiment, memory 270 includes a rendering process module 274 for determining the sequence of a nucleic acid template by removing mutations found in mutant sequence reads. The rendering process module 274 can perform methods disclosed herein, including the process described with respect to Figure 1. Furthermore, memory 270 may include or communicate with a data store 290 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of determining the sequence of a nucleic acid template by removing mutations found in mutant sequence reads in accordance with this disclosure. Such data may include, for example, non-mutant sequence reads, mutant sequence reads, mutation probability information, error probability information, variable mutation rate information, and the most likely sequence and base-by-base precision scores.

[0104] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis functions and sequence data storage devices to a cloud computing environment or cloud-based network. User interaction with sequencing data, genomic data, or other types of biological data may be mediated through a central hub that stores the data and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data, and distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates user modification or annotation of sequence data. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand, or online.

[0105] In some embodiments, software written to perform the methods described herein is stored on several forms of computer-readable media, such as memory, CD-ROMs, DVD-ROMs, memory sticks, flash drives, hard drives, SSD hard drives, servers, and mainframe storage systems.

[0106] In some embodiments, the method may be written in any of several suitable programming languages, such as compiled languages ​​like C, C#, C++, Fortran, and Java. Other programming languages ​​may include scripting languages ​​such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R, and PHP. In some embodiments, the method is written in C, C#, C++, Fortran, Java, Perl, R, or Python. In some embodiments, the method may be a standalone application having data input and data display modules. Alternatively, the method may be a computer software product in which distributed objects may include classes containing applications that include the computation method described herein.

[0107] In some embodiments, the method may be incorporated into existing data analysis software, such as that found in sequencing instruments. Software including the computer implementation method described herein may be installed directly on a computer system or indirectly stored on a computer-readable medium and loaded onto a computer system as needed. Furthermore, the method may be located on a computer that is remote from where the data is generated, such as software found on servers maintained in a different location from where the data is generated, such as those provided by a third-party service provider.

[0108] Assay equipment, desktop computers, laptop computers, or servers may include a processor that makes operational communication with accessible memory, including instructions for implementing the system and method. In some embodiments, the desktop or laptop computer may operably communicate with one or more computer-readable storage media or devices and / or output devices. Assay equipment, desktop computers, and laptop computers may operate under many different computer-based operating languages, such as those used by Apple-based computer systems or PC-based computer systems. Assay equipment, desktop, and / or laptop computer and / or server systems may further provide computer interfaces for creating or modifying experimental definitions and / or conditions, viewing data results, and monitoring experimental progress. In some embodiments, output devices may be computer monitors or computer screens, printers, portable devices such as portable digital assistants (i.e., PDAs, Blackberries, iPhones), tablet computers (e.g., iPads), hard drives, servers, memory sticks, flash drives, or other graphic user interfaces.

[0109] The computer-readable storage device or medium may be any device such as a server, mainframe, supercomputer, or magnetic tape system. In some embodiments, the storage device may be located in close proximity to the assay equipment, for example, adjacent to or in close proximity to the assay equipment. For example, the storage device may be located in relation to the assay equipment in the same room, the same building, an adjacent building, on the same floor of a building, or on a different floor of a building. In some embodiments, the storage device may be located outside or distal to the assay equipment. For example, the storage device may be located in a different part of a city, in a different city, in a different state, or in a different country relative to the assay equipment. In embodiments where the storage device is located distal to the assay equipment, communication between the assay equipment and one or more of the desktop, laptop, or server is typically via an internet connection, either wirelessly or via network cable through an access point. In some embodiments, the storage device may be maintained and managed by an individual or entity directly associated with the assay equipment, while in other embodiments, the storage device may be maintained and managed by a third party, typically distal to the individual or entity associated with the assay equipment. In the embodiments described herein, the output device may be any device for visualizing data.

[0110] Assay instruments, desktops, laptops, and / or server systems may be used to store and / or retrieve computer implementation software programs incorporating computer code for performing and implementing the calculation methods described herein, data for use in the implementation of the calculation methods, etc. One or more of the assay instruments, desktops, laptops, and / or servers may include one or more computer-readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing the calculation methods described herein, data for use in the implementation of the calculation methods, etc. Computer-readable storage media may include, but are not limited to, one or more of hard drives, SSD hard drives, CD-ROM drives, DVD-ROM drives, floppy disks, tapes, flash memory sticks, or cards. Furthermore, networks, including the Internet, may be computer-readable storage media. In some embodiments, computer-readable storage media refers to computing resource storage accessible by a computer network over the Internet or a company network provided by a service provider, rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.

[0111] In some embodiments, a computer-readable storage medium for storing and / or retrieving computer implementation software programs incorporating computer code for performing and implementing the calculation methods described herein, data used for implementing the calculation methods, etc., is operated and maintained by a service provider that operably communicates with assay equipment, desktops, laptops, and / or server systems via an internet connection or network connection.

[0112] In some embodiments, the hardware platform for providing a computing environment includes a processor (i.e., a CPU) where processor time and memory layout, such as random access memory (i.e., RAM), are system considerations. For example, smaller computer systems provide inexpensive, high-speed processors and large memory and storage capacity. In some embodiments, a graphics processing unit (GPU) can be used. In some embodiments, the hardware platform for performing the computing methods described herein includes one or more computer systems having one or more processors. In some embodiments, smaller computers are clustered together to create a supercomputer network.

[0113] In some embodiments, the computation methods described herein are performed on a collection of inter- or intra-inter-connected computer systems (i.e., grid technologies) capable of coordinating the execution of various operating systems. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available from United Devices are an example of the coordination of multiple independent computer systems for handling large amounts of data. These systems may provide a Perl interface for submitting, monitoring, and managing large sequence analysis jobs on clusters in a sequential or parallel configuration. [Examples]

[0114] Some aspects of the embodiments described above are disclosed in more detail in the following examples, but these are not intended to limit the scope of this disclosure. Those skilled in the art will understand that many other embodiments, as described above and in the claims of this specification, also fall within the scope of this disclosure. For example, the notation introduced in some of the following examples assumes the alignment of non-mutant and mutant sequence reads to a reference genome assembly. However, the methods described below can also be used in the direct alignment of non-mutant sequence reads to mutant sequence reads without using a reference sequence as an intermediate.

[0115] In the following examples, the following notation is applicable: “Dream” or “Dream Sequence” refers to a variant sequence read containing mutations that may be removed by rendering.

[0116] A set of mapped "reference reads" covering site i in the reference sequence (non-mutant sequence reads from the reference sequence, such as standard non-mutant shotgun reads) is R i This can be expressed as follows, where r∈R i Let be a read from a set of reads aligned to site i in the reference sequence. The nucleotide vector in read r aligned to site i in the reference sequence is given by r i It is expressed as r. i is a length 4 instruction vector corresponding to one of the four nucleotides (e.g., [1;0;0;0] for A and [0;1;0;0] for C). The probability that the base call at site i in read r is correct (e.g., quality score) is given by q(r). i ) is expressed as.

[0117] d i This is an indicator vector a of length 4. i Regarding this, it represents a vector of length 4 that describes the probabilities of A, C, G, and T at the position in the dream aligned to reference site i, where 1 at that position corresponds to the dream bases A, C, G, and T. As described above, q(a i) represents the probability that the base call at site i of the dream array matches the nucleic acid template. D i is written as follows:

[0118]

Number

[0119] Example 1: The nucleic acid sequence was determined by using a statistical method to remove mutations introduced into the mutant array reads. The details of the statistical method are described below.

[0120] Maximum a posteriori basecalling and Bayes factor quality scores

[0121]

Number

[0122]

Number

[0123] [[ID=4४]] The prior probability of a mutation at any site is denoted as P(u), where u refers to the mutation event. For example, in a typical application of Illumina® Complete Long Read (ICLR) sequencing, P(u) takes values in the range of 0.05 - 0.06.

[0124] z<ひ i shall represent a length-4 vector of the unnormalized posterior probabilities of A, C, G, T at site i. z " i The value of is defined as follows:

[0125]

number

[0126] In the formula, ° represents the Hadamard product (element-wise product) of vectors / matrices, and in the formula, y i is a normalized vector of length 4 describing the probabilities of A, C, G, and T at that position in the ref sequence, and the calculation from the ref read is described above. M is a 4x4 rotation matrix that rotates the elements of the length 4 vector to apply transition mutations:

[0127]

number

[0128] To render the base, the maximum posterior probability value is z i Select from:

[0129]

number

[0130] The quality score is set for part i, and QV(z i ) is expressed as follows, and is equal to:

[0131]

number

[0132] The quality score is calculated based on the base argmax(z i This is a Bayes factor that compares the hypothesis that it is not the case. The molecule has a base argmax(z i This is in response to the hypothesis that (denominator) is...

[0133] Example 2: Extension of MAP-based calling to polymerase error correction The base calling method described above has been extended to allow sequencing when considering changes other than transition mutations, such as changes introduced by the use of PCR during sample preparation and sequencing, or other causes of sequencing errors. Details of this method are described below.

[0134] To account for non-transition polymerase errors,

[0135]

number

[0136] Furthermore, PI is used as the prior probability of polymerase error, for example, 5 × 10 per PCR cycle. -5 Alternatively, in a typical ICLR library, it's 1.6 × 10⁶. -3 The order is expected to be (32 cycles in total). Then, the posterior probability across the bases is redefined as follows:

[0137]

number

[0138] Intuitively, the above formula integrates the probabilities across four scenarios reflected in the four terms of the formula. The first term corresponds to no mutation event and no polymerase error, the second term corresponds to no polymerase error and no mutation, the third term corresponds to both mutation and polymerase error, and the fourth term corresponds to the case where neither mutation nor error occurs. The estimation of base calling and QV is given by z i Except for using the revised definition of , the process proceeds in the same manner as described above in formulas (6) and (7). That is, the output base is argmax(z i Position i is set for the nucleotide corresponding to ), and the quality score QV(z i This is calculated as a Bayes factor that compares the MAP base to three other candidate bases.

[0139] Example 3: Processing of unmapped dream locations If the region in the mutant sequence read was not mapped to the non-mutant sequence read, the sequence was determined using statistical methods. Details of the statistical methods are described below.

[0140] As described above, z i Let P(u) represent a vector of length 4 representing the posterior probability of a base at site i. We define φ as a vector of length 4 representing the probability that any A, C, G, or T base is mutated in any dream. φ extends the concept of P(u) to represent the unequal probability of mutations at sites containing A, C, G, and T. This quantity is reliably calculated by simply counting the number of transition differences at all A, C, G, and T sites in the alignment of all dream sets relative to a reference read, and normalizing this by the total number of A, C, G, and T sites. Formally, φ is defined as follows:

[0141]

number

[0142]

number

[0143] If the Dream base is not aligned,

[0144]

number

[0145] Example 4: Adjustment of local sequence composition and unequal mutation rate When considering local sequence composition and unequal mutation rates, sequences were determined using statistical methods. Details of the statistical methods are described below.

[0146] The variations in local sequence composition and unequal mutation rates at the four bases A, C, G, and T are taken into account by using φ as defined in the section above (see equation (10)). To do so, z i The definition can be further extended as follows:

[0147]

number

[0148] Example 5: Maximizing Expected Value To infer the rendered array, we applied a technique similar to EM. Details of the technique are described below.

[0149] Use the following to determine r for dream

[0150]

number

[0151]

number

[0152]

number

[0153]

number

[0154] The remaining notation is as defined above. The value of φ0 is initialized as described above in equation (10). d i,,0 The value of is d in equation (1). i It is initialized as described above.

[0155] Next, the parameter updates for each iteration of EM are defined as follows: First, x and y are redefined using the likelihood per read and indexed by the EM iterations.

[0156]

number

[0157] The c vector is obtained by summing the likelihoods of the reads used in each part and normalizing the values ​​so that the sum is 1. i,ι It is used in the definition of [the function]. As mentioned above, α is the regularization constant.

[0158] Next, we use the previously defined Z(·) to identify the maximum likelihood base (the most likely base) at each site. This can be expressed as follows:

[0159]

number

[0160] Also,

[0161]

number

[0162]

number

[0163]

number

[0164] Therefore, φ ι is a vector of length 4 (

[0165]

number

[0166]

number

[0167]

number

[0168] Example 6: Practical Considerations Regarding Actual Datasets To account for reads of unequal length, we adjusted the above expectation maximization method using statistical methods, assuming reads of equal length so that the likelihood is calculated based on data of equal dimensions. Details of the statistical methods are described below.

[0169] L represents the maximum read length to be processed. In an exemplary application, L=150. Alternatively, if paired-end 150nt (150 nucleotide) reads are processed as a single entity, L=300 is taken. For a particular read, L represents the number of sites in the read r aligned to the dream.

[0170]

number

[0171]

number

[0172] That is, q r This is the mean likelihood for each aligned region in the lead. r Using this, we can define a fixed-dimensional pseudolikelihood for read r,

[0173]

number

[0174]

number

[0175] In other words, the mean likelihood for each part from the aligned parts is taken as an estimate of how the likelihood for each part will be in any missing part, and the data dimension in the likelihood function is embedded in L parts.

[0176] Finally, the function Δ(·) is defined as counting the number of nucleotide differences between two sequences, as follows:

[0177]

number

[0178] The exemplary method can be expressed as follows with respect to the defined notation and functions:

[0179] [Table 1]

[0180] Example 7: Mutations in a Simple Simulated Dataset The statistical methods described herein were used to render mutations in a simulated dataset without sequencing errors.

[0181] For example, Figure 3 shows an example of rendering site 340 containing a mutation. In alignment 310, the mutant sequence read 320 is compared with several aligned non-mutant sequence reads 330, and the discrepancy is highlighted at site 340. Here, the mutant sequence read has "T" at site 340, while the non-mutant sequence read has "C" at site 340. In the example in Figure 3, the simplest version of the mathematical model of the mutation and sequencing process is used. In this example, the non-mutant sequence read 330 matches the mutant sequence read 320 in supporting the alternative nucleotide, and the difference is explained by a transition mutation. The maximum posterior probability base generated by the model matched the base present in the non-mutant sequence read 330, so the mutation was corrected by the rendering process. In the example in Figure 3, a simple model outlined in equations (6) and (7) was used. The alignment of the non-mutant sequence read 330 to the mutant sequence read 320 is shown in alignment 310. α i , q(a i ), r i , q(r i ), α i The values ​​of , and P(u) are specified as constants in Table 360. i , x i , y i , z i , and QV(z i The use of constant values ​​in the formula for calculating z is shown in Table 370. i The MAP base selected from is C and (0.0506 z) i Corresponding to the maximum value in , the obtained PHRED-scaled QV score is 17.2, corresponding to an estimated probability of approximately 0.98 (the rendered base call is correct). In a typical dataset with higher coverage of non-mutant data (e.g., a larger number of aligned non-mutant sequence reads of 330), the QV score would be higher.

[0182] Example 8: SNPs in a Simple Simulated Dataset The statistical methods described herein were used to render the mutations in the simulated dataset without sequencing errors. The dataset in this embodiment contained at least two haplotypes that were heterozygous for single nucleotide polymorphisms (SNPs).

[0183] Figure 4 shows an example of rendering a site indicating a heterozygous SNP in non-mutant data, where the two involved bases represent the transition difference. In alignment 410, the mutant sequence read 420 matches one of the two bases present at site 440 in the non-mutant sequence read 430, leaving a high degree of uncertainty as to whether site 440 is mutated or not. In Table 470, the MAP estimate generated by the model leaves the bases unchanged and assigns a low QV corresponding to a probability of approximately 0.92 that the bases were correctly called. The low confidence in the base call reflects the prior knowledge that approximately 5% of the sites in the sequence are mutated, due to the nature of sequencing chemistry. In the example in Figure 4, a simple model outlined in equations (6) and (7) was used. Alignment of the non-mutant sequence read to the mutant sequence read is shown in alignment 410. α i , q(a i ), r i , q(r i ), α i The values ​​of , and P(u) are specified as constants in Table 460. i , x i , y i , z i , and QV(z i The use of constant values ​​in the formula for calculating z is shown in Table 470. i The MAP base selected from is G (0.3779 z i The resulting PHRED-scaled QV score (corresponding to the maximum value in ) is 10.9, which corresponds to an estimated probability of approximately 0.92 (the rendered base call is correct).

[0184] Example 9: Accuracy evaluation in a simulated dataset The accuracy of the rendering method was evaluated against a simulated dataset containing mutant sequence read templates, a "reference genome" in which the mutant templates were simulated, and non-mutant "reference reads" (non-mutant sequence reads) similarly simulated from the reference genome. The advantage of evaluating against a simulated dataset was that the accurate rendered sequences (corresponding to the nucleic acid template sequences) were known, and therefore the accuracy of the rendering could be clearly evaluated. The details of the experiment are described below.

[0185] Generation of simulation data The inputs to the simulation process included a reference genome assembly, targeted coverage for ref reads, targeted coverage for long reads, and short read coverage (redundancy) per long read template.

[0186] The simulation process produced several outputs, including a true set of long template sequences (representing nucleic acid template sequences) and their locations in the reference genome assembly, a set of long template sequences with simulated mutations (representing mutant nucleic acid templates), the locations of mutations introduced into the long templates, a set of standard short reads (representing non-mutant sequence reads), a set of short reads simulated from the mutated long template sequences (representing mutant sequence reads), and information on which long template each mutant short read originated from.

[0187] Three datasets were generated using a simulation system. These were referred to as "small," "medium," and "large," and had the characteristics listed in Table 1 below. For the human simulation, a diploid assembly of HG002, which is not available from NCBI, was used. Small contigs were removed from the assembly before data simulation. During analysis, the human dataset used actual ref read sequences rather than simulated reads.

[0188] [Table 2]

[0189] All datasets were simulated using long read lengths derived from a normal distribution with a mean of 6 kbp and a mutation rate of 6%. Mutations were located at randomly and uniformly derived positions throughout each long sequence and consisted strictly of transition mutations (A⇔G, C⇔T).

[0190] Rendering implementation We evaluated four alternative implementations of the rendering method: the original, minihyb, EM, and EM w / cons. These four methods represent implementations of various combinations of the methods presented above. Each method is described below.

[0191] original In the "original" method, non-mutant and mutant sequence reads are mapped to the reference sequence. The first pass of rendering then proceeds using the formula defined in Example 4: Adjustment for local sequence composition and unequal mutation rate above.

[0192] In subsequent passes of rendering, k-mer comparisons are used to further remove any residual errors in the data. Briefly, the k-mers present in each rendered read ("rendered k-mers") are compared to the set of k-mers present in the non-mutant sequence reads ("non-mutant k-mers"). If a rendered read k-mer is adjacent to a rendered k-mer that is not found in the set of non-mutant k-mers but can be found in the set of non-mutant k-mers, an attempt is made to modify the unfound k-mer so that it can be found in the set of non-mutant k-mers. Nucleotides in the first k-mer that are not covered by the second k-mer are corrected by applying a transition mutation, and then, if the resulting k-mer can be found in the set of non-mutant k-mers, the modification is retained; otherwise, it is reverted. This process is repeated for all adjacent k-mers in each rendered read until no more changes are found that increase the number of rendered k-mers found in the set of non-mutant k-mers. The k-mer comparison process is similar to the aforementioned methods for correcting short-read errors, for example, the method described in Greenfield et al., Blue: correcting sequencing errors using consensus and context, Bioinformatics, Volume 30, Issue 19, October 2014, Pages 2723-2732, doi.org / 10.1093 / bioinformatics / btu368.

[0193] minihyb In the "minihyb" method, non-mutant and mutant sequence reads are mapped to a reference sequence. Non-mutant sequence reads with overlapping mapping positions with mutant sequence reads cause gapless realignment of the non-mutant sequence reads to the direct mutant sequence reads. This direct non-mutant:mutant read alignment can produce better alignment in some situations, particularly when the sample DNA sequence branches off from the reference genome sequence. The non-mutant:mutant read alignment is then subjected to soft clipping using the well-known X-drop method. Specifically, if the alignment score drops beyond X from its peak, the alignment following the peak score position is removed.

[0194] The next step in minihyb is to select a set of high-scoring alignments for use in rendering calculations using an ad-hoc method. The strategy applied in this step is to rank the alignments of non-mutant sequence reads by their alignment scores and select the top N alignments in a non-overlapping window of the dream sequence. If the alignment scores are formulated using a permutation matrix and gap penalty approach, the scores may have a probabilistic interpretation as the likelihood that non-mutant and mutant sequence reads were generated by the same original sequence. Selecting the top N alignments effectively imposes a threshold on probability, which is dynamic in that the threshold can be scaled up or down based on the density of mutations in local regions of the dream sequence.

[0195] Next, the selected set of alignments is used for rendering using the formula defined in Example 4: Adjustment of local sequence composition and unequal mutation rate.

[0196] EM In the "EM" method, non-mutant and mutant sequence reads are mapped to a reference sequence. Similar to the minihyb method, non-mutant sequence reads with overlapping mapping positions with mutant sequence reads cause gapless realignment of the non-mutant sequence reads to the direct mutant sequence reads. Further alignments are generated by directly aligning non-mutant sequence reads to mutant sequence reads using a short-read aligner that can multi-map non-mutant sequence reads onto multiple mutant sequence reads. When rendering a particular mutant sequence read, all sets of non-mutant sequence read alignments for that mutant sequence read are collected, and then EM method 1 is applied to perform an expected value maximization rendering.

[0197] EM w / cons The "EM w / cons" method applies expectation maximization with an additional step of generating a consensus sequence for use in non-mutant:mutant sequence read alignment. The consensus sequence is generated as described above in the section titled "Alignment to Reference Sequence and / or Consensus Sequence". Once the set of alignments is collected, the consensus of mutant sequence reads is discarded, and rendering proceeds as described in EM method 1.

[0198] Evaluation of the accuracy of simulation data The simulated dataset was analyzed by providing mutant and non-mutant sequence reads to an analytical statistical method. This approach has the advantage of bypassing the step of assembling mutant sequence reads, which is typically performed in the practical application of this technique, and eliminating errors in the assembled mutant sequence reads from the accuracy of the rendering process, as determined by the inventors' analysis.

[0199] The accuracy of the simulated dataset was evaluated by comparing the rendered long sequence (representing the most likely sequence of the nucleic acid template) with the true non-mutated long sequence (representing the true nucleic acid template sequence) generated by the simulation. The comparison was performed site by site, examining whether each site in the rendered sequence matched the corresponding site in the true sequence.

[0200] If a site in the rendered sequence is identical to the true sequence, the site is classified as either true positive or true negative. A true positive (TP) is a site that has experienced a mutation, which has been accurately identified by the renderer, and the site has been modified to match the original true sequence. A true negative (TN) is a site that has not experienced a simulated mutation and remains unchanged by the rendering process.

[0201] If a site in the rendered sequence differs from its corresponding site in the true sequence, it is classified as either a false positive or a false negative. A false positive (FP) is a site that did not undergo a simulated mutation but was modified by the rendering process so that it no longer matches a site in the true sequence. A false negative (FN) is a site that underwent a simulated mutation but the mutation was not accurately identified and resolved by rendering.

[0202] Using the above definitions of TP, TN, FP, and FN, it becomes possible to calculate various measures of rendering performance, such as accuracy and recall. The accuracy score per base can be further stratified by QV; for example, FP or FN predicted to have a very high or very low QV may be of interest. The results of the accuracy evaluation are shown below.

[0203] Results regarding the simulation data As shown in Table 1, three different simulated datasets were generated. Each dataset was analyzed using the four different rendering methods described above: original, minihyb, EM, and EM w / cons. However, EM w / cons results are not available for medium-sized datasets because the datasets only have 3x coverage of variant sequence reads, which is insufficient to generate consensus sequences for most of the datasets. The evaluation results for the simulated data are shown in Table 2 below.

[0204] [Table 3]

[0205] Several trends emerge from the simulated data. For all datasets, the method using expectation maximization produces significantly better results than simpler rendering methods. Secondly, the use of consensus sequences improves accuracy for small datasets but not for large datasets. This difference could be explained by the difference in how the datasets were handled. Small datasets used the same genome sequence for both the simulation of mutant sequence reads and the mapping of those mutant sequence reads during consensus generation, while large datasets used the HG002 sample assembly for simulation but the hg38 human genome assembly for mutant sequence read mapping and consensus generation. The difference between the hg38 and HG002 assemblies may have resulted in insufficient consensus generation, which in turn may have reduced the accuracy of the results.

[0206] Finally, accuracy for small and medium-sized datasets is slightly higher than accuracy for large datasets. Several factors may contribute to this observation. First, the large datasets used actual sequence data rather than simulated data for the ref reads. Second, as mentioned above, the large datasets were processed using hg38 as the reference rather than the HG002 genome, which was simulated for the reads. Rendering methods that heavily rely on reference genome mapping (original, minihyb) appear to be particularly prone to errors when there is a significant difference between the reference genome and the simulated genome, resulting in relatively lower performance for these methods.

[0207] Example 10: Accuracy evaluation in an actual dataset Next, the rendering method selection described above was applied to an actual human genome dataset. The generation of the actual human genome dataset was as described above. The application of the rendering method was the same as for the simulated data, with simulated non-mutant sequence reads being replaced by standard whole genome sequencing (WGS) non-mutant sequence reads, and simulated mutant sequence reads being replaced by actual mutant sequence reads.

[0208] The advantage of evaluating against actual datasets is that, if reasonably accurate predictions of the original nucleic acid template sequences are possible, actual datasets capture the effects of sample preparation and sequencing biochemistry that are difficult to simulate.

[0209] A total of 22,945,744 variant sequence reads, including 102.2 Gbp of sequence data, were rendered. The N50 variant sequence read in the dataset is 6.2 kbp. In the unrendered variant sequence read data, approximately 5.4% of sites show substitution differences compared to the hg38 reference genome sequence. Some of these differences are true differences between the HG002 sample and the hg38 reference sequence, but the majority are introduced mutations. In the rendered data, only 0.3% of sites show substitution differences compared to the hg38 sample. Therefore, the rendering process significantly reduced the amount of difference relative to the reference.

[0210] Figure 5 shows Display 500 of human genome data before and after the rendering process. In Figure 5, three panels are shown on Display 500. In panels 510, 520, and 530, each sequence read is shown as a gray horizontal strip, and regions that differ from the hg38 reference sequence are marked with vertical bars. Panel 510 shows mutant sequence reads from the HG002 sample before rendering. Numerous differences from the hg38 reference sequence can be seen in the mutant sequence reads in panel 510. Panel 520 shows the same set of mutant sequence reads compared to the reference sequence after rendering using the minihyb method. As shown in panel 520, after rendering, most of the residuals between the rendered sequence and the reference sequence correspond to the true variant in this sample. This true variant is shown in panel 530, which shows standard WGS non-mutant sequence reads from a PCR-free library of the HG002 sample. Thus, before rendering, panel 510 shows that the mutant sequence reads contain numerous differences (mutations) compared to the reference sequence. Rendering removes most of those differences, and the residual matches the non-mutant sequence read (reference read) as a true single nucleotide variant (SNV).

[0211] Other considerations It should be understood that all combinations of the aforementioned concepts and further concepts, which are discussed in more detail below, are intended to be part of the subject matter of the inventions disclosed herein (provided that such concepts do not contradict each other). Specifically, all combinations of claimed subject matter appearing at the end of this disclosure are intended to be part of the subject matter of the inventions disclosed herein. It should also be understood that terms used expressly herein and that may appear in any disclosure incorporated by reference should be given meanings that most coincide with the specific concepts disclosed herein.

[0212] Throughout this specification, references to “one example,” “another example,” and “an example” mean that certain elements (e.g., features, structures, and / or characteristics) described in relation to an example are included in at least one example described herein, and may or may not be present in other examples. In addition, unless explicitly indicated otherwise in the context, it should be understood that elements described in any example may be combined in any preferred manner in various examples.

[0213] It should be understood that the ranges provided herein include the specified ranges and any values ​​or subranges within those specified ranges, as if such values ​​or subranges were explicitly enumerated. For example, the range of approximately 2kbp to approximately 20kbp should be interpreted to include not only the explicitly enumerated limits of approximately 2kbp to approximately 20kbp, but also individual values, e.g., approximately 3.5kbp, approximately 8kbp, 18.2kbp, etc., and subranges, e.g., approximately 5kbp to approximately 10kbp, etc. Furthermore, where "approximately" and / or "substantially" are used to express a value, this means that a small variation (up to ±10%) from the specified value is included.

[0214] While several embodiments have been described in detail, it should be understood that the disclosed examples may be modified. Therefore, the above explanation should be considered non-limiting.

[0215] Although certain embodiments are described, these embodiments are presented only as examples and are not intended to limit the scope of the disclosure. In fact, the novel methods described herein may be embodied in various other forms. Further, various omissions, substitutions, and changes in the methods described herein may be made without departing from the spirit of the disclosure. The appended claims and their equivalents are intended to cover such forms or modifications as fall within the scope and spirit of the disclosure.

[0216] Features, materials, characteristics, or groups described in connection with a particular aspect or embodiment are to be understood as applicable to any other aspect or embodiment described in this section or elsewhere in this specification, unless incompatible therewith. All of this specification (including any appended claims, abstract, and drawings), and / or all steps of any method or process disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The protection is not limited to the details of any of the foregoing embodiments. The protection extends to any novel one or any novel combination of features disclosed in this specification (including any appended claims, abstract, and drawings), or to any novel one or any novel combination of steps of any method or process disclosed as such.

[0217] Certain features described in the context of separate implementations in the present disclosure may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately in multiple implementations or in any suitable partial combination. Further, even if a feature is described as functioning in a particular combination, one or more features from the claimed combination may, in some cases, be deleted from the combination, and the combination may be claimed as a partial combination or a variation of a partial combination.

[0218] Furthermore, the operations may be depicted in the drawings or described herein in a particular order, but such operations need not be performed in the particular order shown or sequentially to achieve the desired result, nor is it necessary that all operations be performed. Other operations not depicted or described can be incorporated into the exemplary methods and processes. For example, one or more additional operations can be performed before, after, concurrently with, or between any of the described operations. Further, the operations may be rearranged or reordered in other implementations. Those skilled in the art will understand that in some embodiments, the actual steps taken in the illustrated and / or disclosed processes may differ from those shown in the figures. Depending on the embodiment, certain steps described above may be removed or others may be added. Additionally, the features and attributes of the specific embodiments disclosed above may be combined in different ways to form additional embodiments, all of which fall within the scope of the present disclosure.

[0219] For purposes of the present disclosure, certain aspects, advantages, and novel features are described herein. Not all such advantages may necessarily be achieved in accordance with a particular embodiment. Thus, for example, those skilled in the art will recognize that the present disclosure may be embodied or carried out in a manner that achieves one advantage or a group of advantages taught herein, even if the present disclosure does not necessarily achieve other advantages that may be taught or suggested herein.

[0220] Unless otherwise specified, conditional language such as “can,” “could,” “might,” or “may,” unless specifically stated otherwise or understood within the context in which they are used, is generally intended to convey that a particular embodiment includes a particular feature, element, and / or step, while other embodiments do not. Therefore, such conditional language is generally not intended to imply that a feature, element, and / or step is required in some way in one or more embodiments, or that one or more embodiments necessarily include logic, with or without user input or prompting, that determines whether these features, elements, and / or steps should be included or performed in any particular embodiment.

[0221] Conjunctions such as “at least one of X, Y, and Z” are understood, unless otherwise specified, in the context in which they are commonly used to indicate that an item, term, etc., may be X, Y, or Z. Therefore, such conjunctions are generally not intended to imply that a particular embodiment requires the presence of at least one of X, at least one of Y, and at least one of Z.

[0222] The terms "approximately," "about," "generally," and "substantially" as used herein refer to values, quantities, or characteristics that are close to the stated values, quantities, or characteristics that still perform the desired function or achieve the desired result.

[0223] The scope of this disclosure is not intended to be limited by any specific disclosure of examples in this section or elsewhere of this Spec. Rather, it may be defined by the claims presented or hereafter presented in this section or elsewhere of this Spec. The language of the claims should be interpreted broadly based on the language used in the claims, and the examples should be interpreted non-exclusively, not limited to the examples described herein or in the proceedings of this application.

Claims

1. A computer implementation method for determining the sequence of a nucleic acid template by removing mutations found in a mutant sequence read, wherein the method is a) Aligning a non-mutant sequence read of a nucleic acid template to a mutant sequence read, wherein the mutant sequence read contains one or more mutations. b) i) The most likely sequence of the nucleic acid template, and ii) For each base of the most likely sequence, determine a precision score for each base that correlates with the probability that the base matches the nucleic acid template. A computer implementation method comprising determining the sequence of the nucleic acid template.

2. The method according to claim 1, wherein the mutant sequence read is a mutant nucleic acid template.

3. The method according to claim 1 or 2, wherein the one or more mutations include mutations introduced during sample preparation.

4. The method according to any one of claims 1 to 3, wherein the one or more mutations include mutations randomly introduced during sample preparation by mutagenesis.

5. The method according to any one of claims 1 to 4, wherein the one or more mutations include mutations that were introduced unintentionally.

6. The method according to any one of claims 1 to 5, wherein the nucleic acid template is longer than the mutant sequence read and the non-mutant sequence read.

7. i) The method according to any one of claims 1 to 6, comprising determining the most likely base for each site related to the alignment of the non-mutant sequence read with respect to the mutant sequence read.

8. i) The method according to any one of claims 1 to 7, comprising determining the most likely sequence of a subset of sites related to the alignment of the non-mutant sequence read with respect to the mutant sequence read.

9. The method according to any one of claims 1 to 8, wherein step b) includes summarizing the probability of each base at each site associated with the non-mutant sequence read, site by site.

10. The method according to any one of claims 1 to 9, wherein step b) includes incorporating information relating to the probability of mutation at each site of the mutant sequence read.

11. The method according to any one of claims 1 to 10, wherein step b) includes incorporating information describing the probability of errors related to the sample preparation or sequencing process.

12. The method according to any one of claims 1 to 11, wherein step b) determines the most likely base for the site of the mutant sequence read that is not aligned to the non-mutant sequence read, based on the statistically estimated probability of mutation at the site.

13. The method according to any one of claims 1 to 12, wherein step b) includes incorporating information describing the variation in mutation rates between mutant sequence reads or the unequal mutation rates of adenine, cytosine, guanosine, and thymine.

14. The method according to any one of claims 1 to 13, wherein the nucleic acid template is derived from a nucleic acid sample containing two or more copies of the repeating region of the nucleic acid template.

15. The method according to claim 14, wherein step b) includes estimating the probability that the non-mutant sequence read originated from a copy of the same repeat region as the mutant sequence read to which the non-mutant sequence read is aligned.

16. The method according to claim 14 or 15, wherein step b) includes applying an optimization process, the optimization process iteratively updates an estimate of the probability that a non-mutant sequence read originated from the same copy of the mutant sequence read to which the non-mutant sequence read is aligned, while updating the inference of the most likely base.

17. The method according to claim 16, wherein the optimization process includes applying an expected value maximization method.

18. The method according to claim 16 or 17, wherein the optimization process incorporates information describing variations in lead length.

19. The method according to any one of claims 14 to 17, wherein the contribution of each non-mutant sequence read to the probability calculation of the most likely base is weighted by a weighting scheme.

20. The method according to claim 19, wherein the weighting scheme relates to the sequence identity of the non-mutant sequence reads and the iterative estimation of the sequence of the nucleic acid template.

21. The method according to claim 14, wherein step b) includes selecting a subset of non-mutant sequence reads from among non-mutant sequence reads derived from different copies of the repeat region.

22. The method according to claim 21, wherein selecting a subset of non-mutant sequence reads includes scoring the alignment of the non-mutant sequence reads to the mutant sequence reads and selecting the top-scoring alignment.

23. The method according to any one of claims 14 to 22, wherein the method comprises determining the sequence of copies of the repeating region of the nucleic acid template.

24. The method according to any one of claims 14 to 23, wherein two or more copies of the repeat region comprise two or more haplotypes, and the method comprises determining the sequence of the haplotypes of the nucleic acid sample.

25. The method according to any one of claims 1 to 24, wherein the precision score for each base includes a quality value (QV) for each base.

26. The method according to any one of claims 1 to 25, further comprising determining an overall accuracy score associated with the most likely sequence.

27. The method according to any one of claims 1 to 26, wherein the method further comprises assembling the non-mutant sequence reads before a), and a) comprises aligning the mutant sequence reads to the assembly of the non-mutant sequence reads.

28. The method according to any one of claims 1 to 27, wherein the method further comprises assembling the mutant sequence reads before a), and a) comprises aligning the non-mutant sequence reads to the assembly of the mutant sequence reads.

29. The method according to any one of claims 1 to 28, further comprising aligning the non-mutant sequence read or the mutant sequence read to a reference sequence before a).

30. The method according to any one of claims 1 to 29, wherein the method further comprises, before a), assembling the mutant sequence reads, generating a consensus of the assembled mutant sequence reads, and aligning the non-mutant sequence reads to the consensus, and further, step a) comprises replacing the consensus with the mutant sequence reads in the alignment of the non-mutant sequence reads to the consensus.

31. The method according to claim 28 or 30, wherein assembling the mutant sequence reads includes aligning the mutant sequence reads to a reference sequence.

32. The method according to any one of claims 1 to 31, wherein the nucleic acid template includes a repeating sequence.

33. The method according to claim 4, wherein the randomly introduced mutation includes substitution.

34. The method according to claim 33, wherein the substitution includes a transition mutation.

35. The method according to any one of claims 1 to 34, further comprising sequencing a non-mutant nucleic acid template molecule and a mutant nucleic acid template molecule to obtain a non-mutant sequence read and a mutant sequence read.

36. The method described above is Aligning the non-mutant sequence read and the mutant sequence read to a reference sequence, The first rendered sequence is determined by determining the most likely base for each site related to the alignment of the non-mutant sequence read with respect to the mutant sequence read, and by incorporating information describing the variation in mutation rates between mutant sequence reads or the unequal mutation rates of adenine, cytosine, guanosine, and thymine. The method according to any one of claims 1 to 15, 25 to 26, or 32 to 35, comprising determining a further rendered sequence by comparing the k-mer from the first rendered sequence with the k-mer from a non-mutant sequence read.

37. The method described above is Aligning the non-mutant sequence read and the mutant sequence read to a reference sequence, Realigning a non-mutant sequence read with a mutant sequence read that has a mapping position that directly overlaps with the mutant sequence read, The alignment of non-mutant sequence reads to the mutant sequence reads is scored, and the alignment with the highest score is selected. The method according to any one of claims 1 to 15, 25 to 26, or 32 to 35, comprising determining the rendered sequence by determining the most likely base for each site related to the alignment of the non-mutant sequence read to the mutant sequence read.

38. The method described above is For each site related to the alignment of the non-mutant sequence read with respect to the mutant sequence read, the most likely base is determined. The method according to any one of claims 1 to 15, 25 to 26, or 32 to 35, wherein the EM optimization process determines a rendered sequence by applying an EM optimization process which iteratively updates an estimate of the probability that a non-mutant sequence read originated from the same copy of the mutant sequence read from which the non-mutant sequence read is aligned, while updating the inference of the most likely base.

39. The method described above is To generate consensus on assembled variant sequence reads, Aligning the non-mutant sequence reads to the consensus, In the alignment of the non-mutant sequence reads with respect to the consensus, the consensus is replaced with the mutant sequence reads, For each site related to the alignment of the non-mutant sequence read with respect to the mutant sequence read, the most likely base is determined. The method according to any one of claims 1 to 15, 25 to 26, or 32 to 35, comprising determining a rendered sequence by applying an expectation maximization (EM) optimization process, the EM optimization process iteratively updates an estimate of the probability that a non-mutant sequence read originated from the same copy of the mutant sequence read from which the non-mutant sequence read is aligned, while updating the inference of the most likely base.

40. It is a method, a) Aligning a non-mutant sequence read of a nucleic acid template to a mutant sequence read of a mutant nucleic acid template, wherein the mutant sequence read contains one or more mutations. b) i) The most likely sequence of the nucleic acid template, and ii) For each base of the most likely sequence, determine a precision score for each base that correlates with the probability that the base matches the nucleic acid template. A system comprising a processor configured to perform a method including determining the sequence of the nucleic acid template.

41. When executed by the processor, a) Aligning a non-mutant sequence read of a nucleic acid template to a mutant sequence read of a mutant nucleic acid template, wherein the mutant sequence read contains one or more mutations. b) i) The most likely sequence of the nucleic acid template, and ii) For each base of the most likely sequence, determine a precision score for each base that correlates with the probability that the base matches the nucleic acid template. A computer-readable medium containing instructions for performing a method, which includes determining the sequence of the nucleic acid template.