Co-occurrence of somatic mutations and aberrantly methylated fragments
Patent Information
- Application Number
- JP2024506251
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-04
- Filing Date
- 2022-08-04
- Publication Date
- 2025-08-01
AI Technical Summary
Existing methods for analyzing nucleic acid sequencing data from liquid biopsies struggle to accurately identify somatic and germline mutations due to the presence of nucleic acid fragments from both healthy and cancerous tissues, leading to challenges such as low prevalence of cancer, class imbalances, and susceptibility to noise, particularly in the absence of matched normal controls.
A method combining methylation data with whole genome sequencing data using a trained binary classifier to identify somatic or germline mutations by binning nucleic acid fragments based on reference and mutant alleles, incorporating methylation status and CpG site indicators, and applying a strand-specific base count set to improve accuracy.
Enhances the identification of somatic and germline mutations, improving diagnostic accuracy for cancer detection, staging, and treatment recommendations by overcoming data imbalance and noise issues, and reducing computational complexity in large-scale sequencing datasets.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Nonprovisional Patent Application No. 17 / 817,421, filed August 4, 2022, which claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 229,797, filed August 5, 2021, which is incorporated by reference in its entirety herein.
[0002] This specification describes technology relating to determining genomic mutations in a subject using sequencing of a nucleic acid sample. [Background technology]
[0003] Increasing knowledge of the molecular basis of cancer and rapid development of next generation sequencing technology has advanced the study of early molecular changes involved in cancer development and detection. Large-scale sequencing technologies such as next generation sequencing (NGS) have provided the opportunity to achieve sequencing at a cost of less than US$1 per million bases, with actual costs of less than US$10 per million bases being realized. As a result, certain genetic and epigenetic changes associated with cancer have been found in biological samples such as plasma, serous fluid, and urine. Such changes can be used as diagnostic biomarkers, for example, to correlate methylation status and other epigenetic modifications with the presence or classification of cancer. For example, DNA methylation plays an important role in regulating gene expression, and aberrant DNA methylation is involved in many disease processes, including certain cancer conditions.
[0004] Thus, the specific patterns of differentially methylated regions and / or allele-specific methylation patterns obtained using methylation sequencing can be useful as molecular markers for non-invasive diagnosis using circulating cell-free DNA (cfDNA). cfDNA found in serous fluid, plasma, urine, and other bodily fluids provides a circulating picture of disease in living subjects, including specific tumor-associated changes such as mutations, methylation, and copy number variations. Analysis of cfDNA in liquid biopsies obtained from subjects with cancerous conditions presents an attractive opportunity for a non-invasive method of screening for various cancers.
[0005] Furthermore, approaches that use deep learning to model and infer complex biological patterns and nonlinearities across genomes can be used to develop clinical and analytical tools for cancer. For example, deep learning strategies using base sequences can be used for a variety of cancer classification, regression, inference and clustering purposes, including Neu-Somatic, DeepVariant, methylation status prediction, and histone denoising. Deep learning approaches aim, in part, to address the rapid and substantial increase in the amount, size and complexity of sequencing datasets that accompany new large-scale sequencing technologies. For example, the assembly and organization of large amounts of high-fidelity nucleic acid sequences into complete genomes, and the analysis and identification of potential diagnostic indicators therein, is a computationally challenging task.
[0006] In addition to the promise and potential of applying deep learning to nucleic acid sequencing data, there are many caveats and dangers to avoid, including, among others, large class imbalance due to the low prevalence of cancer in the general population, an insufficient number of training examples relative to the number of learned parameters, and a tendency to overfit to biological or process-related noise.Similarly, cancer prediction can be approached using a number of modeling techniques (e.g., clustering, outlier, denoising or classification) using various architectures such as autoencoder, recurrent, transformer, wide-and-deep, embedding or convolutional networks, but optimally framing the problem and minimizing data imbalance, noise, overfitting and sparsity for accurate prediction are crucial challenges that require careful consideration.
[0007] For example, the quality and / or purity of samples in the training dataset may vary due to the inclusion of different types of samples (e.g., when using cfDNA from liquid biopsies that may be from multiple cell and / or tissue sources), which may result in poor classifier performance. Thus, for accurate training of a classifier, it is challenging to obtain a sufficient number of high-quality training samples that can be reliably annotated with the state of interest (e.g., cancer, non-cancer, and / or cancer subtype).
[0008] Furthermore, the identification of nucleic acid fragments with tumor-specific mutations in cancer patients remains challenging due to the high proportion of nucleic acid fragments derived from healthy tissues compared to those derived from tumor tissues. Such problems arise especially when using cfDNA fragments obtained from liquid biopsy samples, but can also arise due to clonal heterogeneity in solid tumors.
[0009] In view of the above, there is a need in the art for methods of analyzing genetic information from nucleic acid sequencing data, including data obtained from cfDNA. Summary of the Invention
[0010] The present disclosure addresses the shortcomings noted in the Background Art by providing a robust technique for identifying genomic variants as somatic or germline variants from a biological sample obtained from a subject using nucleic acid data. Combining methylation data with whole genome sequencing data and / or targeted genome sequencing data provides additional diagnostic power over conventional screening methods.
[0011] The present disclosure provides technical solutions (eg, computing systems, methods, and non-transitory computer-readable storage media) for addressing the above problems associated with analyzing datasets.
[0012] The following presents a summary of the invention in order to provide a basic understanding of some aspects of the invention. This summary is not an extensive overview of the invention. It is not intended to identify key / critical elements of the invention or to delineate the scope of the invention. Its sole purpose is to present some concepts of the invention in a simplified form as a prelude to the more detailed description that is presented later.
[0013] One aspect of the present disclosure provides a method for identifying a mutant allele at a genomic location of a test subject as a somatic mutant allele or a germline mutant allele, the method comprising obtaining an identification of a reference allele at the genomic location, obtaining an identification of a mutant allele at the genomic location, and obtaining a methylation state and a respective sequence of each nucleic acid fragment sequence in each of a plurality of nucleic acid fragment sequences in a sequencing dataset (e.g., comprising at least 10^6 nucleic acid fragment sequences) derived from a biological sample obtained from the test subject that are mapped onto the genomic location.
[0014] The identification of the reference allele at the genomic location and the respective sequence of each nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences are used to assign each nucleic acid fragment sequence of each of the plurality of nucleic acid fragment sequences that has the reference allele at the genomic location to a reference subset, and the identification of the variant allele at the genomic location and the respective sequence of each nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences are used to assign each nucleic acid fragment sequence of each of the plurality of nucleic acid fragment sequences that has the variant allele at the genomic location to a variant subset.
[0015] A trained binary classifier (e.g., comprising at least 10 parameters) is applied with at least (i) one or more indicators of methylation state across the methylation states of each nucleic acid fragment sequence in the mutant subset, and (ii) an indicator of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset, thereby obtaining from the trained binary classifier an identification of the mutant allele at this genomic location in the test subject as a somatic mutant allele or a germline mutant allele.
[0016] In some embodiments, the method further includes inputting the reference genome into a computer system including a processor coupled to a non-transitory memory, and determining, using the computer system, that each nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences is mapped to a genomic location by aligning each nucleic acid fragment sequence to the reference genome.
[0017] In some embodiments, a first nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences has a plurality of CpG sites, the first nucleic acid fragment sequence has a corresponding methylation pattern across the plurality of CpG sites, and the methylation state of the first nucleic acid fragment sequence is a p-value, and the method further comprises determining the p-value of the first nucleic acid fragment sequence at least in part by comparing the corresponding methylation pattern of the first nucleic acid fragment sequence to a corresponding distribution of methylation patterns of those nucleic acid fragment sequences in healthy non-cancer cohort datasets, each having a respective plurality of CpG sites.
[0018] In some embodiments, if the mutant allele at the genomic location is determined by the trained binary classifier to be a germline mutant allele, the method further comprises using the mutant allele in the test subject to determine the cancer risk of the test subject. In some embodiments, if the mutant allele at the genomic location is determined by the trained binary classifier to be a germline mutant allele, the method further comprises using the mutant allele in the test subject to predict the ethnicity of the subject. In some embodiments, if the mutant allele at the genomic location is determined by the trained binary classifier to be a somatic mutant allele, the method further comprises using the mutant allele in the test subject to determine the tumor fraction of the subject.
[0019] In some embodiments, the application of the trained binary classifier further applies one or more CpG site indices across the variant subset.
[0020] In some embodiments, the application of the trained binary classifier further applies one or more indices of methylation status across the reference subset.
[0021] In some embodiments, the application to the trained binary classifier further applies one or more CpG site indices across the reference subset.
[0022] In some embodiments, obtaining the identification of the mutant allele at the genomic location includes obtaining a strand-specific base count set for the genomic location, the strand-specific base count set including forward and reverse strand-specific counts of each base in the set of bases (e.g., A, C, T, G) at the genomic location obtained by determining (i) the strand orientation of each base in each nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences and (ii) the identity of each base at the genomic location, where a base at the genomic location in each of the plurality of nucleic acid fragment sequences whose identity may be affected by conversion of a methylated cytosine or an unmethylated cytosine does not contribute to the strand-specific base count set. Using the strand-specific base count set and the sequencing error estimate, a respective forward strand conditional probability and a respective reverse strand conditional probability are calculated for each candidate genotype in the set of candidate genotypes for the genomic location, thereby calculating a plurality of forward strand conditional probabilities and a plurality of reverse strand conditional probabilities. A plurality of likelihoods are calculated, each likelihood in the plurality of likelihoods being a likelihood for a respective candidate genotype in the set of candidate genotypes, the calculation using a combination of (i) a respective forward strand conditional probability for each candidate genotype in the plurality of forward strand conditional probabilities, (ii) a respective reverse strand conditional probability for each candidate genotype in the plurality of reverse strand conditional probabilities, and (iii) a genotype prior probability for each candidate genotype. The plurality of likelihoods are used to identify a mutant allele at the genomic location, thereby obtaining an identification of the mutant allele at the genomic location.
[0023] In some embodiments, the method further comprises repeating the method for each genomic location in the plurality of genomic locations, thereby identifying a plurality of test mutations, and for each mutation in the plurality of mutations, identifying whether the each mutation is a somatic mutation or a germline mutation.
[0024] Another aspect of the disclosure provides a method of training a classifier (e.g., comprising at least 10 parameters) to identify a variant allele at a genomic location of a test subject as a somatic variant allele or a germline variant allele, the method including obtaining an identification of a reference allele at the genomic location and performing the procedure for each genomic location in the plurality of genomic locations in each subject in the plurality of subjects.
[0025] The method includes: i) obtaining an orthogonal call for a mutant allele at each genomic location for each subject as either a somatic mutant allele or a germline mutant allele; ii) obtaining an identification of the mutant allele at each genomic location for each subject; iii) obtaining a methylation state and a respective sequence of each nucleic acid fragment sequence in each of a plurality of nucleic acid fragment sequences in a sequencing dataset (e.g., comprising at least 10^6 nucleic acid fragment sequences) derived from a biological sample obtained from each subject, the sequence being mapped to the respective genomic location; iv) assigning each nucleic acid fragment sequence having the reference allele at the respective genomic location among the respective plurality of nucleic acid fragment sequences to a reference subset using (a) the identification of the reference allele at the respective genomic location and (b) the respective sequence of each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences; and v) assigning each nucleic acid fragment sequence having the mutant allele at the respective genomic location among the respective plurality of nucleic acid fragment sequences to a mutant subset using (a) the identification of the mutant allele at the respective genomic location and (b) the respective sequence of each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences.
[0026] For each genomic location among the plurality of genomic locations in each subject among the plurality of subjects, at least (i) one or more indices of methylation state across the methylation states of each nucleic acid fragment sequence in the mutant subset of each subject at each genomic location, (ii) an indicia of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset of each subject at each genomic location, and (iii) an orthogonal call of the mutant allele at each genomic location of each subject as either a somatic mutant allele or a germline mutant allele is used to train a classifier to identify the mutant allele at the genomic location of the test subject as a somatic mutant allele or a germline mutant allele.
[0027] Another aspect of the present disclosure provides a computing system comprising one or more processors and a memory storing one or more programs executed by the one or more processors, the one or more programs including instructions for performing any of the methods disclosed above, either alone or in combination.
[0028] Yet another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs configured to be executed by a computer, the one or more programs including instructions for performing any of the methods disclosed above, either alone or in combination.
[0029] Various embodiments of the systems, methods, and devices within the scope of the appended claims each have several aspects, no one of which is solely responsible for the desirable attributes described herein. Without limiting the scope of the appended claims, several prominent features are described herein. By considering this description, and particularly by reading the section entitled "Description of the Preferred Embodiments," one will understand how the features of the various embodiments may be used.
[0030] Incorporation by Reference All patents and publications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual patent and publication was specifically and individually indicated to be incorporated by reference. [Brief description of the drawings]
[0031] Implementations disclosed herein are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings in which like reference numerals refer to corresponding parts throughout the several views of the drawings. [Figure 1] FIG. 1 illustrates an exemplary block diagram illustrating a computing device according to some embodiments of the present disclosure. [Figure 2A] 1A-1D collectively depict exemplary flow charts of a method for identifying a mutant allele at a genomic location under test as a somatic mutant allele or a germline mutant allele, according to some embodiments of the present disclosure, where dashed boxes represent optional steps. [Figure 2B] 1A-1D collectively depict exemplary flow charts of a method for identifying a mutant allele at a genomic location under test as a somatic mutant allele or a germline mutant allele, according to some embodiments of the present disclosure, where dashed boxes represent optional steps. [Diagram 3] 1 shows an exemplary flow chart of a method for calling variant alleles according to some embodiments of the present disclosure. [Figure 4A] 1 shows an analysis of the correlation between methylation patterns and somatic mutations according to some embodiments of the present disclosure. [Figure 4B] 1 shows an analysis of the correlation between methylation patterns and somatic mutations according to some embodiments of the present disclosure. [Figure 5A] 1 illustrates exemplary performance measures of methods according to some embodiments of the present disclosure. [Figure 5B] 1 illustrates exemplary performance measures of methods according to some embodiments of the present disclosure. [Figure 6A] 1 illustrates exemplary performance measures of methods according to some embodiments of the present disclosure. [Figure 6B]1 illustrates exemplary performance measures of methods according to some embodiments of the present disclosure. [Figure 7] 1 shows a flowchart of a method for preparing a nucleic acid sample for sequencing according to some embodiments of the present disclosure. [Figure 8] 1 is a graphical representation of a process for obtaining sequence reads according to some embodiments of the present disclosure. [Figure 9] 1 shows an exemplary flowchart of a method for obtaining methylation information in a subject, according to some embodiments of the present disclosure. [Figure 10A] 1 illustrates exemplary performance measures of methods according to some embodiments of the present disclosure. [Figure 10B] 1 illustrates exemplary performance measures of methods according to some embodiments of the present disclosure. [Figure 11A] 1 illustrates exemplary performance measures of methods according to some embodiments of the present disclosure. [Figure 11B] 1 illustrates exemplary performance measures of methods according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0032] Introduction As mentioned above, conventional methods for analyzing nucleic acid sequencing data may not provide accurate determination of cancer-related biomarkers.For example, although the recent development of next-generation sequencing technology and machine learning has brought about advances in the analysis of sequencing data, accurate determination of gene mutations using cfDNA is hindered by the presence of nucleic acid molecules derived from other tissues, such as healthy tissues.Conventional methods may include obtaining and sequencing patient-matched normal (e.g., healthy) control samples, such as white blood cell or tissue biopsy, and performing comparative analysis to determine which mutations observed in liquid biopsy samples may be derived from tumors and which mutations are derived from normal controls.
[0033] In the absence of a matched normal control, determining whether a genomic alteration is a germline or somatic mutation can be difficult, especially for rare or unannotated mutations. However, unlike liquid biopsy samples, matched normal controls may not be routinely obtained in clinical settings. For example, as described herein, the use of bodily fluids advantageously facilitates clinical applications because these bodily fluids are obtained by non-invasive or minimally invasive methods and therefore are easy to collect. This can be in contrast to methods that rely on solid tissue samples, such as biopsies, which often use invasive surgical procedures. Thus, the improved methods described herein can include analyzing nucleic acid sequencing data to accurately identify and classify genetic mutations, such as tumor-specific mutations, in cfDNA. In particular, the improved methods can include identifying mutant alleles as somatic mutant alleles or germline mutant alleles.
[0034] Advantageously, the present disclosure provides methods and systems that provide accurate determination of mutant alleles as somatic mutant alleles or germline mutant alleles. For example, in some embodiments, the methods and systems described herein include using nucleic acid sequencing and methylation sequencing of nucleic acid fragments in a liquid biopsy sample to obtain a plurality of features for input to a binary classifier trained to identify mutant alleles in a subject as somatic mutant alleles or germline mutant alleles. Each nucleic acid fragment that maps to a genomic location of a mutant allele may be binned into a mutant subset if the corresponding sequence reads (e.g., obtained from nucleic acid sequencing) have support for the mutant allele, or into a reference subset if the corresponding sequence reads have support for the reference allele. The features used as input to the classifier may include at least one or more distribution statistics of p-values calculated across methylation vectors (e.g., obtained from methylation sequencing) corresponding to nucleic acid fragments in the mutant subset and the reference subset, respectively. In some embodiments, the features further comprise a count of CpG sites in the nucleic acid fragments assigned to the mutation subset and a count of CpG sites in the nucleic acid fragments assigned to the reference subset, which may result in an output from the trained binary classifier identifying whether a mutation allele at a genomic location of a subject is a somatic mutation allele or a germline mutation allele.
[0035] Accurate identification of mutations as somatic or germline mutations may provide advantages in clinical applications such as diagnosing cancer, determining the stage of cancer, monitoring the progression of cancer, determining prognosis, prescribing or administering treatment, matching or recommending enrollment in clinical trials, monitoring the occurrence of further complications or risks over time, and assessing the efficacy of treatment, among others.
[0036] For example, somatic mutations reflect gene mutations that accumulate over a subject's lifetime through mutagenesis processes (e.g., smoking, drinking, etc.) and are more closely associated with the development of cancer. Potential therapeutic uses of somatic mutation identification may include improving a physician's ability to interpret cancer types and select the most effective treatment options. Thus, accurate identification of genetic mutations as somatic or germline mutations may impact a healthcare provider's ability to determine appropriate treatment recommendations for patients. In addition to cancer risk, monitoring, and treatment, identification of somatic mutations using the methods described herein may also be used for tumor fraction estimation (e.g., to confirm or complement tumor mutation burden calculations obtained using matched normal control samples). Furthermore, somatic mutations may be indicative of other disease types, including clonal hematopoiesis of indeterminate (CHIP), cardiovascular risk, nonalcoholic fatty liver disease (NAFLD or NASH), and other disease conditions.
[0037] In contrast, germline mutations may not be involved in the development of cancer and therefore typically provide less information than somatic mutations regarding the detection and / or identification of cancer. Nevertheless, germline mutations can provide information regarding antecedent cancer risk, either through the identification of annotated cancer-associated germline mutations (e.g., BRCA) or through the calculation of polygenic risk scores (PRS) using genetic information. Furthermore, accurate identification of germline mutations can be used for analytical procedures such as enrichment of somatic mutations in a dataset, or for other applications such as ethnicity prediction.
[0038] Advantageously, the method disclosed herein can overcome the above-mentioned difficulties of identifying somatic mutations in the absence of normal (e.g., healthy) controls by using methylation patterns to improve the quality of mutation calls in nucleic acid sequencing data. The method disclosed herein can improve prior art mutation classification methods that use only nucleic acid sequencing by exploiting the possibility of the co-occurrence of abnormal methylation signals and enrichment of somatic mutations in combination with machine learning algorithms.
[0039] Specifically, p-values and CpG distribution statistics based on methylation sequencing of nucleic acid fragments can be added to the input vector for a trained binary classifier to improve the performance of the classifier compared to the baseline input including reference fragment counts and mutant fragment counts obtained using nucleic acid sequence reads.For example, as reported in Example 6, when methylation fragment p-values and CpG counts are added to the baseline input of reference fragment counts and mutant fragment counts, the performance of logistic regression classifiers and neural network classifiers is improved in terms of area under the curve (AUC), positive predictive value (precision), and sensitivity (recall).Improvements are observed both when using tissue-derived sequencing datasets, as shown in Figures 5A, 5B, 6A, and 6B, and when using cfDNA-derived sequencing datasets, as shown in Figures 10A, 10B, 11A, and 11B.
[0040] Thus, the described methods and systems may improve how treatments are assigned and / or administered due to improved accuracy in identifying mutations as somatic or germline mutations.
[0041] Further benefits Identifying genomic alterations in a patient's cancer genome can be a challenging and computationally demanding problem. For example, the determination of various prognostic metrics useful in clinical practice, including the identification and classification of mutant alleles, uses the analysis of hundreds of millions to billions of sequenced nucleic acid bases. An example of a typical bioinformatics pipeline established for this purpose can include at least five stages of analysis: evaluation of the quality of raw next-generation sequencing data, generation of collapsed nucleic acid fragment sequences and alignment of such sequences to a reference genome, detection of structural variations in the aligned sequence data, annotation of identified mutations, and visualization of the data.
[0042] Furthermore, the method of the present disclosure may include additional processes such as performing methylation sequencing, correlating each methylated fragment sequence to each nucleic acid fragment and its corresponding nucleic acid sequence, binning the plurality of nucleic acid fragments at each mutation position, faceting the nucleic acid fragments based on reference support or alternative support, determining a plurality of features (including, but not limited to, reference fragment count, alternative fragment count, methylation state p-value distribution statistics, and / or CpG site count distribution statistics) for the plurality of fragments binned at each mutation position, and generating a feature vector for input to a binary classifier. In some aspects of the present disclosure, the method may further include training a binary classifier based on a training dataset including a plurality of training subjects to identify mutations as somatic mutations or germline mutations. Each of these steps may be computationally intensive in itself.
[0043] For example, the overall time and space computational complexity of simple global pairwise sequence alignment algorithms and local pairwise sequence alignment algorithms can be quadratic in nature (i.e., quadratic problem) and grow rapidly as a function of the size (n and m) of the nucleic acid sequences to be compared.Specifically, the time and space complexity of these sequence alignment algorithms can be estimated as O(mn), where O is the upper limit of the asymptotic growth rate of the algorithm, n is the number of bases in the first nucleic acid sequence, and m is the number of bases in the second nucleic acid sequence.Considering that the human genome contains more than 3 billion bases, these alignment algorithms can be extremely computationally intensive, especially when used to analyze next-generation sequencing (NGS) data, which can generate more than 3 billion sequence reads per reaction.
[0044] This may be especially true when performed in the context of a liquid biopsy assay, since liquid biological samples may contain a complex mixture of short DNA fragments originating from many different germline (e.g., healthy) and diseased (e.g., cancerous) tissues. Thus, the cellular origin of the sequence reads may be unknown, and sequence signals originating from cancerous cells, which may constitute multiple subclonal populations, may be computationally deconvoluted from signals originating from germline and hematopoietic origins to provide relevant information about the subject's cancer. Thus, in addition to the computationally intensive process used to align sequence reads to the human genome, there may be a computational problem in determining whether one or more sequence reads corresponding to a particular aberrant signal, e.g., a genomic alteration, (i) are not an artifact and (ii) originate from a cancerous source in the subject. This may become increasingly difficult in the early stages of cancer, when treatment is likely most effective and when small amounts of circulating tumor DNA (ctDNA) are diluted by germline and hematopoietic DNA.
[0045] Advantageously, the present disclosure provides various systems and methods for improving computational resolution of genomic alterations (e.g., somatic or germline mutations) in a subject from cfDNA. The methods and systems described herein can solve problems in computing technology, for example, by improving the accuracy of identification of mutations as somatic or germline mutations. As detailed above, classification of mutations is performed using a large sequencing dataset (e.g., at least 1×10 6 The method may involve multiple processes that can be implemented as a bioinformatics pipeline, utilizing multiple sequencing reads (number of sequence reads) and with computational complexity in time and space that increases with the size of the sequencing dataset at a quadratic rate. The large requirements for computational power, including processing time and processing space, may reduce the efficiency of computer-implemented methods. Given these constraints, improvements in such processes may provide a solution to computing technologies by providing more efficient and accurate methods for mutation identification.
[0046] More advantageously, the present disclosure provides various systems and methods for improving computational resolution of genomic alterations (e.g., somatic or germline mutations) in subjects from cfDNA by improving model training and use for more accurate mutation identification. The complexity of a machine learning model can include time complexity (run time, or a measure of the speed of an algorithm for a given input size n), space complexity (space requirements, or the amount of computing power or memory required to run an algorithm for a given input size n), or both. Complexity (and subsequent computational burden) can be applied to both training and prediction by a given model.
[0047] In some examples, computational complexity may be affected by the implementation, the incorporation of additional algorithms or cross-validation methods, and / or one or more parameters (e.g., weights and / or hyperparameters). Nevertheless, computational complexity can generally be expressed as a function of the input size n, where the input data is the number of instances (e.g., number of training samples), the dimension p (e.g., number of features), the number of trees n trees (e.g., for tree-based methods), the number of support vectors, n sv (e.g., for methods based on support vectors), the number of neighbors k (e.g., for k-nearest neighbor algorithms), the number of classes c, and / or the number of neurons in layer i, n i (e.g., for neural networks). Then, with respect to an input size n, an approximation of the computational complexity (e.g., in Big O notation) indicates how the execution time and / or space requirements increase as the input size increases. A function can increase in complexity at a slower or faster rate compared to an increase in the input size. Various approximations of computational complexity include a constant approximation (e.g., O(1)), a logarithmic approximation (e.g., O(log n)), a linear approximation (e.g., O(n)), a log-linear approximation (e.g., O(n log n)), a quadratic approximation (e.g., O(n 2 ), polynomial approximations (e.g., O(n c), exponential approximations (e.g., O(c n ), and / or factorial approximations (e.g., O(n!)). In some examples, simpler functions may have lower levels of computational complexity as the input size increases, such as for constant functions, while more complex functions, such as factorial functions, may experience a substantial increase in complexity in response to small increases in input size.
[0048] The computational complexity of a machine learning model can similarly be expressed by a function (e.g., in Big O notation), where the complexity may vary depending on the type of model, the size of one or more inputs or dimensions, the application (e.g., training and / or prediction), and / or whether time or space complexity is being assessed. For example, the complexity of a decision tree algorithm may be O(n 2 p), prediction is approximated as O(p), and the complexity of the linear regression algorithm is O(p 2 n+p 3 ) for prediction, and O(p) for approximation. For the Random Forest algorithm, the training complexity is O(n 2 p.n. trees ), and the prediction complexity is O(pn trees ) for the gradient boosting algorithm. The complexity is O(npn trees ), and prediction is O(pn trees ) for the kernel support vector machine, the complexity is O(n 2 p+n 3 ), and O(n svFor the Naive Bayes algorithm, the complexity can be approximated as O(np) for training and O(p) for prediction, while for neural networks, the complexity can be approximated as O(pn1+n1n2+...) for prediction. The complexity in the K-nearest neighbors algorithm can be approximated as O(knp) for time and O(np) for space. For the Logistic Regression algorithm, the complexity can be approximated as O(np) for time and O(p) for space. For the Logistic Regression algorithm, the complexity can be approximated as O(np) for time and O(p) for space.
[0049] As mentioned above, for machine learning models, computational complexity may dictate the scalability of the model (e.g., classifier) and changes in model architecture to increase input, feature, and / or class size, and therefore the overall effectiveness and usefulness. In the context of large-scale sequencing technology, the computational complexity of the functions performed on the sequencing dataset (e.g., nucleic acid sequencing data and methylation sequencing data obtained from cfDNA samples) may tax the capacity of many existing systems. Furthermore, as the number of input features (e.g., reference and alternative counts, stratified for reference and alternative subsets, p-value distribution statistics (e.g., mean, minimum, maximum, median, standard deviation), and / or CpG site distribution statistics (e.g., mean, minimum, maximum, median, standard deviation)) and / or the number of instances (e.g., training subjects, test subjects, number of variant alleles, and / or number of genomic locations) increase with the expansion of downstream applications and possibilities, the computational complexity of any given classification model may quickly overwhelm the temporal and spatial capabilities provided by the specifications of the respective system.
[0050] In general (and as defined herein), a parameter (e.g., weight and / or hyperparameter) is a coefficient that changes one or more inputs, outputs, or functions of a model. For example, the value of a parameter can be used to up-weight or down-weight the influence of an input to a model, such as a feature. Thus, a feature can be associated with a parameter of a logistic regression model, an SVM model, a Naive Bayes model, or the like. The value of a parameter can alternatively or additionally be used to up-weight or down-weight the influence of a node (e.g., a node includes one or more activation functions that define the transformation from input to output), a class, or an instance (e.g., of a sample) in a neural network. The assignment of parameters to a particular input, output, function, or feature may be any one paradigm for a given model, but may be used in any suitable model architecture for optimal performance. Nevertheless, reference to coefficients associated with inputs, outputs, functions, or features of a model may be used as an indication of their number, performance, or optimization, such as in the context of the computational complexity of a machine learning algorithm.
[0051] Therefore, a minimum input size (e.g., at least 1×10 6 A machine learning model having a minimum number of sequence reads) and / or parameters (e.g., at least 10, at least 100, or at least 1000 parameters) can refer to a corresponding number of relevant inputs, outputs, functions, or features in the model. The computational complexity of such a model can be increased proportionately, such that the model for the methods of the present disclosure (e.g., identifying somatic or germline mutations from cfDNA in a subject) cannot be used in mental calculations, and the method may be inherently a computational problem.
[0052] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0053] definition As used herein, the term "about" or "approximately" means within an acceptable error range of a particular value as measured by one of ordinary skill in the art, which depends in part on the constraints of how the value is measured or determined, i.e., the measurement system. For example, in some embodiments, "about" means within 1 or more than 1 standard deviation, according to the practice in the art. In some embodiments, "about" means within ±20%, ±10%, ±5%, or ±1% of a given value. In some embodiments, the term "about" or "approximately" means within the same order of magnitude, within 5-fold, or within 2-fold of a value. When a particular value is described in the present application and claims, unless otherwise indicated, the term "about" can be assumed to mean within an acceptable error range for the particular value. The term "about" can have the meaning commonly understood by a person of ordinary skill in the art. In some embodiments, the term "about" refers to ±10%. In some embodiments, the term "about" refers to ±5%.
[0054] When a range of values is provided, it is understood that each intervening value between the upper and lower limits of that range, to one-tenth of the unit of the lower limit, and any other stated or intervening value in that stated range, unless the context clearly dictates otherwise, is encompassed within the invention. The upper and lower limits of these smaller ranges may be independently included in the smaller ranges, and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where a stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included within the invention. For example, as used herein, the term "between" used in a range is intended to include the recited endpoints. For example, a number "between X and Y" can be X, Y, or any value from X to Y.
[0055] As used herein, the term "allele" refers to a specific sequence of one or more nucleotides at a genomic position.For haploid organisms, subjects generally have one allele at every genomic position.For diploid organisms, subjects generally have two alleles at every genomic position.
[0056] As used herein, the term "assay" refers to a technique for determining a characteristic of a substance, such as a nucleic acid, a protein, a cell, a tissue, or an organ. An assay (e.g., a first assay or a second assay) can include a technique for determining copy number variation of a nucleic acid in a sample, a methylation state of a nucleic acid in a sample, a fragment size distribution of a nucleic acid in a sample, a mutation state of a nucleic acid in a sample, or a fragmentation pattern of a nucleic acid in a sample. Any assay can be used to detect any of the nucleic acid characteristics described herein. A nucleic acid characteristic can include a sequence, a genomic identity, a copy number, a methylation state at one or more nucleotide positions, a size of a nucleic acid, the presence or absence of a mutation in a nucleic acid at one or more nucleotide positions, and a fragmentation pattern of a nucleic acid (e.g., a nucleotide position at which a nucleic acid is fragmented). An assay or method can have a particular sensitivity and / or specificity, and their relative utility as a diagnostic tool can be measured using the ROC-AUC statistic.
[0057] As used herein, the term "biological sample" or "sample" refers to any sample taken from a subject (i.e., any type of organism, not just humans) that may reflect a biological state associated with the subject. Examples of biological samples include, but are not limited to, blood, whole blood, plasma, serous fluid, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of a subject. A biological sample may include any tissue or material derived from a living or dead subject. A biological sample may be a cell-free sample and / or may include cell-free DNA. A biological sample may include nucleic acid (e.g., DNA or RNA) or fragments thereof. The term "nucleic acid" may refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or any hybrid or fragment thereof. The nucleic acid in the sample may be cell-free nucleic acid. A sample may be a liquid sample or a solid sample (e.g., a cell sample or a tissue sample). The biological sample may be a bodily fluid, such as blood, plasma, serous fluid, urine, vaginal fluid, fluid from hydrocele (e.g., scrotal), vaginal washing fluid, pleural fluid, peritoneal fluid, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge fluid, aspirated fluid from different parts of the body (e.g., thyroid, breast), etc. The biological sample may be a fecal sample. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) may be cell-free (e.g., more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free). The biological sample may be treated to physically disrupt tissue or cellular structures (e.g., centrifugation and / or cell lysis), thus releasing intracellular components into a solution that may further contain enzymes, buffers, salts, detergents, etc. that may be used to prepare the sample for analysis. Biological samples may be obtained from a subject invasively (eg, by surgical means) or non-invasively (eg, by drawing blood, swabbing, or taking a discharged sample).
[0058] As used herein, the term "cancer" or "tumor" refers to an abnormal mass of tissue in which the growth of the mass exceeds and is uncoordinated with that of normal tissue. Cancers or tumors can be defined as "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation, including morphology and functionality, growth rate, local invasion and metastasis. "Benign" tumors can be well differentiated, characteristically grow slower than malignant tumors, and remain localized at the primary site. Furthermore, in some cases, benign tumors do not have the ability to invade, invade, or metastasize to distant sites. "Malignant" tumors can be poorly differentiated (anaplastic) and have characteristic rapid growth with progressive infiltration, invasion, and destruction of surrounding tissues. Furthermore, malignant tumors can have the ability to metastasize to distant sites.
[0059] As used interchangeably herein, the terms "cancer burden," "tumor burden," "cancer burden," "tumor burden," or "tumor fraction" refer to the concentration or presence of tumor-derived nucleic acid in a test sample. Thus, the terms "cancer burden," "tumor burden," "cancer burden," "tumor burden," and "tumor fraction" are non-limiting examples of cell source fractions in a biological sample. In some embodiments, the tumor fraction is a specific version of the cell source fraction.
[0060] As disclosed herein, the terms "cell-free nucleic acid", "cell-free DNA" and "cfDNA" refer interchangeably to nucleic acid fragments circulating in a subject's body (e.g., in bodily fluids such as bloodstream) and derived from one or more healthy cells and / or one or more cancer cells. Cell-free DNA can be collected from bodily fluids such as blood, whole blood, plasma, serous fluid, urine, cerebrospinal fluid, feces, saliva, sweat, perspiration, tears, pleural fluid, pericardial fluid, or ascites of a subject. Cell-free nucleic acid is used interchangeably with circulating nucleic acid. Examples of cell-free nucleic acid include, but are not limited to, RNA, mitochondrial DNA, or genomic DNA.
[0061] As disclosed herein, the term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments derived from abnormal tissues, such as cells of a tumor or other type of cancer, that may be released into a subject's bloodstream as a result of biological processes, such as apoptosis or necrosis of dying cells, or that may be actively released by viable tumor cells.
[0062] As used herein, the term "classification" refers to any number or other letter associated with a particular characteristic of a sample. For example, a "+" sign (or the word "positive") may indicate that the sample is classified as having a deletion or an amplification. In another example, the term "classification" refers to the amount of tumor tissue in the subject and / or sample, the size of the tumor in the subject and / or sample, the stage of the tumor in the subject, the tumor burden in the subject and / or sample, and the presence of tumor metastases in the subject. In some embodiments, the classification is binary (e.g., positive or negative, somatic mutation or germline mutation, etc.) or has more levels of classification (e.g., a scale of 1-10 or 0-1). In some embodiments, the terms "cutoff" and "threshold" refer to a predefined number used in an operation. In one example, a cutoff size refers to a size above which fragments are excluded. In some embodiments, a threshold is a value above or below which a particular classification is applied. Any of these terms can be used in any of these contexts.
[0063] As used herein, the terms "control sample", "reference sample" and "normal sample" refer to a sample from a subject that does not have a particular condition or is otherwise healthy. In one example, the method disclosed herein can be performed on a subject with a tumor, and the reference sample is a sample taken from the subject's healthy tissue. The reference sample can be obtained from the subject or from a database. The reference sample can be, for example, a reference genome used to map sequence reads obtained from sequencing a sample from the subject. The reference genome can refer to a haploid genome or a diploid genome to which sequence reads from a biological sample and a constituent sample can be aligned and compared. An example of a constituent sample can be the DNA of a white blood cell obtained from a subject. In the case of a haploid genome, there can be one nucleotide at each locus. In the case of a diploid genome, heterozygous loci can be identified, and each heterozygous locus can have two alleles, allowing either allele to match for alignment to the locus.
[0064] As used herein, the term "genomic location" or "locus" refers to a location (e.g., site) within a genome, for example, on a particular chromosome. In some embodiments, a genomic location (e.g., locus) refers to the location of a single nucleotide on a particular chromosome within a genome. In some embodiments, a genomic location refers to a group of nucleotide positions within a genome. In some embodiments, a genomic location refers to one or more genomic coordinates and / or a range of genomic coordinates (e.g., within a reference sequence or genome). For example, in some embodiments, a genomic location is used to indicate or identify a genomic region. In some examples, a genomic location is characterized by a mutation (e.g., a substitution, insertion, deletion, inversion, or translocation) of consecutive nucleotides within a cancer genome. In some examples, a genomic location is a gene, a subgene structure (e.g., a regulatory element, an exon, an intron, or a combination thereof), or a defined range of a chromosome. Because normal mammalian cells have a diploid genome, a normal mammalian genome (e.g., the human genome) generally has two copies of each genomic position (e.g., locus) in the genome, or at least two copies of each genomic position (e.g., locus) located on an autosome, e.g., one copy on the maternal autosome and one copy on the paternal autosome.
[0065] As disclosed herein, the term "genomic region" or "chromosomal region" refers to any contiguous or discontinuous portion of a genome. A genomic region may also be referred to as, for example, a bin, a section, a genomic portion, a portion of a reference genome, a portion of a chromosome, etc. In some embodiments, the genomic region is based on a particular length of a genomic sequence. For example, in some embodiments, the method may include analysis of multiple mapped sequence reads to multiple genomic regions. The genomic regions may be of approximately the same length or of different lengths. In some embodiments, genomic regions of different lengths are adjusted or weighted. In some embodiments, the genomic region is about 3 base pairs (bp) to about 100 bp, about 0.1 kilobases (kb) to about 10 kb, about 10 kb to about 500 kb, about 20 kb to about 400 kb, about 30 kb to about 300 kb, about 40 kb to about 200 kb, and sometimes about 50 kb to about 100 kb. In some embodiments, the genomic region is about 100 kb to about 200 kb. A genomic region is not limited to a contiguous set of sequences. Thus, a genomic region may be composed of contiguous and / or discontinuous sequences. A genomic region is not limited to a single chromosome. In some embodiments, a genomic region includes all or part of one chromosome, or all or part of two or more chromosomes. In some embodiments, a genomic region may span one, two, or more entire chromosomes. Furthermore, a genomic region may span the junction or separation of multiple chromosomes.
[0066] As used herein, the term "measure of central tendency" refers to the center or representative value of a distribution of values. Non-limiting examples of measures of central tendency include the arithmetic mean, weighted average, mid-range, mid-hinge, triangular mean, geometric mean, geometric median, winsorized mean, median, and mode of a distribution of values.
[0067] As used herein, the term "methylation" refers to a modification of deoxyribonucleic acid (DNA) in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. In particular, methylation tends to occur at the dinucleotides of cytosine and guanine, referred to herein as "CpG sites." In other instances, methylation may occur at a cytosine that is not part of a CpG site, or at another nucleotide that is not a cytosine. However, these occur more rarely. In this disclosure, methylation is discussed with reference to CpG sites for clarity. Aberrant cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which may be indicative of a cancerous condition. As is well known in the art, DNA methylation aberrations (compared to healthy controls) can cause a variety of effects that may contribute to cancer.
[0068] Identifying abnormally methylated cfDNA fragments presents various challenges. First, in some cases, determining that a subject's cfDNA is abnormally methylated carries weight compared to a group of control subjects, and as a result, when the number of control groups is small, the determination is unreliable for this small group of controls. In addition, the methylation status may differ between groups of control subjects, which may be difficult to account for when determining that a subject's cfDNA is abnormally methylated. In other respects, in some cases, methylation of cytosines at CpG sites causally affects the methylation at subsequent CpG sites.
[0069] The principles described herein are equally applicable to detection of methylation in non-CpG contexts, including non-cytosine methylation. Furthermore, the methylation state vector may contain elements that are generally vectors of sites that are methylated or unmethylated (even if these sites are not specifically CpG sites). With that substitution, the remainder of the process described herein remains the same, and thus the inventive concepts described herein are applicable to these other forms of methylation.
[0070] In some embodiments, the methylation level of a nucleic acid fragment is provided using a beta value and / or an M value, both of which provide a measure of differential methylation at one or more given CpG sites. For example, a beta value is defined as the ratio of the intensity between the methylated allele (e.g., of a given CpG site) and the sum of all (methylated and unmethylated) alleles. The intensity can be determined by interrogating each CpG site using a methylated probe and an unmethylated probe in a methylation assay (e.g., an Illumina methylation assay). The beta value statistic results in a number between 0 and 1, or between 0% and 100%. Under ideal conditions, a value of 0 indicates that all copies of a CpG site in a sample are completely unmethylated (no methylated molecules were measured), and a value of 1 indicates that all copies of the site are methylated. An M value is defined as the log2 ratio of the intensity between the methylated allele (e.g., of a given CpG site) and the unmethylated allele. The intensities used for M-value estimation can be determined by interrogating each CpG site using a methylated probe and an unmethylated probe in a methylation assay (e.g., an Illumina methylation assay). An M-value close to 0 indicates similarity in intensity between the methylated and unmethylated probes, which generally means that the CpG site is about half methylated. A positive M-value generally means that more fragments are methylated than unmethylated, and a negative M-value means the opposite (more fragments are unmethylated than methylated). In some embodiments, the intensity data is normalized (e.g., by Illumina GenomeStudio or some other external normalization algorithm) before beta-value estimation or M-value estimation. Further details regarding beta-values and M-values are provided in Du et al., "Comparison of Beta-value and M-value methods for quantifying methylation levels by microarray analysis," BMC Bioinformatics 2010, 11:587, which is incorporated herein by reference in its entirety.
[0071] As used herein, the term "methylation index" of each genomic site (e.g., a CpG site, a region of DNA in which a cytosine nucleotide is followed by a guanine nucleotide in the linear sequence of bases along its 5'→3' direction) refers to the proportion of sequence reads that show methylation at that site relative to the total number of reads that cover that site. The "methylation density" of a region can be the number of reads at sites within the region that show methylation divided by the total number of reads that cover sites within the region. These sites can have specific characteristics (e.g., these sites can be CpG sites). The "CpG methylation density" of a region can be the number of reads that show CpG methylation divided by the total number of reads that cover CpG sites within the region (e.g., a particular CpG site, a CpG site within a CpG island, or a larger region). For example, the methylation density of each 100 kb bin in the human genome can be determined from the total number of unconverted cytosines (which may correspond to methylated cytosines) at the CpG site as a percentage of all CpG sites covered by sequence reads mapped to the 100 kb region. In some embodiments, this analysis is performed for other bin sizes, such as 50 kb or 1 Mb. In some embodiments, the region is the entire genome or a chromosome or a portion of a chromosome (e.g., a chromosome arm). The methylation index of a CpG site can be the same as the methylation density of the region if the region contains that CpG site. The "percentage of methylated cytosines" can refer to the number of cytosine sites "C" that are shown to be methylated (e.g., unconverted after bisulfite conversion) over the total number of cytosine residues analyzed, including cytosines outside of CpG contexts in the region. The methylation index, methylation density, and percentage of methylated cytosines are examples of "methylation levels."
[0072] As used herein, the term "methylation pattern" or "methylation state vector" refers to an array of methylation states of one or more CpG sites. Methylation states include, but are not limited to, methylated states (e.g., represented by "M") and unmethylated states (e.g., represented by "U"). For example, a methylation pattern spanning five CpG sites can be represented as "MMMMM" or "UUUUU", with each discrete symbol representing the methylation state at a single CpG site. A methylation pattern may or may not correspond to a particular genomic location and / or a particular CpG site or sites in a reference genome.
[0073] As used interchangeably herein, terms such as "node," "neuron," "unit," "hidden neuron," "hidden unit," and the like refer to a unit of a neural network that accepts inputs and provides outputs via activation functions and one or more parameters (e.g., weights and / or hyperparameters). For example, a node can accept one or more inputs from a previous layer and provide an output that serves as an input for a subsequent layer. In some embodiments, a neural network includes one output node. In some embodiments, a neural network includes multiple output nodes. In general, the output is a prediction, such as a probability or likelihood of a state of interest, such as a cancer state, a binary decision (e.g., presence or absence, a positive or negative result, identification of a somatic or germline mutation, etc.), and / or a label (e.g., a classification). In the case of a single-class classification model, the output can be the likelihood of an input data set (e.g., of a biological sample and / or subject) having a condition (e.g., a label or class). In the case of a multi-class classification model, multiple predictions can be generated, each indicating the likelihood of the input data set for each state of interest. In some embodiments, the nodes are associated with parameters that contribute to the output of the neural network, which are determined based on an activation function. In some embodiments, the nodes are initialized with arbitrary parameters (e.g., randomized weights). In some alternative embodiments, the nodes are initialized with a predefined set of parameters.
[0074] As used herein, the term "normalize" refers to converting a value or set of values into a common reference frame for comparison purposes. For example, when a diagnostic ctDNA level is "normalized" with a baseline ctDNA level, the diagnostic ctDNA level may be compared to the baseline ctDNA level to determine the amount by which the diagnostic ctDNA level differs from the baseline ctDNA level.
[0075] As used interchangeably herein, the terms "nucleic acid" and "nucleic acid molecule" refer to nucleic acids of any compositional form, such as deoxyribonucleic acid (DNA, e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), ribonucleic acid (RNA, e.g., message RNA (mRNA), short inhibitory RNA (siRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA, RNA highly expressed by the fetus or placenta, etc.), and / or DNA or RNA analogs (e.g., including base analogs, sugar analogs, and / or non-natural backbones, etc.), RNA / DNA hybrids, and polyamide nucleic acids (PNAs), all of which may be in single-stranded or double-stranded form. Unless otherwise limited, nucleic acids may include known analogs of natural nucleotides, some of which may function in a manner similar to naturally occurring nucleotides. Nucleic acids may be in any form useful for carrying out the processes herein (e.g., linear, circular, supercoiled, single-stranded, double-stranded, etc.). The nucleic acid in some embodiments may be derived from a single chromosome or a fragment thereof (e.g., a nucleic acid sample may be derived from one chromosome of a sample obtained from a diploid organism). In some embodiments, the nucleic acid comprises a nucleosome, a fragment or portion of a nucleosome, or a nucleosome-like structure. The nucleic acid may comprise proteins (e.g., histones, DNA-binding proteins, etc.). The nucleic acids analyzed by the processes described herein may be substantially isolated and substantially free of association with proteins or other molecules. Nucleic acids also include single-stranded polynucleotides ("sense" or "antisense", "plus" or "minus" strands, "forward" or "reverse" reading frames) and derivatives, variants, and analogs of RNA or DNA synthesized, replicated, or amplified from double-stranded polynucleotides. Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. In the case of RNA, the base cytosine is substituted with uracil, and the sugar 2' position comprises a hydroxyl moiety. Nucleic acid obtained from a subject may be used as a template to prepare the nucleic acid.
[0076] As used herein, the term "nucleic acid fragment sequence" or "nucleic acid fragment" refers to all or part of a polynucleotide sequence of at least three consecutive nucleotides. In the context of sequencing nucleic acid molecules found in a biological sample, the term "nucleic acid fragment sequence" refers to the sequence or a representation thereof (e.g., an electronic representation of a sequence) of a nucleic acid fragment (e.g., a nucleic acid molecule fragment) found in a biological sample. To determine the sequence of a nucleic acid fragment, sequencing data (e.g., raw or corrected sequence reads from whole genome sequencing, targeted sequencing, whole genome bisulfite sequencing, targeted methylation sequencing, etc.) from a unique nucleic acid fragment (e.g., a cell-free nucleic acid molecule) is used. Thus, such sequence reads, which may actually be obtained from sequencing PCR copies of an original nucleic acid fragment, "represent" or "support" that nucleic acid fragment sequence. There may be multiple sequence reads (e.g., PCR copies) each representing or supporting a particular nucleic acid fragment in a biological sample, but there may be one nucleic acid fragment sequence for a particular nucleic acid fragment. In some embodiments, duplicate sequence reads generated for the original nucleic acid fragments are combined or removed (e.g., collapsed into a single sequence, e.g., nucleic acid fragment sequence). Thus, when determining a metric associated with a population of nucleic acid fragments in a sample, each of which encompasses a particular locus (e.g., a metric based on the abundance value of the locus or a feature of the distribution of fragment lengths), the metric can be determined using the nucleic acid fragment sequence of the population of nucleic acid fragments, rather than the supporting sequence reads (which may be generated, e.g., from PCR copies of the nucleic acid fragments in the population). This is because, in such embodiments, one copy of the sequence is used to represent an original (e.g., unique) nucleic acid fragment (e.g., a unique nucleic acid molecule fragment). Note that the nucleic acid fragment sequence of a population of nucleic acid fragments may include several identical sequences, each of which represents a different original nucleic acid fragment, rather than a copy of the same original nucleic acid fragment. In some embodiments, the cell-free nucleic acid is referred to as a nucleic acid fragment.
[0077] As used herein, the term "positive predictive value," "PPV," or "accuracy" refers to the likelihood that an output (e.g., a variant classification) will be correctly called by a predictive algorithm. PPV can be expressed as (number of true positives) / (number of false positives + number of true positives).
[0078] As used herein, the term "reference allele" refers to a sequence of one or more nucleotides at a genomic position that is either the predominant allele represented at that genomic position within a population of a species (e.g., a "wild type" sequence) or a predefined allele within a reference genome for the species.
[0079] As disclosed herein, the term "reference genome" or "genome" refers to any known, sequenced, or characterized genome, whether partial or complete, of any organism or virus that can be used to reference an identified sequence from a subject. Exemplary reference genomes used for human subjects as well as many other organisms are provided in the online genome browsers provided by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz (UCSC). "Genome" refers to the complete genetic information of an organism or virus, represented by nucleic acid sequences. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. A reference genome can be considered a representative example of a gene set for a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).
[0080] The terms "sequence read" or "read", as used interchangeably herein, refer to a nucleotide sequence generated by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment ("single-end read") or from both ends of a nucleic acid (e.g., paired-end read, double-end read). The length of a sequence read is often related to a particular sequencing technology. For example, high-throughput methods provide sequence reads that can vary in size from tens of base pairs (bp) to hundreds of base pairs (bp). In some embodiments, the sequence reads have an average, median, or average length of about 15 bp to 900 bp long (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, the sequence reads have an average, median, or average length of about 1000 bp or more. For example, nanopore sequencing can provide sequence reads that vary in size from tens, hundreds, to thousands of base pairs. Illumina parallel sequencing can provide sequence reads with a lower degree of variability (e.g., most sequence reads are about 200 bp or less in length). A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a series of nucleotides). For example, a sequence read can correspond to a series of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, a series of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of an entire nucleic acid fragment.Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes (e.g., in hybridization arrays or capture probes), or using amplification techniques such as polymerase chain reaction (PCR) or linear or isothermal amplification using a single primer.
[0081] As disclosed herein, the terms "sequencing," "sequencing," and the like generally refer to any biochemical process that can be used to determine the order of a biological polymer, such as a nucleic acid or a protein. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as a DNA fragment.
[0082] As used herein, the term "sensitivity", "recall", or "true positive rate" (TPR) refers to the number of true positives divided by the number of true positives plus the number of false negatives. Sensitivity can characterize the ability of an assay or method to accurately identify the proportion of a population that truly has a condition. For example, sensitivity can characterize the ability of a method to accurately identify the number of subjects in a population that have cancer. In another example, sensitivity can characterize the ability of a method to accurately identify one or more markers that are indicative of cancer.
[0083] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the number of true negatives plus the number of false positives. Specificity can characterize the ability of an assay or method to accurately identify the proportion of a population that is truly free of a condition. For example, specificity can characterize the ability of a method to accurately identify the number of subjects in a population that are free of cancer. In another example, specificity characterizes the ability of a method to accurately identify one or more markers that are indicative of cancer.
[0084] As disclosed herein, the terms "subject," "reference subject," "training subject," or "test subject" refer to any living or non-living organism, including, but not limited to, a human (e.g., male, female, fetus, pregnant woman, child, etc.), a non-human animal, a plant, a bacterium, a fungus, or a protist. Any human or non-human animal may serve as a subject, including, but not limited to, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cows), equines (e.g., horses), caprines (and ovines (e.g., sheep, goats), porcines (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursidae (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The terms "subject" and "patient" are used interchangeably herein and refer to a human or non-human animal known to have or potentially have a medical condition or disorder, such as, for example, cancer. In some embodiments, a subject is male or female (e.g., man, woman, or child) of any age.
[0085] The subject from whom a sample is taken or treated by any of the methods or compositions described herein may be of any age, and may be an adult, infant, or child. Optionally, the subject, e.g., patient, may be 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 109, 101, 102 , 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99 years of age, or within a range therein (e.g., from about 2 years to about 20 years of age, from about 20 years to about 40 years of age, or from about 40 years to about 90 years of age). A particular class of subjects, e.g., patients, that can benefit from the methods of the present disclosure are subjects, e.g., patients over the age of 40.
[0086] Another particular class of subjects, e.g., patients who may benefit from the methods of the present disclosure, are pediatric patients, who may be at higher risk for chronic heart disease. Additionally, subjects, e.g., patients from whom samples are taken or treated by any of the methods or compositions described herein, may be male or female.
[0087] As used herein, the term "tissue" corresponds to a group of cells grouped together as a functional unit. More than one type of cell can be found in a single tissue. Different types of tissues can consist of different kinds of cells (e.g., liver cells, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (maternal vs. fetal) or healthy vs. tumor cells. The term "tissue" can generally refer to any group of cells found in the human body (e.g., cardiac tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some embodiments, the term "tissue" or "tissue type" can be used to refer to the tissue from which the cell-free nucleic acid is derived. In one example, the viral nucleic acid fragments can be derived from blood tissue. In another example, the viral nucleic acid fragments can be derived from tumor tissue.
[0088] As used herein, the term "tumor mutation burden" (TMB) refers to a measure of mutations in a cancer per unit of a patient's genome (e.g., a measure of mutations carried by tumor cells). For example, tumor mutation burden can be expressed as a measure of central tendency (e.g., average) of the number of somatic mutations per million base pairs in the genome. In some embodiments, tumor mutation burden refers to a measure of one or more types of possible mutations, e.g., one or more of SNVs, MNVs, indels, or genomic rearrangements. In some embodiments, tumor mutation burden refers to a subset of one or more types of possible mutations, such as nonsynonymous mutations (e.g., mutations that change the amino acid sequence of an encoded protein). In other embodiments, for example, tumor mutation burden refers to the number of one or more types of mutations that occur in a protein coding sequence (e.g., whether or not they change the amino acid sequence of an encoded protein). As an example, in some embodiments, tumor mutation burden is calculated by dividing the number of mutations (e.g., all mutations and / or nonsynonymous mutations) identified in the sequencing data by the size of the capture probe panel used for targeted sequencing (e.g., the size in megabases of the electronic file). Other methods for calculating tumor mutation burden in liquid biopsy and / or solid tissue samples are known in the art.
[0089] As used herein, the term "tumor fraction" refers to the fraction of nucleic acid molecules in a sample derived from a subject's cancerous tissue, rather than from a non-cancerous tissue (e.g., germline or hematopoietic tissue). Tumor fraction can be measured using solid tissue samples or liquid biopsy samples. For example, as used herein, the term "circulating tumor fraction" refers to the fraction of acellular nucleic acid molecules in a liquid biopsy sample derived from a subject's cancerous tissue, rather than from a non-cancerous tissue. However, estimating tumor fraction from liquid biopsy samples can be difficult, as such samples generally have a lower tumor fraction compared to solid tumor samples, and the target panels used for liquid biopsy sequencing are typically small.
[0090] Software packages for calculating tumor fraction include, for example, PureCN, which is designed to estimate tumor purity from targeted short-read sequencing data of solid tumor samples, and FACETS, which is designed to estimate tumor fraction from sequencing data of solid tumor samples.In addition, the ichorCNA package applies a probabilistic model to normalized read coverage from ultra-low-range whole genome sequencing data of cell-free DNA to estimate tumor fraction in liquid biopsy samples.Tumor fraction can also be determined using a maximum likelihood model based on the copy number of alleles in the sample and the mutant allele frequency in paired control samples.
[0091] Methods for determining tumor fraction and tumor mutation burden are described in further detail in U.S. Patent Application No. 17 / 185885, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 25, 2021, and PCT Application No. PCT / US2021 / 019746, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 2021, each of which is incorporated by reference in its entirety herein.
[0092] As used herein, the term "untrained classifier" refers to a classifier that is not trained on a target data set. For example, consider the case of a first canonical set of methylation state vectors and a second canonical set of methylation state vectors described below. Each canonical set of methylation state vectors, together with each referent cell source represented by the first canonical set of methylation state vectors (hereinafter "main training data set"), is applied as a collective input to the untrained classifier to train the untrained classifier on the cell source, thereby obtaining a trained classifier. Furthermore, it will be understood that the term "untrained classifier" does not exclude the possibility that transfer learning techniques are used in such training of the untrained classifier. When transfer learning is used, the untrained classifier described above is provided with additional data in addition to the data of the main training data set. That is, in a non-limiting example of an embodiment of transfer learning, an untrained classifier receives (i) a canonical set of methylation state vectors and a cell source label of each of the reference objects represented by the canonical set of methylation state vectors ("main training dataset"), and (ii) additional data. Typically, this additional data is in the form of coefficients (e.g., regression coefficients) learned from another auxiliary training dataset. Furthermore, although a description of a single auxiliary training dataset is disclosed, it will be understood that there is no limit to the number of auxiliary training datasets that may be used to complement the main training dataset in training an untrained classifier in this disclosure. For example, in some embodiments, two or more auxiliary training datasets, three or more auxiliary training datasets, four or more auxiliary training datasets, or five or more auxiliary training datasets are used to complement the main training dataset through transfer learning, each such auxiliary dataset being different from the main training dataset. In such an embodiment, any method of transfer learning can be used. For example, consider the case where, in addition to the main training dataset, there is a first auxiliary training dataset and a second auxiliary training dataset.The coefficients learned from the first auxiliary training dataset (by applying a classifier such as a regression to the first auxiliary training dataset) can be applied to the second auxiliary training dataset using transfer learning techniques (e.g., the two-dimensional matrix multiplication described above), resulting in a trained intermediate classifier whose coefficients are applied to the main training dataset, which, together with the main training dataset itself, is applied to the untrained classifier. Alternatively, a first set of coefficients learned from the first auxiliary training dataset (by applying a classifier such as a regression to the first auxiliary training dataset) and a second set of coefficients learned from the second auxiliary training dataset (by applying a classifier such as a regression to the second auxiliary training dataset) may each be applied separately (e.g., by separate and independent matrix multiplications) to separate instances of the main training dataset, and both such applications of these coefficients to separate instances of the main training dataset may then be applied to the untrained classifier, together with the main training dataset itself (or some reduced form of the main training dataset, such as principal components or regression coefficients learned from the main training set), to train the untrained classifier. In both examples, knowledge of cell source (e.g., cancer type, etc.) derived from the first and second auxiliary training datasets is used in conjunction with the cell source-labeled main training dataset to train an untrained classifier.
[0093] As used herein, the term "mutation" or "mutation" refers to a detectable change in the genetic material of one or more cells. Mutation or mutation can refer to various types of changes in the genetic material of a cell, including changes in the primary genomic sequence at a single nucleotide or multiple nucleotide positions, such as single nucleotide variants (SNVs), multi-nucleotide variants (MNVs), indels (e.g., insertions or deletions of nucleotides), DNA rearrangements (e.g., inversions or translocations of a portion of a chromosome or a chromosome), copy number variations (CNVs) of a locus (e.g., an exon, a large span of a gene or chromosome), partial or complete changes in the ploidy of a cell, and / or changes in the epigenetic information of the genome, such as changes in DNA methylation patterns. For example, a single nucleotide variant or "SNV" refers to the substitution of one nucleotide with a different nucleotide at a position (e.g., site) of a nucleotide sequence, e.g., a sequence read from an individual. A substitution of a first nucleobase X with a second nucleobase Y can be represented as "X>Y". For example, a cytosine to thymine SNV may be represented as "C>T". In some embodiments, a mutation is a change in the genetic information of a cell relative to one or more "normal" or "reference" alleles found in a particular reference genome or population of a species of interest. In some embodiments, a mutation is a change in the genetic information of a cell compared to a reference cell or tissue, such as a "normal" or "healthy" tissue of a subject. In some embodiments, a mutation is a germline mutation or a somatic mutation.
[0094] In some instances, the mutation refers to a cancer metric derived from nucleic acid sequencing data. In some instances, the mutation refers to tumor mutational burden, microsatellite instability (MSI) status, ploidy, or tumor fraction. In some instances, the mutation refers to a fusion, amplification, and / or isoform.
[0095] As used herein, the term "mutant allele" refers to a sequence of one or more nucleotides at a genomic location that is not the predominant allele represented at that genomic location within a population of a species (e.g., is not a "wild type" sequence) or is not a predefined allele in a reference genome for the species.
[0096] As used herein, the term "parameter" refers to any coefficient, or similarly any value (e.g., weights and / or hyperparameters) of an internal or external element of a model, classifier, or algorithm that can affect (e.g., modify, adapt, and / or tune) one or more inputs, outputs, and / or functions of the model, classifier, or algorithm. For example, in some embodiments, a parameter refers to any coefficient, weight, and / or hyperparameter that can be used to control, modify, adapt, and / or tune the behavior, learning, and / or performance of a model. In some embodiments, a parameter has a fixed value. In some embodiments, the value of a parameter is manually and / or automatically adjustable. In some embodiments, the value of a parameter is modified by a classifier validation and / or training process (e.g., by error minimization and / or backpropagation, as described elsewhere herein).
[0097] Some aspects are described below with reference to illustrative example applications. It should be understood that numerous specific details, relationships, and methods are described to provide a thorough understanding of the features described herein. However, one skilled in the art will readily recognize that the features described herein can be practiced without one or more of the specific details, or by using other methods. The features described herein are not limited to the illustrated order of operations or events, since some operations may occur in different orders and / or simultaneously with other operations or events. Furthermore, not all illustrated operations or events may be used to implement a method in accordance with the features described herein.
[0098] Exemplary System Embodiments Details of an exemplary system are now described in relation to Figure 1. Figure 1 is a block diagram illustrating a system 100 according to some implementations. The system 100 in some implementations includes one or more processing units CPU 102 (also called processors or processing cores), one or more network interfaces 104, a user interface 106, a non-persistent memory 111, a persistent memory 112, and one or more communication buses 114 interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes called a chipset) that interconnects and controls communication between the system components. The non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, while the persistent memory 112 typically includes CD-ROM, digital versatile disk (DVD) or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid-state storage device. The persistent memory 112 optionally includes one or more storage devices located remotely from the CPU 102. The persistent memory 112, and the non-volatile memory devices within the non-persistent memory 112, include non-transitory computer-readable storage media. In some implementations, the non-persistent memory 111 or alternatively the non-transitory computer-readable storage media, possibly together with the persistent memory 112, stores the following programs, modules and data structures, or a subset thereof: optional instructions, programs, data, or information associated with an optional operating system 116, including procedures for handling various basic system services and performing hardware-dependent tasks; instructions, programs, data, or information associated with an optional network communications module (or instructions) 118 for connecting the system 100 to other devices or communications networks; instructions, programs, data, or information associated with an allele set 122 that stores an identification of a reference allele 126 (e.g., 126-1-1) and an identification of a variant allele 128 (e.g., 128-1-1) for a genomic location 124 (optionally for each genomic location in a plurality of genomic locations 124-1...124-Y); a biological sample obtained from the test subject (e.g., a sequencing dataset 130 derived from a liquid biological sample), comprising a respective set of nucleic acid fragments mapped onto genomic locations 132 (optionally a respective set of fragments for each genomic location in the plurality of genomic locations 132-1...132-Y), and a respective methylation state 136 (e.g., 136-1-1) of each nucleic acid fragment 134 (e.g., 134-1-1...134-1-N) in the set of nucleic acid fragments, and a respective sequence (e.g., 138-1-1) of a nucleic acid fragment 138; a reference subset 140 comprising each nucleic acid fragment 132 in each set of nucleic acid fragments 134 having a reference allele at a genomic location 124, each nucleic acid fragment being assigned to the reference subset using the identification of the reference allele 126 at that genomic location and each sequence 138 of the nucleic acid fragment; a mutant subset 142 comprising each nucleic acid fragment 134 in each set of nucleic acid fragments 132 having a mutant allele at a genomic location 124, each nucleic acid fragment being assigned to a mutant subset using the identification of the mutant allele 128 at that genomic location and the respective sequence 138 of the nucleic acid fragment; a classification module 144 for applying to the trained binary classifier at least (i) one or more indices of methylation state across the methylation states 136 of each nucleic acid fragment sequence in the mutation subset, and (ii) an indicia of the number of nucleic acid fragment sequences in the reference subset 140 relative to the number of nucleic acid fragment sequences in the mutation subset 142, thereby obtaining from the trained binary classifier an identification of the mutant allele at this genomic location in the test subject as a somatic mutant allele or a germline mutant allele; and Optionally, a classifier training module 146 for training a binary classifier used for identification of the variant allele at this genomic location.
[0099] In some implementations, one or more of the above elements are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for performing the functions described above. The above modules, data, or programs (e.g., sets of instructions) may not be implemented as separate software programs, procedures, data sets, or modules, and thus in various implementations, various subsets of these modules and data may be combined or otherwise reconfigured. In some implementations, the non-persistent memory 111 optionally stores a subset of the above modules and data structures. Additionally, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above elements are stored in a computer system other than that of the visualization system 100 that is addressable by the visualization system 100 such that the visualization system 100 can retrieve all or a portion of such data.
[0100] Although FIG. 1 illustrates "system 100," the diagram is intended as a functional description of various features that may be present in a computer system, rather than as a structural schematic of the implementations described herein. In practice, items shown separately may be combined and some items may be separate. Additionally, while FIG. 1 illustrates some data and modules in non-persistent memory 111, some or all of these data and modules may be in persistent memory 112.
[0101] Having disclosed a system according to the present disclosure with reference to Figure 1, a method according to the present disclosure will now be described in detail with reference to Figures 2A, 2B, and 3. Any of the methods of the present disclosure can utilize any of the assays or algorithms disclosed in U.S. Patent Application No. 15 / 793830, filed October 25, 2017, and / or International Patent Publication No. WO2018 / 081130, entitled "Methods and Systems for Tumor Detection," each of which is incorporated herein by reference, to determine a cancerous condition in a test subject or the likelihood that the subject has a cancerous condition. For example, any of the methods of the present disclosure can work in conjunction with any of the methods or algorithms disclosed in U.S. Patent Application No. 15 / 793830, filed October 25, 2017, and / or International Patent Publication No. WO2018 / 081130, entitled "Methods and Systems for Tumor Detection."
[0102] Identification of mutant alleles With reference to Figures 2A and 2B, provided herein is a method 200 for identifying a mutant allele at a test genomic location as a somatic mutant allele or a germline mutant allele.
[0103] Subjects and samples In some embodiments, the test subject is a mammal. In some embodiments, the test subject is a human. In some embodiments, the test subject is a patient with cancer.
[0104] In some embodiments, the method includes obtaining a biological sample from a test subject. In some embodiments, the biological sample is one of multiple biological samples obtained from the test subject (e.g., multiple replicates and / or multiple samples including matched tumor samples and matched normal samples). In some embodiments, the multiple biological samples are obtained from the test subject simultaneously or spaced apart over a period of time (e.g., for serial analysis). For example, in some such embodiments, the period between obtaining biological samples from the test subject is at least 1 day, at least 2 days, at least 1 week, at least 2 weeks, at least 1 month, at least 2 months, at least 3 months, at least 4 months, at least 6 months, or at least 1 year.
[0105] In some embodiments, a biological sample is obtained from any tissue, organ or bodily fluid from a subject.
[0106] In some embodiments, the biological sample is a liquid biological sample (e.g., a liquid biopsy sample). In some embodiments, the liquid biological sample comprises blood, whole blood, plasma, serous fluid, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the test subject. In some embodiments, the liquid biological sample consists of blood, whole blood, plasma, serous fluid, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the test subject.
[0107] In some embodiments, the biological sample is a tissue sample. In some embodiments, the tissue sample is a tumor sample from the test subject. In some embodiments, the tumor sample is a tumor sample from an allogeneic tumor. In some embodiments, the tumor sample is a tumor sample from a xenogeneic tumor.
[0108] In some embodiments, the biological sample comprises a plurality of nucleic acid fragments. In some embodiments, each of the plurality of nucleic acid fragments comprises cell-free nucleic acid fragments (e.g., cfDNA). In some embodiments, each of the plurality of nucleic acid fragments comprises cell-free nucleic acid fragments (e.g., cfDNA). In some embodiments, the nucleic acid fragments in the plurality of nucleic acid fragments comprise any of the embodiments of nucleic acid disclosed herein (see, e.g., definition: nucleic acid).
[0109] In some embodiments, the biological sample comprises a mixture of nucleic acid molecules derived from diseased cells and healthy cells. For example, in some embodiments, the biological sample is a blood sample comprising cfDNA (e.g., ctDNA) derived from tumor cells, cfDNA derived from normal cells, and / or normal cells (e.g., white blood cells).
[0110] In some embodiments, the biological sample is processed to extract nucleic acids in preparation for sequencing analysis. As a non-limiting example, in some embodiments, cell-free nucleic acid fragments are extracted from a liquid biological sample (e.g., a blood sample) collected from a subject in a K2 EDTA tube. If the biological sample is blood, as a non-limiting example, the sample is processed within 2 hours of collection, first by a double spin at 1000g for 10 minutes of the biological sample, then a spin at 2000g for 10 minutes of the obtained plasma. The plasma is then stored in 1 mL aliquots at -80°C. In this way, an appropriate amount of plasma (e.g., 1-5 mL) is prepared from the biological sample for the purpose of cell-free nucleic acid extraction. In some embodiments, cell-free nucleic acid is extracted using a QIAamp Circulating Nucleic Acid kit (Qiagen) and eluted in DNA suspension buffer (Sigma). In some embodiments, the purified cell-free nucleic acid is stored at -20°C until use.
[0111] Other equivalent methods can be used to prepare and / or extract nucleic acid fragments (e.g., cell-free nucleic acid fragments) from biological samples for sequencing purposes, and all such methods are within the scope of the present disclosure.
[0112] In some embodiments, each plurality of nucleic acid fragments (e.g., cell free nucleic acid fragments) from a test subject includes 100 or more nucleic acid fragments, 1,000 or more nucleic acid fragments, 10,000 or more nucleic acid fragments, 20,000 or more nucleic acid fragments, 50,000 or more nucleic acid fragments, 100,000 or more nucleic acid fragments, 200,000 or more nucleic acid fragments, 500,000 or more nucleic acid fragments, 1,000,000 or more nucleic acid fragments, 2,000,000 or more nucleic acid fragments, 5,000,000 or more nucleic acid fragments, 10,000,000 or more nucleic acid fragments, or 50,000,000 or more nucleic acid fragments. In some embodiments, the nucleic acid fragments (e.g., cell free nucleic acid fragments) from the test subject include no more than 50,000,000, no more than 10,000,000, no more than 5,000,000, no more than 2,000,000, no more than 1,000,000, no more than 500,000, no more than 200,000, no more than 100,000, no more than 50,000, no more than 20,000, no more than 10,000, or no more than 1,000 nucleic acid fragments. In some embodiments, the nucleic acid fragments (e.g., cell-free nucleic acid fragments) from the test subject include 100-1,000, 1,000-10,000, 10,000-100,000, 100,000-1,000,000, 1,000,000-10,000,000, or 10,000,000-50,000,000 nucleic acid fragments. In some embodiments, the nucleic acid fragments (e.g., cell-free nucleic acid fragments) from the test subject fall within another range beginning with 100 or more nucleic acid fragments and ending with 50,000,000 or less nucleic acid fragments.
[0113] In some embodiments, the nucleic acid fragments obtained from the biological sample are cell-free nucleic acids derived from tumor cells (e.g., ctDNA). In some embodiments, the nucleic acid fragments obtained from the biological sample are cell-free nucleic acids derived from normal cells. In some embodiments, the nucleic acid fragments obtained from the biological sample are obtained directly from tumor cells (e.g., solid tumor biopsy). In some embodiments, the nucleic acid fragments obtained from the biological sample are obtained directly from normal cells (e.g., healthy tissue and / or white blood cells).
[0114] In some embodiments, the nucleic acid fragments obtained from the biological sample are any form of nucleic acid as defined herein (e.g., cell-free nucleic acid fragments), or combinations thereof (see, e.g., Definition: Nucleic Acid). For example, in some embodiments, the nucleic acid obtained from the biological sample is a mixture of RNA and DNA (e.g., cell-free RNA and / or cell-free DNA).
[0115] In some embodiments, the method comprises obtaining a plurality of nucleic acid fragment sequences by sequencing a plurality of nucleic acid molecules in a biological sample obtained from a test subject. For example, in some embodiments, the biological sample is a liquid biological sample, and each nucleic acid fragment sequence in each plurality of nucleic acid fragment sequences represents all or a portion of each cell-free nucleic acid molecule in a population of cell-free nucleic acid molecules in the liquid biological sample. In some embodiments, alternatively or additionally, the biological sample is a tissue sample, and each nucleic acid fragment sequence in each plurality of nucleic acid fragment sequences represents all or a portion of each nucleic acid molecule in a population of nucleic acid molecules in the tissue sample. Non-limiting embodiments of the method for obtaining nucleic acid fragment sequences are detailed in the following section (see "Obtaining Nucleic Acid Fragment Sequences").
[0116] Reference and variant alleles Referring to blocks 202 and 204, the method further includes obtaining an identification of a reference allele at the genomic location and obtaining an identification of a variant allele at the genomic location.
[0117] In some embodiments, the variant allele is an insertion, deletion, single nucleotide variation (SNV) or single nucleotide polymorphism (SNP). In some embodiments, the variant allele is any variation or mutation as defined herein (see definition: mutation).
[0118] In some embodiments, the genomic location is any genomic location or locus as defined herein (see definition: genomic location). For example, in some embodiments, the genomic location is a single base position and the mutation is a single nucleotide variation (SNV) or single nucleotide polymorphism (SNP). In some embodiments, the genomic location is two or more base positions and the mutation is an insertion or deletion. In some embodiments, the genomic location is a portion or region of a reference genome.
[0119] In some embodiments, the genomic location is associated with a clinically actionable mutation. For example, in some embodiments, the genomic location indicates a genomic mutation associated with an increased risk of a cancer condition, such as increased severity, likelihood of progression, and / or an indication of the type of cancer (e.g., KRAS mutation in lung cancer). In some such embodiments, the presence and / or identification of the respective genomic mutation may affect clinical decision-making, such as treatment recommendations, clinical trial enrollment, and other physician actions. In some embodiments, the clinically actionable mutation is a somatic mutation or a germline mutation. In some embodiments, the clinically actionable mutation is associated with a gene.
[0120] In some embodiments, the genomic location includes all or part of a gene or is characterized by a mutation in the gene. In some embodiments, the gene is, for example, a cancer gene, where malfunction of the gene is associated with cancer. Non-limiting examples of malfunction include genomic alterations (e.g., mutations and / or mutant alleles), dysregulation, altered activity, altered expression, and / or altered epigenetic modifications such as methylation. In some embodiments, the cancer gene includes known cancer genes, candidate cancer genes, cancer genes, tumor suppressor genes, and / or tissue-specific genes (e.g., genes associated with a particular cancer type). In some embodiments, the cancer gene is obtained based on annotations from sequencing screens, manual curation by experts, and / or experimental data. In some embodiments, the cancer genes are obtained from databases such as the Network of Cancer Genes (NCG), the International Cancer Genome Consortium (ICGC), the Cancer Genome Atlas (TCGA), COSMIC, DoCM, DriverDB, Cancer Genome Interpreter, OncoKB, cBIOPortal, Cancer Gene Census (CGC), ONGene, TSGene, and / or CoReCG.
[0121] These are A1CF, ABI1, ABL1, ABL2, ACKR3, ACSL3, ACSL6, and ACVR1. ACVR1B, ACVR2A, AFDN, AFF1, AFF3, AFF4, AKAP9, AKT1, AKT2, AKT3, ALDH2, AL K, AMER1, ANK1, APC, APOBEC3B, AR, ARAF, ARHGAP26, ARHGAP5, ARHGEF10, AR HGEF10L, ARHGEF12, ARID1A, ARID1B, ARID2, ARNT, ASPSCR1, ASXL1, ASXL2, A TF1, ATIC, ATM, ATP1A1, ATP2B3, ATR, ATRX, AXIN1, AXIN2, B2M, BAP1, BARD1 、BAX、BAZ1A、BCL10、BCL11A、BCL11B、BCL2、BCL2L12、BCL3、BCL6、BCL7A、BCL 9. BCL9L, BCLAF1, BCOR, BCORL1, BCR, BIRC3, BIRC6, BLM, BMP5, BMPR1A, BRA F BRCA1 BRCA2 BRD3 BRD4 BRIP1 BTG1 BTK BUB1B C15orf65 CACNA1DC ALR, CAMTA1, CANT1, CARD11, CARS, CASP3, CASP8, CASP9, CBFA2T3, CBFB, CB L, CBLB, CBLC, CCDC6, CCNB1IP1, CCNC, CCND1, CCND2, CCND3, CCNE1, CCR4, CC R7, CD209, CD274, CD28, CD74, CD79A, CD79B, CDC73, CDH1, CDH10, CDH11, CD H17, CDK12, CDK4, CDK6, CDKN1A, CDKN1B, CDKN2A, CDKN2C, CDX2, CEBPA, CEP8 9 CHCHD7, CHD2, CHD4, CHEK2, CHIC2, CHST11, CIC, CIITA, CLIP1, CLP1, CLT C. CLTCL1, CNBD1, CNBP, CNOT3, CNTNAP2, CNTRL, COL1A1, COL2A1, COL3A1, CO X6C, CPEB3, CREB1, CREB3L1, CREB3L2, CREBBP, CRLF2, CRNKL1, CRTC1, CRTC 3. CSF1R, CSF3R, CSMD3, CTCF, CTNNA2, CTNNB1, CTNND1, CTNND2, CUL3, CUX1CXCR4, CYLD, CYP2C8, CYSLTR2, DAXX, DCAF12L2, DCC, DCTN1, DDB2, DDIT3, D DR2, DDX10, DDX3X, DDX5, DDX6, DEK, DGCR8, DICER1, DNAJB1, DNM2, DNMT1, D NMT3A, DROSHA, EBF1, ECT2L, EED, EGFR, EIF1AX, EIF3E, EIF4A2, ELF3, ELF4 ELK4, ELL, ELN, EML4, EP300, EPAS1, EPHA3, EPHA7, EPS15, ERBB2, ERBB3, ER BB4, ERC1, ERCC2, ERCC3, ERCC4, ERG, ESR1, ETNK1, ETV1, ETV4, ETV5, ETV6 EWSR1, EXT1, EXT2, EZH2, EZR, FAM131B, FAM135B, FAM46C, FAM47C, FANCA, FA NCC, FANCD2, FANCE, FANCF, FANCG, FAS, FAT1, FAT3, FAT4, FBLN2, FBXO11, F BXW7, FCGR2B, FCRL4, FEN1, FES, FEV, FGFR1, FGFR1OP, FGFR2, FGFR3, FGFR4 FH, FHIT, FIP1L1, FKBP9, FLCN, FLI1, FLNA, FLT3, FLT4, FNBP1, FOXA1, FOXL 2. FOXO1, FOXO3, FOXO4, FOXP1, FOXR1, FSTL3, FUBP1, FUS, GAS7, GATA1, GATA 2. GATA3, GLI1, GMPS, GNA11, GNAQ, GNAS, GOLGA5, GOPC, GPC3, GPC5, GPHN, G RIN2A, GRM3, H3F3A, H3F3B, HERPUD1, HEY1, HIF1A, HIP1, HIST1H3B, HIST1H4 I、HLA-A、HLF、HMGA1、HMGA2、HNF1A、HNRNPA2B1、HOOK3、HOXA11、HOXA13、HO XA9, HOXC11, HOXC13, HOXD11, HOXD13, HRAS, HSP90AA1, HSP90AB1, ID3, IDH1 IDH2, IGF2BP2, IKBKB, IKZF1, IL2, IL21R, IL6ST, IL7R, IRF4, IRS4, ISX, I TGAV, ITK, JAK1, JAK2, JAK3, JAZF1, JUN, KAT6A, KAT6B, KAT7, KCNJ5, KDM5A.KDM5C, KDM6A, KDR, KDSR, KEAP1, KIAA1549, KIF5B, KIT, KLF4, KLF6, KLK2, K MT2A, KMT2C, KMT2D, KNL1, KNSTRN, KRAS, KTN1, LARP4B, LASP1, LCK, LCP1, L EF1, LEPROTL1, LHFPL6, LIFR, LMNA, LMO1, LMO2, LPP, LRIG3, LRP1B, LSM14A LYL1, LZTR1, MAF, MAFB, MALT1, MAML2, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MA P3K13, MAPK1, MAX, MB21D2, MDM2, MDM4, MDS2, MECOM, MED12, MEN1, MET, MGM T, MITF, MKL1, MLF1, MLH1, MLLT1, MLLT10, MLLT11, MLLT3, MLLT6, MN1, MNX1 MPL, MSH2, MSH6, MSI2, MSN, MTCP1, MTOR, MUC1, MUC16, MUC4, MUTYH, MYB, MY C, MYCL, MYCN, MYD88, MYH11, MYH9, MYO5A, MYOD1, N4BP2, NAB2, NACA, NBEA, N BN, NCKIPSD, NCOA1, NCOA2, NCOA4, NCOR1, NCOR2, NDRG1, NF1, NF2, NFATC2 NFE2L2, NFIB, NFKB2, NFKBIE, NIN, NKX2-1, NONO, NOTCH1, NOTCH2, NPM1, NR4 A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1, NTRK1, NTRK3, NUMA1, NUP2 14. NUP98, NUTM1, NUTM2A, NUTM2B, OLIG2, OMD, P2RY8, PABPC1, PAFAH1B2, PA LB2, PATZ1, PAX3, PAX5, PAX7, PAX8, PBRM1, PBX1, PCBP1, PCM1, PDCD1LG2 DGFB, PDGFRA, PDGFRB, PER1, PHF6, PHOX2B, PICALM, PIK3CA, PIK3CB, PIK3R1 、PIM1、PLAG1、PLCG1、PML、PMS1、PMS2、POLD1、POLE、POLG、POLQ、POT1、POU2 AF1, POU5F1, PPARG, PPFIBP1, PPM1D, PPP2R1A, PPP6C, PRCC, PRDM1, PRDM16.<h2 style=";text-align:left;direction:ltr">PRDM2、PREX2、PRF1、PRKACA、PRKAR1A、PRKCB、PRPF40B、PRRX1、PSIP1、PTCH 1、PTEN、PTK6、PTPN11、PTPN13、PTPN6、PTPRB、PTPRC、PTPRD、PTPRK、PTPRT、 PWWP2A、QKI、RABEP1、RAC1、RAD17、RAD21、RAD51B、RAF1、RALGDS、RANBP2、R AP1GDS1、RARA、RB1、RBM10、RBM15、RECQL4、REL、RET、RFWD3、RGPD3、RGS7、RH OA、RHOH、RMI2、RNF213、RNF43、ROBO2、ROS1、RPL10、RPL22、RPL5、RPN1、RSP O2、RSPO3、RUNX1、RUNX1T1、S100A7、SALL4、SBDS、SDC4、SDHA、SDHAF2、SDHB、 SDHC、SDHD、SEPT5、SEPT6、SEPT9、SET、SETBP1、SETD1B、SETD2、SF3B1、SFPQ 、SFRP4、SGK1、SH2B3、SH3GL1、SHTN1、SIRPA、SIX1、SIX2、SKI、SLC34A2、SLC4 5A3、SMAD2、SMAD3、SMAD4、SMARCA4、SMARCB1、SMARCD1、SMARCE1、SMC1A、SM O、SND1、SNX29、SOCS1、SOX2、SOX21、SOX9、SPECC1、SPEN、SPOP、SRC、SRGAP3 、SRSF2、SRSF3、SS18、SS18L1、SSX1、SSX2、SSX4、STAG1、STAG2、STAT3、STAT 5B、STAT6、STIL、STK11、STRN、SUFU、SUZ12、SYK、TAF15、TAL1、TAL2、TBL1XR1 、TBX3、TCEA1、TCF12、TCF3、TCF7L2、TCL1A、TEC、TERT、TET1、TET2、TFE3、TF EB、TFG、TFPT、TFRC、TGFBR2、THRAP3、TLX1、TLX3、TMEM127、TMPRSS2、TNC、TN FAIP3、TNFRSF14、TNFRSF17、TOP1、TP53、TP63、TPM3、TPM4、TPR、TRAF7、TRI M24、TRIM27、TRIM33、TRIP11、TRRAP、TSC1、TSC2、TSHR、U2AF1、UBR5、USP44、Selected from USP6, USP8, VAV1, VHL, VTI1A, WAS, WDCP, WIF1, WNK2, WRN, WT1, WWTR1, XPA, XPC, XPO1, YWHAE, ZBTB16, ZCCHC8, ZEB1, ZFHX3, ZMYM2, ZMYM3, ZNF331, ZNF384, ZNF429, ZNF479, ZNF521, ZNRF3, and ZRSR2.
[0122] Cancer genes are further described in Repana et al., 2019, "The Network of Cancer Genes (NCG): a comprehensive catalogue of known and candidate cancer genes from cancer sequencing screens," Genome Biology 20:1, doi:10.1186 / s13059-018-1612-0, which is incorporated by reference in its entirety.
[0123] In some embodiments, the genomic location is selected from a plurality of genomic locations. For example, in some embodiments, the systems and methods disclosed herein can be used to identify a plurality of mutant alleles at a corresponding plurality of genomic locations as somatic mutant alleles or germline mutant alleles. In some embodiments, the plurality of genomic locations includes at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, or at least 20,000 genomic locations. In some embodiments, the plurality of genomic locations comprises 20,000 or less, 10,000 or less, 5000 or less, 4000 or less, 3000 or less, 2000 or less, 1000 or less, 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, 200 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, or 20 or less genomic locations. In some embodiments, the plurality of genomic locations is 10-50, 50-100, 100-500, 500-1000, 1000-5000, 5000-10,000, or 10,000-20,000 genomic locations. In some embodiments, the plurality of genomic locations is in another range beginning with 10 or more genomic locations and ending with 20,000 or fewer genomic locations.
[0124] In some embodiments, each genomic location in the plurality of genomic locations is associated with a respective clinically actionable mutation (e.g., a cancer gene). In some embodiments, each genomic location in the plurality of genomic locations is associated with a respective clinically actionable mutation (e.g., a cancer gene). In some embodiments, the plurality of genomic locations is a panel of clinically actionable mutations (e.g., a cancer gene of interest).
[0125] Mutation calling Referring again to blocks 202 and 204, in some embodiments, the identification of the reference allele at the genomic location is obtained from a reference genome, which may include any of the embodiments disclosed herein (see definition: reference genome).
[0126] In some embodiments, obtaining an identification of a variant allele at a genomic location comprises determining that each of the plurality of nucleic acid fragments supports a variant allele call at the genomic location.
[0127] For example, in some embodiments, obtaining the identification of a mutant allele at a genomic location is performed by a method that determines the likelihood that the genomic location has each genotype in a plurality of candidate genotypes from a plurality of nucleic acid fragments. Selection of each genotype from the plurality of candidate genotypes can be determined based on a comparison of the calculated likelihoods (e.g., by ranking the genotypes by their corresponding likelihoods and / or by applying a likelihood threshold to the estimated likelihoods). In general, the mutant allele can be identified as the candidate genotype that is most likely not to be a reference genotype (e.g., a reference allele obtained from a reference genome). In some embodiments, the reference genotype of the genomic location is homozygous (e.g., A / A, T / T, G / G, C / C).
[0128] In some embodiments, obtaining the identification of the variant allele at the genomic location is performed using a Bayesian likelihood model (e.g., variant calling). An exemplary method 320 for variant calling in a test subject can be described with reference to FIG.
[0129] Referring to block 328, in some embodiments, the method 320 for variant calling is performed by deriving a prior probability (e.g., in an electronic format) of each genotype at a genomic location for each candidate genotype in the set of candidate genotypes using nucleic acid data obtained from a reference population (e.g., a population of multiple reference subjects of a given species (e.g., human)). In some embodiments, the reference population includes at least 100 reference subjects. In some embodiments, the reference population includes at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 reference subjects.
[0130] In some embodiments, each candidate genotype in the set of genotypes is of the form X / Y, where X is the identity of a base in the set of bases {A,C,T,G} at a genomic location of the reference genome, and Y is the identity of a base in the set of bases {A,C,T,G} at a genomic location being tested. In other words, in some embodiments, each candidate genotype in the set of genotypes represents a respective diploid genotype, where the paternal and maternal alleles at a genomic location are denoted by X and Y, respectively.
[0131] At the single nucleotide level, in some embodiments there are 10 possible genotypes for each autosomal position. In some embodiments, the set of candidate genotypes consists of 2-10 genotypes in the set {A / A, A / C, A / G, A / T, C / C, C / G, C / T, G / G, G / T, and T / T}. In some embodiments, the set of candidate genotypes includes at least 2, 4, 5, 6, 7, 8, or 9 genotypes in the set {A / A, A / C, A / G, A / T, C / C, C / G, C / T, G / G, G / T, and T / T}. In some embodiments, the set of candidate genotypes consists of the entire set {A / A, A / C, A / G, A / T, C / C, C / G, C / T, G / G, G / T, and T / T}.
[0132] Referring to block 334, in some embodiments, the method 320 for variant calling continues by obtaining, for a genomic location, a forward and reverse strand-specific base count set including a respective forward strand base count and a respective reverse strand base count of each base in the set {A, T, C, G} at the genomic location based on (i) the strand orientation and (ii) the identity determination of each base at the genomic location in each nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences mapped to the genomic location. For example, in some embodiments, each of the plurality of nucleic acid fragment sequences is obtained from a plurality of nucleic acid molecules in the liquid biological sample of the test subject by nucleic acid sequencing and / or methylation sequencing. Details regarding obtaining each of the plurality of nucleic acid fragment sequences and mapping the nucleic acid fragment sequences to genomic locations are further disclosed below, for example, in the section entitled "Obtaining Nucleic Acid Fragment Sequences." In some embodiments, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 50 or more, or 100 or more nucleic acid fragment sequences are mapped to genomic locations and included in the strand-specific base counts. In some embodiments, bases at genomic locations in each of the plurality of nucleic acid fragment sequences whose identity can be affected by conversion of methylated or unmethylated cytosines do not contribute to the strand-specific base count set.
[0133] In some embodiments, the forward direction is a F1R2 read (sense) orientation and the reverse direction is a F2R1 (antisense) read orientation. The pair of orientations can refer to whether the respective nucleic acid fragment sequence is derived from the 5' strand or the 3' strand of the fragment at a given genomic location. For example, the F1R2 read orientation refers to sequence reads derived from the positive (sense) strand of the nucleic acid fragment and the F2R1 read orientation refers to sequence reads derived from the negative (antisense) strand of the nucleic acid fragment. In some embodiments, the forward direction is a F1R2 or R2F1 read (sense) orientation and the reverse direction is a F2R1 or R1F2 (antisense) read orientation.
[0134] In some embodiments, a strand-specific base count set is used to account for bisulfite conversion. Methylation sequencing can inherently provide strand-specific chemistry that affects the detection of C and T alleles at genomic locations. For example, bisulfite conversion results in the conversion of C to T on the forward strand of a nucleic acid fragment and the conversion of A to G on the corresponding reverse strand. Since A and G alleles are not directly affected by bisulfite conversion, the allele counts on the positive strand can be resolved, and the C and T alleles on the positive strand are identified by the A and G alleles on the negative strand. As a validation, the sum of the C and T allele counts cannot be affected by bisulfite conversion.
[0135] Referring to block 340, in some embodiments the method 320 for variant calling further includes calculating a respective forward strand conditional probability and a respective reverse strand conditional probability for each candidate genotype in the set of candidate genotypes for the genomic location using the strand-specific base count set and the sequencing error estimate, thereby calculating a plurality of forward strand conditional probabilities and a plurality of reverse strand conditional probabilities for the genomic location.
[0136] In some embodiments, the sequencing error estimate is between 0.01 and 0.0001. In some embodiments, the sequencing error estimate is less than 0.01, less than 0.009, less than 0.008, less than 0.007, less than 0.006, less than 0.005, less than 0.004, less than 0.003, less than 0.002, less than 0.001, less than 0.00075, less than 0.0005, or less than 0.0075. In some embodiments, for each candidate genotype in the set of candidate genotypes, a respective sequencing error estimate is used. In some embodiments, for each candidate genotype in the set of candidate genotypes, the same sequencing error estimate is used. In some embodiments, one or more of the candidate genotypes have a corresponding sequencing error estimate that is different from the sequencing error estimate used for the remaining candidate genotypes in the set of candidate genotypes. In some embodiments, a symmetric error estimate is assumed for each genotype. In some embodiments, the sequencing error is fixed or variable.
[0137] Referring to block 344, in some embodiments, the method 320 for variant calling further includes calculating a plurality of likelihoods for the genomic location. Each likelihood in the plurality of likelihoods is a likelihood for each candidate genotype in the set of candidate genotypes. In some embodiments, the plurality of likelihoods is calculated using a combination of (i) a respective forward strand conditional probability for each candidate genotype in the plurality of forward strand conditional probabilities, (ii) a respective reverse strand conditional probability for each candidate genotype in the plurality of reverse strand conditional probabilities, and (iii) a genotype prior probability for each candidate genotype.
[0138] In some embodiments, Bayes' theorem is used to calculate the likelihood of observing each genotype. In some embodiments, the prior likelihood of each genotype is calculated using the observed allele frequencies. In some embodiments, each candidate genotype in the set of candidate genotypes for a genomic position is ranked in order of its Bayes probability.
[0139] In some embodiments, the likelihood of each of the candidate genotypes in the set of candidate genotypes has the following form: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(G) In the formula, Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) is the conditional probability of each candidate genotype, and Pr(R AG ,R C ,R T │R ACGT , genotype,ε) is the reverse strand conditional probability for each candidate genotype, Pr(G) is the prior probability of the genotype at the genomic position for each candidate genotype, ε is the sequencing error estimate, genotype is each candidate genotype, and F A is the forward base count of base A at a genomic position across each of the multiple nucleic acid fragment sequences in the strand-specific base count set, and F G is the forward base count of the base G at the genomic position across each of the multiple nucleic acid fragment sequences in the strand-specific base count set, and F CT is the sum of (i) the forward base count of the base C and (ii) the forward base count of the base T at a genomic position across each of the multiple nucleic acid fragment sequences in the strand-specific base count set, and R C is the reverse base count of base C at a genomic position across each of the multiple nucleic acid fragment sequences in the strand-specific base count set, and R T is the reverse base count of the base T at the genomic position across each of the multiple nucleic acid fragment sequences in the strand-specific base count set, and R AGis the sum of (i) the reverse base count of the base A and (ii) the reverse base count of the base G at a genomic position across each of the multiple nucleic acid fragment sequences in the strand-specific base count set.
[0140] In some embodiments, this multiplication depends on the assumption of symmetric sequencing error estimates for each candidate genome. In some embodiments, the likelihood is the log-likelihood, which is determined by taking the logarithm of the formula defined above.
[0141] In some embodiments, each candidate genotype G is A / A, and each likelihood of A / A is: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(A / A) Calculating includes calculating:
[0142]
number
[0143] In some embodiments, each candidate genotype G is A / A, and each likelihood of A / A is: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(A / A) Computing σ includes computing the log-likelihood:
[0144]
number
[0145] In some embodiments, each candidate genotype G is A / C, and each likelihood of A / C: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(A / C) Calculating includes calculating:
[0146]
number
[0147] In some embodiments, each candidate genotype G is A / C, and each likelihood of A / C: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(A / C) Computing σ includes computing the log-likelihood:
[0148]
number
[0149] In some embodiments, each candidate genotype G is A / G, and each likelihood of A / G: Pr(F A ,F G ,F CT│F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(A / G) Calculating includes calculating:
[0150]
number
[0151] In some embodiments, each candidate genotype G is A / G, and each likelihood of A / G: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(A / G) Computing σ includes computing the log-likelihood:
[0152]
number
[0153] In some embodiments, each candidate genotype G is A / T and each likelihood of A / T: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(A / T) Calculating includes calculating:
[0154]
number
[0155] In some embodiments, each candidate genotype G is A / T and each likelihood of A / T: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(A / T) Computing σ includes computing the log-likelihood:
[0156]
number
[0157] In some embodiments, each candidate genotype G is C / C, and each likelihood of C / C: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(C / C) Calculating includes calculating:
[0158]
number
[0159] In some embodiments, each candidate genotype G is C / C, and each likelihood of C / C: Pr(FA ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(C / C) Computing σ includes computing the log-likelihood:
[0160]
number
[0161] In some embodiments, each candidate genotype G is C / G, and each likelihood of C / G: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(C / G) Calculating includes calculating:
[0162]
number
[0163] In some embodiments, each candidate genotype G is C / G, and each likelihood of C / G: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(C / G) Computing σ includes computing the log-likelihood:
[0164]
number
[0165] In some embodiments, each candidate genotype G is C / T, and each likelihood of C / T: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(C / T) Calculating includes calculating:
[0166]
number
[0167] In some embodiments, each candidate genotype G is C / T, and each likelihood of C / T: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(C / T) Computing σ includes computing the log-likelihood:
[0168]
number
[0169] In some embodiments, each candidate genotype G is G / G, and each likelihood of G / G is: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(G / G) Calculating includes calculating:
[0170]
number
[0171] In some embodiments, each candidate genotype G is G / G, and each likelihood of G / G is: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(G / G) Computing σ includes computing the log-likelihood:
[0172]
number
[0173] In some embodiments, each candidate genotype G is G / T, and each likelihood of G / T: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,RT │R ACGT ,genotype,ε) * Pr(G / T) Calculating includes calculating:
[0174]
number
[0175] In some embodiments, each candidate genotype G is G / T, and each likelihood of G / T: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(G / T) Computing σ includes computing the log-likelihood:
[0176]
number
[0177] In some embodiments, each candidate genotype G is T / T, and each likelihood of T / T: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(T / T) Calculating includes calculating:
[0178]
number
[0179] In some embodiments, each candidate genotype G is T / T, and each likelihood of T / T: Pr(F A ,F G ,F CT │F ACGT ,genotype,ε) * Pr(R AG ,R C ,R T │R ACGT ,genotype,ε) * Pr(T / T) Computing σ includes computing the log-likelihood:
[0180]
number
[0181] In some embodiments, one or more respective likelihood calculations further include the corresponding bisulfite conversion rate before accounting for the apparent discrepancy between the counts of C on the corresponding forward and reverse strands. For example, if a greater number of C bases are observed on the forward strand, this suggests that T / T is less likely than C / T for the final C / C genotype. Examples of likelihood calculations that account for bisulfite conversion rates, base quality scores, and other sequencing information are known in the art.
[0182] Referring to block 346, in some embodiments, the method 320 for variant calling further includes determining whether the multiple likelihoods (e.g., calculated in block 344) support a variant call at the genomic location. In some embodiments, this includes determining whether any likelihoods among the multiple likelihoods of any of the proposed genotypes (e.g., including reference genotypes) at the genomic location meet a variant threshold. In some embodiments, if the likelihoods of any of the proposed genotypes (e.g., including reference genotypes) at the genomic location meet the variant threshold, the variant at the genomic location is considered identified. Thus, if the likelihood of a variant allele meets the threshold among the multiple likelihoods corresponding to the multiple different variant alleles, the variant allele is called among the multiple different variant alleles. If three or more variant alleles meet the threshold, the variant allele with the highest likelihood of meeting the threshold is called. If none of the variant alleles meets the threshold, the variant allele is not called.
[0183] In some embodiments, the likelihood is expressed as a log likelihood (e.g., an unnormalized likelihood) and the mutation threshold is met when the log likelihood of the reference genotype at the genomic location is less than -10. In some embodiments, the mutation threshold is met when the log likelihood of the reference genotype at the genomic location is less than -1, less than -5, less than -10, less than -25, less than -50, or less than -100. In some embodiments, the likelihood is expressed as a log likelihood and the mutation threshold is met when the log likelihood of the reference genotype at the genomic location is between -25 and -5. In some embodiments, the likelihood is expressed as a log likelihood and the mutation threshold is met when the log likelihood of the reference genotype at the genomic location is between -10 and -1, -10 and -5, -25 and -1, -25 and -10, -25 and -15, -50 and -1, -50 and -5, -50 and -10, or -50 and -25.
[0184] In some embodiments, the method 320 further includes, once a mutation at a genomic location has been called, determining the identity of the mutation by selecting, among the set of candidate genotypes for the genomic location, the candidate genotype with the best likelihood among the multiple likelihoods as the mutation. In some embodiments, this determination may be by ranking the candidate genotypes by their corresponding likelihood or log-likelihood. In some embodiments, a single identity of the mutation is called by selecting the highest ranked genotype for the mutation. In some embodiments, at least two, at least three, or at least four identities of the mutation are called by selecting the top two, top three, or top four ranked genotypes of the mutation, respectively.
[0185] In some embodiments, method 320 further includes repeating the method for each genomic location in the plurality of genomic locations of the test subject (eg, thereby obtaining a plurality of variant calls for the test subject).
[0186] In some embodiments, the plurality of variant calls comprises 200 variant calls. In some embodiments, the plurality of variant calls comprises at least 10 variant calls, at least 20 variant calls, at least 30 variant calls, at least 40 variant calls, at least 50 variant calls, at least 60 variant calls, at least 70 variant calls, at least 80 variant calls, at least 90 variant calls, at least 100 variant calls, at least 200 variant calls, at least 300 variant calls, at least 400 variant calls, at least 500 variant calls, at least 600 variant calls, at least 700 variant calls, at least 800 variant calls, at least 900 variant calls, at least 1000 variant calls, at least 2000 variant calls, at least 3000 variant calls, at least 4000 variant calls, between 10 and 10,000 variant calls, between 50 and 5000 variant calls, or between 100 and 4500 variant calls for the test subject using sequencing data obtained from the biological sample of the test subject. In some embodiments, the number of variant calls obtained in the plurality of variant calls corresponds to the number of genomic positions in the plurality of genomic positions.
[0187] In some embodiments, a number of variant calls are filtered, e.g., in some embodiments, variant calls obtained using any of the methods disclosed herein do not meet one or more filtering criteria and are not retained for further analysis (e.g., to identify the variant allele as a somatic variant allele or a germline variant allele).
[0188] In some embodiments, if a variant call is determined to be a germline variant call using a sequencing dataset obtained from a matched germline sample from the test subject, the variant call is removed from further analysis. For example, in some embodiments, the method further includes obtaining a second plurality of variant calls using a second plurality of nucleic acid fragment sequences in electronic form obtained from sequencing a second plurality of nucleic acid fragments in a second biological sample from the test subject, the second biological sample being a matched germline sample from the subject (e.g., a normal tissue sample), and removing from the plurality of variant calls each variant call that is also in the second plurality of variant calls (e.g., removing a germline variant call). In some embodiments, a variant allele is identified as a germline variant if a variant calling algorithm such as FreeBayes, VarDict, MuTect, MuTect2, MuSE, FreeBayes, VarDict, and / or MuTect (e.g., for the test subject using a sample-matched sequencing assay) identifies the variant as a germline variant.
[0189] In some embodiments, a variant call is removed from further analysis if it is a germline variant call obtained from a list of known germline variants (e.g., gnomad, dbSNP). GnomAD and dbSNP refer to reference databases of known germline variants. In some embodiments, any other known germline variants are removed from the first plurality of variant calls.
[0190] In some embodiments, a variant call is removed from further analysis if it is found in tissue samples from subjects other than the test subject (e.g., a recurrent variant tissue blacklist). For example, in some embodiments, some portions of the reference genome are determined to be more informative (e.g., more informative in variant identification or downstream analysis).
[0191] In some embodiments, variant calls are removed from further analysis if they do not meet a quality metric (e.g., minimum allele fraction, maximum allele fraction, quality of base calls (e.g., Phred score), minimum depth, etc.).
[0192] In some embodiments, the quality metric is the minimum variant allele fraction in each of the plurality of nucleic acid fragment sequences in electronic form that are mapped to the genomic location of each variant call. In some embodiments, the minimum variant allele fraction is 10 percent. In some embodiments, the minimum variant allele fraction is less than 1 percent, less than 2 percent, less than 3 percent, less than 4 percent, less than 5 percent, less than 6 percent, less than 7 percent, less than 8 percent, less than 9 percent, less than 10 percent, less than 15 percent, or less than 20 percent.
[0193] In some embodiments, the quality metric is the maximum variant allele fraction in each of the plurality of nucleic acid fragment sequences in electronic form that are mapped to the genomic location of each variant call. In some embodiments, the maximum variant allele fraction is 90 percent. In some embodiments, the maximum variant allele fraction is at least 55 percent, at least 60 percent, at least 70 percent, at least 80 percent, at least 90 percent, at least 95 percent, or at least 99 percent.
[0194] In some embodiments, the quality metric is the minimum depth of each of the plurality of nucleic acid fragment sequences in electronic form that are mapped to the genomic location of each variant call. In some embodiments, the minimum depth is 10. In some embodiments, the minimum depth is at least 5, at least 10, at least 50, at least 100, or at least 200.
[0195] In some embodiments, variant calls are removed from further analysis if they are listed in a blacklist of known noisy genomic locations. In some embodiments, such sites are based on a set of 642 samples from the CCGA-1 method described in Example 5 below. In some embodiments, the blacklist is all or a portion of the ENCODE blacklist.
[0196] In some embodiments, mutation calling is performed using a matched normal control sample (e.g., using cfDNA from a liquid biological sample and a patient-matched normal tissue sample). In some embodiments, mutation calling is performed without a matched normal control sample (e.g., using cfDNA from a liquid biological sample).
[0197] Alternative methods for variant calling may be contemplated. Suitable variant calling methods include methods for calling SNVs and indels (e.g., FreeBayes, GATK HaplotypeCaller, Platypus, Samtools / BCFtools, etc.), methods for calling somatic mutations (e.g., deepSNV, MuSE, MuTect2, SomaticSniper, Strelka2, VarDict, VarScan2, etc.), methods for calling copy number variations (e.g., cn.MOPS, CONTRA, CoNVEX, ExomeCNV, ExomeDepth, XHMM, etc.), methods for calling structural variations (e.g., DELLY, Lumpy, Manta, Pindel, SVMerge, etc.), and / or methods for calling gene fusions (RNA-seq) (e.g., fusionCatcher, fusionMap, mapSplice, SOAPfuse, STAR-Fusion, TopHat-Fusion, etc.). In some embodiments, mutation calling is performed using any of the methods disclosed herein, or any substitution, modification, addition, deletion, and / or combination thereof.
[0198] Methods for variant calling are described in more detail in U.S. patent application Ser. No. 17 / 185885, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 25, 2021, and PCT application Ser. No. PCT / US2021 / 019746, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 2021, each of which is incorporated by reference in its entirety.
[0199] Obtaining nucleic acid fragment sequences Referring to block 206 of FIG. 2A, the method includes obtaining a sequencing dataset (e.g., at least 1×10 nucleotides) derived from a biological sample (e.g., a liquid biological sample) obtained from the test subject that is mapped onto the genomic location. 6 Pieces, at least 2 x 10 6 Pieces, at least 3 × 10 6 Pieces, at least 4 × 10 6 Pieces, at least 5 × 10 6 Pieces, at least 6 x 10 6 Pieces, at least 7 x 10 6 Pieces, at least 8 x 10 6 Pieces, at least 9 x 10 6 Pieces, at least 1 × 10 7 or at least 1 × 10 8 The method further comprises obtaining a methylation state and a respective sequence of each nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences in the plurality of nucleic acid fragment sequences (including each of the plurality of nucleic acid fragment sequences).
[0200] In some embodiments, the biological sample is prepared for sequencing using any suitable method (see "Subjects and Samples" above). In some embodiments, preparing the biological sample includes obtaining a respective plurality of nucleic acid fragments (e.g., nucleic acid molecules) for the test subject. In some embodiments, each of the plurality of nucleic acid fragments obtained from the biological sample are cell-free nucleic acid fragments.
[0201] After obtaining a plurality of nucleic acid fragments from a biological sample, in some embodiments, the nucleic acid fragments are sequenced. In some embodiments, the sequencing is methylation sequencing. In some embodiments, the methylation sequencing is whole genome methylation sequencing. In some embodiments, the methylation sequencing is targeted DNA methylation sequencing using a plurality of nucleic acid probes. In some embodiments, the plurality of nucleic acid probes comprises 100 or more probes. In some embodiments, the plurality of nucleic acid probes comprises 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 2000 or more, 3000 or more, 4000 or more, 5000 or more, 6000 or more, 7000 or more, 8000 or more, 9000 or more, 10,000 or more, 25,000 or more, or 50,000 or more probes. In some embodiments, the plurality of nucleic acid probes includes 50,000 or less, 250,000 or less, 10,000 or less, 9000 or less, 8000 or less, 7000 or less, 6000 or less, 5000 or less, 4000 or less, 3000 or less, 2000 or less, 1000 or less, 900 or less, 800 or less, 700 or less, 600 or less, or 500 or less probes. In some embodiments, the plurality of nucleic acid probes includes 100 to 500, 500 to 1000, 1000 to 2000, 1000 to 5000, 100 to 5000, 5000 to 10,000, or 10,000 to 50,000 probes. In some embodiments, the plurality of nucleic acid probes falls within another range beginning with 100 or more probes and ending with 50,000 or less probes. In some embodiments, some or all of the probes uniquely map to genomic regions described in International Patent Publication No. WO2020154682A3, entitled "Detecting Cancer, Cancer Tissue or Origin, or Cancer Type," which is incorporated by reference herein, including the sequence listing referenced therein.In some embodiments, some or all of the probes are uniquely mapped to genomic regions described in International Patent Publication No. WO2020 / 069350A1, entitled "Methylated Markers and Targeted Methylation Probe Panel," which is incorporated by reference herein, including the sequence listing referenced therein. In some embodiments, some or all of the probes are uniquely mapped to genomic regions described in International Patent Publication No. WO2019 / 195268A2, entitled "Methylated Markers and Targeted Methylation Probe Panels," which is incorporated by reference herein, including the sequence listing referenced therein.
[0202] In some embodiments, methylation sequencing detects one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in each nucleic acid fragment in each of the plurality of nucleic acid fragments. In some embodiments, methylation sequencing comprises converting one or more unmethylated cytosines or one or more methylated cytosines in each of the plurality of nucleic acid fragments to one or more corresponding uracils. In some embodiments, one or more uracils are converted during amplification and detected as one or more corresponding thymines during methylation sequencing. In some embodiments, the conversion of one or more unmethylated cytosines or one or more methylated cytosines comprises chemical conversion, enzymatic conversion, or a combination thereof.
[0203] In some embodiments, prior to sequencing, a plurality of nucleic acid fragments are treated to convert unmethylated cytosines to uracil. In some embodiments, the methylation sequencing is bisulfite sequencing. For example, in some embodiments, the method uses bisulfite treatment of DNA (e.g., cfDNA) to convert unmethylated cytosines to uracil without converting methylated cytosines. For example, in some embodiments, a commercially available kit is used for bisulfite conversion, such as, for example, EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, or EZ DNA Methylation™-Lightning kit (available from Zymo Research Corp, Irvine, CA). In some embodiments, the conversion of unmethylated cytosines to uracil is achieved using an enzymatic reaction. For example, the conversion can use a commercially available kit for the conversion of unmethylated cytosines to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).
[0204] In some embodiments, the methylation sequencing is whole genome bisulfite sequencing. In some embodiments, the whole genome bisulfite sequencing assay looks for variations in methylation patterns in the genome. See U.S. Patent Application Publication No. US 2019-0287652 A1, entitled "Anomalous Fragment Detection and Classification."
[0205] From the cell-free nucleic acid fragments converted, prepare a sequencing library.Optionally, the sequencing library is enriched for cell-free nucleic acid fragments or genomic regions with high information value for cellular origin using multiple hybridization probes, such as any combination of regions disclosed in International Patent Publication No. WO2020154682A3 entitled "Detecting Cancer, Cancer Tissue or Origin, or Cancer Type", International Patent Publication No. WO2020 / 069350A1 entitled "Methylated Markers and Targeted Methylation Probe Panel", and / or International Patent Publication No. WO2019 / 195268A2 entitled "Methylated Markers and Targeted Methylation Probe Panels", each of which is incorporated herein by reference. In some embodiments, the hybridization probes are short oligonucleotides that hybridize to specifically designated cell-free nucleic acid fragments or target regions, and enrich those fragments or regions for subsequent sequencing and analysis, for example as disclosed in International Patent Publication No. WO2020154682A3 entitled "Detecting Cancer, Cancer Tissue or Origin, or Cancer Type," International Patent Publication No. WO2020 / 069350A1 entitled "Methylated Markers and Targeted Methylation Probe Panel," and / or International Patent Publication No. WO2019 / 195268A2 entitled "Methylated Markers and Targeted Methylation Probe Panels," each of which is incorporated herein by reference. In some embodiments, the hybridization probes are used to perform targeted deep analysis of a set of specific CpG sites that are highly informative about the cellular origin. Once prepared, the sequencing library or a portion thereof can be sequenced to obtain multiple sequence reads (e.g., nucleic acid fragment sequences).
[0206] In some embodiments, any form of sequencing can be used to obtain sequence reads (e.g., nucleic acid fragment sequences) from multiple nucleic acid fragments derived from biological samples of test subjects.Examples of sequencing methods include, but are not limited to, high-throughput sequencing systems, such as Roche 454 platform, Applied Biosystems SOLID platform, Helicos True Single Molecule DNA sequencing technology, Affymetrix Inc.'s hybridization sequencing platform, Pacific Biosciences' single molecule real-time (SMRT) technology, 454 Life Sciences, Illumina / Solexa and Helicos Biosciences' synthesis sequencing platform, and Applied Biosystems' ligation sequencing platform.ION TORRENT technology and nanopore sequencing from Life technologies can also be used to obtain sequence reads from multiple nucleic acid fragments obtained from biological samples.
[0207] In some embodiments, sequencing by synthesis and reversible terminator-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 2500 (Illumina, San Diego CA)) are used to obtain sequence reads from multiple nucleic acid fragments (e.g., cell-free nucleic acid fragments) from a biological sample. In some such embodiments, millions of nucleic acid fragments (e.g., cfDNA fragments) are sequenced in parallel. One example of this type of sequencing technology uses a flow cell that includes an optically clear slide with eight individual lanes with oligonucleotide anchors (e.g., adapter primers) attached to its surface. A flow cell is often a solid support configured to hold a reagent solution over the bound analytes and / or allow for the orderly passage of a reagent solution. In some examples, the flow cell is planar in shape, optically clear, typically millimeter or sub-millimeter scale, and often has channels or lanes where analyte / reagent interactions occur. In some embodiments, a sample that includes multiple nucleic acid fragments (e.g., cfDNA fragments) can include a signal or tag to facilitate detection. In some such embodiments, obtaining sequence reads from nucleic acid fragments includes obtaining quantitative information of signals or tags through various techniques such as, for example, flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene chip analysis, microarrays, mass spectrometry, cytofluorimetric analysis, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electric field suspension, sequencing, and combinations thereof.
[0208] In some embodiments, the sequencing comprises whole genome methylation sequencing (e.g., whole genome bisulfite sequencing (WGBS)) and / or whole genome sequencing (e.g., whole genome sequencing (WGS) or whole exome sequencing (WES)), where the sequencing is used to sequence at least a portion of the genome of the test subject. In some embodiments, the portion of the genome is at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent, 90 percent, 95 percent, 99 percent, 99.9 percent, or all of the genome (e.g., a human reference genome). In some embodiments, the sequencing comprises whole genome methylation sequencing and / or whole genome sequencing, where the sequencing obtains at least 1x, at least 2x, at least 3x, at least 4x, at least 5x, at least 10x, at least 15x, at least 20x, at least 25x, at least 30x, at least 50x, at least 100x, at least 200x, at least 300x, at least 400x, at least 500x, or at least 1000x sequencing coverage of the portion of the genome across the entire sequenced portion of the genome. In some embodiments, the sequencing obtains at least 5x, at least 10x, at least 15x, at least 20x, at least 25x, at least 30x, at least 50x, at least 100x, at least 200x, at least 300x, at least 400x, at least 500x, or at least 1000x sequencing coverage across the entire genome.
[0209] In some embodiments, the sequencing is targeted sequencing (e.g., targeted methylation sequencing), which obtains at least 5x, at least 10x, at least 15x, at least 20x, at least 25x, at least 30x, at least 50x, at least 100x, at least 250x, at least 500x, or at least 1000x sequencing coverage (e.g., sequencing depth) of a targeted portion of the genome of a test subject (e.g., a panel of genes to which one or more probes are mapped). In some embodiments, targeted sequencing obtains at least 100x, at least 200x, at least 500x, at least 1,000x, at least 2,000x, at least 3,000x, at least 4,000x, at least 5,000x, at least 10,000x, at least 15,000x, at least 20,000x, at least 25,000x, at least 30,000x, at least 40,000x, at least 50,000x, at least 60,000x, or at least 70,000x sequencing coverage across the entire target region of the genome.
[0210] In some embodiments, the plurality of sequence reads, e.g., nucleic acid fragment sequences, obtained from sequencing of the biological sample comprises at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1,000,000, at least 2,000,000, at least 3,000,000, at least 4,000,000, at least 5,000,000, at least 6,000,000, at least 7,000,000, at least 8,000,000, at least 9,000,000 or more sequence reads in a sequencing dataset. In some embodiments, the plurality of sequence reads comprises at least 1×10 7 Pieces, at least 2 x 107 Pieces, at least 3 × 10 7 Pieces, at least 4 × 10 7 Pieces, at least 5 × 10 7 Pieces, at least 6 x 10 7 Pieces, at least 7 x 10 7 Pieces, at least 8 x 10 7 Pieces, at least 9 x 10 7 Pieces, at least 1 × 10 8 Pieces, at least 2 x 10 8 Pieces, at least 3 × 10 8 Pieces, at least 4 × 10 8 Pieces, at least 5 × 10 8 Pieces, at least 6 x 10 8 Pieces, at least 7 x 10 8 Pieces, at least 8 x 10 8 Pieces, at least 9 x 10 8 Pieces, at least 1 × 10 9 In some embodiments, the plurality of sequence reads comprises 5×10 or more sequence reads in a sequencing dataset. 7 pcs or less, 1×10 7 Less than or equal to 5×10 6 pcs or less, 4×10 6 pcs or less, 3×10 6 pcs or less, 2×10 6 pcs or less, 1×10 6or less, 500,000 or less, 100,000 or less, 50,000 or less, 30,000 or less, 20,000 or less, 10,000 or less, 9000 or less, 8000 or less, 7000 or less, 6000 or less, 5000 or less, 4000 or less, 3000 or less, 2000 or less, 1000 or less, or less sequence reads. In some embodiments, the plurality of sequence reads includes 1000-5000, 1000-10,000, 2000-20,000, 5000-50,000, 10,000-100,000, 100,000-500,000, 10,000-500,000, 500,000-1,000,000, 1,000,000-30,000,000, 30,000,000-80,000,000, or 10,000,000-500,000,000 sequence reads in a sequencing dataset. In some embodiments, the plurality of sequence reads includes 1000 or more sequence reads, 1×10 9 within another range ending with fewer than 1 sequence read.
[0211] In some embodiments, obtaining a respective sequence for each nucleic acid fragment sequence in the plurality of nucleic acid fragment sequences further comprises mapping each nucleic acid fragment sequence in the sequencing dataset to a reference sequence (e.g., a human reference genome). In some embodiments, the method comprises mapping all or a portion of the sequencing dataset comprising the plurality of nucleic acid fragment sequences to a reference sequence.
[0212] For example, for each genomic location, in some embodiments the method further includes inputting a reference genome (e.g., a human reference genome) into a computer system including a processor coupled to a non-transitory memory, and determining, using the computer system, that each nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences is mapped to a genomic location by aligning each nucleic acid fragment sequence to the reference genome.
[0213] In some embodiments, the mapping is performed using Smith-Waterman gapped alignment, e.g., as implemented in Arioc, or Burrows-Wheeler transformation, e.g., as implemented in Bowtie. Other suitable alignment programs can include, but are not limited to, BarraCUDA, BBMap, BFAST, BigBWA, BLASTN, BLAT, BWA, BWA-PSSM, CASHX. In some embodiments, the mapping tolerates mismatches. In some embodiments, the mapping includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more than 10 mismatches. Other methods of mapping sequence reads to a reference sequence can be used.
[0214] In some embodiments, mapping the nucleic acid fragment sequences in the sequencing dataset to the reference sequence includes using a CpG index. For example, in some embodiments, the CpG index includes a list of each CpG site in a plurality of CpG sites (e.g., CpG 1, CpG 2, CpG 3, etc.) in a reference sequence (e.g., a human reference genome). The CpG index may further include, for each CpG site in the CpG index, a corresponding genomic position in the corresponding reference sequence. Thus, each CpG site in each nucleic acid sequence fragment can be indexed to a specific position in each reference sequence, which can be determined using the CpG index. In some embodiments, the reference sequence is obtained in an electronic format.
[0215] In some embodiments, for each genomic location, the method includes mapping all or a portion of a sequencing dataset comprising the plurality of nucleic acid fragment sequences to at least a portion of a reference sequence that comprises the genomic location.
[0216] In some embodiments, each nucleic acid fragment sequence in the plurality of nucleic acid fragment sequences that are mapped to a genomic location is determined by mapping to overlap all or a portion of the genomic location.
[0217] In some embodiments, the plurality of nucleic acid fragment sequences mapped to genomic locations comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, at least 20,000, or at least 30,000 nucleic acid fragment sequences mapped to genomic locations. In some embodiments, the plurality of nucleic acid fragment sequences mapped to genomic locations include 70,000 or less, 50,000 or less, 30,000 or less, 10,000 or less, 5000 or less, 2000 or less, 1000 or less, 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, 200 or less, 100 or less, 50 or less, or 30 or less nucleic acid fragment sequences mapped to genomic locations. In some embodiments, the plurality of nucleic acid fragment sequences mapped to genomic locations include 5 to 20, 20 to 50, 50 to 100, 100 to 500, 500 to 1000, 500 to 5000, 2000 to 10,000, or 10,000 to 70,000 nucleic acid fragment sequences mapped to genomic locations. In some embodiments, the plurality of nucleic acid fragment sequences that are mapped to a genomic location are within another range beginning with 10 or more nucleic acid fragment sequences and ending with 70,000 or less nucleic acid fragment sequences. In some embodiments, the plurality of nucleic acid fragment sequences that are mapped to a genomic location are determined at least in part based on the sequencing coverage (e.g., sequencing depth) of the sequencing method used.
[0218] In some embodiments, where the method is performed for each of a plurality of genomic locations, mapping comprises mapping a plurality of nucleic acid fragment sequences to a region of a reference sequence (e.g., a reference genome) that includes at least the plurality of genomic locations.
[0219] In some embodiments, obtaining the methylation state of each nucleic acid fragment sequence in the sequencing data set comprises determining the corresponding methylation state of each CpG site in each nucleic acid fragment sequence. For example, in some embodiments, each nucleic acid fragment sequence can have one or more CpG sites, and each CpG site in the nucleic acid fragment sequence is determined to have a corresponding methylation state by methylation sequencing.
[0220] In some embodiments, the methylation state of each CpG site in the corresponding one or more CpG sites in each nucleic acid fragment sequence is methylated if the respective CpG site is determined to be methylated by methylation sequencing, and is unmethylated if the respective CpG site is determined to be unmethylated by methylation sequencing. In some embodiments, the methylation state is designated as "M" and the unmethylation state is designated as "U".
[0221] Other methylation states may be possible. For example, in some embodiments, if methylation sequencing cannot call the methylation state of each CpG site as methylated or unmethylated, the methylation state is "other". In some embodiments, possible methylation states further include, but are not limited to, ambiguous methylation states (e.g., meaning that the underlying CpG is not covered by any of the fragment sequences in the multiple fragment sequences), variant methylation states (e.g., meaning that the fragment sequence does not match the CpG present at its expected position based on the reference sequence, which may be caused by an actual mutation or sequence error at that site), or conflict states (e.g., when two or more fragment sequences both overlap a CpG site but have inconsistent methylation states). See, for example, U.S. Provisional Patent Application No. 62 / 948,129, entitled "Cancer classification using patch convolutional neural networks," filed December 13, 2019, which is incorporated herein by reference in its entirety.
[0222] In some embodiments, obtaining the methylation state of each nucleic acid fragment sequence in the sequencing dataset comprises determining a methylation state vector of the nucleic acid fragment sequence. In some embodiments, the methylation state vector is a sequence of methylation states indicating the methylation state of all CpG sites contained in each nucleic acid fragment. The methylation state vector is further described in accordance with any of the techniques disclosed in, for example, U.S. Patent Application No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification," filed March 13, 2019, or U.S. Provisional Patent Application No. 62 / 847,223, entitled "Model-Based Featurization and Classification," filed May 13, 2019, each of which is incorporated herein by reference.
[0223] Methods for sequencing nucleic acid fragments obtained from a test subject biological sample, including processing the biological sample, extracting nucleic acid fragments from the biological sample, processing the nucleic acid fragments for methylation sequencing, preparing a sequencing library, enriching target nucleic acids, hybridization probes, obtaining sequence reads, mapping fragment sequences to reference sequences, and / or generating methylation state vectors, are described in further detail below in Examples 1, 2, and 4, with reference to Figures 7, 8, and 9. Other methods for obtaining nucleic acid fragment sequences are contemplated, including processing the biological sample, extracting nucleic acid fragments from the biological sample, processing the nucleic acid fragments for methylation sequencing, preparing a sequencing library, enriching target nucleic acids, hybridization probes, obtaining sequence reads, mapping fragment sequences to reference sequences, and / or generating methylation state vectors.
[0224] Subset Allocation Referring to block 208, the method further includes assigning each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences that has the reference allele at the genomic location to a reference subset using (i) the identification of the reference allele at the genomic location and (ii) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences. The method also includes assigning each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences that has the mutant allele at the genomic location to a mutant subset using (i) the identification of the variant allele at the genomic location and (ii) the respective sequence of each nucleic acid fragment sequence in the respective plurality of nucleic acid fragment sequences.
[0225] In some embodiments, the assignment of each nucleic acid fragment sequence to a reference subset includes, for each nucleic acid fragment sequence in the sequencing dataset, determining whether the respective nucleic acid fragment sequence has a reference allele at the genomic location based on a comparison of the nucleic acid fragment sequence obtained by sequencing to a nucleic acid sequence of a reference allele (identified as described above with reference to block 202; see "Reference Alleles and Mutant Alleles"). In some embodiments, the comparison is performed using a lookup table.
[0226] In some embodiments, assigning each nucleic acid fragment sequence to a mutant subset includes, for each nucleic acid fragment sequence in the sequencing dataset, determining whether each nucleic acid fragment sequence has a mutant allele at a genomic location based on a comparison of the nucleic acid fragment sequence obtained by sequencing to the nucleic acid sequence of the mutant allele (identified as described above with reference to block 204; see "Reference Alleles and Mutant Alleles").
[0227] In some embodiments, the method comprises obtaining a count of the number of nucleic acid fragment sequences assigned to the reference subset.
[0228] In some embodiments, the method comprises obtaining a count of the number of nucleic acid fragment sequences assigned to the mutation subset.
[0229] In some embodiments, the plurality of nucleic acid fragment sequences in the sequencing dataset are filtered using one or more filters. In some embodiments, the filtering is performed prior to the assignment of the nucleic acid fragment sequences to the reference and variant subsets. In some embodiments, the filtering is performed after the assignment of the nucleic acid fragment sequences to the reference and variant subsets. In some embodiments, the filtering is performed using a count of the nucleic acid fragment sequences assigned to the reference and variant subsets. In some embodiments, the filtering comprises removing one or more nucleic acid fragment sequences from each of the plurality of nucleic acid fragment sequences at each genomic location that do not meet the filtering criteria. In some embodiments, when the method is performed on a plurality of genomic locations, the filtering comprises removing one or more genomic locations from the plurality of genomic locations that do not meet the filtering criteria. In some embodiments, when the method is performed on a plurality of genomic locations, the filtering comprises removing genomic locations from the plurality of genomic locations if at least a threshold amount of nucleic acid fragment sequences mapped to the respective genomic location do not meet the filtering criteria.
[0230] For example, in some embodiments, the nucleic acid fragment sequences in the sequencing dataset are filtered based on the ratio of fragments containing mutant alleles to fragments containing reference alleles at the genomic location. In some embodiments, when the method is performed on a plurality of genomic locations, the filtering comprises removing genomic locations where the ratio of mutant allele fragments to reference allele fragments is less than a threshold ratio. In some embodiments, when the method is performed on a plurality of genomic locations, the filtering comprises removing genomic locations where the count of mutant allele fragments in the mutant subset is less than a threshold count.
[0231] In some embodiments, the threshold count of mutant allele fragments in the mutant subset is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 200, at least 300, at least 400, at least 500, or at least 1000 nucleic acid fragments from a test subject that maps to the genomic region of the mutant allele and has the mutant allele.
[0232] In some embodiments, the one or more filters include a minimum mutant allele frequency, a maximum mutant allele frequency, a minimum sequencing depth for each allele, a blacklist of germline mutations from the test subject (e.g., marked by freebayes), a blacklist of custom databases (e.g., a recurrent tissue blacklist), or a blacklist of germline mutations from a reference database (e.g., from the gnomad and / or dbSNP databases).
[0233] In some embodiments, one or more filters is a minimum variant allele frequency (minimum VAF). In some such embodiments, the minimum allele frequency is at least 3%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% of the nucleic acid fragments from the test subject.
[0234] In some embodiments, one or more filters is a maximum variant allele frequency (max VAF), in some embodiments, the maximum allele frequency is 95% or less, 90% or less, 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, or 50% or less of the nucleic acid fragments from the test subject.
[0235] In some embodiments, one or more filters is a minimum sequencing depth (e.g., of all nucleic acid fragment sequences at a genomic location, including the reference and variant subsets). In some embodiments, the minimum sequencing depth is at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 200, at least 300, at least 400, at least 500, or at least 1000 nucleic acid fragments from a test subject that are mapped to a genomic location.
[0236] Other filters may also be contemplated. For example, in some embodiments, the plurality of nucleic acid fragment sequences are filtered for, e.g., depth, minimum mapping quality (MAPQ), overlapping fragments, uncalled fragments, unconverted fragments, ambiguous calls, variant calls, competing calls, minimum or maximum fragment length, minimum or maximum number of base pairs, minimum or maximum CpG count, and / or p-value (described in more detail below).
[0237] Further, in some embodiments, the sequencing dataset is further processed by any suitable method, such as a bioinformatics pipeline. For example, in some embodiments, the multiple nucleic acid fragment sequences are further normalized to account for, for example, pull-down, amplification, background copy number (e.g., duplication), and / or sequencing bias (e.g., mappability, GC bias, etc.).
[0238] Entering indicators Referring to block 210, the method further includes applying at least (i) one or more indices of methylation state across methylation states of each nucleic acid fragment sequence in the mutant subset, and (ii) an indicator of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset to a trained binary classifier (e.g., comprising at least 10 parameters), thereby obtaining from the trained binary classifier an identification of the mutant allele at the genomic location in the test subject as a somatic mutant allele or a germline mutant allele.
[0239] In some embodiments, (i) one or more measures of methylation status across the methylation status of each nucleic acid fragment in the mutant subset is a p-value. In some embodiments, the p-value indicates whether the respective nucleic acid fragment is aberrantly methylated compared to a healthy reference.
[0240] Thus, referring to block 212 of FIG. 2B, in an exemplary embodiment, a first nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences has a plurality of CpG sites, the first nucleic acid fragment sequence has a corresponding methylation pattern across the plurality of CpG sites, and the methylation state of the first nucleic acid fragment sequence is a p-value, and the method further includes determining the p-value of the first nucleic acid fragment sequence at least in part by comparing the corresponding methylation pattern of the first nucleic acid fragment sequence to a corresponding distribution of methylation patterns of those nucleic acid fragment sequences in healthy non-cancer cohort datasets, each having a respective plurality of CpG sites.
[0241] P-value determination is further described in Example 5 of International Patent Application PCT / US2020 / 034317, entitled "Systems and Methods for Determining Whether a Subject has a Cancer Condition Using Transfer Learning," filed May 22, 2020, and U.S. Patent Application Serial No. 16 / 352,602, entitled "Anomalous fragment detection and classification," filed March 13, 2019, and currently published as U.S. Patent Application Publication No. 2019 / 0287652, each of which is incorporated herein by reference in its entirety. The goal of p-value determination can be to measure aberrant methylation in nucleic acid fragment sequences based on their corresponding methylation state vectors. For example, for each nucleic acid fragment in a biological sample, the methylation state vector corresponding to the fragment is used to determine (e.g., via analysis of sequence reads derived from the fragment) whether the fragment is aberrantly methylated compared to an expected methylation state vector (e.g., the expected methylation state vector is determined from sequence analysis of a cohort of healthy subjects). Generation of such methylation state vectors of nucleic acid fragments (e.g., cell-free nucleic acid fragments) is disclosed above and in, e.g., U.S. Patent Application Publication No. 2019 / 0287652, which is incorporated by reference in its entirety.
[0242] In some embodiments, the healthy cohort comprises at least 20 subjects, and the plurality of nucleic acid fragment sequences comprises at least 10,000 different corresponding methylation patterns. In some embodiments, the healthy cohort comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 subjects. In some embodiments, the healthy cohort comprises 1-10, 10-50, 50-100, 100-500, 500-1000, or more than 1000 subjects. In some embodiments, the plurality of nucleic acid fragment sequences comprises 1-1000, 1000-2000, 2000-4000, 4000-6000, 6000-8000, 8000-10,000, 10,000-20,000, 20,000-50,000, or more than 50,000 different corresponding methylation patterns.
[0243] In some embodiments, aberrant fragments are identified as fragments having more than a threshold number of CpG sites, and either more than a threshold percentage of CpG sites that are methylated (hypermethylated) or more than a threshold percentage of CpG sites that are unmethylated (hypomethylated). In some embodiments, the threshold percentage of methylated and / or unmethylated CpG sites is at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, or at least 95%. In some embodiments, the threshold percentage of methylated and / or unmethylated CpG sites is between 50% and 100%.
[0244] In some embodiments, a Markov model (e.g., a hidden Markov model "HMM") is used to determine the probability that a set of methylation states (e.g., including methylated "M" and / or unmethylated "U") may be observed for each nucleic acid fragment sequence, given a set of probabilities that determine for each state in the methylation pattern of each nucleic acid fragment sequence the likelihood of observing the next state in the sequence. In some embodiments, the set of probabilities is obtained by training the HMM. In some embodiments, such training involves calculating statistics (e.g., the probability that a first state transitions to a second state (transition probability) and / or the probability that a given methylation state is observed for each CpG site (output probability)) given an initial training data set of observed methylation state sequences (e.g., methylation patterns) obtained from a cohort of non-cancer subjects. In some embodiments, the HMM is trained using supervised training (e.g., using samples where the underlying sequence and observed states are known). In some alternative embodiments, the HMM is trained using unsupervised training (e.g., Viterbi learning, maximum likelihood estimation, expectation-maximization training, and / or Baum-Welch training). For example, expectation-maximization algorithms such as the Baum-Welch algorithm estimate transition and output probabilities from observed sample sequences and generate a parameterized probabilistic model that best explains the observed sequences. Such algorithms iteratively calculate the likelihood function until the expected number of correctly predicted states is maximized.
[0245] In some embodiments, the p-value of each nucleic acid fragment sequence is determined by a method other than a Markov model or a hidden Markov model. In some embodiments, the p-value of each nucleic acid fragment sequence is determined using a mixed model. For example, the mixed model can detect abnormal methylation patterns in nucleic acid fragment sequences by determining the likelihood of a methylation state vector (e.g., a methylation pattern) of each nucleic acid methylation fragment based on the number of possible methylation state vectors at the same length and the same corresponding genomic location. This can be performed by generating multiple possible methylation states for a vector of a specified length at each genomic location in a reference sequence (e.g., a human reference genome). The multiple possible methylation states can be used to determine the number of all possible methylation states, and then the probability of each predicted methylation state at the genomic location. The likelihood of the sample nucleic acid methylation fragment corresponding to the genomic location in the reference sequence can then be determined by matching the sample nucleic acid fragment sequence with a predicted (e.g., possible) sequence and taking the calculated probability of the predicted methylation state. An abnormal methylation score is then calculated based on the probability of the sample nucleic acid fragment sequence.
[0246] In some embodiments, the p-value of each nucleic acid methylation fragment is determined using the learned representation. As will be apparent to one of skill in the art, any other suitable method of determining a p-value is contemplated.
[0247] In some embodiments, the p-value (e.g., determined by any of the methods disclosed herein) is used as a filter to remove nucleic acid fragment sequences that are not sufficiently aberrant to be used as input (e.g., for a model) in the systems and methods for identifying mutant alleles disclosed herein.
[0248] In some such embodiments, nucleic acid fragment sequences having a p-value below the threshold are retained for further use in the method (e.g., as input to a model for identifying mutant alleles as somatic mutant alleles or germline mutant alleles). For example, in some embodiments, the plurality of nucleic acid fragment sequences are filtered by removing each nucleic acid fragment sequence having a p-value whose corresponding methylation pattern (e.g., methylation state vector) across the corresponding plurality of CpG sites in each fragment does not meet the p-value threshold.
[0249] In some embodiments, the p-value threshold is between 0.001 and 0.20. In some embodiments, the threshold is 0.01 (e.g., in such embodiments, p can be <0.01). In some embodiments, the threshold is 0.001, 0.005, 0.01, 0.015, 0.02, 0.05, or 0.10. In some embodiments, the threshold is between .0001 and 0.20. In some embodiments, the p-value threshold for a methylation pattern from a subject is met if a methylation pattern corresponding to each cell-free fragment in the plurality of cell-free fragments has a p-value of 0.10 or less, 0.05 or less, or 0.01 or less.
[0250] Referring again to block 210, in some embodiments, (i) each index in the one or more indices of methylation states across the methylation states of each nucleic acid fragment sequence in the mutation subset is a measure of central tendency of methylation state p-values across the mutation subset, a minimum methylation state p-value across the mutation subset, a maximum methylation state p-value across the mutation subset, or a measure of the spread of methylation state p-values across the mutation subset.
[0251] For example, in some embodiments, one of the one or more indices of methylation status across the mutation subsets is a measure of central tendency of the methylation status p-values across the mutation subsets, where the measure of central tendency is the arithmetic mean, weighted average, mid-range, mid-hinge, triangular mean, winsorized mean, average, or mode of the methylation status p-values across the mutation subsets. In some embodiments, one of the one or more indices of methylation status across the mutation subsets is a measure of spread of the methylation status p-values across the mutation subsets, where the measure of spread is the standard deviation, variance, range, or interquartile range of the methylation status p-values across the mutation subsets.
[0252] In some embodiments, the one or more indices of methylation state across the mutation subsets are a plurality of indices of methylation state across the mutation subsets, including at least two, at least three, or all four of a measure of central tendency of methylation state p-values across the mutation subsets, a minimum methylation state p-value across the mutation subsets, a maximum methylation state p-value across the mutation subsets, and a measure of spread of methylation state p-values across the mutation subsets.
[0253] In some embodiments, the one or more indices of methylation state across the mutation subsets are a plurality of indices of methylation state across the mutation subsets, including the mean p-value, median p-value, minimum p-value, maximum p-value, and standard deviation of p-values across the mutation subsets.
[0254] In some embodiments, the one or more indices of methylation status across the mutation subsets are the highest ranked (e.g., most significant) set of p-values from the mutation subsets. For example, in some embodiments, the one or more indices of methylation across the mutation subsets include at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 of the highest ranked (e.g., most significant) p-values from the mutation subsets. In some embodiments, the one or more indices of methylation across the mutation subsets include the top 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or top 1% of the most highly ranked (e.g., most significant) p-values from the mutation subset.
[0255] In some embodiments, the one or more indices of methylation state across the methylation states of each nucleic acid fragment in the mutation subset comprise a methylation state vector and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the mutation subsets, a minimum across the mutation subsets, a maximum across the mutation subsets, and a measure of spread across the mutation subsets).
[0256] In some embodiments, the one or more indices of methylation state across the methylation states of each nucleic acid fragment in the mutation subset include a beta value and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the mutation subsets, a minimum across the mutation subsets, a maximum across the mutation subsets, and a measure of spread across the mutation subsets).
[0257] In some embodiments, the one or more indices of methylation state across the methylation states of each nucleic acid fragment in the mutation subset include an M value and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the mutation subsets, a minimum across the mutation subsets, a maximum across the mutation subsets, and a measure of spread across the mutation subsets).
[0258] In some embodiments, the one or more indices of methylation status across the methylation status of each nucleic acid fragment in the mutation subset comprise an aberrant methylation score and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the mutation subsets, a minimum across the mutation subsets, a maximum across the mutation subsets, and a measure of spread across the mutation subsets).
[0259] In some embodiments, the one or more indices of methylation state across the methylation states of each nucleic acid fragment in the mutation subset include a mutual information score and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the mutation subsets, a minimum across the mutation subsets, a maximum across the mutation subsets, and a measure of spread across the mutation subsets). Further details regarding mutual information scores are disclosed in U.S. Provisional Patent Application No. 62 / 948,129, entitled "Cancer Classification using Patch Convolutional Neural Networks," filed December 13, 2019, which is incorporated herein by reference in its entirety.
[0260] In some embodiments, the measure of central tendency is the arithmetic mean, weighted average, midrange, midhinge, triad mean, winsorized mean, average, or mode of the methylation state p-values across the mutation subsets, in some embodiments, the measure of spread is the standard deviation, variance, range, or interquartile range of the methylation state p-values across the mutation subsets.
[0261] In some embodiments, the one or more indicators of methylation states across the methylation states of each nucleic acid fragment in the mutation subset comprise at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 500, at least 800, or at least 1000 indicators of methylation states across the mutation subset. In some embodiments, the one or more indices of methylation states across the methylation states of each nucleic acid fragment in the mutation subset comprises 2000 or less, 1000 or less, 500 or less, 200 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 40 or less, 30 or less, 20 or less, or 10 or less indices of methylation states across the mutation subset. In some embodiments, the one or more indices of methylation states across the methylation states of each nucleic acid fragment in the mutation subset comprises 3-10, 5-20, 10-50, 20-100, 50-200, 100-500, 300-1000, or 500-2000 indices of methylation states across the mutation subset. In some embodiments, one or more indices of methylation state in the mutation subset are within another range starting from 3 or more indices of methylation state across the mutation subset and ending with 2000 or fewer indices.
[0262] Referring to block 214, in some embodiments, the method further includes applying (iii) one or more CpG site indices across the variant subset to the trained binary classifier.
[0263] In some embodiments, the indicator of CpG sites is a CpG count. For example, in some embodiments, the CpG count is obtained by counting the number of CpG sites in the nucleic acid fragment based on the nucleic acid fragment sequence. In some embodiments, each nucleic acid fragment sequence in the mutation subset has the same CpG count. In some embodiments, two or more nucleic acid fragment sequences in the mutation subset have different CpG counts. In some embodiments, each nucleic acid fragment sequence in the mutation subset has at least a minimum number of CpG sites (e.g., a plurality of nucleic acid fragment sequences at each genomic location are filtered using a minimum CpG count or a maximum CpG count).
[0264] In some embodiments, the minimum number of CpG sites is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites. In some embodiments, the minimum number of CpG sites is 1-10, 10-20, 20-30, 30-40, 40-50, or more than 50 CpG sites.
[0265] In some embodiments, an index of the one or more CpG site indices across the mutation subsets includes a measure of central tendency of CpG counts across the mutation subsets, a minimum CpG count across the mutation subsets, a maximum CpG count across the mutation subsets, and a measure of spread of CpG counts across the mutation subsets.
[0266] For example, in some embodiments, one of the one or more CpG site indices across the mutation subsets is a measure of central tendency of the CpG counts across the mutation subsets, where the measure of central tendency is the arithmetic mean, weighted mean, midrange, midhinge, trimean, winsorized mean, average, or mode of the CpG counts across the mutation subsets. In some embodiments, one of the one or more CpG site indices across the mutation subsets is a measure of spread of the CpG counts across the mutation subsets, where the measure of spread is the standard deviation, variance, range, or interquartile range of the CpG counts across the mutation subsets.
[0267] In some embodiments, the one or more CpG index across the mutation subset is a plurality of CpG site indexes across the mutation subset including at least two, at least three, or all four of a measure of central tendency of CpG counts across the mutation subset, a minimum CpG count across the mutation subset, a maximum CpG count across the mutation subset, and a measure of spread of CpG counts across the mutation subset.
[0268] In some embodiments, the one or more CpG index across the mutation subset is a plurality of CpG site indexes across the mutation subset including a CpG count across the mutation subset, a median CpG count, a minimum CpG count, a maximum CpG count, and a standard deviation of the CpG counts.
[0269] In some embodiments, the one or more CpG indices across the mutation subsets include genomic locations of CpG sites and / or one or more distribution statistics thereof. In some embodiments, the one or more CpG indices across the mutation subsets include CpG density and / or one or more distribution statistics thereof. In some embodiments, the one or more CpG indices across the mutation subsets include genomic distance between two or more CpG sites and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the mutation subsets, a minimum across the mutation subsets, a maximum across the mutation subsets, and a measure of spread across the mutation subsets).
[0270] In some embodiments, the one or more CpG index across the mutation subsets comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 CpG index across the mutation subsets. In some embodiments, the one or more CpG indices across the mutation subsets include 200 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 40 or less, 30 or less, 20 or less, or 10 or less CpG indices across the mutation subsets. In some embodiments, the one or more CpG indices across the mutation subsets include 3-10, 5-20, 10-50, 20-100, or 50-200 CpG indices across the mutation subsets. In some embodiments, the one or more CpG indices in the mutation subsets are in another range starting with 3 or more CpG indices across the mutation subsets and ending with 200 or less CpG indices.
[0271] Referring to block 216, in some embodiments, applying to the trained binary classifier further applies one or more indices of methylation status across the reference subset.
[0272] In some embodiments, the one or more indices of methylation status across the reference subset are p-values, hi some embodiments, the p-values of the reference subsets are obtained using any of the methods disclosed herein, or any suitable substitutions, modifications, additions, deletions, and / or combinations thereof.
[0273] In some embodiments, each index in the one or more indices of methylation state across the reference subset is a measure of central tendency of methylation state p-values across the reference subset, a minimum methylation state p-value across the reference subset, a maximum methylation state p-value across the variant references, or a measure of the spread of methylation state p-values across the reference subset.
[0274] For example, in some embodiments, one of the one or more indices of methylation status across the reference subsets is a measure of central tendency of the methylation status p-values across the reference subsets, where the measure of central tendency is the arithmetic mean, weighted average, mid-range, mid-hinge, triad mean, winsorized mean, average, or mode of the methylation status p-values across the reference subsets. In some embodiments, one of the one or more indices of methylation status across the reference subsets is a measure of spread of the methylation status p-values across the reference subsets, where the measure of spread is the standard deviation, variance, range, or interquartile range of the methylation status p-values across the reference subsets.
[0275] In some embodiments, applying to the trained binary classifier further applies multiple measures of methylation state across the reference subset, including at least two, at least three, or all four of: a measure of central tendency of methylation state p-values across the reference subset, a minimum methylation state p-value across the reference subset, a maximum methylation state p-value across the reference subset, and a measure of spread of methylation state p-values across the reference subset.
[0276] In some embodiments, the one or more indices of methylation state across the subsets are a plurality of indices of methylation state across the reference subsets, including a mean p-value, a median p-value, a minimum p-value, a maximum p-value, and a standard deviation of the p-values across the reference subsets.
[0277] In some embodiments, the one or more indices of methylation status across the reference subset are the highest ranked (e.g., most significant) set of p-values from the reference subset. For example, in some embodiments, the one or more indices of methylation across the reference subset include at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 of the highest ranked (e.g., most significant) p-values from the reference subset. In some embodiments, the one or more indices of methylation across the reference subset include the top 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or top 1% of the most highly ranked (e.g., most significant) p-values from the reference subset.
[0278] In some embodiments, the one or more indices of methylation state across the methylation states of each nucleic acid fragment in the reference subset comprise a methylation state vector and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the reference subset, a minimum across the reference subset, a maximum across the reference subset, and a measure of spread across the reference subset).
[0279] In some embodiments, the one or more indices of methylation state across the methylation states of each nucleic acid fragment in the reference subset include a beta value and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the reference subset, a minimum value across the reference subset, a maximum value across the reference subset, and a measure of spread across the reference subset).
[0280] In some embodiments, the one or more indices of methylation state across the methylation states of each nucleic acid fragment in the reference subset include an M value and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the reference subset, a minimum value across the reference subset, a maximum value across the reference subset, and a measure of spread across the reference subset).
[0281] In some embodiments, the one or more indices of methylation status across the methylation status of each nucleic acid fragment in the reference subset comprise an aberrant methylation score and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the reference subset, a minimum across the reference subset, a maximum across the reference subset, and a measure of spread across the reference subset).
[0282] In some embodiments, the one or more indices of methylation state across the methylation states of each nucleic acid fragment in the reference subset include a mutual information score and / or one or more distribution statistics thereof (e.g., a measure of central tendency across the reference subsets, a minimum across the reference subsets, a maximum across the reference subsets, and a measure of spread across the reference subsets). Further details regarding mutual information scores are disclosed in U.S. Provisional Patent Application No. 62 / 948,129, entitled "Cancer Classification using Patch Convolutional Neural Networks," filed December 13, 2019, which is incorporated herein by reference in its entirety.
[0283] In some embodiments, the one or more indicators of methylation states across the methylation states of each nucleic acid fragment in the reference subset comprise at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 500, at least 800, or at least 1000 indicators of methylation states across the reference subset. In some embodiments, the one or more indices of methylation states across the methylation states of each nucleic acid fragment in the reference subset comprise 2000 or less, 1000 or less, 500 or less, 200 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 40 or less, 30 or less, 20 or less, or 10 or less indices of methylation states across the reference subset. In some embodiments, the one or more indices of methylation states across the methylation states of each nucleic acid fragment in the reference subset comprise 3-10, 5-20, 10-50, 20-100, 50-200, 100-500, 300-1000, or 500-2000 indices of methylation states across the reference subset. In some embodiments, one or more indices of methylation status in the reference subset are within another range starting from 3 or more indices of methylation status across the entire reference subset and ending with 2000 or fewer indices.
[0284] Referring to block 218, in some embodiments, applying to the trained binary classifier further applies one or more CpG site indices across the reference subset, in some embodiments, the CpG site indices are CpG counts (e.g., as described above).
[0285] In some embodiments, each nucleic acid fragment sequence in the reference subset has the same CpG count. In some embodiments, two or more nucleic acid fragment sequences in the reference subset have different CpG counts. In some embodiments, each nucleic acid fragment sequence in the reference subset has at least a minimum number of CpG sites (e.g., a plurality of nucleic acid fragment sequences at each of the genomic locations are filtered using a minimum or maximum CpG count). In some embodiments, the minimum number of CpG sites is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 CpG sites. In some embodiments, the minimum number of CpG sites is 1-10, 10-20, 20-30, 30-40, 40-50, or more than 50 CpG sites.
[0286] In some embodiments, an index of the one or more CpG site indices across the reference subset includes a measure of central tendency of CpG counts across the reference subset, a minimum CpG count across the reference subset, a maximum CpG count across the reference subset, and a measure of spread of CpG counts across the reference subset.
[0287] For example, in some embodiments, one of the one or more CpG site indices across the reference subset is a measure of central tendency of the CpG counts across the reference subset, where the measure of central tendency is the arithmetic mean, weighted mean, midrange, midhinge, trimean, winsorized mean, average, or mode of the CpG counts across the reference subset. In some embodiments, one of the one or more CpG site indices across the reference subset is a measure of spread of the CpG counts across the reference subset, where the measure of spread is the standard deviation, variance, range, or interquartile range of the CpG counts across the variant subset.
[0288] In some embodiments, the application to the trained binary classifier further applies a plurality of CpG site indices across the reference subset, the plurality of CpG site indices across the reference subset comprising at least two, at least three, or all four of a measure of central tendency of CpG counts across the reference subset, a minimum CpG count across the reference subset, a maximum CpG count across the reference subset, and a measure of spread of CpG counts across the reference subset.
[0289] In some embodiments, the one or more CpG indexes across the reference subset are a plurality of CpG site indexes across the reference subset, including a CpG count across the reference subset, a median CpG count, a minimum CpG count, a maximum CpG count, and a standard deviation of the CpG counts.
[0290] In some embodiments, the one or more CpG indices across the reference subset include at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 CpG indices across the reference subset. In some embodiments, the one or more CpG indices across the reference subset include 200 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 40 or less, 30 or less, 20 or less, or 10 or less CpG indices across the reference subset. In some embodiments, the one or more CpG indices across the reference subset include 3-10, 5-20, 10-50, 20-100, or 50-200 CpG indices across the reference subset. In some embodiments, the one or more CpG indices in the reference subset are in another range starting with 3 or more CpG indices across the reference subset and ending with 200 or less CpG indices.
[0291] Referring again to block 210, in some embodiments, (ii) the indication of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset comprises a count of nucleic acid fragment sequences in the reference subset. In some embodiments, the indication of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset comprises a count of nucleic acid fragment sequences in the mutant subset. In some embodiments, the indication of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset comprises a ratio of the count of nucleic acid fragment sequences in the mutant subset compared to the count of nucleic acid fragment sequences in the reference subset.
[0292] In some embodiments, the indices for application to the trained binary classifier (e.g., one or more indices of methylation state of the variant subset, one or more indices of methylation state of the reference subset, an index of the number of nucleic acid fragment sequences in the reference subset relative to the variant subset, one or more CpG indices of the variant subset, and / or one or more CpG indices of the reference subset) are pooled (e.g., variant subset and reference subset) and binned into an input vector of genomic locations. In some embodiments, the pooled indices in the input vector are labeled as variant and / or reference.
[0293] In some embodiments, indices for application to a trained binary classifier are faceted such that indices corresponding to a variant subset are binned into a first input vector of the variant subset of genomic locations, and indices corresponding to a reference subset are binned into a second input vector of the reference subset of genomic locations.
[0294] In some cases, the indices in the input vector are applied as features to a trained binary classifier.
[0295] In some embodiments, the input vector has a fixed length. In some embodiments, the input vector has a variable length. In some embodiments, each genomic location in the plurality of genomic locations has an input vector of the same length or a different length.
[0296] In some embodiments, the input vector for each genomic location includes at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 500, at least 800, at least 1000, at least 2000, or at least 5000 indices (e.g., features). In some embodiments, the input vector for each genomic location includes 10,000 or less, 5000 or less, 2000 or less, 1000 or less, 500 or less, 200 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 40 or less, 30 or less, 20 or less, or 10 or less indices (e.g., features). In some embodiments, the input vector for each genomic location includes 3-10, 5-20, 10-50, 20-100, 50-200, 100-500, 300-1000, 500-2000, or 1000-10,000 indices. In some embodiments, the input vector for each genomic location includes a plurality of indices (e.g., features) in another range starting from 3 or more indices and ending with 10,000 or less indices.
[0297] Thus, in an exemplary implementation, identifying a mutant allele at each genomic location of a subject as a somatic mutant allele or a germline mutant allele includes providing one or more input vectors to a trained binary classifier, where the genomic locations are for candidate mutant alleles of the subject (e.g., identified as described above with reference to block 204), and the one or more input vectors include a plurality of features (e.g., indices) for each genomic location. The plurality of features may include, for example, (i) one or more p-values and / or distribution statistics thereof, (ii) an indicator of the number of mutant nucleic acid fragment sequences relative to a reference nucleic acid fragment sequence, and (iii) one or more CpG counts and / or distribution statistics thereof, obtained for a plurality of nucleic acid fragment sequences mapped to the genomic locations. The trained classifier may then provide as an output a determination of whether the mutation is a somatic mutation or a germline mutation based on the plurality of indices in the input vector.
[0298] classifier In some embodiments, the trained classifier is a trained logistic regression classifier or a multi-layer perceptron classifier.
[0299] In some embodiments, the trained classifier is a trained decision tree classifier, a trained random forest classifier, a trained support vector machine classifier, a trained k-nearest neighbor classifier, a trained nearest centroid classifier, a trained neural network classifier, or a trained Naive Bayes classifier. In some embodiments, the trained classifier is any of the classifiers disclosed in Example 3 below.
[0300] In some embodiments, the trained classifier includes a corresponding number of parameters (e.g., weights; see, e.g., Definition: Parameters).
[0301] In some embodiments, the trained classifier comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 parameters. In some embodiments, the trained classifier comprises at least 100, at least 500, at least 800, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 15,000, at least 20,000, or at least 30,000 parameters. In some embodiments, the trained classifier comprises 30,000 or less, 20,000 or less, 15,000 or less, 10,000 or less, 9000 or less, 8000 or less, 7000 or less, 6000 or less, 5000 or less, 4000 or less, 3000 or less, 2000 or less, 1000 or less, 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, 200 or less, 100 or less, or 50 or less parameters. In some embodiments, the trained classifier includes 2-20, 2-200, 2-1000, 10-50, 10-200, 20-500, 100-800, 50-1000, 500-2000, 1000-5000, 5000-10,000, 10,000-15,000, 15,000-20,000, or 20,000-30,000 parameters. In some embodiments, the trained classifier includes multiple parameters in different ranges starting from 2 or more parameters and ending with 30,000 or less parameters.
[0302] In some embodiments, the trained classifier is a neural network that includes multiple hidden layers and multiple hidden neurons. For example, in some embodiments, the trained classifier is a neural network and the multiple hidden layers include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 hidden layers. In some embodiments, the plurality of hidden layers includes 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, 40 or less, 30 or less, 20 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, or 5 or less hidden layers. In some embodiments, the plurality of hidden layers includes 1-5, 1-10, 1-20, 10-50, 2-80, 5-100, 10-100, 50-100, or 3-30 hidden layers. In some embodiments, the plurality of hidden layers falls within another range starting from 1 or more and ending with 100 or less hidden layers.
[0303] In some embodiments, the trained classifier is a neural network, and each hidden neuron in the plurality of hidden neurons is associated with one or more corresponding parameters (e.g., weights) in a corresponding plurality of parameters for the trained classifier. For example, in some embodiments, the plurality of hidden neurons includes 2-20, 2-200, 2-1000, 10-50, 10-200, 20-500, 100-800, 50-1000, 500-2000, 1000-5000, 5000-10,000, 10,000-15,000, 15,000-20,000, or 20,000-30,000 parameters. In some embodiments, the plurality of hidden neurons includes at least as many hidden neurons as there are parameters in the corresponding plurality of parameters for the classifier.
[0304] In some embodiments, the trained classifier is a neural network, and each hidden neuron in the plurality of hidden neurons is associated with a first activation function type and / or a second activation function type.
[0305] In some embodiments, the first activation function and / or the second activation function (e.g., for each hidden neuron) are selected from the group consisting of all or a combination of tanh, sigmoid, softmax, logistic, Gaussian, Boltzmann weighted average, absolute value, linear, rectified linear unit (ReLU), leaky ReLU, exponential linear unit (eLU), bounded rectified linear, soft rectified linear, parameterized rectified linear, mean, maximum, minimum, sign, square, square root, polyquadratic surface, inverse quadratic surface, inverse polyquadratic surface, polyharmonic spline, and thin plate spline.
[0306] In some embodiments, the disclosure provides methods for training a classifier (e.g., an untrained or partially untrained model) to identify a mutant allele at a test genomic location as a somatic mutant allele or a germline mutant allele.
[0307] The classifier training can be performed by obtaining an identification of a reference allele at a genomic location. For each genomic location in the plurality of genomic locations in each subject in the plurality of subjects, a procedure can be performed that includes obtaining an orthogonal call for a mutant allele at each genomic location in each subject as either a somatic mutant allele or a germline mutant allele, and obtaining an identification of the mutant allele at each genomic location in each subject. The method includes obtaining an identification of the mutant allele at each genomic location in each subject from a biological sample (e.g., at least 1×10 6 The method may further comprise obtaining a methylation state and a respective sequence of each nucleic acid fragment sequence in a respective plurality of nucleic acid fragment sequences in a sequencing dataset (including each nucleic acid fragment sequence).
[0308] Each nucleic acid fragment sequence having a reference allele at a respective genomic location among each of the plurality of nucleic acid fragment sequences can be assigned to a reference subset using (a) the identification of a reference allele at a respective genomic location and (b) the respective sequence of each nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences. Further, each nucleic acid fragment sequence having a mutant allele at a respective genomic location among each of the plurality of nucleic acid fragment sequences can be assigned to a mutant subset using (a) the identification of a mutant allele at a respective genomic location and (b) the respective sequence of each nucleic acid fragment sequence in each of the plurality of nucleic acid fragment sequences.
[0309] The method may further include training a classifier to identify a mutant allele at a genomic location in a test subject as a somatic mutant allele or a germline mutant allele using, for each genomic location in the plurality of genomic locations in each subject in the plurality of subjects, at least: (i) one or more indices of methylation state across methylation states of each nucleic acid fragment sequence in the mutant subset of each subject at each genomic location; (ii) an indices of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset of each subject at each genomic location; and (iii) an orthogonal call of the mutant allele at each genomic location in each subject as either a somatic mutant allele or a germline mutant allele.
[0310] For example, in some embodiments, the method can include applying at least (i) one or more indices of methylation state, (ii) an indicia of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset, and (iii) an orthogonal call of a mutant allele as a somatic mutant allele or a germline mutant allele to an untrained or partially untrained model, thus training a classifier to identify a mutant allele at a genomic location under test as a somatic mutant allele or a germline mutant allele.
[0311] In some embodiments, the untrained or partially untrained model comprises any of the classifiers disclosed herein (eg, above and / or in Example 3 below).
[0312] In some embodiments, the untrained or partially untrained model includes at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 parameters. In some embodiments, the untrained or partially untrained model includes at least 100, at least 500, at least 800, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 15,000, at least 20,000, or at least 30,000 parameters. In some embodiments, an untrained or partially untrained model includes 30,000 or less, 20,000 or less, 15,000 or less, 10,000 or less, 9000 or less, 8000 or less, 7000 or less, 6000 or less, 5000 or less, 4000 or less, 3000 or less, 2000 or less, 1000 or less, 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, 200 or less, 100 or less, or 50 or less parameters. In some embodiments, the untrained or partially untrained model includes 2-20, 2-200, 2-1000, 10-50, 10-200, 20-500, 100-800, 50-1000, 500-2000, 1000-5000, 5000-10,000, 10,000-15,000, 15,000-20,000, or 20,000-30,000 parameters.In some embodiments, the untrained or partially untrained model includes a number of parameters in another range starting from 2 or more parameters and ending with 30,000 or fewer parameters.
[0313] In some embodiments, the plurality of training subjects includes at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 subjects. In some embodiments, the plurality of training subjects includes at least 100, at least 500, at least 800, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, or at least 20,000 subjects. In some embodiments, the plurality of training objects includes 20,000 or less, 10,000 or less, 5000 or less, 4000 or less, 3000 or less, 2000 or less, 1000 or less, 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, or 200 or less objects. In some embodiments, the plurality of training objects includes 20-500, 100-800, 50-1000, 500-2000, 1000-5000, or 5000-10,000 objects. In some embodiments, the plurality of training objects is in another range starting with 20 or more objects and ending with 20,000 or less objects.
[0314] In some embodiments, training the classifier comprises using a training dataset for a plurality of training subjects. In some embodiments, the training dataset comprises a plurality of nucleic acid fragment sequences for each of the plurality of training subjects in electronic form. In some embodiments, obtaining the plurality of nucleic acid fragment sequences is performed for each of the plurality of training subjects using any of the methods disclosed herein and / or any suitable substitutions, modifications, additions, deletions, and / or combinations thereof.
[0315] In some embodiments, the method includes obtaining a plurality of biological samples for each training subject in the plurality of training subjects, and each biological sample in the plurality of biological samples of each subject is used to obtain a respective plurality of nucleic acid fragment sequences. For example, in some embodiments, a first plurality of nucleic acid fragment sequences can be obtained from a first biological sample (e.g., cell-free nucleic acid from a liquid biological sample), and a second plurality of nucleic acid fragment sequences can be obtained from a second matched biological sample (e.g., a healthy tissue sample or a solid tumor sample) from the same respective training subject.
[0316] In some embodiments, the method includes, for each training subject among the plurality of training subjects, sequencing each biological sample obtained from each training subject using a plurality of sequencing methods, each sequencing method generating a respective plurality of nucleic acid fragment sequences. For example, in some embodiments, a first plurality of nucleic acid fragment sequences can be obtained from a first sequencing method (e.g., WGS) of each biological sample obtained from each training subject, and a second plurality of nucleic acid fragment sequences can be obtained from a second sequencing method (e.g., WGBS and / or targeted methylation) of each biological sample obtained from each training subject.
[0317] In some embodiments, any number of matched samples and / or matched sequencing assays can be carried out for each training subject in a plurality of training subjects.For example, in some embodiments, a first plurality of nucleic acid fragment sequences can be obtained using a first sequencing method of a first biological sample of each training subject (e.g., WGS on healthy tissue sample), and a second plurality of nucleic acid fragment sequences can be obtained using a second sequencing method other than the first sequencing method of a second biological sample different from the first biological sample from each training subject (e.g., target methylation on cfDNA in liquid biological sample).
[0318] In some embodiments, the classifier is trained using a training dataset obtained from the same biological sample type as the sequencing dataset of the test subject. For example, in some embodiments, the classifier is trained using nucleic acid fragment sequences derived from solid tissue samples from multiple training subjects, and the method of identifying mutations as somatic or germline mutations using the trained classifier is performed using nucleic acid fragment sequences derived from solid tissue samples from the test subject. In some embodiments, the classifier is trained using a training dataset obtained from a different biological sample type as the sequencing dataset of the test subject. For example, in some forms, the classifier is trained using nucleic acid fragment sequences derived from solid tissue samples from multiple training subjects, and the method of identifying mutations as somatic or germline mutations using the trained classifier is performed using cell-free nucleic acid fragment sequences derived from liquid biological samples from the test subject.
[0319] Alternatively or additionally, in some embodiments, the classifier is trained using a training dataset obtained by the same sequencing method as that used for the test subject. For example, in some embodiments, the classifier is trained using nucleic acid fragment sequences obtained from whole genome sequencing (WGS) of tissue samples from multiple training subjects, and using the trained classifier to identify mutations as somatic mutations or germline mutations is performed using nucleic acid fragment sequences obtained from whole genome sequencing (WGS) of tissue samples from the test subject. In some embodiments, the classifier is trained using a training dataset obtained by a different sequencing method than that used for the test subject. For example, in some embodiments, the classifier is trained using nucleic acid fragment sequences obtained from whole genome sequencing (WGS) of tissue samples from multiple training subjects, and using the trained classifier to identify mutations as somatic mutations or germline mutations is performed using nucleic acid fragment sequences obtained from targeted methylation of cell-free nucleic acids in a liquid biological sample from the test subject.
[0320] In some embodiments, the training dataset further comprises a tumor fraction and / or a tumor mutation burden for each training subject in the plurality of training subjects.
[0321] As defined above, tumor fraction can refer to the fraction of nucleic acid molecules in a sample derived from a subject's cancerous tissue compared to non-cancerous tissue (see definition: "Tumor Fraction"). The tumor fraction can be expressed as a value between 0 and 1 or can be converted to a percentage (e.g., 0 to 100). In some embodiments, the tumor fraction is 10 -6 In some embodiments, the tumor fraction is 10 -5 In some embodiments, the tumor fraction is 10 -4 In some embodiments, the tumor fraction is 0.001 to 0.999. In some embodiments, the tumor fraction is 0.01 to 0.99. In some embodiments, the tumor fraction is 10 -5 ~0.04, 10-4 In some embodiments, the tumor fraction is 0.3 or less, 0.2 or less, 0.1 or less, 0.09 or less, 0.08 or less, 0.07 or less, 0.06 or less, 0.05 or less, 0.04 or less, 0.03 or less, 0.02 or less, 0.01 or less, 0.009 or less, 0.008 or less, 0.007 or less, 0.006 or less, 0.005 or less, 0.004 or less, 0.003 or less, 0.002 or less, 0.001 or less, 10 -4 Less than or equal to 10 -5 In some embodiments, the tumor fraction is at least 10 -4 , at least 0.001, at least 0.005, at least 0.01, at least 0.05, at least 0.1, at least 0.2, at least 0.3, or at least 0.5. In some embodiments, the tumor fraction is 10 -6 It is within another range that starts greater than or equal to 0.999 and ends less than or equal to 0.999.
[0322] As defined above, tumor mutation burden refers to a measure of mutations in a cancer per unit of a patient's genome (see definition: "Tumor Mutation Burden"). In some embodiments, tumor mutation burden is measured in number of mutations per megabase (Mb) (e.g., of a patient's genome and / or coding sequence). In some embodiments, tumor mutation burden is 0.0001-5, 0.001-5, 0.001-1, or 0.1-5 mutations per Mb. In some embodiments, tumor mutation burden is 5-10 mutations per Mb. In some embodiments, tumor mutation burden is 10-20, 10-30, 10-50, or 10-100 mutations per Mb. In some embodiments, the tumor mutation burden is 50 or less, 30 or less, 20 or less, 10 or less, 9 or less, 8 or less, 7 or less, 6 or less, 5 or less, 4 or less, 3 or less, 2 or less, 1 or less, 0.5 or less, 0.1 or less, 0.05 or less, 0.01 or less, 0.005 or less, 0.001 or less, 0.0005 or less, or 0.0001 or less mutations per Mb. In some embodiments, the tumor mutation burden is at least 0.001, at least 0.005, at least 0.01, at least 0.05, at least 0.1, at least 0.5, at least 1, at least 5, or at least 10 mutations per Mb. In some embodiments, the tumor mutation burden is in another range starting at 0.0001 or more mutations per Mb and ending at 100 or less mutations per Mb.
[0323] In some embodiments, the training dataset includes weighting factors and / or dilution factors (e.g., to account for differences in sample type and / or tumor fraction) for one or more training subjects in the plurality of training subjects.
[0324] In some embodiments, the training dataset is filtered (e.g., using any of the filters disclosed herein, e.g., see the section above entitled "Assigning Subsets"); in some embodiments, the filtering comprises removing genomic locations from the plurality of genomic locations across all training subjects in the plurality of training subjects.
[0325] In some embodiments, filtering comprises removing a training object from the plurality of training objects. For example, in some embodiments, if all of the genomic positions in the plurality of genomic positions of each training object cannot meet the filtering criteria (e.g., all of the genomic positions of the training object are removed from the dataset), the corresponding plurality of nucleic acid fragment sequences of each training object are removed from the dataset.
[0326] Any suitable sample type, tissue type, sampling, sequencing method, processing and / or bioinformatics analysis can be used to obtain a training dataset for one or more training subjects for the test subjects and / or any substitutions, modifications, additions, deletions, and / or combinations thereof disclosed herein.
[0327] In some embodiments, other aspects such as training of the classifier (e.g., for each genomic location in the plurality of genomic locations in each subject in the plurality of subjects), inclusion of subjects and samples, obtaining identification of mutant and reference alleles, sequencing (e.g., methylation sequencing), processing of nucleic acid fragment sequences, obtaining methylation states, assigning reference and mutant subsets, obtaining features, etc., are performed using any of the methods disclosed herein for systems and methods for identifying mutant alleles as somatic mutant alleles or germline mutant alleles (including, e.g., inclusion of subjects and samples, obtaining identification of mutant and reference alleles, sequencing (e.g., methylation sequencing), processing of nucleic acid fragment sequences, obtaining methylation states, assigning reference and mutant subsets, obtaining features, etc.), and / or using any suitable substitutions, modifications, additions, deletions, and / or combinations thereof.
[0328] As noted above, in some embodiments, training the classifier includes, for each genomic location in the plurality of genomic locations in each subject in the plurality of subjects, obtaining an orthogonal call for a mutant allele at each genomic location as either a somatic mutant allele or a germline mutant allele. Thus, the training dataset includes, for each genomic location of a mutation of interest in each subject, a corresponding label that the mutation is a somatic mutation or a germline mutation.
[0329] In some embodiments, the orthogonal call for the mutant allele is determined using a comparison between the abnormal sample and the reference sample. For example, as described in Example 6 below, in some embodiments, the orthogonal call for the mutant allele is determined using an analysis between the patient-matched tumor sample and a normal tissue reference. The orthogonal call (e.g., somatic label or germline label) is then used as input together with the multiple indices of each training subject to train a classifier.
[0330] In general, training a classifier (e.g., a logistic regression model, a neural network, and / or another suitable model) involves updating multiple parameters of each classifier by backpropagation (e.g., gradient descent). First, forward propagation is performed in which input data is accepted into an untrained or partially untrained model and outputs are calculated based on a selected activation function and an initial set of parameters (e.g., weights). Next, a backward pass can be performed by calculating an error gradient for each parameter, and the error for each parameter is determined by calculating a loss (e.g., error) based on the output (e.g., predicted value) and the input data (e.g., expected value or true label).
[0331] The parameters can then be updated by adjusting the values (e.g., small adjustments vs. large adjustments) based on the calculated loss as measured by a predefined learning rate hyperparameter that specifies the degree or severity to which the parameters are updated, thereby training the untrained or partially untrained model.
[0332] For example, in some common embodiments of machine learning, backpropagation is a method of training an untrained or partially untrained model that includes multiple parameters (e.g., embeddings). The output of the untrained or partially untrained model (e.g., identification of a mutation as a somatic mutation or a germline mutation) can be generated using an arbitrarily selected set of initial parameters. The output is then compared to the original input (e.g., orthogonal calls of each training subject's mutant allele at each genomic location) by evaluating an error function (e.g., using a loss function) to calculate an error. The parameters can then be updated to minimize the error (e.g., according to the loss function). In some embodiments, any one of a variety of backpropagation algorithms and / or methods is used to update the multiple parameters.
[0333] In some embodiments, the error is calculated using an error function (e.g., a loss function). In some embodiments, the loss function is mean squared error, quadratic loss, mean absolute error, mean bias error, hinge, multi-class support vector machine, and / or cross entropy. In some embodiments, training the untrained or partially untrained model includes calculating the error according to a gradient descent algorithm and / or a minimization function.
[0334] In some embodiments, the error function is used to update one or more parameters of an untrained or partially untrained model by adjusting the value of the one or more parameters by an amount proportional to the calculated loss, thereby training the model. In some embodiments, the amount by which the parameters are adjusted (e.g., a smaller or larger adjustment) is scaled by a pre-defined learning rate that dictates the extent or severity to which the parameters are updated. In some embodiments, the learning rate is a hyperparameter that can be selected by a physician.
[0335] In some embodiments, training the untrained or partially untrained model forms a trained classifier according to a first evaluation of the error function. In some such embodiments, training the untrained or partially untrained model forms a trained classifier according to a first update of one or more parameters based on the first evaluation of the error function. In some alternative embodiments, training the untrained or partially untrained model forms a trained classifier according to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 500, at least 1000, at least 10,000, at least 50000, at least 100,000, at least 200,000, at least 500,000, or at least 1 million evaluations of the error function. In some such embodiments, training an untrained or partially untrained model includes increasing or decreasing at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 500, at least 1000, at least 10,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, or at least 1 In one embodiment, a trained classifier is formed according to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 500, at least 1000, at least 10,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, or at least 1 million updates of one or more parameters based on 1 million evaluations.
[0336] In some embodiments, training the untrained or partially untrained model forms a trained classifier if the model meets a minimum performance requirement. For example, in some embodiments, training the untrained or partially untrained model forms a trained classifier if the error calculated for the trained classifier meets an error threshold according to an evaluation of an error function across one or more training datasets for each one or more training subjects. In some embodiments, the error calculated by the error function across one or more training datasets for each one or more training subjects meets the error threshold if the error is less than 20 percent, less than 18 percent, less than 15 percent, less than 10 percent, less than 5 percent, or less than 3 percent.
[0337] In some embodiments, the minimum performance requirement is met based on validation training, hi some embodiments, the validation training is performed by K-fold cross-validation.
[0338] In some embodiments, classifier training is performed on multiple machines (e.g., computers and / or systems). In some embodiments, use of the classifier for mutant alleles at test genomic locations as somatic mutant alleles or germline mutant alleles is performed on multiple machines (e.g., computers and / or systems).
[0339] In some embodiments, the classifier training further includes fixing (e.g., freezing) one or more parameters in the plurality of parameters, thereby obtaining a corresponding trained classifier that can be used to determine and / or classify (e.g., a mutant allele at a genomic location as a somatic mutant allele or a germline mutant allele).
[0340] As will be apparent to one of ordinary skill in the art, any other model parameters and architectures suitable for training are contemplated.
[0341] Purpose Referring to block 220, in some embodiments, if the mutant allele at the genomic location is determined by the trained binary classifier to be a germline mutant allele, the method further includes determining a risk of cancer in the test subject using the mutant allele in the test subject. For example, in some embodiments, the genomic location is the BRCA1 locus or the BRCA2 locus, and the mutant allele at the genomic location is determined by the trained binary classifier to be a germline mutant allele, and the method further includes determining that the test subject is at risk for breast cancer.
[0342] Referring to block 222, in some embodiments, if the mutant allele at the genomic location is determined to be a germline mutant allele by the trained binary classifier, the method further includes predicting the ethnicity of the subject using the mutant allele in the test subject. For example, germline mutations in cancer genes have been reported to be ethnic-specific, with different mutant alleles at a given locus being overrepresented in various ethnic groups. Thus, for each subject, the mutant allele at the locus of the cancer gene (e.g., BRCA1 or BRCA2) can be used to determine the ethnicity and assess the cancer risk for each ethnicity.
[0343] In some embodiments, if the mutant allele at the genomic location is determined to be a somatic mutant allele by the trained binary classifier, the method further comprises using the mutant allele in the test subject to make a clinical determination of the disease. In some implementations, the clinical determination of the disease is diagnosis, determining the stage of the disease, monitoring the progression, prognosis, prescribing or administering a treatment, matching or recommending enrollment in a clinical trial, monitoring the occurrence of further complications or risks over time, and / or evaluating the efficacy of a treatment. In some embodiments, the disease is cancer. In some embodiments, the disease is indeterminate clonal hematopoiesis (CHIP), cardiovascular risk, non-alcoholic fatty liver disease (NAFLD), and / or non-alcoholic steatohepatitis (NASH).
[0344] For example, in some embodiments, the genomic location is a KRAS locus and the mutant allele at the genomic location is determined to be a somatic mutant allele by a trained binary classifier, and the method further includes using the mutant allele to diagnose the patient as having cancer (e.g., pancreatic cancer, colorectal cancer, and / or lung cancer).
[0345] In some embodiments, if the mutant allele at the genomic location is determined to be a somatic mutant allele by the trained binary classifier, the method further comprises using the mutant allele in the test subject to determine the tumor mutation burden of the subject (e.g., a normalized count of somatic mutations per base pair unit). Typical methods for calculating tumor mutation burden generally utilize a tumor sample and a normal control sample (e.g., a normal reference). In some embodiments, the method provides a complementary method (e.g., using a liquid biological sample) for determining the tumor mutation burden of the subject using the mutant allele in the test subject.
[0346] Referring to block 224, in some embodiments, if the mutant allele at the genomic location is determined to be a somatic mutant allele by the trained binary classifier, the method further includes using the mutant allele in the test subject to determine the subject's tumor fraction. For example, in some embodiments, when the biological sample of each test subject is derived from cell-free nucleic acid, the cell-free nucleic acid may indicate a significant tumor fraction. In some embodiments, the corresponding tumor fraction in each test subject is at least 2 percent, at least 5 percent, at least 10 percent, at least 15 percent, at least 20 percent, at least 25 percent, at least 50 percent, at least 75 percent, at least 90 percent, at least 95 percent, or at least 98 percent. In some embodiments, the corresponding tumor fraction in each test subject is 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, 1% or less, or 0.1% or less. In some such embodiments, such tumor fraction estimates are used to detect cancer in the subject, as described in Example 3 below.
[0347] The tumor fraction and / or tumor mutation burden, in some implementations, can be used for further diagnostic applications, for example, the tumor fraction and / or tumor mutation burden can be used to evaluate or monitor the effectiveness of cancer treatments (e.g., chemotherapy, immunotherapy, etc.).
[0348] In some embodiments, the method includes obtaining a tumor fraction estimate for the test subject at a first time point and a second time point, and the diagnosis of the test subject is changed if the subject's tumor fraction estimate is observed to change by a threshold amount between the first time point and the second time point. For example, in some embodiments, the diagnosis is changed from having cancer to being in remission. As another example, in some embodiments, the diagnosis is changed from not having cancer to having cancer. As another example, in some embodiments, the diagnosis is changed from being in a first stage of cancer to being in a second stage of cancer. As another example, in some embodiments, the diagnosis is changed from being in a second stage of cancer to being in a third stage of cancer. As yet another example, in some embodiments, the diagnosis is changed from being in a third stage of cancer to being in a fourth stage of cancer. As yet another example, in some embodiments, the diagnosis is changed from having cancer that has not metastasized to having cancer that has metastasized.
[0349] In some embodiments, when the subject's tumor fraction estimate is observed to change by a threshold amount between the first time point and the second time point, the prognosis of the test subject is changed.For example, in some embodiments, the prognosis comprises life expectancy, and the prognosis is changed from a first life expectancy to a second life expectancy, and the first life expectancy and the second life expectancy are different in duration.In some embodiments, the change in prognosis increases the subject's life expectancy.In some embodiments, the change in prognosis decreases the subject's life expectancy.
[0350] In some embodiments, if the subject's tumor fraction estimate is observed to change by a threshold amount between the first and second time points, the treatment of the test subject is altered, in some embodiments, the treatment alteration comprises initiating a cancer therapeutic, increasing the dosage of the cancer therapeutic, ceasing the cancer therapeutic, or decreasing the dosage of the cancer therapeutic.
[0351] In some embodiments, a treatment regimen is administered to the subject based at least in part on the value of the test subject's tumor fraction estimate and / or the identification of the mutation at the genomic location as a somatic mutation or a germline mutation. For example, in some embodiments, the method further comprises administering a first treatment to the test subject if the mutant allele at the genomic location is determined by the trained binary classifier to be a somatic mutant allele, and administering a second treatment to the test subject if the mutant allele at the genomic location is determined by the trained binary classifier to be a germline mutant allele.
[0352] In some embodiments, the treatment regimen comprises administering an anti-cancer drug to the test subject. In some embodiments, the anti-cancer drug is a hormone, an immunotherapy, a radiography, or a cancer treatment drug. In some embodiments, the anti-cancer drug is lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, human papillomavirus tetravalent (types 6, 11, 16, and 18), a vaccine, pertuzumab, pemetrexed, nilotinib, denosumab, abiraterone acetate, eltrombopag, imatinib, everolimus, palbociclib, erlotinib, bortezomib, bortezomib, or a generic equivalent thereof.
[0353] In some embodiments, the test subject has been treated with an agent for cancer, and the tumor fraction estimate of the test subject and / or the identification of the mutation at the genomic location as a somatic mutation or a germline mutation is used to assess the subject's response to the agent for cancer, details of which are provided elsewhere herein.
[0354] In some embodiments, the test subject is being treated with a drug for cancer, and the tumor fraction estimate of the test subject and / or the identification of the mutation at the genomic location as a somatic mutation or a germline mutation is used to determine whether to enhance or discontinue the drug for cancer in the test subject. For example, in some embodiments, the observation of at least a tumor fraction estimate (e.g., greater than 0.05, greater than 0.10, greater than 0.15, greater than 0.20, greater than 0.25, or greater than 0.30, etc.) is used as a basis for enhancing the drug for cancer in the test subject (e.g., increasing the dose, increasing the radiation level in radiotherapy, etc.). In some embodiments, the observation of a threshold tumor fraction estimate (e.g., less than 0.30, less than 0.25, less than 0.20, less than 0.15, less than 0.10, less than 0.05, or less than 0.01, etc.) is used as a criterion for ceasing the use of the drug for cancer in the test subject.
[0355] In some embodiments, the test subject has undergone a surgical intervention to address cancer, and the test subject's tumor fraction estimate and / or the identification of the mutation at the genomic location as a somatic or germline mutation is used to assess the test subject's status in response to the surgical intervention. In some embodiments, the status is a metric based on the tumor fraction estimate and / or the identification of the mutation at the genomic location as a somatic or germline mutation using the methods provided in the present disclosure.
[0356] Methods for determining tumor fraction and tumor mutation burden are described in further detail in U.S. Patent Application No. 17 / 185885, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 25, 2021, and PCT Application No. PCT / US2021 / 019746, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 2021, each of which is incorporated by reference in its entirety herein.
[0357] In some embodiments, the systems and methods of the present disclosure include detecting contamination using identification of a mutation at a genomic location of a test subject as a somatic mutation or a germline mutation. For example, in some embodiments, techniques disclosed in U.S. Patent Application No. 15 / 900645, entitled "Detecting cross-contamination in sequencing data using regression techniques," filed February 20, 2018 and published as U.S. Patent Application Publication No. 2018 / 0237838, entitled "Detecting cross-contamination in sequencing data," filed June 26, 2018 and published as U.S. Patent Application Publication No. 2018 / 0373832, and / or U.S. Patent Application No. 63 / 080,670, entitled "Detecting cross-contamination in sequencing data," filed September 18, 2020, are used to detect cross-contamination using identification of a mutation at a genomic location of a test subject as a somatic mutation or a germline mutation.
[0358] Additional Embodiments Referring to block 226, in some embodiments, the method further includes repeating the method for each genomic location in the plurality of genomic locations, thereby identifying a plurality of test mutations, and for each mutation in the plurality of mutations, identifying whether the each mutation is a somatic mutation or a germline mutation.
[0359] In some embodiments, the plurality of mutations comprises 200 mutations.
[0360] In some embodiments, the plurality of mutations comprises at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, or at least 20,000 mutations. In some embodiments, the plurality of mutations comprises 20,000 or less, 10,000 or less, 5000 or less, 4000 or less, 3000 or less, 2000 or less, 1000 or less, 900 or less, 800 or less, 700 or less, 600 or less, 500 or less, 400 or less, 300 or less, 200 or less, 100 or less, 90 or less, 80 or less, 70 or less, 60 or less, 50 or less, or 20 or less mutations. In some embodiments, the plurality of mutations is 10-50, 50-100, 100-500, 500-1000, 1000-5000, 5000-10,000, or 10,000-20,000 mutations. In some embodiments, the plurality of mutations falls within another range beginning with 10 or more mutations and ending with 20,000 or less mutations.
[0361] In some embodiments, each mutation in the plurality of mutations is a clinically actionable mutation (e.g., an oncogene). Suitable embodiments for clinically actionable mutations can include any of the embodiments disclosed herein (see, e.g., the section entitled "Reference Alleles and Mutant Alleles" above). In some embodiments, the plurality of mutations is a panel of clinically actionable mutations (e.g., an oncogene of interest).
[0362] In some embodiments, the plurality of mutations are filtered. Suitable methods for filtering the plurality of mutations include any of the embodiments for filtering mutation calls, genomic locations, and / or nucleic acid fragment sequences disclosed in detail herein (see, e.g., the above sections entitled "Mutation Calling," "Subset Assignment," and "Indicator Input"), or any substitutions, modifications, additions, deletions, and / or combinations thereof, as will be apparent to one of skill in the art.
[0363] In some embodiments, the method further comprises removing each mutation from the plurality of mutations if the each mutation does not satisfy the quality metric.
[0364] In some embodiments, the quality metric is the minimum variant allele fraction in each of the plurality of nucleic acid fragment sequences in electronic form that are mapped to the genomic location of each variant call. In some embodiments, the minimum variant allele fraction is 10 percent.
[0365] In some embodiments, the quality metric is the maximum mutant allele fraction in each of the plurality of nucleic acid fragment sequences in electronic form that are mapped to the genomic location of each mutation. In some embodiments, the maximum mutant allele fraction is 90 percent.
[0366] In some embodiments, the quality metric is the minimum depth of each of the plurality of nucleic acid fragment sequences that are mapped to the genomic location of each mutation. In some embodiments, the minimum depth is 10.
[0367] Additional embodiments of quality metrics contemplated for use in the present disclosure include the quality metrics described in the section "Mutation Calling" above.
[0368] Another aspect of the present disclosure provides a computing system comprising one or more processors and a memory storing one or more programs executed by the one or more processors, the one or more programs including instructions for performing any of the methods disclosed above, either alone or in combination.
[0369] Yet another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs configured to be executed by a computer, the one or more programs including instructions for performing any of the methods disclosed above, either alone or in combination.
[0370] Additional Exemplary Embodiments Example 1 - Obtaining multiple sequence reads 7 is a flow chart of a method 700 for preparing a nucleic acid sample for sequencing according to some embodiments of the present disclosure. Method 700 includes, but is not limited to, the following steps. For example, any step of method 700 may include a quantification substep for quality control or any other laboratory assay procedure.
[0371] Referring to block 702, a nucleic acid sample (DNA or RNA) is extracted from a subject. The sample may be any subset of the human genome, including the entire genome. The sample may be extracted from a subject known to have cancer or suspected of having cancer. The sample may include blood, plasma, serous fluid, urine, feces, saliva, other types of bodily fluids, or any combination thereof. In some embodiments, methods for taking blood samples (e.g., syringe or finger prick) may be less invasive than procedures for obtaining tissue biopsies, which may use surgery. The extracted sample may include cfDNA and / or ctDNA. In healthy individuals, the human body can naturally remove cfDNA and other cellular debris. If the subject has cancer or disease, ctDNA in the extracted sample may be present at detectable levels for diagnosis.
[0372] Referring to block 704, a sequencing library was prepared. During library preparation, unique molecular identifiers (UMIs) were added to nucleic acid molecules (e.g., DNA molecules) by adaptor ligation. UMIs are short nucleic acid sequences (e.g., 4-10 base pairs) that are added to the ends of DNA fragments during adaptor ligation. In some embodiments, the UMIs were degenerate base pairs that served as unique tags that could be used to identify sequence reads derived from specific DNA fragments. During PCR amplification after adaptor ligation, the UMIs were replicated along with the bound DNA fragments. This provided a way to identify sequence reads derived from the same original fragment in downstream analysis.
[0373] Referring to block 706, target DNA sequences were enriched from the library. During enrichment, hybridization probes (also referred to herein as "probes") were used to target and pull down nucleic acid fragments that are highly informative for the presence or absence of cancer (or disease), the state of cancer, or the classification of cancer (e.g., the class of cancer or tissue of origin). For a given workflow, in some embodiments, the probes were designed to anneal (or hybridize) to a target (complementary) strand of DNA. In some embodiments, each probe was 8-5000 bases long, 12-2500 bases long, or 15-1225 bases long. In some embodiments, the target strand has a "positive" strand (e.g., the strand that is transcribed into mRNA and subsequently translated into protein) or a complementary "negative" strand. In some embodiments, the probes may range in length from tens, hundreds, or thousands of base pairs.
[0374] In some embodiments, the probes were designed based on a panel of methylation sites.
[0375] In some embodiments, the probes were designed based on a panel of target genes and / or genomic regions to analyze specific mutations or target regions of the genome (e.g., of a human or another organism) suspected to correspond to a particular cancer or other type of disease. For example, in some embodiments, each of the probes was uniquely mapped to a genomic region described in International Patent Publication Nos. WO2020154682A3, WO2020 / 069350A1, or WO2019 / 195268A2, each of which is incorporated herein by reference.
[0376] In some embodiments, the probes covered overlapping portions of the target region. Referring to block 708, in some embodiments, the probes were used to generate sequence reads for the nucleic acid sample.
[0377] FIG. 8 is a graphical representation of a process for obtaining sequence reads, according to one embodiment. FIG. 8 shows an example of a nucleic acid segment 800 from a sample. The nucleic acid segment 800 can be a single-stranded nucleic acid segment. In some embodiments, the nucleic acid segment 800 was a double-stranded cfDNA segment. The illustrated example shows three regions 805A, 805B, and 805C of the nucleic acid segment that can be targeted by different probes. Specifically, each of the three regions 805A, 805B, and 805C includes an overlapping position on the nucleic acid segment 800. An exemplary overlapping position is shown in FIG. 8 as a cytosine ("C") nucleotide base 802. The cytosine nucleotide base 802 is located near a first edge of region 805A, in the center of region 805B, and near a second edge of region 805C.
[0378] In some embodiments, one or more (or all) of the probes are designed based on a gene panel or methylation site panel to analyze a specific mutation or target region of the genome (e.g., of a human or another organism) suspected to correspond to a particular cancer or other type of disease. Method 800 may be used to increase the sequencing depth of the target region by using a target gene panel or methylation site panel rather than sequencing all expressed genes of the genome, also known as "whole exome sequencing", where depth refers to the count of the number of times a given target sequence in a sample is sequenced. Increasing the sequencing depth reduces the input amount of nucleic acid sample used. For example, in some embodiments, the target gene panel or methylation site panel includes multiple probes, each of which uniquely maps to a genomic region described in International Patent Publication Nos. WO2020154682A3, WO2020 / 069350A1, or WO2019 / 195268A2, each of which is incorporated herein by reference.
[0379] Hybridization of the nucleic acid sample 800 using one or more probes results in the understanding of a target sequence 870. As shown in FIG. 8, the target sequence 870 is a nucleotide base sequence of a region 805 targeted by a hybridization probe. The target sequence 870 may also be referred to as a hybridized nucleic acid fragment. For example, target sequence 870A corresponds to region 805A targeted by a first hybridization probe, target sequence 870B corresponds to region 805B targeted by a second hybridization probe, and target sequence 870C corresponds to region 805C targeted by a third hybridization probe. Assuming that a cytosine nucleotide base 802 is located at a different position within each region 805A-C targeted by a hybridization probe, each target sequence 870 includes a nucleotide base that corresponds to a cytosine nucleotide base 802 at a particular position on the target sequence 870.
[0380] After the hybridization step, the hybridized nucleic acid fragments can be captured and amplified using PCR. For example, the target sequence 870 can be enriched to obtain enriched sequences 880 that can then be sequenced. In some embodiments, each enriched sequence 880 was replicated from the target sequence 870. The enriched sequences 880A and 880C amplified from the target sequences 870A and 870C, respectively, also contain a thymine nucleotide base located near the edge of each sequence read 880A or 880C. As used below, a mutated nucleotide base (e.g., a thymine nucleotide base) in the enriched sequence 880 that is mutated relative to a reference allele (e.g., a cytosine nucleotide base 802) was considered as an alternative allele. Additionally, each enriched sequence 880B amplified from the target sequence 870B contained a cytosine nucleotide base located near or at the center of each enriched sequence 880B.
[0381] Referring again to block 708 of FIG. 7, sequence reads are generated from the enriched DNA sequences, such as, for example, the enriched sequences 880 shown in FIG. 8. Sequencing data can be obtained from the enriched DNA sequences. For example, method 800 can include next-generation sequencing (NGS) techniques, including synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing was performed using sequencing by synthesis with reversible dye terminators.
[0382] In some embodiments, the sequence reads are aligned to the reference genome using methods known in the art to determine alignment position information. The alignment position information may indicate the start and end positions of a region in the reference genome that corresponds to the start and end nucleotide bases of a given sequence read. The alignment position information may also include the sequence read length, which may be determined from the start and end positions. The region in the reference genome may be associated with a gene or a segment of a gene.
[0383] In some embodiments, the average sequence read length of the corresponding multiple sequence reads obtained by methylation sequencing for each fragment was between 140 and 280 nucleotides.
[0384] In various embodiments, the sequence read is composed of a read pair, denoted as R1 and R2. For example, the first read R1 may be sequenced from a first end of the nucleic acid fragment, and the second read R2 may be sequenced from a second end of the nucleic acid fragment. Thus, the nucleotide base pairs of the first read R1 and the second read R2 may be aligned (e.g., in opposite orientation) with the nucleotide bases of the reference genome. The alignment position information derived from the read pair R1 and R2 may include a start position (e.g., R1) in the reference genome corresponding to the end of the first read and an end position (e.g., R2) in the reference genome corresponding to the end of the second read. In other words, the start and end positions in the reference genome represent positions in the reference genome to which the nucleic acid fragment is likely to correspond. An output file in SAM (sequence alignment map) format or BAM (binary) format may be generated and output for further analysis, such as methylation status determination.
[0385] Example 2 - Generation of methylation state vectors according to some embodiments of the present disclosure FIG. 9 is a flow chart illustrating a process 900 for sequencing fragments of cfDNA to obtain a methylation state vector, according to one embodiment of the present disclosure.
[0386] Referring to block 902, cfDNA fragments were obtained from a biological sample. Referring to block 920, the cfDNA fragments were treated to convert unmethylated cytosines to uracil. In some embodiments, the cfDNA was subjected to bisulfite treatment to convert unmethylated cytosines of cfDNA fragments to uracil without converting methylated cytosines. For example, in some embodiments, a commercially available kit such as EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, or EZ DNA Methylation™-Lightning kit (available from Zymo Research Corp, Irvine, CA) was used for bisulfite conversion. In other embodiments, the conversion of unmethylated cytosines to uracil was achieved using an enzymatic reaction. For example, the conversion can use a commercially available kit for converting unmethylated cytosines to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).
[0387] A sequencing library is prepared from the converted cfDNA fragments (Block 930). Optionally, the sequencing library is enriched for cfDNA fragments or genomic regions that are highly informative about the cancer state using multiple hybridization probes (Block 935). Hybridization probes are short oligonucleotides that can hybridize to specific cfDNA fragments or target regions and enrich those fragments or regions for subsequent sequencing and analysis. Hybridization probes can be used to perform targeted deep analysis of a specific set of CpG sites of interest to a researcher. Once prepared, the sequencing library or a portion thereof can be sequenced to obtain multiple sequence reads (Block 940). The sequence reads can be in a computer-readable digital format for processing and interpretation by computer software.
[0388] From the sequence reads, the location and methylation state of each of the CpG sites was determined based on an alignment of the sequence reads to the reference genome (Block 950). A methylation state vector for each fragment (Block 960) was generated that identifies the location of the fragment in the reference genome (e.g., identified by the location of the first CpG site in each fragment, or another similar metric), the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment.
[0389] Example 3 - Ability to detect cancer as a function of cfDNA fraction In some embodiments, the method further includes training a classifier to determine the subject's cancer status or the likelihood that the subject will have a cancer status using at least the tumor fraction estimation information associated with the plurality of variant calls (e.g., based at least in part on one or more respective called variants identified as somatic and / or germline variants for one or more corresponding allelic positions in the subject).
[0390] For example, in some embodiments, an untrained classifier is trained on a training set that includes one or more reference plurality of mutation calls (e.g., identified as somatic and / or germline mutations), where each reference plurality of mutation calls is associated with corresponding tumor fraction estimation information.
[0391] In some embodiments, the classifier is a logistic regression. In some embodiments, the classifier is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted trees algorithm, a random forest algorithm, a decision tree algorithm, a multinomial logistic regression algorithm, a linear model, or a linear regression algorithm.
[0392] Classifiers for use in some embodiments are described in further detail, for example, in U.S. Patent Application No. 17 / 119,606, filed December 11, 2020, and U.S. Patent Application Publication No. 2019-0385813 A1, entitled "Systems and Methods for Estimating Cell Source Fractions Using Methylation Information," filed December 18, 2020, each of which is incorporated by reference in its entirety into this specification.
[0393] In some embodiments, the classifier is based on a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, or a logistic regression algorithm, a mixture model, or a hidden Markov model. In some embodiments, the trained classifier is a polynomial classifier.
[0394] In some embodiments, the classifier utilized the B-score classifier described in U.S. Patent Application Publication No. 2019-0287649 A1, filed March 13, 2019, entitled "Method and System for Selecting, Managing, and Analyzing Data of High Dimensionality," which is incorporated herein by reference.
[0395] In some embodiments, the classifier utilized the M-score classifier described in U.S. Patent Application Publication No. 2019-0287652 A1, entitled "Methylation Fragment Anomaly Detection," filed March 13, 2019, which is incorporated herein by reference.
[0396] In some embodiments, the classifier was a neural network or a convolutional neural network. For a disclosure of convolutional neural networks that can be used to classify methylation patterns according to the present disclosure, see U.S. Patent Application No. 62 / 679,746, entitled "Convolutional Neural Network Systems and Methods for Data Classification," filed June 1, 2018, which is incorporated herein by reference.
[0397] In some embodiments, the classifier was a support vector machine (SVM). When used for classification, an SVM separates a given set of labeled binary data with a hyperplane that is maximally far from the labeled data. When linear separation is not possible, an SVM can work in combination with the technique of "kernels" that automatically achieve a nonlinear mapping to the feature space. The hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space.
[0398] In some embodiments, the classifier was a decision tree. Tree-based methods divide the feature space into a set of rectangles and then fit a model (such as a constant) to each. In some embodiments, the decision tree was a random forest regression. One particular algorithm that may be used is a classification and regression tree (CART). Other particular decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forest.
[0399] In some embodiments, the classifier was an unsupervised clustering model. In some embodiments, the classifier is a supervised clustering model. The clustering problem is described as one of finding natural classifications in a data set. To identify natural classifications, two problems are addressed. First, a method for measuring the similarity (or dissimilarity) between two samples is determined. This metric (e.g., a similarity measure) is used to ensure that samples in one cluster are more similar to each other than to samples in other clusters. Second, a mechanism is determined for splitting the data into clusters using the similarity measure. One way to start the clustering search is to define a distance function and calculate a matrix of distances between all pairs of samples in the training set. If distance is a good similarity measure, the distance between reference entities in the same cluster will be significantly smaller than the distance between reference entities in different clusters. Clustering does not require the use of a distance metric. For example, a non-metric similarity function s(x,x') can be used to compare two vectors x and x'. Traditionally, s(x,x') is a symmetric function that is large if x and x' are "similar" in some way. Once a method for measuring "similarity" or "dissimilarity" between points in a dataset is selected, clustering requires a criterion function that measures the clustering quality of any partition of the data. The partition of the dataset that extremizes the criterion function is used to cluster the data. Certain exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor, farthest neighbor, average linkage, centroid, or sum of squares algorithms), k-means clustering, fuzzy k-means clustering algorithm, and Jarvis-Patrick clustering. In some embodiments, clustering includes unsupervised clustering (e.g., without a pre-determination of a pre-expected number of clusters and / or cluster assignments).
[0400] In some embodiments, the classifier is a regression model, such as a multi-category logit model. In some embodiments, the classifier utilizes a regression model.
[0401] In some embodiments, the classifier was a Naive Bayes algorithm. In some embodiments, the classifier was a nearest neighbor algorithm, such as a non-parametric method. In some embodiments, the classifier is a mixture model. In some embodiments, particularly those including a temporal component, the classifier was a hidden Markov model.
[0402] In some embodiments, the classifier was an A-score classifier. The A-score classifier was a tumor mutation burden classifier based on targeted sequencing analysis of nonsynonymous mutations. For example, a logistic regression on tumor mutation burden data can be used to calculate a classification score (e.g., "A-score"), and an estimate of tumor mutation burden for each individual is obtained from a targeted cfDNA assay. In some embodiments, tumor mutation burden can be estimated as the total number of mutations per individual that are called candidate mutations in cfDNA and that pass noise modeling and joint calling and / or are found as nonsynonymous in any gene annotation that overlaps with the mutation. The tumor mutation counts of the training set can be fed into a penalized logistic regression classifier to determine the cutoff at which 95% specificity is achieved using cross-validation.
[0403] In some embodiments, the classifier was a B-score classifier. The B-score classifier is described in US Patent Application Publication No. 2019-0287649 A1, entitled "Method and System for Selecting, Managing, and Analyzing Data of High Dimensionality," which is incorporated herein by reference. According to the B-score method, a first set of sequence reads of nucleic acid samples from healthy subjects in a reference group of healthy subjects is analyzed for regions of low variability. Thus, each sequence read in the first set of sequence reads of nucleic acid samples from each healthy subject is aligned to a region in the reference genome. From this, a training set of sequence reads from sequence reads of nucleic acid samples from subjects in a training group is selected. Each sequence read in the training set is aligned to one of the regions of low variability in the reference genome identified from the reference set. The training set includes sequence reads of nucleic acid samples from healthy subjects as well as sequence reads of nucleic acid samples from diseased subjects known to have cancer. The nucleic acid samples from the training group are of the same or similar type of nucleic acid sample as the type of nucleic acid sample from the reference group of healthy subjects. From this, quantities derived from the sequence reads of the training set are used to determine one or more metrics reflecting the difference between the sequence reads of the nucleic acid samples from the healthy subjects and the sequence reads of the nucleic acid samples from the diseased subjects in the training group. A test set of sequence reads associated with a nucleic acid sample comprising cell-free nucleic acid fragments from a test subject with an unknown cancer status is then received, and a likelihood that the test subject has cancer is determined based on the one or more metrics.
[0404] In some embodiments, the classifier was an M-score classifier, which is described in U.S. Patent Application Publication No. US 2019-0287652 A1, entitled "Anomalous Fragment Detection and Classification," which is incorporated herein by reference.
[0405] Example 4 - Whole Genome Bisulfite Sequencing (WGBS) WGBS is described in U.S. Patent Application Publication No. US 2019-0287652 A1, entitled "Anomalous Fragment Detection and Classification," which is incorporated herein by reference.
[0406] Example 5 - Cell-Free Genome Atlas Study (CCGA) cohort In the examples of the present disclosure, subjects from CCGA [NCT02889978] were used. CCGA is a prospective, multicenter, observational cfDNA-based early cancer detection study that enrolled 15,254 demographically balanced participants at 141 centers. Blood samples were collected from 15,254 enrolled participants (56% cancer, 44% non-cancer) from subjects with newly diagnosed, treatment-naïve cancer (C, cases) and participants without a cancer diagnosis defined at enrollment (non-cancer [NC], controls).
[0407] In the first cohort (prespecified substudy) (CCGA-1), plasma cfDNA extracts were obtained from 3,583 CCGA and STRIVE participants (CCGA: 1,530 cancer subjects and 884 non-cancer subjects; STRIVE 1,169 non-cancer participants). STRIVE is a multicenter prospective cohort study of women undergoing screening mammography (99,259 participants enrolled). For plasma cfDNA extraction, blood was collected from 984 CCGA participants with newly diagnosed untreated cancer (20 tumor types, all stages) and 749 participants (controls) without a cancer diagnosis (n=1,785). This pre-planned substudy included 878 cases, 580 controls, and 169 assay controls (n=1627) across 20 tumor types and all clinical stages.
[0408] Three sequencing assays were performed on blood collected from each participant: 1) paired cfDNA and white blood cell (WBC) targeted sequencing (60000X, 507-gene panel) for single-base variants / indels (ART sequencing assay), which removed WBC-derived somatic variants and residual technical noise by joint collation; 2) paired cfDNA and WBC whole-genome sequencing (WGS; 35X) for copy number variants, which generated cancer-associated signal scores by novel machine learning algorithms, and shared events were identified by joint analysis; and 3) paired cfDNA whole-genome bisulfite sequencing (WGBS; 34X) for methylation, which used aberrantly methylated fragments to generate a normalized score. In addition, tissue samples were collected from participants with cancer, and 4) paired tumor and WBC gDNA whole-genome sequencing (WGS; 30X) was performed for identification of tumor variants for comparison.
[0409] Within the context of the CCGA-1 study, several methods were developed for estimating the tumor fraction of cfDNA samples. See International Patent Publication No. WO2019 / 204360, entitled "SYSTEMS AND METHODS FOR DETERMINING TUMOR FRACTION IN CELL-FREE NUCLEIC ACID," International Patent Publication No. WO2020 / 132148, entitled "SYSTEMS AND METHODS FOR ESTIMATING CELL SOURCE FRACTIONS USING METHYLATION INFORMATION," and U.S. Patent Application Publication No. US 2020-0340064 A1, entitled "SYSTEMS AND METHODS FOR TUMOR FRACTION ESTIMATION FROM SMALL VARIANTS," each of which is incorporated herein by reference.
[0410] In the second prespecified substudy (CCGA-2), we used a targeted bisulfite sequencing assay rather than a whole genome sequencing assay to develop a cancer vs. non-cancer and tissue of origin classifier based on a targeted methylation sequencing approach. For CCGA-2, 3,133 training participants and 1,354 validation samples (775 with cancer and 579 without cancer as determined at enrollment before confirmation of cancer vs. non-cancer status) were used. Plasma cfDNA was subjected to a bisulfite sequencing assay (COMPASS assay) targeting the most informative regions of the methylome identified from the intrinsic methylation database and previous prototype whole genome and targeted sequencing assays to identify cancer- and tissue-defining methylation signals. Of the original 3,133 samples reserved for training, 1,308 samples were deemed clinically evaluable and analyzable. Analyses were performed on a primary analysis population of n=927 (654 cancer and 273 non-cancer) and a secondary analysis population of n=1027 (659 cancer and 373 non-cancer). Finally, genomic DNA from formalin-fixed paraffin-embedded (FFPE) tumor tissues and isolated cells from tumors was subjected to whole-genome bisulfite sequencing (WGBS) to generate a large database of cancer-defining methylation signals for use in panel design and training to optimize performance.
[0411] These data demonstrate the feasibility of achieving >99% specificity for invasive cancer and support the promise of cfDNA assays for early cancer detection. See, e.g., Klein et al., 2018, "Development of a comprehensive cell-free DNA (cfDNA) assay for early detection of multiple tumor types: The Circulating Cell-free Genome Atlas (CCGA) study," J. Clin. Oncology 36(15), 12021-12021; doi:10.1200 / JCO.2018.36.15_suppl.12021, and Liu et al., 2019, "Genome-wide cell-free DNA (cfDNA) methylation signatures and effect on tissue of origin (TOO) performance," J. Clin. Oncology 37(15), 3049-3049; doi:10.1200 / JCO.2019.37.15_suppl.3049, each of which is incorporated by reference in its entirety.
[0412] Within the context of the CCGA-2 study, multiple methods were developed to estimate the tumor fraction of a cfDNA sample based on methylation data (obtained by targeted methylation or WGBS) (see, for example, International Publication No. WO 2020 / 132148, entitled "SYSTEMS AND METHODS FOR ESTIMATING CELL SOURCE FRACTIONS USING METHYLATION INFORMATION," and U.S. Provisional Patent Application No. 62 / 983443, filed February 28, 2020, entitled "Identifying Methylation Patterns that Discriminate or Indicate a Cancer Condition," each of which is incorporated herein by reference in its entirety). In an exemplary approach, nucleic acid samples from formalin-fixed paraffin-embedded (FFPE) tumor tissue were analyzed by whole genome bisulfite sequencing (WGBS). Somatic mutations identified based on the sequencing data were analyzed against matched cfDNA WGBS sequencing data from the same patient and used to determine tumor fraction estimates.
[0413] Example 6 - Co-occurrence of abnormal methylation patterns and somatic mutations Experiment 1 An initial experiment to determine whether a correlation exists between hypermethylated and mutated fragments was performed by simulated pull-down of hypermethylated fragments using Fisher's exact test and assessment for enrichment of somatic mutations.
[0414] A dataset of 220 tissue samples sequenced using WGBS was subset to select methylation-enriched regions. The dataset further contained approximately 13,500 somatic mutations annotated using patient-matched tissues sequenced using WGS. Somatic mutations were called based on the analysis including patient-matched normal tissue references and were therefore considered as ground truth. For each somatic mutation, the dataset was divided into "reference" or "alternative" fragments based on whether each fragment in the dataset corresponding to the mutation position supports the reference allele or the alternative allele. Each fragment was further determined to be hypomethylated or hypermethylated by calculating the methylation fraction (beta value) of each fragment. For example, fragments with a beta value greater than 0.5 were determined to be hypermethylated, while fragments with a beta value less than or equal to 0.5 were determined to be hypomethylated. For each somatic mutation, the correlation between hypermethylation and mutated fragments was evaluated using Fisher's exact test according to the matrix shown below.
[0415] [Table 1]
[0416] Hypermethylated and hypomethylated mutations were aggregated and plotted, respectively. 6.6% of mutations were found to be significantly associated with hypermethylation (FDR<0.05), indicating that hypermethylated fragments did not significantly enrich for somatic mutations in the isolates. Figure 4A shows these results using a distribution plot of probability density of alternative fragments plotted against fragment beta values (x-axis) across mutations.
[0417] Another approach was used to determine whether fragment-level methylation fractions, rather than mutation-level, could be correlated with somatic mutations. All fragments in the dataset were faceted into reference support and alternative support, and aggregated together across mutations. Methylation fractions (beta values) were calculated for each fragment. Figure 4B shows a distribution plot of the probability density of alternative and reference fragments plotted against beta values (x-axis), further indicating that alternative fragments were not significantly enriched in the hypermethylated fraction.
[0418] Experiment 2 Experiments were performed to determine whether tumor-derived fragments marked by methylation could be highly informative for somatic mutation detection, especially in the presence of nearby CpG sites.
[0419] A dataset of 238 tissue samples sequenced using WGBS from the CCGA-1 substudy (see Example 5) was subset to select methylation-enriched regions. A simplified variant calling workflow was performed using a Bayesian likelihood filter, the database of single nucleotide polymorphisms (dbSNP; NCBI), and tissue recurrence blacklists disclosed in U.S. Patent Application No. 17 / 185885, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 25, 2021, and PCT Application No. PCT / US2021 / 019746, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 2021, each of which is incorporated herein by reference in its entirety. The dataset included 12,928 somatic mutations and 49,083 germline mutations obtained using patient-matched tissues sequenced using WGS. For each candidate mutation, the fragments were grouped into a "reference" or "alternative" bin based on whether each fragment supports either the reference or alternative allele. For each candidate mutation, the p-value distribution statistics (e.g., mean, minimum, maximum, median, and standard deviation) across the reference and alternative bins, respectively, were calculated. Additionally, for each candidate mutation, the distribution statistics (e.g., mean, minimum, maximum, median, and standard deviation) for the number of CpG sites across all fragments in the reference and alternative bins, respectively, were calculated. The reference and alternative counts, p-values, number of CpG sites, and their distribution statistics were determined according to some embodiments of the present disclosure as disclosed herein. For each candidate mutation, the obtained features (e.g., reference and alternative fragment counts, p-values, and / or CpG sites) were binned together into a fixed-length vector for each mutation and used as input to train and evaluate a classifier to determine whether the candidate mutation is a somatic or germline mutation. The classifiers were trained and evaluated using an 80 / 20 training-test mutation split.
[0420] 5A and 5B show the performance of a baseline binary classification model using reference and alternative fragment counts as inputs. FIG. 5A is a receiver operating characteristic (ROC) curve showing an evaluation of the performance of a logistic regression classifier to determine whether a candidate mutation is a somatic or germline mutation. Similar performance was observed for both the training and testing datasets (training: AUC=0.70; testing: AUC=0.69). FIG. 5B shows the precision-recall curve of the logistic regression classifier, which achieves a sensitivity (recall) of 20% with a positive predictive value (PPV or precision) of 50%. As defined above, positive predictive value (PPV) refers to the proportion of mutations that are correctly classified as somatic or germline mutations (e.g., the number of true positives divided by the number of true positives plus the number of false positives).
[0421] In contrast, Figures 6A and 6B show the performance of a binary classification model using extended feature inputs including reference and alternative fragment counts, p-value distribution statistics (e.g., mean, minimum, maximum, median, and standard deviation), and distribution statistics for the number of CpG sites across all fragments for each of the reference and alternative bins (e.g., mean, minimum, maximum, median, and standard deviation). Figure 6A is a ROC curve showing an evaluation of the performance of a multi-layer perceptron (MLP) neural network classifier to determine whether a candidate mutation is a somatic or germline mutation. Similar performance is observed for both the training and testing datasets (training: AUC=0.80; testing: AUC=0.80), further improving compared to previous models utilizing reference and alternative fragment counts as inputs. In addition, Figure 6B shows the precision-recall curve for the MLP classifier, where the sensitivity (recall) achieved at a positive predictive value (PPV or precision) of 50% is 60% compared to 20% in the previous model.
[0422] Experiment 3 Additional experiments were performed to determine whether tumor-derived fragments marked by methylation could be highly informative for somatic mutation detection in cfDNA samples. A dataset of 148 cfDNA samples sequenced using targeted methylation was subset to select methylation-enriched regions. The dataset included 404 somatic mutations and 62,575 germline mutations that were annotated using WGS and filtered to remove mutations with zero read support in fragments sequenced from cfDNA samples (e.g., filtering mutations with non-zero alternate support depth). Classifiers were trained and evaluated using an 80 / 20 training-test mutation split.
[0423] Figure 10A and Figure 10B show the performance of the baseline binary classification model using reference fragment counts and alternative fragment counts as input. Figure 10A shows the ROC curves that show the evaluation of the performance of the logistic regression classifier to determine whether a candidate mutation is a somatic or germline mutation. Similar performance was observed for both the training and testing datasets (training: AUC=0.63; testing: AUC=0.63). Figure 10B shows the precision-recall curves of the logistic regression classifier, showing poor resolution of mutations, as shown by the low accuracy obtained by the model (likely due to low tumor signal and high proportion of noise from normal-derived fragments in cfDNA samples compared to tissue samples).
[0424] In contrast, Figures 11A and 11B show the performance of the model using extended feature inputs including reference and alternative fragment counts, p-value distribution statistics (e.g., mean, minimum, maximum, median, and standard deviation), and distribution statistics (e.g., mean, minimum, maximum, median, and standard deviation) for the number of CpG sites across all fragments for each of the reference and alternative bins, respectively. Figure 11A shows the ROC curves that show an evaluation of the performance of the logistic regression model, with similar performance observed for both the training and testing datasets (training: AUC=0.86; testing: AUC=0.85), revealing an improvement over the model that utilizes reference and alternative fragment counts as inputs (training: AUC=0.63; testing: AUC=0.63). In addition, Figure 11B shows the precision-recall curves of the logistic regression model, showing that the PPV is improved, achieving a sensitivity of about 30% with a PPV of about 10%.
[0425] Conclusion The data show that abnormal methylation patterns co-occur with somatic mutations when CpG sites are present in the vicinity of the mutation. For example, in WGBS tissue, this relationship can be used to achieve a similar PPV (50%) as the filtering method used in the aforementioned tumor fraction estimation method using WGS cfDNA, although the sensitivity is reduced by 40%. See, for example, U.S. Patent Application No. 17 / 185885, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 25, 2021, and PCT Application No. PCT / US2021 / 019746, entitled "Systems and Methods for Calling Variants using Methylation Sequencing Data," filed February 2021, each of which is incorporated herein by reference in its entirety.
[0426] In targeted methylated cfDNA, the above experiments revealed an increase in PPV for somatic mutation detection when using expanded feature input. In some cases, larger training datasets and methods for reducing class balance can be used to offset the differences between somatic and germline mutations in cfDNA (e.g., more closely approximate class balance in tissues), which can further improve PPV and sensitivity.
[0427] conclusion The terms used herein are for the purpose of describing particular cases only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" are intended to include the plural, unless the context clearly dictates otherwise. Additionally, the term "and / or" as used in this disclosure shall be understood to refer to and include any and all possible combinations of one or more of the associated listed terms. Furthermore, it shall be understood that the terms "comprise" and / or "comprising" as used herein indicate the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Additionally, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof are used in any of the detailed description and / or claims, such terms are intended to be inclusive in the same manner as the term "comprising."
[0428] Multiple instances may be provided for components, operations, or structures described herein as a single instance. Finally, boundaries between various components, operations, and data stores are somewhat arbitrary, with particular operations being illustrated in the context of specific exemplary configurations. Other allocations of functionality are contemplated and may be included within the scope of the implementations. In general, structures and functions presented as separate components in an exemplary configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements are included within the scope of the implementations.
[0429] In this specification, terms such as first, second, etc. may be used to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, a first object can be called a second object, and similarly, a second object can be called a first object, without departing from the scope of the present disclosure. Although a first object and a second object are both objects, they are not the same object.
[0430] The term "if" as used herein shall be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (a stated condition or event) is detected" shall be interpreted to mean "upon determining" or "in response to determining," or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)," depending on the context.
[0431] The foregoing description has included exemplary systems, methods, techniques, instruction sequences, and computer program products embodying exemplary implementations. For purposes of explanation, numerous specific details have been set forth to provide an understanding of various implementations of the inventive subject matter. However, it will be apparent to those skilled in the art that implementations of the inventive subject matter may be practiced without these specific details. In general, well-known instruction instances, protocols, structures, and techniques have not been shown in detail.
[0432] The foregoing has been described with reference to specific implementations for purposes of explanation. However, the exemplary discussion above is not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. These embodiments have been chosen and described in order to best explain the principles and their practical application, thereby enabling others skilled in the art to best utilize these implementations and various implementations with various modifications suited to the particular uses contemplated.
Claims
1. A method for identifying a mutant allele at a genomic position of a test subject as a somatic mutant allele or a germline mutant allele, comprising: obtaining an identification of a reference allele at the genomic position; obtaining an identification of the mutant allele at the genomic position; Obtaining, for each nucleic acid fragment sequence in each of a plurality of nucleic acid fragment sequences in a sequencing data set derived from a liquid biological sample obtained from the test subject, which is mapped on the genomic position, the methylation state and each sequence, wherein the sequencing data set includes at least 1×10 6 individual nucleic acid fragment sequences, and obtaining; using (i) the identification of the reference allele at the genomic position and (ii) the respective sequences of each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences, assigning each nucleic acid fragment sequence having the reference allele at the genomic position among the respective plurality of nucleic acid fragment sequences to a reference subset; using (i) the identification of the mutant allele at the genomic position and (ii) the respective sequences of each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences, assigning each nucleic acid fragment sequence having the mutant allele at the genomic position among the respective plurality of nucleic acid fragment sequences to a mutant subset; applying to a trained binary classifier at least (i) one or more indicators of the methylation state over the entire methylation state of each nucleic acid fragment sequence in the mutant subset and (ii) an indicator of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset, wherein the trained binary classifier includes at least 10 parameters, whereby obtaining an identification of the mutant allele at the genomic position in the test subject as a somatic mutant allele or a germline mutant allele from the trained binary classifier; A method comprising the above.
2. The method further comprises: inputting a reference genome into a computer system including a processor coupled to a non-transitory memory; using the computer system to determine that each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences is mapped to the genomic position by aligning each nucleic acid fragment sequence to the reference genome; The method according to claim 1, further comprising the above.
3. The first nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences has a plurality of CpG sites; The first nucleic acid fragment sequence has a corresponding methylation pattern over the plurality of CpG sites; The methylation state of the first nucleic acid fragment sequence is a p-value; The method Determining at least in part the p-value of the first nucleic acid fragment sequence by comparing the corresponding methylation pattern of the first nucleic acid fragment sequence with the corresponding distribution of the methylation patterns of those nucleic acid fragment sequences in a healthy non-cancer cohort dataset each having the respective plurality of CpG sites The method according to claim 1, comprising:
4. If the mutant allele at the genomic position is determined by the trained binary classifier to be a germline mutant allele, the method comprises: Using the mutant allele in the test subject to determine the cancer risk of the test subject The method according to claim 1, further comprising:
5. Each of the one or more indicators of the methylation state across the entire mutant subset is A measure of the central tendency of the methylation state p-values across the entire mutant subset, The minimum methylation state p-value across the entire mutant subset, The maximum methylation state p-value across the entire mutant subset, or A measure of the spread of the methylation state p-values across the entire mutant subset The method according to claim 1, wherein:
6. The application to the trained binary classifier further comprises (iii) applying one or more CpG site indicators across the entire mutant subset. The method according to claim 1
7. The obtaining of the identification of the mutant allele at the genomic position includes determining that each of the plurality of nucleic acid fragments supports a mutant allele call at the genomic position. The method according to claim 1
8. A computing system, comprising: One or more processors; and A memory storing one or more programs executed by the one or more processors, the one or more programs comprising: Obtaining an identification of a reference allele at the genomic position; Obtaining an identification of the mutant allele at the genomic position; Obtaining the methylation state and respective sequences of each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences in a sequencing dataset derived from a liquid biological sample obtained from the test subject and mapped onto the genomic position, wherein the sequencing dataset includes at least 10^6 nucleic acid fragment sequences (i) the identification of the reference allele at the genomic position, and (ii) using the respective sequences of each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences, assigning each nucleic acid fragment sequence having the reference allele at the genomic position among the respective plurality of nucleic acid fragment sequences to a reference subset; (i) the identification of the mutant allele at the genomic position, and (ii) using the respective sequences of each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences, assigning each nucleic acid fragment sequence having the mutant allele at the genomic position among the respective plurality of nucleic acid fragment sequences to a mutant subset; applying to a trained binary classifier at least (i) one or more indicators of the methylation state over the entire methylation state of each nucleic acid fragment sequence in the mutant subset, and (ii) an indicator of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset, wherein the trained binary classifier includes at least 10 parameters, whereby obtaining, from the trained binary classifier, an identification as a somatic mutant allele or a germline mutant allele of the mutant allele at the genomic position in the subject under test; A method comprising: a memory including instructions for calling a mutation at a genomic position of a subject under test; A computing system comprising:
9. A non-transitory computer-readable storage medium storing one or more programs for calling a mutation at a genomic position of a subject under test, the one or more programs being configured to be executed by a computer, the one or more programs comprising: obtaining an identification of a reference allele at the genomic position; obtaining an identification of the mutant allele at the genomic position; obtaining the methylation state and the respective sequences of each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences in a sequencing data set derived from a liquid biological sample obtained from the subject under test and mapped onto the genomic position, the sequencing data set including at least 10^6 nucleic acid fragment sequences; (i) the identification of the reference allele at the genomic position, and (ii) using the respective sequences of each nucleic acid fragment sequence among the plurality of nucleic acid fragment sequences, assign each nucleic acid fragment sequence having the reference allele at the genomic position among the plurality of nucleic acid fragment sequences to a reference subset, (i) the identification of the mutant allele at the genomic position, and (ii) using the respective sequences of each nucleic acid fragment sequence among the plurality of nucleic acid fragment sequences, assign each nucleic acid fragment sequence having the mutant allele at the genomic position among the plurality of nucleic acid fragment sequences to a mutant subset, Apply to a trained binary classifier at least (i) one or more indicators of the methylation state over the entire methylation state of each nucleic acid fragment sequence in the mutant subset, and (ii) an indicator of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset. The trained binary classifier includes at least 10 parameters, whereby an identification as a somatic mutant allele or a germline mutant allele of the mutant allele at the genomic position in the test subject is obtained from the trained binary classifier. A non-transitory computer-readable storage medium containing instructions for
10. A method of training a classifier to identify a mutant allele at a genomic position of a test subject as a somatic mutant allele or a germline mutant allele, A) obtaining an identification of a reference allele at the genomic position; B) for each genomic position among a plurality of genomic positions in each subject among a plurality of subjects, i) obtaining an orthogonal call as either a somatic mutant allele or a germline mutant allele for the mutant allele at the respective genomic position of the respective subject; ii) obtaining an identification of the mutant allele at the respective genomic position of the respective subject; iii) Obtaining the methylation state and each sequence of each nucleic acid fragment sequence among a plurality of nucleic acid fragment sequences in a sequencing data set derived from a liquid biological sample obtained from each of the respective targets, which are mapped on the respective genomic positions, wherein the sequencing data set includes at least 1 × 10 6 nucleic acid fragment sequences, and obtaining; iv) (a) the identification of the reference allele at the respective genomic position, and (b) using the respective sequences of each nucleic acid fragment sequence among the plurality of nucleic acid fragment sequences, assign each nucleic acid fragment sequence having the reference allele at the respective genomic position among the plurality of nucleic acid fragment sequences to a reference subset; v) (a) identifying the mutant alleles at each of the respective genomic positions; and (b) using the respective sequences of each nucleic acid fragment sequence among the respective plurality of nucleic acid fragment sequences, assigning each nucleic acid fragment sequence having the mutant allele at the respective genomic position among the respective plurality of nucleic acid fragment sequences to a mutant subset; C) for each genomic position among the plurality of genomic positions in each of the plurality of subjects, at least: (i) one or more indicators of the methylation state over the entire methylation state of each nucleic acid fragment sequence in the mutant subset of the respective subject at the respective genomic position; (ii) an indicator of the number of nucleic acid fragment sequences in the reference subset relative to the number of nucleic acid fragment sequences in the mutant subset of the respective subject at the respective genomic position; and (iii) using the orthogonal call as either a somatic mutant allele or a germline mutant allele for the mutant allele at the respective genomic position of the respective subject, training the classifier to identify the mutant allele at the genomic position of the test subject as a somatic mutant allele or a germline mutant allele, wherein the classifier includes at least 10 parameters; performing a procedure comprising; A method comprising.