Tumor fraction estimation using methylation variants
Patent Information
- Application Number
- JP2024544960
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-20
- Filing Date
- 2023-02-15
- Publication Date
- 2026-02-24
AI Technical Summary
Existing cancer detection methods have problems with low detection rates and high false positive rates, making it difficult to detect cancer early, and traditional methods require separate screening for different types of cancer, and the detection efficiency of abnormal methylation signals in cellular free DNA (cfDNA).
By developing a computational model, the identification and quantification of cancer signals is performed using abnormal methylation signatures in cell free DNA (cfDNA) released from tumor cells. The computational model constructs a database that records the methylation patterns of healthy individuals and determines the possibility of abnormal methylation patterns through statistical methods to identify cancer signals.
It realizes efficient identification and quantification of very small amounts of cancer signals, and can screen multiple cancer types simultaneously from a single cfDNA sample, improving the sensitivity and specificity of traditional methods and reducing dependence on high-risk individuals.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 311,402, filed February 17, 2022, U.S. Provisional Patent Application No. 63 / 408,412, filed September 20, 2022, U.S. Provisional Patent Application No. 63 / 432,461, filed December 14, 2022, and U.S. Provisional Patent Application No. 63 / 480,859, filed January 20, 2023, all of which are incorporated by reference in their entireties herein.
[0002] FIELD OF THEINVENTION The present disclosure relates generally to early cancer detection via computational models for predicting tumor fraction from nucleic acid samples. [Background technology]
[0003] Cancer is a leading cause of death worldwide. Cancer mortality is driven by the fact that cancer is usually detected at a late stage, limiting the effectiveness of treatment options for long-term survival. Current detection methods are generally cancer type specific, i.e., each cancer type is screened individually. Each individual screening process is tailored to the cancer type. For example, mammography scans are utilized for breast cancer detection, while colonoscopic or fecal tests are useful for colorectal cancer detection. Each of the various screening methods is generally not cross-applicable to other cancer types. Furthermore, current screening methods are hampered by low detection rates or high false positive rates. Low detection rates often result in the inability to detect early stages of cancer as they develop. High positive rates result in cancer-free individuals being misdiagnosed as positive for cancer status. As a result, most screening tests are only practical when used to test individuals with a high risk of developing the screened cancer, and have limited ability to detect cancer in the general population.
[0004] New research suggests aberrant DNA methylation in many disease processes, including cancer. DNA methylation plays a role in suppressing gene expression. Thus, aberrant DNA methylation can cause problems in normal gene expression pathways, thereby resulting in cancer or other diseases. For example, specific patterns of differentially methylated regions can be useful as molecular markers for various disease states. Nevertheless, even such models face some challenges. Early cancer detection is particularly difficult because the ratio of tumor cells to non-cancerous cells in a subject is very small. The minimum ratio can be about 1:1000, 1:10,000, or even 1:100,000. This creates a challenge to detect small amounts of cancer signals among healthy signals. Furthermore, normal cfDNA can be released by blood cells, which can contain age-related genetic mutations that often resemble cancerous aberrant methylation. These aberrantly methylated fragments released from blood cells can often increase the apparent cancer signal.
[0005] After a cancer patient is diagnosed, further challenges arise. First, healthcare providers need to understand the cancer status to tailor and individualize treatment for the patient. Factors that can be evaluated in the treatment decision-making process may include the stage of the cancer, whether the cancer is benign or malignant, whether the cancer has metastasized, the cancer's original tissue, etc. These factors are typically observed from cumbersome traditional screening processes. Second, during and / or after treatment, healthcare providers rely once again on those traditional screening processes to evaluate the effectiveness of the treatment. These traditional screening processes can be burdensome to the patient. For these reasons, there remains a critical need in the art for accurate and precise quantification of cancer signals (e.g., tumor fraction) as a valuable post-diagnosis tool. Summary of the Invention
[0006] The invention described herein in this disclosure provides improvements to cancer detection and treatment, and in particular provides valuable insights to patients being screened for cancer. The invention can quantify cancer signals in individuals, which can inform prognosis, cancer staging, cancer progression, monitoring of minimal residual disease (MRD), evaluation of treatment efficacy, and the like. The invention includes screening for cancer signals in cell-free deoxyribonucleic acid (cfDNA) samples of subjects. Such cfDNA samples may contain hundreds of thousands, if not millions, of cfDNA fragments, which may result in a similar order of sequence reads output by a sequencer, or even multiples of such order based on the sequencing depth of the sample. Each sequence read for cfDNA fragments can vary in length, for example, up to 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000bp long.These next-generation sequencing techniques greatly increase the amount of fragments that can be sequenced and analyzed, thereby allowing such models to distinguish even very small amounts of cancer signals in samples.The present invention can screen for cancer in general or for multiple cancer types from a single sample.This improves on the traditional screening methods that are tailored for each cancer type by providing a single comprehensive screen that can screen for various cancer types from a single cfDNA sample.
[0007] The present invention implements a computer model that identifies and quantifies cancer signals based on cfDNA released from tumor cells. The computer model can identify cancer signals from cfDNA fragments that contain an aberrant methylation signature from sequence reads of cfDNA. The computer model can identify aberrantly methylated fragments by building a database of counts of methylation patterns from a healthy population (i.e., subjects with no prior diagnosis of disease and / or cancer). The computer model can utilize the database to determine whether a methylation pattern associated with a cfDNA fragment from a test sample is aberrantly methylated. The computer model can apply statistical methods to determine the likelihood of observing a fragment in a normal subject, even if the fragment has a methylation pattern not yet observed in the database. Thus, the computer model can identify aberrant methylation patterns, i.e., the proverbial needle in a haystack.
[0008] From the aberrant methylation patterns, the computer model may implement a trained cancer classifier to characterize the aberrant methylation patterns of the sample and generate a cancer prediction. The cancer prediction may be a binary prediction and / or a multi-class prediction. A binary prediction may be the likelihood of the presence of cancer. A multi-class prediction may be the likelihood of a particular cancer type. By training a cancer classifier that can screen across multiple cancer types, medical professionals can utilize a single comprehensive screen rather than multiple different screens.
[0009] The computer model may further train a probabilistic model for one or more cancer types to determine the tissue of origin of the anomalously methylated fragments. Each probabilistic model may input a methylation pattern and output a likelihood that the methylation pattern is from a particular cancer type (or more generally, a disease type). The computer model may characterize methylation patterns that exceed a threshold likelihood to further quantify the cancer signal in the individual. The ability to quantify the cancer signal based on the fragments released from the tumor cells improves upon conventional screening methods in that conventional methodologies have relied primarily on visual observation by medical professionals. Such visual observation leads to subjectivity and human error. This quantification may be practically applied to predict the stage of the cancer, which may inform the individual of available treatment options.
[0010] In other exemplary applications, quantification of cancer signals can inform a combination of prognosis, assessment of cancer progression, and assessment of treatment efficacy. For example, in informing cancer progression, quantification of cancer signals can indicate reduced tumors, benign tumors, aggressive tumors, and / or anything in between. Quantification of cancer signals can also be used in staging of cancer, for example, a low signal can indicate stage I, or a high signal can indicate a later stage (e.g., stage III or IV). Furthermore, quantification of cancer signals can indicate metastasis of tumors to other tissues in the body. For example, the present invention may first identify a cancer signal present in a first tissue of a subject, but then identify a cancer signal present in a different tissue of the subject, which is indicative of a metastatic tumor. In addition, quantification of cancer signals before and during (or after) treatment can inform treatment efficacy in an individual. For example, if before treatment, the present invention predicts the cancer signal to be at a first level, but during and / or after treatment, the present invention predicts the cancer signal to be at a different level, the difference in levels can indicate the effectiveness of the treatment. Armed with that knowledge, a medical professional can make better informed decisions in treating a subject, for example, whether to continue a current treatment if the cancer signal is decreasing, or to switch treatment if the cancer signal remains the same or increases.
[0011] A computer-implemented method for generating a tumor fraction estimate from a cell-free deoxyribonucleic acid (cfDNA) sample of a subject is disclosed, the computer-implemented method includes receiving a dataset of methylated sequence reads from the subject's cfDNA sample; partitioning the dataset into a plurality of variants, each variant comprising a methylation pattern across one or more CpG sites; filtering the plurality of variants based on a bank of reference sequence reads to generate a subset of filtered variants, the bank comprising reads generated from a non-cancer cfDNA sample and a biopsy sample of a plurality of tissues of a reference individual; for each variant in the filtered subset, determining a count of methylated sequence reads that comprise the variant; inputting the counts of methylated sequence reads for variants of the filtered subset into a model trained based on recurrence rates of the plurality of variants; and using the model to generate a tumor fraction prediction for the cfDNA sample.
[0012] A computer-implemented method is disclosed in which the recurrence rates of multiple variants are determined based on reference sequence reads in a bank.
[0013] A computer-implemented method is disclosed, in which filtering the plurality of variants to generate a subset of filtered variants based on a reference sequence read includes filtering out one or more variants whose abundance in the non-cancer sample exceeds a threshold.
[0014] A computer-implemented method is disclosed in which a particular recurrence rate of a particular variant corresponds to the observation rate of the particular variant among reference sequence reads in a bank.
[0015] A computer-implemented method is disclosed in which the tumor fraction prediction is a distribution of the probability of the fraction of fragments in a cfDNA sample that are tumor-derived.
[0016] A computer-implemented method is disclosed in which the tumor fraction prediction is the fraction of fragments in a cfDNA sample that are tumor-derived.
[0017] A computer-implemented method is disclosed in which the model includes at least one probability model, the probability model including a Poisson distribution for a particular variant, the Poisson distribution being weighted by the recurrence rate of the particular variant.
[0018] A computer-implemented method is disclosed in which the model includes a plurality of probability distributions, each probability distribution corresponding to a particular variant and parameterized based on a site-specific noise rate for the particular variant and a sequencing depth per site for the particular variant.
[0019] Computer-implemented methods are disclosed in which each probability distribution corresponding to a particular variant is further parameterized based on at least one of the depth of the cfDNA sample, the targeted panel pull-down efficiency of the cfDNA sample, and the estimated tumor fraction of the cfDNA sample.
[0020] A computer-implemented method is disclosed in which the count for each variant of the filtered subset includes a count of methylated sequence reads of the cfDNA sample that contain a methylation pattern across one or more CpG sites of the variant.
[0021] A computer-implemented method is disclosed in which a particular variant comprising multiple adjacent CpG sites is encoded by a series of binary values, the series corresponding to the adjacent CpG sites, a first binary value at a particular CpG site representing observed methylation and a second binary value at a particular CpG site representing observed unmethylation.
[0022] A computer-implemented method is disclosed in which the tumor fraction prediction includes multiple fractions for subsets of tissue.
[0023] A computer-implemented method is disclosed in which each fraction represents a percentage of the fragments of the cfDNA sample derived from each tissue of a subset of tissues.
[0024] A computer-implemented method is disclosed in which the model is a binomial mixed model that assumes independence between variants in the filtered set.
[0025] A computer-implemented method is disclosed in which the model includes a plurality of methylation sub-models, each methylation sub-model associated with a variant in the filtered set and parameterized by a recurrence rate and an estimated tumor fraction of the variant across a subset of tissues, and each methylation sub-model configured to calculate a likelihood of observing a methylated sequence read count based on the methylated sequence read count.
[0026] A computer-implemented method is disclosed in which the model comprises a weighted sum of methylation sub-models that enumerate all possibilities for tissues within the subset that are activated or inactivated for the variant, with each methylation sub-model representing the likelihood of the possibility and including a weight parameterized by the recurrence rate of the activated tissue.
[0027] A computer-implemented method is disclosed in which each methylation sub-model is Poisson distributed.
[0028] A computer-implemented method is disclosed in which the model generates a tumor fraction prediction by identifying the estimated tumor fraction that has the greatest likelihood as calculated by the methylation submodel.
[0029] A computer-implemented method is disclosed in which the maximum likelihood is determined via a grid search.
[0030] A computer-implemented method is disclosed in which a subset of tissues is selected for a plurality of tissues.
[0031] A computer-implemented method is disclosed in which the model generates tumor fraction predictions by identifying an estimated tumor fraction that has a maximum likelihood as calculated by a methylation sub-model across two or more subsets of tissues, each subset of tissue being a different combination of tissues selected from a plurality of tissues from the other subsets of tissues.
[0032] A computer-implemented method, wherein the subset of tissues is of one size.
[0033] A computer-implemented method is disclosed in which the size of the tissue subset is selected from 2, 3, or 4.
[0034] A computer-implemented method is disclosed in which the model includes a cancer classifier that generates a prediction of one or more tissues from which the fragments originate, and the subset of tissues includes the one or more tissues predicted by the cancer classifier.
[0035] A computer-implemented method is disclosed in which methylated sequence reads in a dataset are generated by a targeted methylation assay.
[0036] A computer-implemented method is disclosed, wherein a plurality of tissues of a reference individual used to generate reference sequence reads are selected from the group including breast tissue, thyroid tissue, lung tissue, bladder tissue, cervical tissue, small intestine tissue, colorectal tissue, esophageal tissue, stomach tissue, tonsil tissue, liver tissue, ovarian tissue, fallopian tube tissue, pancreatic tissue, prostate tissue, kidney tissue, and uterine tissue.
[0037] A computer-implemented method is disclosed in which tumor fraction estimates are generated without the use of matching biopsy samples.
[0038] A computer-implemented method is disclosed in which a cfDNA sample is obtained from a bodily fluid without an invasive biopsy.
[0039] A computer-implemented method is disclosed in which one or more reference individuals in the bank are determined to have cancer or a type of cancer that is breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis and ureter, non-urothelial renal cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal squamous cell carcinoma, non-squamous esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from hepatocytes, hepatobiliary carcinoma arising from cells other than hepatocytes, pancreatic cancer, human papillomavirus-associated head and neck cancer, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung cancer, lung cancer other than squamous cell carcinoma, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, or leukemia.
[0040] A computer-implemented method is disclosed, wherein the model is a machine learning model.
[0041] A computer-implemented method is disclosed, wherein the machine learning model is one or more of a constant model, a binomial model, an independent sites model, a neural network model, or a Markov model.
[0042] A computer-implemented method is disclosed in which a machine learning model is trained by: identifying, for each variant of the filtered variants, for each reference sample in the bank comprising a non-cancer cfDNA sample and a biopsy sample, a count of reads containing the variant; determining, for each variant of the filtered variants, a non-cancer recurrence rate based on the count of reads for the variant in the non-cancer sample; determining, for each variant of the filtered variants, a cancer recurrence rate based on the count of reads for the variant in the biopsy sample; and training a model with the non-cancer recurrence rate and the cancer recurrence rate, wherein the model is configured to predict a tumor fraction prediction based on the count of reads of the filtered variant in a given sample.
[0043] Computer-implemented methods are disclosed in which a cfDNA sample is used to perform one or more of cancer surveillance for previously diagnosed cancer and early cancer screening for multiple cancer types.
[0044] A computer-implemented method is disclosed, wherein the subject's cfDNA sample is a liquid biopsy collected after initiation of a treatment for cancer, the computer-implemented method further comprising: determining a confidence score for the tumor fraction prediction based on a count of methylation sequence reads that comprise the subset of filtered variants; in response to determining that the confidence score is below a confidence threshold, sequencing a tissue sample collected after initiation of the treatment for cancer; receiving a second dataset of methylation sequence reads from the subject's tissue sample; partitioning the second dataset into a second plurality of variants; filtering the second plurality of variants based on a bank of reference sequence reads to generate a second subset of filtered variants; for each variant in the filtered subset, determining a second count of methylation sequence reads that comprise the variant; inputting the second count of methylation sequence reads for variants of the second filtered subset into a model; and using the model to generate a second tumor fraction prediction for the tissue sample.
[0045] A computer-implemented method is disclosed, further including returning the tumor fraction prediction with a second tumor fraction prediction and a confidence score.
[0046] A computer-implemented method is disclosed, wherein the subject's cfDNA sample is a liquid biopsy sample collected after initiation of a treatment for cancer, the computer-implemented method further comprising: determining that a tumor fraction prediction of the liquid biopsy sample is below a threshold signal; and in response to determining that the tumor fraction prediction of the liquid biopsy sample is below a threshold signal, sequencing a tissue sample collected after initiation of a treatment for cancer; receiving a second dataset of methylated sequence reads from the subject's tissue sample; partitioning the second dataset into a second plurality of variants; filtering the second plurality of variants based on a bank of reference sequence reads to generate a second subset of filtered variants; for each variant in the filtered subset, determining a second count of methylated sequence reads that include the variant; inputting the second count of methylated sequence reads for variants of the second filtered subset into a model; and using the model to generate a second tumor fraction prediction for the tissue sample.
[0047] A computer-implemented method is disclosed, further comprising returning the second tumor fraction prediction and the tumor fraction prediction below the threshold signal.
[0048] A computer-implemented method is disclosed, further including determining a first fraction of the first tissue that is apoptotic non-cancerous tissue, determining whether the first fraction of the first tissue is above a threshold, and in response to determining that the first fraction of the first tissue is above the threshold, returning an indication that the test subject has a higher propensity for uncontrolled inflammation as a side effect to an immune checkpoint inhibitor.
[0049] A computer-implemented method is disclosed in which the cfDNA sample is derived from the subject's urine.
[0050] A computer-implemented method is disclosed, wherein the tumor fraction prediction is for bladder cancer, prostate cancer, renal cancer, or any combination thereof.
[0051] Disclosed are computer-implemented methods that further include determining a poor prognosis for the subject in response to determining that the tumor fraction prediction for the cfDNA sample is above a threshold.
[0052] A computer-implemented method is disclosed in which a poor prognosis indicates that the disease in the subject is progressive.
[0053] A computer-implemented method is disclosed in which a poor prognosis indicates that the subject is at higher risk of disease recurrence.
[0054] A computer-implemented method is disclosed that further includes providing a list of recommended treatments based on the subject's poor prognosis.
[0055] Disclosed is a computer-implemented method, wherein a cfDNA sample is collected from a subject after initiation of treatment, the method further including determining the presence of residual disease in the subject in response to determining that a tumor fraction prediction of the cfDNA sample is above a threshold.
[0056] A computer-implemented method is disclosed that further includes providing a list of recommended treatments based on the presence of residual disease in the subject.
[0057] A computer-implemented method is disclosed, in which a cfDNA sample is collected from a subject following initiation of a treatment for a disease in the subject, the method further comprising evaluating the treatment based on the tumor fraction prediction.
[0058] A computer-implemented method is disclosed in which evaluating the treatment includes determining that the treatment is efficacious in response to determining that a tumor fraction prediction of a cfDNA sample collected after initiation of the treatment is less than an initial tumor fraction prediction of an initial cfDNA sample collected before initiation of the treatment.
[0059] Computer-implemented methods are disclosed in which evaluating the treatment includes determining that the treatment is ineffective in response to determining that a tumor fraction prediction of a cfDNA sample collected after initiation of the treatment is substantially equal to or greater than an initial tumor fraction prediction of an initial cfDNA sample collected before initiation of the treatment.
[0060] A computer-implemented method is disclosed that further includes, in response to determining that the treatment is ineffective, providing a list of alternative treatments to exclude the treatment.
[0061] Disclosed is a non-transitory computer-readable medium configured to store computer code including instructions for generating a tumor fraction estimate from a cell-free deoxyribonucleic acid (cfDNA) sample of a subject, the instructions, when executed by one or more processors, causing the one or more processors to perform any of the methods disclosed above.
[0062] A system is disclosed that includes one or more processors and a memory configured to store computer code including instructions for generating a tumor fraction estimate from a cell-free deoxyribonucleic acid (cfDNA) sample of a subject, the instructions, when executed by the one or more processors, cause the one or more processors to perform any one of the methods disclosed above. [Brief description of the drawings]
[0063] [Figure 1] 1 is an exemplary flowchart illustrating the overall workflow of cancer classification of a sample, according to one or more embodiments. [Figure 2A] 1 illustrates a flow chart of a device for sequencing a nucleic acid sample, according to one embodiment. [Figure 2B] FIG. 1 is a block diagram of an analytical system for processing sequence reads, in accordance with various embodiments. [Figure 3A] 1 is a flow chart illustrating a process for sequencing a nucleic acid, according to various embodiments. [Figure 3B] 1 is an illustration of a process for obtaining methylation information and a methylation state vector according to various embodiments. [Figure 4A] 1 illustrates the generation of a data structure for a control group, according to various embodiments. [Figure 4B] 1 illustrates a flow chart describing a process for determining abnormally methylated fragments from a sample, according to various embodiments. [Diagram 5] 1 is an illustration of a block of a reference genome, according to various embodiments. [Figure 6] 1 is a flowchart of a method for generating a classifier for predicting a disease state, according to various embodiments. [Figure 7] FIG. 1 is a conceptual diagram illustrating how a tumor fraction estimation model can be trained using a bank of reference sequence reads obtained from a reference individual, according to various embodiments. [Figure 8] 1 is an exemplary plot of the number of methylation variants found per participant, according to an exemplary implementation. [Figure 9] 1 is an exemplary plot of recurrence rates of various methylation variants found in different participants, according to an exemplary implementation. [Figure 10] 1 is a flowchart depicting a process for generating a tumor fraction estimate from a subject's cfDNA sample, according to some embodiments. [Figure 11] 1 includes plots of the number of methylation variants per million bases in targeted methylation panels and WGBS panels, according to exemplary implementations. [Figure 12] 1 is a plot illustrating average methylation variant cfDNA allele fraction calibration according to an exemplary implementation. [Figure 13] According to an exemplary implementation, data is provided comparing estimated tumor fraction in urine and plasma samples from subjects with or without any stage of bladder cancer. [Figure 14]1 provides data comparing estimated tumor fraction in urine and plasma samples from subjects with or without stage III or IV prostate cancer, according to an exemplary implementation. [Figure 15] 1 provides data comparing estimated tumor fraction in urine and plasma samples from subjects with and without renal cancer, according to an exemplary implementation. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0064] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. It should be noted that wherever possible, similar or like reference numbers may be used in the figures and may indicate similar or like functionality. It should also be noted that the contents of all published materials (such as patent applications, patents, papers, and conference proceedings) referenced herein are incorporated herein by reference in their entirety.
[0065] I. Overview IA definition Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. As used herein, the following terms have the following meanings ascribed to them:
[0066] The term "individual" refers to a human, animal, or any other multi-cellular organism. The term "healthy individual" refers to an individual who is suspected to be free of cancer or disease.
[0067] The term "subject" refers to an individual whose DNA is being analyzed. A subject may be a test subject whose DNA is evaluated using whole genome sequencing or a targeted panel as described herein to assess whether the person has a disease state (e.g., cancer, a type of cancer, or a tissue of origin of the cancer). A subject may be part of a control group known to not have cancer or another disease. A subject may also be part of a cancer or other disease group known to have cancer or another disease. Control groups and cancer / disease groups may be used to aid in the design or validation of a targeted panel.
[0068] The term "biological sample" or "sample" refers to a specimen obtained from an individual that contains the individual's genetic material. Examples of biological samples include, but are not limited to, blood, whole blood, plasma, serous fluid, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of a subject. A biological sample may include any tissue or material from a living or dead subject. A biological sample may be a cell-free sample. A biological sample may include nucleic acid (e.g., DNA or RNA) or fragments thereof. The term "nucleic acid" may refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or any hybrid or fragment thereof. The nucleic acid in the sample may be a cell-free nucleic acid. A sample may be a liquid sample or a solid sample (e.g., a cell sample or a tissue sample). The biological sample may be a bodily fluid, such as blood, plasma, serous fluid, urine, vaginal fluid, fluid from edema (e.g., of the scrotum), vaginal washing fluid, pleural fluid, peritoneal fluid, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, bronchial lavage fluid, nipple discharge fluid, aspirated fluid from different parts of the body (e.g., thyroid, breast), etc. The biological sample may be a fecal sample. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) may be cell-free (e.g., more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free). The biological sample may be treated to physically disrupt tissue or cellular structures (e.g., centrifugation and / or cell lysis), thus releasing intracellular components into a solution that may further contain enzymes, buffers, salts, detergents, etc. that may be used to prepare the sample for analysis.
[0069] The term "reference sample" refers to a sample obtained from a subject with a known disease state.
[0070] The term "training sample" refers to a sample obtained from a known disease state. The training sample can be applied to a probabilistic model to generate features that can be utilized for disease state classification.
[0071] The term "test sample" refers to a sample that may have an unknown disease state.
[0072] The term "sequence read" refers to a nucleotide sequence read from a sample obtained from an individual. A sequence read may be generated from a nucleic acid fragment in a sample. A sequence read may be a folded sequence read generated from multiple sequence reads derived from multiple amplicons from a single original nucleic acid molecule. In some embodiments, a sequence read may be a de-duplicated sequence read. A sequence read may be obtained through various methods known in the art.
[0073] The term "disease state" refers to the presence or absence of a disease, the type of disease, and / or the tissue of disease origin. For example, in one embodiment, the present disclosure provides methods, systems, and non-transitory computer-readable media for detecting cancer (i.e., the presence or absence of cancer), the type of cancer, or the tissue of disease origin.
[0074] The term "tissue of origin" or "TOO" refers to an organ, group of organs, body area, or cell type in which a disease state may originate or occur. For example, identification of the tissue of origin or cancer cell type typically allows one to identify appropriate next steps to further diagnosis, stages, and treatment decisions.
[0075] The term "methylation," as used herein, refers to the chemical process by which a methyl group is added to a DNA molecule. Two of the four bases of DNA, cytosine ("C") and adenine ("A"), can be methylated. For example, the hydrogen atom on the pyrimidine ring of the cytosine base can be converted to a methyl group to form 5-methylcytosine. Methylation tends to occur at the dinucleotides of cytosine and guanine, referred to herein as "CpG sites." In other instances, methylation can occur at a cytosine that is not part of a CpG site, or at another nucleotide that is not a cytosine. However, these occur more rarely. In this disclosure, methylation is discussed with reference to CpG sites for clarity. However, the principles described herein are equally applicable to the detection of methylation in non-CpG contexts, including non-cytosine methylation. For example, adenine methylation has been observed in bacterial, plant, and mammalian DNA, but has received little attention.
[0076] In such embodiments, the wet lab assays used to detect methylation may differ from those described herein, as is well known in the art. Furthermore, the methylation state vector may contain elements that are generally vectors of sites that are methylated or unmethylated (even if these sites are not specifically CpG sites). With that substitution, the remainder of the process described herein remains the same, and thus the inventive concepts described herein are applicable to those other forms of methylation.
[0077] The term "CpG site" refers to a region of a DNA molecule in which a cytosine nucleotide is followed by a guanine nucleotide in a linear sequence of bases along its 5' to 3' direction. "CpG" is an abbreviation for 5'-C-phosphate-G-3', where the cytosine and guanine are separated by only one phosphate group, and the phosphate links together any two nucleotides in the DNA. The cytosine in a CpG dinucleotide can be methylated to form 5-methylcytosine.
[0078] The term "methylation site" refers to a single site on a DNA molecule to which a methyl group can be added. Although "CpG" sites are the most common methylation sites, methylation sites are not limited to CpG sites. For example, DNA methylation can occur at cytosine in CHG and CHH, where H is adenine, cytosine, or thymine. Cytosine methylation in the form of 5-hydroxymethylcytosine (see, for example, WO 2010 / 037001 and WO 2011 / 127136, which are incorporated herein by reference), and its features, can also be evaluated using the methods and procedures disclosed herein. The terms "hypomethylated" or "hypermethylated" refer to the methylation status of a DNA molecule that contains multiple (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) CpG sites, where a high percentage (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50%-100%) of the CpG sites are unmethylated or methylated, respectively.
[0079] The term "modified methylation pattern" refers to a methylation pattern that falls within a predetermined number of CpG sites and meets one or more selection criteria. In this disclosure, the term "modified methylation pattern" is used interchangeably with the term "QMP" unless otherwise specified. In some embodiments, the modified methylation pattern corresponds to one or more CpG sites (e.g., a span or section of CpG sites) indexed to one or more specific sites in a reference genome. For example, when a modified methylation pattern is identified in one or more fragments in each of a plurality of fragments aligned to a reference genome, the modified methylation pattern includes one or more CpG sites, each respective CpG site includes a respective methylation state, and is indexed to a specific site in the reference genome. Thus, in some such embodiments, the modified methylation pattern refers to a specific sequence of methylation states at specific locations in a reference genome that meet one or more selection criteria. A modified methylation pattern (e.g., a representation of a respective sequence of methylation states for a modified methylation pattern such as "MMMMM" or "UUUUU") may be identified in one or more fragments of each of a plurality of fragments aligned to a reference genome, the fragment methylation pattern of each of the plurality of fragments being represented by an interval map, by matching a query methylation pattern to the representation of each fragment methylation pattern at each node of the interval map and determining whether the matched methylation pattern satisfies one or more selection criteria. In some embodiments, the modified methylation pattern does not correspond to either a specific CpG site or a specific position in the reference genome (e.g., when the genomic position of one or more CpG sites in the modified methylation pattern is unknown and / or when the sequence of methylation states in the modified methylation pattern occurs at multiple positions throughout the reference genome).
[0080] The terms "cell-free deoxyribonucleic acid," "cell-free DNA," or "cfDNA" refer to deoxyribonucleic acid fragments that circulate in bodily fluids, such as blood, sweat, urine, or saliva, and originate from one or more healthy cells and / or one or more cancer cells.
[0081] The term "circulating tumor DNA" or "ctDNA" refers to deoxyribonucleic acid fragments originating from tumor cells or other types of cancer cells, and which may be released into an individual's bodily fluids, such as blood, sweat, urine, or saliva, as a result of biological processes such as apoptosis or necrosis of dying cells, or may be actively released by viable tumor cells.
[0082] The term "tumor fraction" refers to the fractional contribution of tumor material to the cfDNA present in a sample. Tumor fraction estimates may be further subdivided into the amount of contribution from the underlying tissue and / or cell type.
[0083] The term "methylation variant" refers to the pattern and status of adjacent CpG sites that distinguish DNA from one biological source from another. Additionally, the term "methylation variant allele fraction" (MVAF) is an estimate for measuring the proportion of tumor-derived cfDNA in a sample that is aberrantly methylated.
[0084] IB Cancer Prediction Workflow FIG. 1 is an exemplary flow chart illustrating an overall workflow 100 for cancer classification of a sample, according to one or more embodiments. The workflow 100 is by one or more entities, including, for example, a healthcare provider, a sequencing device, an analysis system, etc. The purpose of the workflow includes detecting and / or monitoring cancer in an individual. From a health management perspective, the workflow 100 can serve to complement other existing cancer diagnostic tools. The workflow 100 can serve to provide early detection of cancer and / or regular cancer monitoring to better inform treatment plans for individuals diagnosed with cancer. The overall workflow 100 may include additional / fewer steps than those shown in FIG. 1.
[0085] A healthcare provider performs sample collection 110. An individual to undergo cancer classification visits a healthcare provider. The healthcare provider collects a sample to perform cancer classification. Examples of biological samples include, but are not limited to, a subject's tissue biopsy, blood, whole blood, plasma, serous fluid, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or ascites. The sample contains genetic material belonging to the individual, which can be extracted and sequenced for cancer classification. Once the sample is collected, the sample is provided to a sequencing device. Along with the sample, the healthcare provider may collect other information about the individual, such as biological sex, age, ethnicity, smoking status, any previous diagnoses, etc.
[0086] The sequencing device performs sample sequencing 120. A laboratory clinician may perform one or more processing steps on the sample in preparation for sequencing. Once prepared, the clinician loads the sample into the sequencing device. Examples of devices utilized in sequencing are further described in conjunction with Figures 2A and 2B. The sequencing device generally extracts and isolates fragments of the nucleic acid to be sequenced to determine the sequence of the nucleic acid bases corresponding to the fragments. Sequencing may also include amplification of the nucleic acid material. Different sequencing processes include Sanger sequencing, fragment analysis, and next generation sequencing. Sequencing may be whole genome sequencing or targeted sequencing using a targeted panel. In the context of DNA methylation, bisulfite sequencing (e.g., as further described in Figures 3A and 3B) can determine methylation status via bisulfite conversion of unmethylated cytosines at CpG sites. Sample sequencing 120 generates sequences for multiple nucleic acid fragments in the sample. In one or more embodiments, the sequence may include methylation state vectors, each methylation state vector describing the methylation status of CpG sites on the fragment.
[0087] The analytical system performs pre-analysis processing 130. An exemplary analytical system is illustrated in FIG. 10B. Pre-analysis processing 130 may include, but is not limited to, de-duplication of sequence reads, determining metrics related to coverage, determining whether a sample is contaminated, removing contaminated fragments, calling sequencing errors, and the like.
[0088] The analysis system performs one or more analyses 140. The analyses are statistical analyses or application of one or more trained models to predict at least the cancer status of the individual from whom the sample is derived. Different genetic features may be evaluated and taken into account, such as methylation of CpG sites, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), and other types of genetic mutations. In the context of methylation, the analyses 140 may include identification of aberrant methylation 142 (e.g., as further described in Figures 4A and 4B), feature extraction 144 (e.g., as further described in Figures 6, 7, and 10), and application of a cancer classifier 146 to determine a cancer prediction (e.g., as further described in Figures 6, 7, and 10). The cancer classifier 146 inputs the extracted features to determine a cancer prediction. The cancer prediction may be a label or a value. The label may indicate a particular cancer state, e.g., a binary label may indicate the presence or absence of cancer, and a multi-class label may indicate one or more cancer types from a plurality of cancer types screened for a stage of cancer (e.g., stage I, II, III, or IV). The value may indicate the likelihood of a particular cancer state, e.g., the likelihood of cancer, and / or the likelihood of a particular cancer type. The prediction 150 may further indicate a quantification of the cancer signal, which may include a quantification of one or more particular tissue of origin signals. For example, the quantification of the cancer signal may be expressed as a tumor burden calculated based on the percentage of sequence reads determined to be derived from tumor cells.
[0089] The analysis system returns a prediction 150 to the healthcare provider. The healthcare provider may establish or adjust a treatment plan based on the cancer prediction. Optimization of treatment is further described in Section VI.C. Treatment.
[0090] IC Exemplary Sequencer and Analysis System 2A is a flow chart of a system and device for sequencing a nucleic acid sample, according to one embodiment. This illustrative flow chart includes devices such as a sequencer 270 and an analysis system 200. The sequencer 270 and the analysis system 200 may operate in conjunction to perform one or more steps in the processes described herein.
[0091] In various embodiments, the sequencer 270 receives the enriched nucleic acid sample 260. As shown in FIG. 2A, the sequencer 270 can include a graphical user interface 275 that allows user interaction with a particular task (e.g., start sequencing or end sequencing), as well as another loading station 280 for loading a sequencing cartridge containing the enriched fragment sample and / or for loading buffers required to perform a sequencing assay. Thus, once a user of the sequencer 270 provides the necessary reagents and sequencing cartridge to the loading station 280 of the sequencer 270, the user can initiate sequencing by interacting with the graphical user interface 275 of the sequencer 270. Once initiated, the sequencer 270 performs sequencing and outputs sequence reads of the enriched fragments from the nucleic acid sample 260.
[0092] In some embodiments, the sequencer 270 is communicatively coupled to the analysis system 200. The analysis system 200 includes several computing devices used to process sequence reads for various applications, such as determining methylation status at one or more CpG sites, variant calling, or quality control. The sequencer 270 may provide the sequence reads to the analysis system 200 in a BAM file format. The analysis system 200 may be communicatively coupled to the sequencer 270 through wireless communication techniques, wired communication techniques, or a combination of wireless communication techniques and wired communication techniques. In general, the analysis system 200 is configured to include a processor and a non-transitory computer-readable storage medium that stores computer instructions that, when executed by the processor, cause the processor to process sequence reads or perform one or more steps of any of the methods or processes disclosed herein.
[0093] In some embodiments, the sequence reads may be aligned to the reference genome using methods known in the art to determine alignment position information. The alignment position may generally describe the start and end positions of a region in the reference genome that corresponds to the starting and ending nucleotide bases of a given sequence read. Corresponding to methylation sequencing, the alignment position information may be generalized to indicate the first and last CpG sites included in the sequence read according to the alignment to the reference genome. The alignment position information may further indicate the methylation status and positions of all CpG sites in the given sequence read. The regions in the reference genome may be associated with genes or segments of genes. Thus, the analysis system 200 may label the sequence read with one or more genes that align to the sequence read. In one embodiment, the fragment length (or size) is determined from the start and end positions.
[0094] In various embodiments, for example, when a paired-end sequencing process is used, the sequence read is composed of a read pair shown as R_1 and R_2. For example, the first read R_1 may be sequenced from a first end of a double-stranded DNA (dsDNA) molecule, and the second read R_2 may be sequenced from a second end of the double-stranded DNA (dsDNA). Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 may be aligned (e.g., in opposite orientation) with the nucleotide bases of the reference genome. The alignment position information derived from the read pair R_1 and R_2 may include a start position (e.g., R_1) in the reference genome corresponding to the end of the first read and an end position (e.g., R_2) in the reference genome corresponding to the end of the second read. In other words, the start position and the end position in the reference genome represent the positions in the reference genome to which the nucleic acid fragment is likely to correspond. In one embodiment, the read pair R_1 and R_2 may be assembled into fragments, which are used for subsequent analysis and / or classification. An output file having a SAM (Sequence Alignment Map) format or a BAM (Binary) format may be generated and output for further analysis.
[0095] Reference is now made to Figure 2B, which is a block diagram of an analysis system 200 for processing a DNA sample, according to one embodiment. The analysis system implements one or more computing devices for use in analyzing a DNA sample. The analysis system 200 includes a sequence processor 210, a sequence database 215, a model database 225, one or more probabilistic models 230 and / or one or more classifiers 240, and a parameter database 235. In some embodiments, the analysis system 200 performs one or more steps in a method or process disclosed herein.
[0096] The sequence processor 210 generates methylation state vectors for fragments from the sample. For each CpG site on the fragment, the sequence processor 210 generates a methylation state vector for each fragment via process 360 of FIG. 4A that specifies the location of the fragment in the reference genome, the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment, i.e., methylated, unmethylated, or uncertain. The sequence processor 210 may store the methylation state vectors for the fragments in a sequence database 215. The data in the sequence database 215 may be organized such that the methylation state vectors from the samples are related to each other.
[0097] Additionally, multiple different models 230 may be stored in model database 225 or retrieved for use with a test sample. In one example, a model is a trained cancer classifier 240 to determine a cancer prognosis for a test sample using feature vectors derived from anomalous fragments. The training and use of cancer classifiers is discussed elsewhere herein. Analysis system 200 may train one or more models 230 and / or one or more classifiers 240 and store various trained parameters in parameter database 235. Analysis system 200 stores models 230 and / or classifiers along with functions in model database 225.
[0098] During inference, the machine learning engine 220 uses one or more models 230 and / or classifiers 240 and returns outputs. The machine learning engine accesses the models 230 and / or classifiers 240 in the model database 225 along with the trained parameters from the parameter database 235. According to each model, the machine learning engine 220 receives appropriate inputs for the model and calculates outputs based on the received inputs, parameters, and functions of each model that relate inputs and outputs. In some use cases, the machine learning engine 220 further calculates metrics that correlate with confidence in the calculated outputs from the models. In other use cases, the machine learning engine 220 calculates other intermediate values for use in the models.
[0099] II. Assay Protocol 3A is a flow chart illustrating a process 300 for sequencing nucleic acids, according to one embodiment. In some embodiments, the process 300 is performed to generate sequence reads as part of the sample sequencing 120 of the method 100 of FIG.
[0100] In step 310, a nucleic acid sample (e.g., DNA or RNA) was extracted from the subject. In this disclosure, unless otherwise indicated, DNA and RNA can be used interchangeably. That is, the embodiments described herein may be applicable to both DNA and RNA types of nucleic acid sequences. However, the examples described herein may focus on DNA for clarity and illustration. The sample may include nucleic acid molecules derived from any subset of the human genome, including the entire genome. The sample may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. In some embodiments, methods for collecting blood samples (e.g., syringe or finger prick) may be less invasive than procedures for obtaining tissue biopsies, which may require surgery. The extracted sample may include cfDNA and / or ctDNA. If the subject has a disease condition, such as cancer, the cell-free nucleic acid (e.g., cfDNA) in the sample extracted from the subject generally includes detectable levels of nucleic acid that can be used to determine the disease condition.
[0101] In step 315, the extracted nucleic acid (e.g., including cfDNA fragments) is treated to convert unmethylated cytosines to uracil. In some embodiments, method 300 uses bisulfite treatment of the sample, which converts unmethylated cytosines to uracil without converting methylated cytosines. For example, a commercially available kit such as EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, or EZ DNA Methylation™-Lightning kit (available from Zymo Research Corp, Irvine, CA) is used for bisulfite conversion. In another embodiment, the conversion of unmethylated cytosines to uracil is achieved using an enzymatic reaction. For example, the conversion can use a commercially available kit for the conversion of unmethylated cytosines to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).
[0102] In step 320, a sequencing library is prepared. In some embodiments, the preparation includes at least two steps. In the first step, a ssDNA ligation reaction is used to add ssDNA adaptors to the 3'-OH ends of the bisulfite-converted ssDNA molecules. In some embodiments, the ssDNA ligation reaction uses CircLigase II (Epicentre) to ligate ssDNA adaptors to the 3'-OH ends of the bisulfite-converted ssDNA molecules, the 5' ends of the adaptors being phosphorylated, and the bisulfite-converted ssDNA being dephosphorylated (i.e., the 3' ends have hydroxyl groups). In another embodiment, the ssDNA ligation reaction uses Thermostable 5'AppDNA / RNA ligase (available from New England BioLabs, Ipswich, Mass.) to ligate ssDNA adaptors to the 3'-OH ends of the bisulfite-converted ssDNA molecules. In this example, the first UMI adaptor is adenylated at the 5' end and blocked at the 3' end. In another embodiment, the ssDNA ligation reaction uses T4 RNA ligase (available from New England BioLabs) to ligate the ssDNA adaptor to the 3'-OH end of the bisulfite converted ssDNA molecule.
[0103] In the second step, the second strand DNA is synthesized in an extension reaction.For example, an extension primer that hybridizes to the primer sequence contained in the ssDNA adaptor is used in a primer extension reaction to form a double-stranded bisulfite-converted DNA molecule.Optionally, in some embodiments, the extension reaction uses an enzyme that can read through the uracil residue in the bisulfite-converted template strand.
[0104] Optionally, in a third step, dsDNA adaptors are added to the double-stranded bisulfite converted DNA molecules. The double-stranded bisulfite converted DNA can then be amplified to add sequencing adaptors. For example, PCR amplification using a forward primer containing a P5 sequence and a reverse primer containing a P7 sequence is used to add P5 and P7 sequences to the bisulfite converted DNA. Optionally, during library preparation, unique molecular identifiers (UMIs) can be added to nucleic acid molecules (e.g., DNA molecules) via adaptor ligation. UMIs are short nucleic acid sequences (e.g., 4-10 base pairs) that are added to the ends of DNA fragments during adaptor ligation. In some embodiments, UMIs are degenerate base pairs that function as unique tags that can be used to identify sequence reads originating from specific DNA fragments. During PCR amplification after adaptor ligation, the UMIs are replicated along with the attached DNA fragments, which provides a method of identifying sequence reads originating from the same original fragment in downstream analysis.
[0105] In optional step 325, the nucleic acids (e.g., fragments) can be hybridized. Hybridization probes (also referred to herein as "probes") can be used to target and pull down nucleic acid fragments that are informative for disease states. For a given workflow, the probes can be designed to anneal (or hybridize) with a target (complementary) strand of DNA or RNA. The target strand can be the "positive" strand (e.g., the strand that is transcribed into mRNA and then translated into protein), or the complementary "negative" strand. Probes can range in length from 10s, 100s, or 1000s of base pairs. Additionally, the probes can cover overlapping portions of the target region.
[0106] In optional step 330, the hybridized nucleic acid fragments can be captured and concentrated, e.g., amplified using PCR. In some embodiments, the target DNA sequence can be enriched from the library. This is used, for example, when a target panel assay is performed on the sample. For example, the target sequence can be enriched to obtain an enriched sequence that can then be sequenced. In general, any method known in the art can be used to isolate and enrich the probe-hybridized target nucleic acid. For example, as is well known in the art, a biotin moiety can be added to the 5' end of the probe (i.e., biotinylated) to facilitate isolation of the probe-hybridized target nucleic acid using a streptavidin-coated surface (e.g., streptavidin-coated beads).
[0107] In step 335, sequence reads are generated from the nucleic acid sample, e.g., the enriched sequences. Sequencing data can be obtained from the enriched DNA sequences by means known in the art. For example, the method can include next generation sequencing (NGS) techniques, including synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing by synthesis with reversible dye terminators.
[0108] In step 340, the sequence processor 210 can generate methylation information using the sequence reads. The methylation information determined from the sequence reads can then be used to generate a methylation state vector.
[0109] FIG. 3B is an illustration of a process 360 for obtaining methylation information and a methylation state vector, according to one embodiment. The process 360 may be part of the process 300 described in FIG. 3A. As an example, the analysis system receives a cfDNA molecule 312, which in this example contains three CpG sites. As shown, the first and third CpG sites of the cfDNA molecule 312 are methylated 314. During a treatment step 315, the cfDNA molecule 312 is converted to generate a converted cfDNA molecule 322. During treatment 315, the second CpG site, which was not methylated, has its cytosine converted to uracil. However, the first and third CpG sites were not converted.
[0110] After conversion, the sequencing library 330 is prepared and sequenced to generate sequence reads 342. The analysis system aligns the sequence reads 342 to a reference genome 344 (not shown). The reference genome 344 provides context as to where in the human genome the fragment cfDNA originates. In this simplified example, the analysis system aligns the sequence reads 342 such that the three CpG sites correlate with CpG sites 23, 24, and 25 (arbitrary reference identifiers used for convenience of illustration). Thus, the analysis system generates information about both the methylation status of all CpG sites on the cfDNA molecule 312 and the location in the human genome to which the CpG sites map. As shown, the CpG sites on the sequence reads 342 that are methylated are read as cytosines. In this example, cytosines appear only at the first and third CpG sites in the sequence reads 342, allowing one to infer that the first and third CpG sites in the original cfDNA molecule were methylated. Meanwhile, the second CpG site is read as a thymine (U is converted to T during the sequencing process), and therefore it can be inferred that the second CpG site was unmethylated in the original cfDNA molecule. Using these two pieces of information, namely methylation status and location, the analysis system generates 200 a methylation state vector 352 for the fragment cfDNA 312. In this example, the resulting methylation state vector 352 is <M 23 ,U 24 ,M 25 >, where M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscripts correspond to the position of each CpG site in the reference genome.
[0111] In some embodiments, one or more of the methylation state vectors 352 each have multiple adjacent methylation sites. The number of methylation sites may be fixed for some of the methylation state vectors 352. For example, the number may be five in some embodiments, but other numbers may also be used. Adjacent methylation sites may refer to sites at the same locus. For example, the sites may be consecutive CpG sites. In some embodiments, adjacent or consecutive sites do not necessarily mean that the two sites are immediately next to each other at the DNA sequence level. Instead, in some embodiments, the next adjacent site may refer to the next methylation state found in the sequence. The two adjacent sites may be 10 or 100 bases apart.
[0112] III. Identifying Anomalous Fragments In some embodiments, the analysis system uses the methylation state vector of the sample to determine the anomalous fragments of the sample. For example, for each nucleic acid molecule or fragment in the sample, the analysis system uses the methylation state vector corresponding to the nucleic acid molecule to determine (through analysis of sequence reads derived therefrom) whether the nucleic acid molecule or fragment is an anomalous methylated molecule or fragment, as compared to an expected methylation state vector from a non-cancer sample. In one embodiment, the analysis system calculates a p-value score for each methylation state vector (e.g., as described in U.S. Patent Application Publication No. 2019 / 0287652, incorporated herein by reference), which describes the probability of observing that methylation state vector, or the probability of observing other methylation state vectors that are less likely in a non-cancer control group. The process for calculating the p-value score is also discussed below in Section III. AP Value Filtering. The analysis system may determine, and optionally filter out, sequence reads of nucleic acid molecules or fragments that have a methylation state vector with a p-value score below a threshold to be anomalous fragments. In another embodiment, the analysis system further labels fragments having at least some number of CpG sites with methylation or unmethylation above some threshold percentage as hypermethylated and hypomethylated fragments, respectively. Hypermethylated and hypomethylated fragments may also be referred to as unusual fragments with extreme methylation (UFXM). In other embodiments, the analysis system may implement various other probabilistic models for determining anomalous molecules or fragments. Examples of other probabilistic models include mixture models, deep probabilistic models, and the like. In some embodiments, the analysis system may use any combination of the processes described below to identify anomalous fragments. With the identified anomalous fragments, the analysis system may filter a set of methylation state vectors for the sample for use in other processes, for example, for use in training and deploying a cancer classifier.
[0113] III.AP value filtering In one embodiment, the analysis system calculates a p-value score for each methylation state vector that is compared to the methylation state vectors from fragments in a non-cancer control group. The non-cancer control group may include healthy samples (e.g., from healthy individuals not diagnosed with any disease) or other non-cancer samples (e.g., from individuals not diagnosed with cancer but may include other diagnoses). The p-value score describes the probability of observing a nucleic acid molecule in the non-cancer control group that has a methylation status consistent with that methylation state vector. To determine that a DNA fragment is anomalously methylated, the analysis system uses a non-cancer control group in which the majority of fragments are normally methylated. When performing this probabilistic analysis to determine anomalous fragments, the determination carries a weight compared to the group of control subjects that constitute the non-cancer control group. To ensure robustness in the non-cancer control group, the analysis system may select some threshold number of healthy individuals to source samples containing DNA fragments. Figure 4A below describes how the analysis system generates a data structure for the non-cancer control group from which the p-value score can be calculated. Figure 4B describes how the generated data structure is used to calculate the p-value score.
[0114] 4A is a flow chart illustrating process 400 for generating a data structure for a non-cancer control group, according to one embodiment. To create the non-cancer control group data structure, the analysis system receives a plurality of DNA fragments (e.g., cfDNA) from a plurality of healthy and / or non-cancer individuals. A methylation state vector is identified for each fragment, e.g., via process 300 and / or process 360.
[0115] Using the methylation state vector of each fragment, the analysis system subdivides 405 the methylation state vector into strings of CpG sites. In one embodiment, the analysis system subdivides 405 the methylation state vector such that the resulting strings are all shorter than a given length. For example, a methylation state vector of length 11 may be subdivided into strings of three or less lengths, resulting in nine strings of length 3, ten strings of length 2, and eleven strings of length 1. In another example, a methylation state vector of length 7 may be subdivided into strings of four or less lengths, resulting in four strings of length 4, five strings of length 3, six strings of length 2, and seven strings of length 1. If the methylation state vector is shorter than or equal to the specified string length, the methylation state vector may be converted into a single string that includes all of the CpG sites of the vector.
[0116] The analysis system 200 tallies 410 the strings by counting the number of strings present in the control group that have the CpG site designated as the first CpG site in the string and have the possible methylation state, for each possible CpG site in the vector and possible methylation state. For example, at a given CpG site, given a string length of 3, there are 2^3 or 8 possible string configurations. At that given CpG site, for each of the 8 possible string configurations, the analysis system records 410 how many times each possible methylation state vector occurs in the control group. Continuing with the example, this means that for each starting CpG site x in the reference genome, the following quantities are tallied: <M x ,M x+1 ,M x+2 >, <M x ,M x+1 ,U x+2 >,..., x ,U x+1 ,U x+2 The analysis system creates 415 a data structure that stores the tabulated counts for each start CpG site and string possibility.
[0117] There are several advantages to setting an upper limit on string length. First, depending on the maximum length of the string, the size of the data structure created by the analysis system can increase dramatically in size. For example, a maximum string length of 4 means that all CpG sites have at least 2^4 numbers to count for strings of length 4. Increasing the maximum string length to 5 means that all CpG sites have an additional 2^4 or 16 numbers to count compared to the previous string length, doubling the number to count (and the computer memory required) compared to the previous string length. Reducing the string size helps keep data structure creation and execution (e.g., for use for later access as described below) reasonable in terms of computation and storage. Second, the statistical consideration of limiting the maximum string length is to avoid overfitting of downstream models that use string counts. If a long string of CpG sites does not biologically have a strong effect on the outcome (e.g., predicting an anomaly that predicts the presence of cancer), calculating a probability based on a large string of CpG sites can be problematic because it requires a significant amount of data that may not be available and therefore becomes too sparse for the model to operate properly. For example, calculating the probability of an anomaly / cancer conditional on the previous 100 CpG sites requires a count of strings in a data structure of length 100, ideally some that match exactly with the methylation state of the previous 100. If only a sparse count of strings of length 100 is available, there is insufficient data to determine whether a given string of length 100 in a test sample is anomaly.
[0118] 4B is a flow chart illustrating process 420 for identifying abnormally methylated fragments from an individual, according to one embodiment. In process 420, the analysis system generates methylation state vectors from the subject's cfDNA fragments, e.g., by process 300 and / or process 360. The analysis system treats each methylation state vector as follows:
[0119] For a given methylation state vector, the analysis system enumerates 430 all possibilities for methylation state vectors that have the same starting CpG site and the same length (i.e., set of CpG sites) as in the methylation state vector. Since each methylation state is generally either methylated or unmethylated, there are effectively two possible states for each CpG site, and therefore the count of different possibilities for a methylation state vector is calculated by dividing the nth possible methylation state vector by the nth possible methylation state vector. n The probability of a methylation state vector depends on a power of 2 so that it is associated with a probability of 1. In the case of having a methylation state vector that includes an uncertain state for one or more CpG sites, the analysis system may enumerate 430 the possibilities of the methylation state vector considering only CpG sites that have an observed state.
[0120] The analysis system calculates 440 the probability of observing each possible methylation state vector for the identified starting CpG site and methylation state vector length by accessing the non-cancer control group data structure. In one embodiment, calculating the probability of observing a given possibility models the joint probability calculation using Markov chain probability. In other embodiments, calculation methods other than Markov chain probability are used to determine the probability of observing each possible methylation state vector.
[0121] The analysis system uses the calculated probability for each possibility to calculate 450 a p-value score for the methylation state vector. In one embodiment, this involves identifying a calculated probability that corresponds to the possibility of matching the methylation state vector in question. Specifically, this is the possibility of having the same set of CpG sites as the methylation state vector, or also the same starting CpG site and length. The analysis system sums the calculated probabilities of all possibilities that have a probability less than or equal to the identified probability to generate a p-value score.
[0122] This p-value represents the probability of observing the methylation state vector of the fragment, or other methylation state vectors that are less likely in the non-cancer control group. Thus, a low p-value score generally corresponds to a methylation state vector that is rare in healthy or non-cancer individuals and causes the fragment to be labeled as anomalously methylated relative to the non-cancer control group. A high p-value score generally relates to a methylation state vector that is expected to be present in a non-cancer individual in a relative sense. For example, if the non-cancer control group is a non-cancer group, a low p-value indicates that the fragment is anomalously methylated relative to the non-cancer group, and thus likely indicates the presence of cancer in the test subject.
[0123] As described above, the analysis system calculates a p-value score for each of a plurality of methylation state vectors, each of which represents a cfDNA fragment in a test sample. To identify which of the fragments are anomaly methylated, the analysis system may filter 460 the set of methylation state vectors based on their p-value scores. In one embodiment, the filtering is performed by comparing the p-value scores to a threshold and retaining only those fragments that are below the threshold. This threshold p-value score may be on the order of 0.1, 0.01, 0.001, 0.0001, or the like.
[0124] According to an exemplary result from the process, the analysis system yielded a median number of aberrant methylation patterns of 2,800 fragments (ranging from 1,500 to 12,000 fragments) for participants without cancer who participated in the training, and a median number of aberrant methylation patterns of 3,000 fragments (ranging from 1,200 to 220,000 fragments) for participants with cancer who participated in the training. These filtered sets of fragments with aberrant methylation patterns can be used for downstream analysis as described below.
[0125] In one embodiment, the analysis system uses 455 a sliding window to determine the probabilities of the methylation state vector and to calculate p-values. Instead of enumerating the probabilities and calculating p-values for the entire methylation state vector, the analysis system enumerates the probabilities and calculates p-values only over a window of contiguous CpG sites, where the window is shorter in length (of CpG sites) than at least some fragment (otherwise the window serves no purpose). The window length may be static, user-determined, dynamic, or otherwise selected.
[0126] When calculating p-values for a methylation state vector larger than a window, the window identifies a set of contiguous CpG sites from the vector within the window, starting from the first CpG site in the vector. The analysis system calculates a p-value score for the window that includes the first CpG site. The analysis system then "slides" the window to a second CpG site in the vector and calculates another p-value score for the second window. Thus, for a window size of l and a methylation vector length of m, each methylation state vector generates m-l+1 p-value scores. After completing the p-value calculations for each portion of the vector, the lowest p-value score from all sliding windows is taken as the overall p-value score for the methylation state vector. In another embodiment, the analysis system aggregates the p-value scores of the methylation state vectors to generate an overall p-value score.
[0127] The use of a sliding window helps to reduce the number of possibilities of the methylation state vector to be enumerated and their corresponding probability calculations that would otherwise need to be performed. To give a realistic example, a fragment can have more than 54 CpG sites. Instead of calculating probabilities for 2^54 (approximately 1.8×10^16) possibilities to generate a single p-score, the analysis system can instead use a window of size 5 (for example), which results in 50 p-value calculations for each of the 50 windows of the methylation state vector for the fragment. Each of the 50 calculations enumerates 2^5 (32) possibilities of the methylation state vector, which in total results in 50×2^5 (1.6×10^3) probability calculations. This results in a significant reduction in the calculations that are performed without meaningful hits for the correct identification of anomalous fragments.
[0128] In embodiments with uncertain states, the analysis system may calculate a p-value score that sums out the CpG sites with uncertain states in the methylation state vector of the fragment. The analysis system identifies all possibilities that have a match with all methylation states in the methylation state vector, excluding the uncertain states. The analysis system may assign a probability to the methylation state vector as the sum of the probabilities of the identified possibilities. By way of example, the analysis system may<M1,M2,U3> and<M1,U2,U3> The methylation state vector<M1,I2,U3> The method calculates the probability that methylation states for CpG sites 1 and 3 are observed and match the methylation states of the fragments at CpG sites 1 and 3. The method of summing out CpG sites with uncertain states uses a calculation of probabilities of up to 2^i possibilities, where i denotes the number of uncertain states in the methylation state vector. In additional embodiments, a dynamic programming algorithm may be implemented to calculate the probability of a methylation state vector having one or more uncertain states. Advantageously, the dynamic programming algorithm operates in linear computation time.
[0129] In one embodiment, the computational burden of calculating the probability and / or p-value scores may be further reduced by caching at least some of the calculations. For example, the analysis system may cache in temporary or persistent memory the probability calculations for the likelihood of a methylation state vector (or a window thereof). Caching the likelihood probabilities allows for efficient calculation of p-score values without needing to recalculate the underlying likelihood probabilities if other fragments have the same CpG sites. Equivalently, the analysis system may calculate a p-value score for each of the possibilities of a methylation state vector associated with a set of CpG sites from the vector (or a window thereof). The analysis system may cache the p-value scores for use in determining the p-value scores of other fragments containing the same CpG sites. In general, the p-value scores of the possibilities of methylation state vectors with the same CpG sites may be used to determine the p-value scores of different ones of the possibilities from the same set of CpG sites.
[0130] III.B. Hypermethylated and Hypomethylated Fragments In some embodiments, the analysis system determines anomalous fragments as those having more than a threshold number of CpG sites, more than a threshold percentage of CpG sites being methylated, or more than a threshold percentage of CpG sites being unmethylated, and the analysis system identifies such fragments as hypermethylated or hypomethylated. Exemplary thresholds for fragment (or CpG site) length include greater than 3, greater than 4, greater than 5, greater than 6, greater than 7, greater than 8, greater than 9, greater than 10, etc. Exemplary percentage thresholds for methylation or unmethylation include greater than 80%, greater than 85%, greater than 90%, or greater than 95%, or any other percentage in the range of 50%-100%.
[0131] III.C. Reference Genome Blocks 5 is an illustration of blocks of a reference genome, according to one embodiment. The sequence processor 210 can partition the reference genome (or a subset of the reference genome) in one or more stages, for use cases including, for example, targeted methylation assays. For example, the sequence processor 210 separates the reference genome into blocks of methylation sites (e.g., CpG sites). Each block is defined when there is a separation between two adjacent CpG sites that exceeds a threshold value, e.g., 200 base pairs (bp), 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or more than 1,000 bp, among other values. Thus, the blocks can vary in size in base pairs. For each block, the sequence processor 210 can subdivide the block into windows of a certain length, e.g., 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1,100 bp, 1,200 bp, 1,300 bp, 1,400 bp, or 1,500 bp, among other values. In other embodiments, the windows can be 200 bp to 10 kilobase pairs (kbp) in length, 500 bp to 2 kbp, or about 1 kbp. The windows (e.g., contiguous ones) can overlap by a few base pairs or a percentage of the length, e.g., 10%, 20%, 30%, 40%, 50%, or 60%, among other values. The window can be separated between two adjacent CpG sites by greater than a threshold value, for example, 200 base pairs (bp), 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1,000 bp, among other values.
[0132] The sequence processor 210 can analyze sequence reads derived from DNA fragments using a windowing process. In particular, the sequence processor 210 scans the block window by window and reads the fragments in each window. The fragments can originate from tissues and / or high signal cfDNA. The high signal cfDNA samples can be determined by a binary classification model, by cancer stage, or by another metric. By partitioning the reference genome (e.g., using blocks and windows), the sequence processor 210 can facilitate computational parallelization. Furthermore, the sequence processor 210 can reduce computational resources for processing the reference genome by targeting sections of base pairs that contain CpG sites while skipping other sections that do not contain CpG sites.
[0133] IV. Cancer Classification Model FIG. 6 is a flow chart of a method 600 for identifying a plurality of features for generating a classifier for predicting a disease state (e.g., presence or absence of a disease, a type of disease, and / or a disease tissue of origin), according to various embodiments. In some embodiments, the analysis system 200 executes the method 600 to process sequence reads of fragments from a nucleic acid sample. The method 600 includes, but is not limited to, the following steps: generating 610 sequence reads; training 620 variant models associated with each of a plurality of different disease states (e.g., different cancer types); applying 630 the variant models to determine a value based on a probability that the sequence read originates from a sample associated with each of a plurality of disease states associated with each variant model; identifying 640 features by determining a count of sequence reads having a value above a threshold; training 650 a classifier using the features, and optionally applying 660 the classifier to predict a disease state and / or a tissue of origin associated with the disease state. Each of which is described with respect to components of the analysis system 200.
[0134] The analysis system 200 generates 610 a first set of sequence reads from a plurality of samples, each having a known or suspected disease state, such as the presence or absence of disease, the type of disease, and / or the disease tissue of origin. For example, in some embodiments, the plurality of samples can include any number of cancer samples from individuals known to have cancer, and / or non-cancer samples from individuals who do not have cancer. In addition, the samples can include any of acellular nucleic acid samples (e.g., cfDNA), solid tumor samples, and / or other types of samples.
[0135] Analysis system 200 trains 620 variant models, each associated with a different disease state. Analysis system 200 isolates a cohort of known disease states to train the variant models. To train the variant models, analysis system 200 may construct a distribution based on the counts of methylation variants across genomic regions in the cohort for a given disease state. The trained variant models may be configured to input methylation sequence reads and output the probability that the methylation sequence reads result from a disease state. In one or more embodiments, there may be two disease states, i.e., non-cancer and cancer, resulting in two variant models. In other embodiments, there may be more disease states, i.e., non-cancer, a first cancer type, a second type, etc. Analysis system 200 may train each variant model based on the methylation sequence reads of each cohort.
[0136] Analysis system 200 identifies 640 features by determining a count of methylated sequence reads having a value above a threshold. For example, a cancer variant model inputs methylated sequence reads and outputs a probability that the methylated sequence read originates from a cancer tissue or cell. Analysis system 200 may set a threshold probability before counting methylated sequence reads for a feature. In some example implementations, the threshold is selected as 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%.
[0137] Analysis system 200 trains 650 a cancer classifier using the features. Analysis system 200 may train the cancer classifier using a cancer cohort of cancer training samples and a non-cancer cohort of non-cancer training samples. Analysis system 200 may characterize each training sample resulting in a feature vector according to step 640. Analysis system 200 then trains the cancer classifier to determine a cancer prediction based on the input feature vector. Analysis system 200 may train the cancer classifier as a machine learning model. In one or more embodiments, the cancer training samples may have known cancer type such that analysis system 200 can train the cancer classifier to further predict tissue of origin in the cancer prediction.
[0138] Analysis system 200 may utilize the trained cancer classifier to predict 660 the disease state of the test sample and / or tissue of origin associated with the disease state by applying the cancer classifier. Analysis system 200 may use the variant model to characterize the test sample according to step 640, resulting in a test feature vector. Analysis system 200 may apply the cancer classifier to the test feature vector to determine a cancer prediction, i.e., input the test feature vector to the cancer classifier, which outputs a cancer prediction. The cancer prediction may be a binary prediction, i.e., a positive disease state or a negative disease state. The cancer prediction may additionally or alternatively be a multi-class prediction, i.e., no disease state, a first disease state, a second disease state, etc.
[0139] In some embodiments, the analysis system 200 can train the cancer classifier according to the type of sample used. For example, the analysis system 200 can train the cancer classifier in one way according to blood samples and in another way according to urine samples. Training in different ways according to the type of sample used ensures that the variation in the level of cell-free DNA present in different types of samples does not bias the cancer prediction. For example, a first type of sample may have a higher level of cfDNA compared to a second type of sample. A cancer classifier trained on a first type of sample and then applied to a second type of sample may skew the cancer prediction.
[0140] In some embodiments, the analysis system 200 may train a cancer classifier to incorporate other types of features in addition to the methylation features described above. Other exemplary features include fragment length, endpoint information to better understand tumor-derived or non-tumor-derived fragments, protein readouts or mutations, other genetic material including RNA, covariates, or any combination thereof. Other features may be incorporated as inputs to the cancer classifier with the methylation features as input. In other embodiments, other features may be used before and / or after classification. Before classification, the other features may be used to classify samples into one or more categories, for example based on covariates. For example, the other features may be used to separate samples according to predicted ethnicity, predicted age group, predicted smoking status, etc. After classification, the analysis system 200 may train a first classifier based on the methylation features as input and train a second classifier whose input consists of the output from the first classifier and the other features fed to the second classifier. The first classifier and / or the second classifier may be a machine learning model (e.g., gradient boosting).
[0141] In some embodiments, the analysis system 200 trains a first classifier (e.g., a cancer classifier) to generate a cancer prediction and trains a second classifier to predict tumor fraction, methylation variant allele fraction, or some combination thereof. The first classifier may be trained utilizing methylation features and may further include other features. The first classifier inputs the features and outputs a cancer prediction. The cancer prediction may be binary (presence or absence of cancer, and / or likelihood thereof) and / or multi-class (presence of one or more cancer types, and / or likelihood thereof). Based on the cancer prediction, the second classifier inputs the methylation features and may further include other features to predict tumor fraction, methylation variant allele fraction, or some combination thereof. In an exemplary embodiment, the cancer classifier ranks the cancer type (or tissue of origin) for a particular sample. Based on the ranked cancer types, the analysis system can utilize a separately trained classifier for each ranked cancer type to predict tumor fraction, methylation variant allele fraction, or some combination thereof. Each classifier can be trained independently with different training parameters and metrics (e.g., sensitivity, specificity, etc.).
[0142] IV.A. Tumor Fraction Models and Banks FIG. 7 is a conceptual diagram illustrating how a tumor fraction estimation model 710 may be trained using a bank 720 of reference sequence reads obtained from a reference individual, according to various embodiments. The model 710 may be used to estimate a subject's tumor fraction using the subject's methylated sequence reads. Tumor fraction, in this context, may refer to the proportion of fragments (e.g., DNA fragments) in the subject's biological sample that are tumor-derived. In some embodiments, the model 710, or an adjusted version of the model, may also be used to estimate a subject's allele fraction. In this context, allele fraction may refer to the proportion of fragments that are tumor-derived and contain a variant corresponding to an allele.
[0143] In some embodiments, the model 710 used to estimate the tumor fraction may be trained based on sequence reads of biological samples of various reference individuals with known disease states. A collection of these sequence reads may be referred to as a bank of reference sequence reads 720. In some embodiments, the reference individuals may include individuals diagnosed with one or more cancers and healthy individuals (e.g., not diagnosed with cancer). In some embodiments, different biological samples may be obtained to build the bank. For example, some reference biological samples are tissue biopsy samples of individuals, and other reference biological samples are cfDNA samples obtained from various types of bodily fluids. In some embodiments, one or more reference individuals may provide more than one type of biological sample, while other reference individuals may provide only a single biological sample. For example, a first reference individual may provide bodily fluid samples from different bodily tissues and various biopsy samples, while a second reference individual may provide only a bodily fluid sample or a single biopsy sample. The bank 720 may include a mixture of various individuals with different biological samples provided.
[0144] Non-limiting examples of tissue samples include breast tissue, thyroid tissue, lung tissue, bladder tissue, cervical tissue, small intestine tissue, colorectal tissue, esophageal tissue, stomach tissue, tonsil tissue, liver tissue, ovarian tissue, fallopian tube tissue, pancreatic tissue, prostate tissue, kidney tissue, and uterine tissue. In some embodiments, tissue samples may be assigned to one or more groups that include similar tissues. For example, lung adenocarcinoma and lung squamous cell carcinoma may be separate or may be grouped into a single lung cancer group. Non-limiting examples of bodily fluids include blood, sweat, urine, and saliva. Bank 720 may include sequence reads generated from two, three, four, or all of the different types of tissues. Bank 720 may also include sequence reads of cfDNA samples generated from one or more types of bodily fluids.
[0145] In some embodiments, the sequence reads in bank 720 may be generated using various sequencing methods. For example, for methylation sequencing, whole genome bisulfite sequencing (WGBS) and targeted methylation (TM) assays may be used. Other methods for methylation sequencing, including those disclosed herein and / or any modifications, substitutions, or combinations thereof, may also be used to generate methylation patterns. The sequence reads in bank 720 may also include control sequence reads for both WGBS and TM assays. The control samples may be sequence reads obtained from randomly generated sequences using one or more predefined ratios of methylated and unmethylated sites (e.g., 50:50, 40:60, x 60:40, etc.). For example, the control sample may be a 50:50 mixture of fully methylated and fully unmethylated sheared genomic DNA. Although bank 720 is discussed primarily with methylated sequencing reads, DNA nucleotide reads may also be included in the bank. For example, in some cases, a methylation sequence read obtained from a biological sample may also have a corresponding DNA nucleotide read obtained from the sample. The corresponding nucleotide reads may also be stored in bank 720.
[0146] In some embodiments, the bank 720 includes a collection of methylation variants 730 that can be generated by filtering biopsy WGBS reference samples 732 and non-cancer cfDNA WGBS reference samples 734. In some embodiments, WGBS can take the form of a sequencing process in which a nucleic acid sample undergoes bisulfite treatment before the converted nucleic acid molecules are evaluated for sequencing information and methylation status on a genome-wide basis. In some embodiments, a whole genome bisulfite sequencing assay looks for variations in methylation patterns in a genome. An example of a WGBS process is illustrated in FIG. 3. cfDNA fragments are obtained from a sample. From the fragments, the location and methylation state of each of the CpG sites are determined based on alignment of the nucleic acid fragments to a reference genome. The methylation state vector for each fragment can be used to represent information such as the location of the fragment in the reference genome (e.g., specified by the location of the first CpG site in each fragment, or another similar metric), the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment. Each methylation state vector, which may include multiple potential methylation sites, may be referred to as a methylation variant. In some embodiments, the term whole genome is used, but WGBS may determine the methylation sequence reads of a portion of a genome without analyzing the entire genome. For example, WGBS may include the analysis of a large portion of a genome.
[0147] In some embodiments, the various biopsy WGBS reference samples 732 and non-cancer cfDNA WGBS reference samples 734 are obtained from different reference individuals. The methylation variants 730 are generated by filtering out methylation variants in the biopsy WGBS reference sample 732 that have a noise rate higher than a threshold rate. For example, only methylation variants in the biopsy WGBS reference sample 732 that have a noise rate lower than 1 / 10,000 (or another suitable threshold rate) are retained after passing through the filter 740. The noise rate represents an estimated background noise rate for a particular variant in the non-cancer cfDNA data (e.g., using both TM and WGBS non-cancer data). The noise rate may be defined in any suitable statistical manner. In some embodiments, the filtering 740 may be performed by determining, for each methylation state vector, a metric that describes the probability of observing that methylation state vector or other methylation state vectors that are less likely in the non-cancer control group. The value of the metric is compared to a threshold value (e.g., 1 / 10,000).
[0148] In some embodiments, the bank 720 also includes targeted methylation (TM) sequence reads 750 of non-cancer cfDNA samples from healthy reference individuals. TM may take the form of an assay that examines the methylation status of a distinct set of targets in a genome. In various embodiments, TM sequencing can be performed in different ways. Combinations with different enzyme and chemical treatments can convert either methylated or unmethylated cytosines. For example, in some embodiments, targeted DNA methylation sequencing detects one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in a plurality of nucleic acids. As another example, targeted DNA methylation sequencing can include converting one or more unmethylated cytosines or one or more methylated cytosines in a plurality of nucleic acids to the corresponding one or more uracils. As another example, in some embodiments, the targeted DNA methylation sequencing can include converting one or more unmethylated cytosines in the plurality of nucleic acids to one or more corresponding uracils, and the DNA methylation sequence reads out the one or more uracils as one or more corresponding thymines. In some embodiments, the targeted DNA methylation sequencing includes converting one or more methylated cytosines in the plurality of nucleic acids to one or more corresponding uracils, and the DNA methylation sequence reads out the one or more 5mC or 5hmC as one or more corresponding thymines.
[0149] In the TM sequencing process, probes are used to enrich nucleic acid samples. In some embodiments, the probes can be designed to bind to sequences after cytosines at methylated or unmethylated CpG sites have been converted (e.g., in a chemical or enzymatic conversion process). In embodiments in which methylation sequencing is used, the sequence of the probes can be complementary to the sequence of the converted DNA fragments rather than to the corresponding genomic sequence.
[0150] The TM sequence reads 750 of non-cancer cfDNA samples can be used to improve noise rate estimates. The TM assay can enrich recurrent methylation variants. The number of methylation variants per block or window of methylation sites in the TM assay can be an order of magnitude larger than the number of methylation variants in the WGBS. Thus, by including the TM sequence reads 750 of non-cancer cfDNA samples in the bank 720, the signal-to-noise ratio of the target methylation sites can be improved, thereby improving the noise rate estimate of the methylation variants 730.
[0151] In some embodiments, the bank 720 also includes control TM and WGBS samples 760 to estimate the pull-down bias of various methylation variants. The pull-down bias can be specific to a particular methylation variant. The pull-down bias is typically introduced through the use of probes with a particular allelic pattern, expressed relative to the exclusion of alternative allelic patterns at a qualifying methylation pattern (QMP) genomic site. In some embodiments, the pull-down bias is estimated by the bias of the QMP genomic site i(bias i ), and (bias i ) is the pull-down bias at QMP genomic site i, as follows:
[0152]
number
[0153] This above pull-down bias is corrected for the pull-down bias in the targeted methylation sequencing at QMP genomic site i using WGBS control data and TM control data. In particular, such control data is used to calculate alpha. For example, to calculate alpha, an aberration count at each site in a plurality of QMP genomic sites (study controls) from the WGBS control is obtained ("control (WGBS count) aberration count"). Thus, there are a plurality of WGBS aberration counts for different QMP genomic sites, each obtained using the WGBS control. There is no specific requirement regarding the cancer state of this WGBS control. In other words, the WGBS control may or may not have a specific cancer state. In some embodiments, the WGBS control is an engineered cell line with a predetermined known percentage of methylated genomic DNA sequenced using WGBS. In some embodiments, the WGBS control is a mixture of 0% methylation and 100% methylated genomic DNA of a predetermined composition (e.g., a 50 / 50 or 40 / 60 or 30 / 70 mixture of 0% and 100% methylated genomic DNA). Further, an aberration count at each site in the plurality of QMP genomic sites from the targeted methylation sequencing is obtained ("TM control (TM count) aberration count"). In some embodiments, the source of DNA for the TM control is the same as that for the WGBS control, the only difference being that for the TM control, the control DNA is sequenced using targeted sequencing with the pull-down probes used in the TM, rather than with WGBS. The quantity alpha in such embodiments may represent the slope of a line fitted to a scatter plot of the control (WGBS count) aberration count / TM control (TM count) aberration count. Each respective point in the scatter plot is for a different QMP genomic site j in the plurality of QMP genomic sites under study, the x coordinate for each point is the (WGBS count) aberration count at genomic site j, and the y coordinate for each point is the (TM count) aberration count at genomic site j.Furthermore, as shown in the formula for alpha, in an exemplary embodiment, only data from the 75th quantile of WGBS controls (WGBS counts) abnormal counts and only data from the 75th quantile of TM controls (TM counts) are used in the scatter plot from which alpha is computed. The quantity alpha is the slope of the line fitted to the scatter plot data. The use of the 75th quantile is exemplary and can be adjusted upwards (e.g., 85th quantile) or downwards (e.g., 65th quantile) depending on the application. For example, it can be treated as a hyperparameter that is optimized as part of the optimization of the downstream classifier. Furthermore, rather than performing a quantile cut, other methods for removing outliers can instead be used before using the scatter plot to compute alpha.
[0154] Although bank 720 has been described as having methylation variants 730, TM sequence reads 750, and control TM and WGBS samples 760, in various examples bank 720 may include fewer or additional types of samples. For example, in some embodiments bank 720 may include only methylation variants 730. In other embodiments bank 720 may include other types of samples that are used to improve various parameters in model 710, including noise estimates, pulldown estimates, depth estimates, and other suitable parameters.
[0155] Biopsy samples typically have more than 1,000 methylation variants. Figure 7 is an exemplary plot of the number of methylation variants found per participant, according to some embodiments. Each dot in the plot 700 represents an individual. The median number of methylation variants found per individual is about 2,635, according to one experiment. Thus, for a potential subject, the number of methylation variants found in any biological sample of the subject can be expected to range into the thousands, regardless of whether the sample is a biopsy or a cfDNA sample. The large number of methylation variants in the subject allows the methylation variants stored in the bank 720 to be used to build a model 710 that is used to estimate the tumor fraction.
[0156] Each methylation variant 730 may be associated with a recurrence rate. The recurrence rate for a given methylation variant may refer to the fraction of the methylation variant present in other samples in the bank 720. For example, the recurrence rate of a given methylation variant may be determined based on the number of samples containing one or more supporting fragments that include the methylation variant divided by the total number of samples. In some embodiments, methylation variants are retained using a per-participant threshold for the recurrence rate to be counted. For example, for a particular individual's cancer tissue WGBS sequencing, the analysis system identifies methylation variants that exceed a threshold occurrence in this one sample. Methylation variants that pass this threshold in at least one reference sample are then retained in the multiple variants. According to various embodiments, the recurrence rate of the methylation variant is expected to be high. Figure 8 is an exemplary plot of the recurrence rate of various methylation variants found in different participants, according to some embodiments. Each dot in the plot 800 represents a methylation variant. As shown in plot 800, the majority of methylation variants have a recurrence rate greater than 0.75. The median recurrence rate is 0.868. Thus, the study according to one embodiment shows that methylation variants are recurrent. In other words, methylation variants associated with a disease, such as cancer, found in one biological sample are expected to be found in other biological samples associated with the disease.
[0157] In some embodiments, the methylation variants with a relatively high recurrence rate and the number of recurrent methylation variants are used to determine the tumor fraction of the subject. The recurrence characteristics are specific to the methylation variants and are not observed in other types of variants such as single base variants, insertions, or deletions. In other words, the methylation variants determined by performing sequencing of a biopsy sample of a specific tissue are expected to be found in other tissues of the subject. Thus, the presence or absence of the methylation variants can be used to determine the cancer status of the subject. It is not necessary to determine the methylation variants from the tissue that is the primary source of the cancer, since the methylation variants are expected to be observed in other tissues as well. The methylation variants are also expected to be observed from the cfDNA sample. Based on the high recurrence rate for the majority of the methylation variants found in biological samples, the methylation variants in the cfDNA sample are correlated with the methylation variants in the biopsy sample of the tissue suspected to be the primary source of the cancer. Thus, in some embodiments, the tumor fraction of the subject can be estimated using the model 710 using only the cfDNA sample of the subject without a matching biopsy, while in other embodiments, a matching biopsy can also be provided. Using only cfDNA samples to determine tumor fraction is feasible using methylation variants, but using other variants such as single nucleotide variants (SNVs) can be difficult, because SNVs often have low recurrence rates.For example, cancer-related single nucleotide variants found in tissues may not necessarily be expected to be found in cfDNA samples, because they have low recurrence rates.Based on the large number of methylation variants and the high recurrence rates expected to be found in individuals, the tumor fraction of a subject can be determined based on a model 710 that considers the recurrence rates of different methylation variants.
[0158] For example, in some embodiments, the model 710 may be trained based on the recurrence rate of the methylation variants 730 in the reference sequence reads stored in the bank 720. Some methylation variants 730 may be determined to be likely cancer-associated. The presence or absence of such cancer-associated methylation variants in the biological sample is used in the model 710 to determine the tumor fraction of the sample. The relative weight of each methylation variant may be discounted based on the recurrence rate of the methylation variant. For example, the model 710 may associate a lower weight with a particular methylation variant if the particular methylation variant has a low recurrence rate. Because the presence or absence of a methylation variant in a sequenced biological sample does not necessarily mean the presence or absence of the same methylation variant in another tissue. The training of the model 710 takes into account the recurrence rate of the methylation variants 730. One or more statistical models and distributions may be used to model the correlation between the methylation variants 730 and the tumor fraction. The accuracy of the model 710 can be further refined using non-cancer cfDNA TM sequence reads 750 to improve noise rate estimates, and using control TM and WGBS samples 760 to improve pull-down bias estimates.
[0159] The model 710 may be a machine learning model or an algorithmic model that includes a set of rules and one or more statistical models used to determine an estimate of tumor fraction. In various embodiments, the model 710 may take different forms depending on the implementation. For example, the model 710 may include a constant model, a binomial model, a Poisson model, an independent site model, a neural network model, and / or a Markov model. In some embodiments, the model 710 is trained based on the recurrence rates of various methylation variants 730 in the reference sequence reads bank 720. In some embodiments, at least a portion of the model 710 may be represented as a probability distribution that may be expressed as follows:
[0160]
number
[0161] The above formula expresses the conditional probability of a tumor fraction (tf) value given a subject's data (e.g., a cfDNA sample) as a distribution of multiples of the various methylation variant-specific Poisson distributions. i represents the count of differentially methylated variants at site i in a subject's cfDNA sample, and λ i represents the Poisson lambda parameter for site i. The Poisson lambda parameter for i can be expressed as:
[0162]
number
[0163] In the above formula, vaf i represents the variant allele fraction at site i in the biopsy, and noise i represents the site-specific noise rate in the cfDNA sample, and depth i represents a depth estimate of site i in the cfDNA sample, including those not pulled down by the probe. The depth estimate at site i may be affected by the pull-down bias of site i, which may be determined based on the control TM and WGBS samples 760 in bank 720.
[0164] In some embodiments, the probability of observing a methylation variant count given a tumor fraction value may be expressed as a probability distribution, which can be expressed as follows:
[0165]
number
[0166] In the above formula,
[0167]
number
[0168]
number
[0169]
number
[0170] In some embodiments, the probability of observing a methylation variant count given a tumor fraction value may also be assigned a uniform distribution or may be biased. For example, the probability of observing a methylation variant count given a tumor fraction value may be assigned a term α × Uniform(0,d i ), where α represents the fraction of the likelihood of assigning a probability to a uniform distribution, and d i represents the fractional count of fragments overlapping with methylation variant i.
[0171] In some embodiments, the analysis system may filter methylation variants based on a set of non-cancer samples of the subject. In such embodiments, non-cancer tissue data is used to filter out common methylation variants present in non-cancer samples and / or healthy samples. This is advantageous for complementing non-cancer sample types that may be vulnerable to low levels of signal-to-noise. For example, the true noise for a particular variant in a non-cancer urine sample may be 1:10, but a low signal may lead to concluding that the variant has a noise of less than 1:1000. Knowing that background cfDNA in non-cancer urine is likely to come from bladder, kidney, and / or prostate tissue, healthy tissue data may be used to further filter methylation variants to ensure the identification of informative methylation variants.
[0172] IV.B Primary tissue fraction estimation In some embodiments, the model 710 determines tumor fractions from multiple primary tissues. The model 710 may be a machine learning model. In some embodiments, the model 710 includes multiple methylation submodels and a maximum likelihood function that identifies tumor fraction predictions with maximum likelihood based on the methylation submodels. The methylation submodels may be trained based on reference samples (also referred to as training samples) in a bank. The methylation submodels may be associated with specific methylation variants. For example, a first methylation submodel may be trained to predict the likelihood that a first methylation variant originates from a specific tissue. The methylation submodels may further be tailored to subsets of primary tissues. For example, the methylation model may estimate the probability of a fragment originating from a specific cancer tissue with an alternative methylation pattern and also estimate the probability of a fragment having an alternative methylation pattern in non-cancerous and / or healthy tissues.
[0173] In one or more embodiments, the model 710 utilizes a binomial mixture model with a maximum likelihood function to predict tumor fraction from multiple primary tissues. The binomial mixture model assumes independence between methylation variant sites. In one or more embodiments, the binomial mixture model can be expressed as follows:
[0174]
number
[0175] In one or more embodiments, the model 710 may identify a tumor fraction that maximizes the likelihood of a count given a methylation sub-model for one or more subsets of k primary tissues. The number of primary tissues, k, may be any integer greater than 1 and less than or equal to the total number of primary tissues considered. For example, if the total number of primary tissues is M, k may be any one of 2, 3, 4, 5, ..., M-1, and M. The model 710 may calculate the likelihood of an observed count of a methylation variant at each site using the methylation sub-models generated based on the training samples. Types of methylation sub-models that may be implemented include, but are not limited to, Bernoulli, binomial, normal, Cauchy, t, Weibull, etc.
[0176] In one or more embodiments, the model 710 uses a Poisson distribution to calculate the likelihood of observing a methylation variant count at a given site. The Poisson distribution is one implementation of the methylation submodel and represents the probability that a given number of events will occur at a fixed interval when the events occur at a known constant average rate and are independent of each other. The Poisson distribution serves as an approximation to the binomial distribution when the number of events is large and the success probability of each event is small, but the product of two is not too large or too small. In the context of methylation variants, the likelihood of observing a methylation variant count at a given site can be modeled using a Poisson distribution parameterized by a weighted average of the variant recurrence rates across a subset of k primary tissues.
[0177]
number
[0178] A hypothetical example of average parameter calculation is as follows: At a site, the primary lung tissue sample has a methylation variant fraction of 0.4, the liver fraction of 0.7, and the healthy tissue fraction of 0.1. Assume estimated tumor fractions of 0.1 for lung, 0.2 for liver, and 0.7 for healthy tissue. The average parameters for the estimated tumor fractions and observed tissue fractions are:
[0179]
number
[0180] In one or more embodiments, the likelihood of observing a count of a methylation variant at site i further takes into account the likelihood that the methylation variant is activating or inactivating for a particular sample. A methylation variant may be "activating" for a particular tissue if the methylation variant is present in the cfDNA due to tumor shedding of the methylation variant. A methylation variant may be "inactivating" for a particular tissue if there is no tumor shedding that causes the absence of the methylation variant from that particular tissue in the cfDNA. One way to calculate the average parameter is to sum the various possibilities that the primary tissue in a subset of k primary tissues is activating or inactivating. For example, for lung and liver, there are four possibilities (or permutations): (1) lung is activating, liver is activating, (2) lung is activating, liver is inactivating, (3) lung is inactivating, liver is activating, and (4) lung is inactivating, liver is inactivating. The likelihood that methylation is activating or inactivating for a particular primary tissue at a particular site may be observed across training samples. One way to calculate the likelihood of an estimated tumor fraction is as follows.
[0181]
number
[0182] In one or more embodiments, the model 710 performs a base grid search across all vectors of tumor fractions to identify the estimated tumor fraction that maximizes the likelihood. The estimated tumor fraction that maximizes the likelihood may be returned by the model 710 as an estimated tumor fraction based on the count of methylation variants across sites. The model 710 identifies all combinations of primary tissues evaluated and performs a base grid search to identify the tumor fraction for each subset that maximizes the likelihood. The best tumor fractions for all subsets may then be evaluated to identify the best of the best tumor fractions that maximize the likelihood. In one or more embodiments, the subset size of k is 2. The model 710 identifies all combinations of subsets of two primary tissues. For example, if there are a total of 10 primary tissues screened, there are a total of 45 combinations of two primary tissues. The model 710 performs a first step of searching for the best tumor fraction in each subset (e.g., each of the 45 combinations). Model 710 then compares the best tumor fractions from each subset to identify the best of the best tumor fractions.
[0183] In other embodiments, the model 710 may perform a multi-step search. In a first step, the model 710 predicts the first primary tissue from which the sample is most likely to originate. This single primary tissue prediction may be generated by a multinomial cancer classifier or other machine learning algorithm capable of predicting primary tissue. In some embodiments, the single primary tissue prediction may be diagnosed by a physician or other healthcare provider. The model 710 may then select a combination of two primary tissues that includes the predicted primary tissue. The combinations that include the predicted primary tissue may be grid searched to estimate the tumor fraction between at least two primary tissues (including at least the primary tissue predicted by the classifier). Upon identifying the tumor fraction between the two primary tissues that maximizes the likelihood, the model 710 may further increase the subset of primary tissues that are evaluated to determine whether a tumor fraction across three primary tissues results in a greater likelihood than a tumor fraction across two primary tissues. The model 710 may refine the search in each step within some partition of the estimated tumor fraction of the previous step.
[0184] IV.C Machine Learning Implementation In some embodiments, the model 710 may additionally or alternatively be trained by one or more machine learning techniques. The bank 720 may accumulate a large number of biological samples and tumor fraction estimates for those samples. The biological samples in the bank 720 may serve as a set of training samples for iteratively training one or more machine learning models that can be used to estimate the tumor fraction of a new cfDNA sample of a subject. Additionally or alternatively, in some embodiments, one or more parameter values may be determined by one or more machine learning models. For example, a recurrence rate, an error rate, a pull-down bias, or a depth estimate may be determined using a machine learning model trained using data in the bank 720. For example, each biological sample may be associated with one or more labels, and the one or more labels may be estimates of the tumor fraction, the recurrence rate, the error rate, the pull-down bias, etc. The biological samples may be used as training samples for iteratively training a machine learning model in a supervised manner.
[0185] In various embodiments, a wide variety of machine learning techniques may be used. Examples include unsupervised learning, clustering, embedding, support vector regression (SVR) models, random forest classifiers, support vector machines (SVMs), such as kernel SVMs, gradient boosting, linear regression, logistic regression, and other forms of regression. Deep learning techniques such as neural networks, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory networks (LSTMs), may also be used. Each biological sample may be converted into a feature vector that includes different dimensions. As an example, methylation variants at different sites may be represented as values in different dimensions of the feature vector. The feature vectors of various training samples may be input to the machine learning model to iteratively train the machine learning model.
[0186] In various embodiments, the training technique for training the machine learning model can be supervised, semi-supervised, or unsupervised. In supervised training, the machine learning algorithm can be trained using a set of labeled training samples. For example, the biological samples in the bank 720 can be labeled with their estimated tumor fraction. The label for each training sample can be a binary value, a multi-class value, or a continuous variable. In some cases, unsupervised learning techniques can be used. The samples used for training are not labeled. Various unsupervised learning techniques, such as clustering, can be used. In some cases, the training can be semi-supervised, using a training set with a mixture of labeled and unlabeled samples.
[0187] The machine learning model is associated with an objective function that generates a metric value that describes the objective goal of the training process. For example, the training is intended to reduce the error rate of the model in determining tumor fraction estimates or parameter values (e.g., recurrence rate, pull-down bias, etc.) in one or more statistical models mentioned above. In such a case, the objective function may monitor the error rate of the machine learning model. For example, in one machine learning model, the objective function may be the training error rate in predicting tumor fraction in a training set. In another machine learning model, the objective function may be the training error rate in predicting the recurrence rate of methylation variants. Such objective functions may be referred to as loss functions. Other forms of objective functions may also be used, especially for unsupervised learning models where the error rate is not easily determined due to the lack of labels. In various embodiments, the error rate may be measured as cross-entropy loss, L1 loss (e.g., absolute distance between predicted and actual values), L2 loss (e.g., root mean square distance).
[0188] Machine learning models may take various suitable structures. For example, in a neural network, the neural network may receive inputs and generate outputs. The neural network may include different types of layers, such as convolutional layers, pooling layers, recurrent layers, fully connected layers, and custom layers. A convolutional layer convolves the input of the layer with one or more kernels to generate convolutional features. Each convolution result may be associated with an activation function. A convolutional layer may be followed by a pooling layer that selects the maximum (max pooling) or average value (average pooling) from the portion of the input covered by the kernel size. The pooling layer reduces the spatial size of the extracted features. In some embodiments, the pair of convolutional and pooling layers 540 may be followed by a recurrent layer that includes one or more feedback loops. The recurrent layer may be gated in the case of LSTM. The feedback may be used to explain the location relationship between methylation variants. A neural network may also include multiple fully connected layers with nodes connected to each other. The fully connected layer may be used for classification and regression. In some embodiments, one or more custom layers may be presented to generate a specific format of output. The order of layers and the number of layers in the neural network may vary. In some embodiments, the neural network includes one or more convolutional layers, but may or may not include any pooling or recurrent layers. If a pooling layer is present, not all convolutional layers need to be followed by a pooling layer. Recurrent layers may also be arranged differently in other positions of the neural network. For each convolutional layer, the size of the kernel (e.g., 3×3, 5×5, 7×7, etc.) and the number of kernels allowed to be trained may be different from other convolutional layers.
[0189] A machine learning model includes certain layers, nodes, kernels, and / or coefficients. Training a machine learning model includes forward and backward propagation iterations. Each layer in a neural network may include one or more nodes that may be fully or partially connected to other nodes in adjacent layers. In forward propagation, the neural network performs operations in a forward direction based on the output of the previous layer. The behavior of a node may be defined by one or more functions. The functions that define the behavior of a node may include various arithmetic operations such as convolution of data with one or more kernels, pooling, recurrent loops in RNNs, various gates in LSTMs, etc. The functions may also include activation functions that adjust the weights of the node's output. Nodes in different layers may be associated with different functions.
[0190] Each of the functions in the neural network may be associated with different coefficients (e.g., weights and kernel coefficients) that are adjustable during training. In addition, some nodes in the neural network may each be associated with an activation function that determines the weight of the node's output in forward propagation. Common activation functions may include step functions, linear functions, sigmoid functions, hyperbolic tangent functions (tanh), and rectified linear unit functions (ReLU). After an input is provided to the neural network and passed forward through the neural network, the results may be compared to training labels or other values in the training set to determine the performance of the neural network. The process of prediction may be repeated for other transactions in the training set to compute the value of the objective function in a particular training round. The neural network then performs backpropagation by using gradient descent, such as stochastic gradient descent (SGD), to adjust the coefficients in various functions to improve the value of the objective function.
[0191] Multiple iterations of forward and backward propagation may be performed. Training may be completed when the objective function is sufficiently stable (e.g., when the machine learning model has converged) or after a predetermined number of rounds for a particular set of training samples. In some embodiments, the trained machine learning model may be used as model 710 to predict tumor fraction using the target methylation results of cfDNA samples as input. In some embodiments, the trained machine learning model, used as a sub-model of del 710, estimates one or more parameters of the statistical distribution used in model 710.
[0192] IV. Exemplary Processes FIG. 10 is a flowchart depicting a process for generating a tumor fraction estimate from a subject's cfDNA sample, according to some embodiments. Process 1000 may be a computer-implemented process. For example, the flowchart may represent a portion of a software algorithm for generating a tumor fraction estimate, according to some embodiments. The software algorithm may be stored as computer instructions executable by one or more general-purpose processors (e.g., CPU, GPU). The instructions, when executed by the processor, cause the processor to perform various steps described in process 1000. In various embodiments, one or more steps in process 1000 may be skipped or modified. Process 1000 may be performed by analysis system 200. In general, a computer-implemented process may be performed using a computer.
[0193] In step 1010, a dataset of methylated sequence reads from a cfDNA sample of a subject is received. The dataset of methylated sequence reads may be generated using process 300 described in FIG. 3A and / or process 360 described in FIG. 3B. In some embodiments, the dataset included in estimating the tumor fraction may include methylated sequence reads determined from the cfDNA sample only. Thus, the process 1000 may be used to determine the amount of cancer signal in a cfDNA sample of a subject without a reference biopsy. For example, the process 1000 may be used as an early diagnosis technique where a cfDNA sample is obtained from a bodily fluid without an invasive biopsy. In some cases, the subject has not been diagnosed with cancer or the tissue origin of the cancer is still unknown. Thus, no biopsy sample is provided to determine the tumor fraction. Although the process 1000 may be performed using only a cfDNA sample, in some embodiments, the dataset may also include other sequence reads generated from other biological samples of the subject. For example, a matching biopsy sample may also be presented in the process 1000 to provide additional data points.
[0194] In some embodiments, the methylation sequence reads in the dataset can be generated based on targeted methylation (TM) assay. The TM assay can select methylation variant sites that are expected to have high recurrence rates. Using the TM assay, the sequencing depth and coverage of the methylation variant sites of interest can also be improved compared to WGBS. Thus, the noise estimate of the targeted methylation sites can be improved.
[0195] In step 1020, the dataset is divided into a number of variants. Each variant may be a methylation variant that includes one or more CpG sites. For example, at least one of the variants includes multiple adjacent CpG sites. For example, the sequence reads may be divided by the pattern of n (e.g., 5) adjacent CpGs and their states that distinguish DNA from one biological source (e.g., cancer) from another (e.g., non-cancer cfDNA). The dataset may be stored in a map of intervals where each key defines a different reference chromosome. The dataset may be referred to as a match tree, which may take the form of a serializable data structure that stores fragment counts for unique fragment-level methylation patterns. The underlying interval map keys may take the form of intervals defined by the fragment's first and last CpG indexes. The data stored in each interval map entry may be the number of fragment observations for each unique observed set of methylation states in the interval. In some embodiments, each variant has the same length of CpG sites. For example, each variant contains a length of 5 CpG sites. In other embodiments, the length of the CpG sites in the various variants may vary.
[0196] In step 1030, the methylation status of multiple variants is determined. In some embodiments, a particular methylation variant may include multiple CpG sites (e.g., adjacent CpG sites). The methylation variant may be encoded by a series of binary values. The series corresponds to the CpG sites. For each site, a first binary value (e.g., 1) represents that methylation is observed, and a second binary value (e.g., 0) represents that methylation is not observed. The methylation status of a methylation variant may be referred to as a methylation state vector, as illustrated in Figures 3A and 3B. For example, a methylation state vector of 11111 represents that all sites in the variant are methylated. Similarly, a methylation state vector of 00000 represents that all sites in the variant are unmethylated. The methylation state of a variant may also be any of 11111 to 00000. Determining the methylation status of a variant may be referred to as a variant calling process. Variants can be called from an input fragment file representing sequenced fragments from biological samples and a reference non-cancer match tree that is specifically constructed from WGBS non-cancer cfDNA. The variant calling process can also include fragment filtering. For example, fragments are filtered to remove duplicates, unconverted fragments, and uncalled fragments. Fragments that are too short, such as those with a minimum number of CpG sites that are less than the typical variant size of 5 CpG sites, can be filtered out. Fragments can also be filtered based on a minimum mapping quality threshold.
[0197] In step 1040, the plurality of variants are filtered based on a bank of reference sequence reads to generate a subset of filtered variants. The bank may be bank 720 as described in FIG. 7, and includes reads generated from non-cancer and biopsy samples of a plurality of tissues of a reference individual. The filtering to generate the subset of filtered variants may include filtering out one or more methylation variants whose prevalence in non-cancer samples exceeds a threshold. For example, if a methylation variant is also commonly present in the non-cancer samples in bank 720, the methylation variant may be associated with a high noise rate and may be filtered out as a result. Filtering of methylation variants may be based on other suitable criteria. In some embodiments, methylation variants are filtered to retain those that are not in the top 5% (or another suitable percentage threshold) of raw sample counts. In some embodiments, filtering may retain only methylation variants that are in autosomes. In some embodiments, only certain patterns of methylation status are retained after filtering. For example, in one embodiment, only methylation variants that are fully methylated or fully unmethylated are retained after filtering. In some embodiments, the selected methylation variants are those that have at least one matching fragment in the input TM reference data. In some embodiments, the selected methylation variants are those that have a TM estimation noise ratio below a certain threshold, such as 1 / 10,000. In some embodiments, the selected methylation variants are those that have a WGBS noise ratio below a certain threshold, such as 1 / 10,000.
[0198] In step 1050, the methylation state counts of the variants in the filtered subset are determined. The counts of a particular methylation state corresponding to a particular variant may represent the number of variant reads in the dataset that have a particular methylation state. For example, the counts at each site may be x i, which represents the count of the differential methylation variant at site i in the cfDNA sample of the subject. Each methylation variant in the cfDNA sample can be associated with a specific count. The counts for different methylation variants can depend on the sequencing depth. The counts can take the form of fractional counts of fragments containing the differential methylation variant, taking into account the total number of fragments sequenced. depth, which represents the depth estimate of site i in the cfDNA sample i Other parameters such as depth, sigma, and sigma may also be determined. Depth may also take the form of a fraction count, which represents fragments that overlap with a variant but do not necessarily contain the specific methylation state of the variant. After the counts of the various methylation variants in the filtered subsets have been determined, the counts may be stored as a vector representing the input of a model for determining a tumor fraction estimate.
[0199] In step 1060, the counts of methylation status are input to a model that is trained based on the recurrence rates of multiple variants in the reference sequence reads in the bank. For example, the model can be model 710 of FIG. 7. The model takes into account the recurrence rates of various methylation variants. A particular recurrence rate of a particular variant can correspond to the observation rate of a particular variant among the reference sequence reads in the bank. In some embodiments, the variant recurrence can be estimated as the number of tissue biopsy samples (e.g., samples with one or more supporting fragments) for a particular label (e.g., breast cancer) that have some evidence of the variant. The recurrence rate can be estimated as the number of samples in the bank 720 that contain one or more supporting fragments that include a methylation variant with a methylation status divided by the total number of samples.
[0200] The model used to determine the tumor fraction estimate may consider one or more recurrence rates for various methylation variants. In some embodiments, the model includes a probabilistic model. The probabilistic model includes a Poisson distribution for various methylation variants. A particular Poisson distribution for a particular methylation variant may be factored by the recurrence rate of the particular variant. In some embodiments, the model includes multiple probability distributions. Each probability distribution may correspond to a particular variant and may be parameterized based on the site-specific noise rate of the particular variant and the site depth of the particular variant. In some embodiments, the model includes a machine learning model. Details of various implementations of the model 710 are described in FIG. 7.
[0201] The model may be trained using data from the bank 720. Some reference individuals in the bank 720 are healthy, while other reference individuals in the bank are determined to have cancer or a type of cancer. The cancer or type of cancer may be breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis and ureter, non-urothelial renal cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal squamous cell carcinoma, non-squamous esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from hepatocytes, hepatobiliary carcinoma arising from cells other than hepatocytes, pancreatic cancer, human papillomavirus-associated head and neck cancer, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung cancer, lung cancer other than squamous cell carcinoma, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, or leukemia.
[0202] In step 1070, a tumor fraction estimate is generated for the cfDNA sample using the model. In some embodiments, the tumor fraction estimate may be an output of the model. In various embodiments, the output may take different forms. For example, in some embodiments, the tumor fraction estimate is a distribution of the probability of the fraction of fragments in the cfDNA sample being tumor-derived. In some embodiments, instead of a distribution, the tumor fraction estimate is a score that determines the fraction of fragments that are likely to be tumor-derived. In some embodiments, the tumor fraction estimate may be a binary decision (e.g., the subject is predicted to have cancer or not have cancer), or a multi-class decision (e.g., no cancer, likely to have cancer, likely to have cancer, further decisions are required, etc.). In some embodiments, the tumor fraction estimate may be an estimated fraction across multiple tissues of origin. The tumor fraction estimate may be used as a signal to estimate the amount of cancer in the subject's cfDNA sample in the absence of a reference biopsy sample. The tumor fraction estimate may also have other uses, such as cancer classification and determining the tissue of origin of the cancer.
[0203] V. Additional Improvements VA Example 1 - Enrichment of methylation variants using a targeted assay In some embodiments, methylation variants can be enriched using a targeted methylation assay. Figure 10 includes a plot of the number of methylation variants per million bases in a targeted methylation panel and a WGBS panel. The plot shows that the TM panel has a significantly higher number of methylation variants, which is an order of magnitude higher than the number of methylation variants in the WGBS panel. The results show that the TM assay can enrich all methylation variants, including recurrent methylation variants. This is not necessarily true for other variants, such as single base variants. In some embodiments, in step 910 of process 900, a TM assay is used to generate a dataset of methylation sequence reads from a cfDNA sample of a subject.
[0204] In some embodiments, tumor fraction estimates can also be used to generate allele fraction estimates. Allele fraction can refer to the proportion of molecules that are tumor-derived and contain variants. Figure 11 is a plot illustrating average methylation variant cfDNA allele fraction calibration according to some embodiments. The average regulated allele fraction of methylation variants on the y-axis can be determined based on allele fraction estimates from tumor fraction multiplied by the average biopsy allele fraction. The average allele fraction of both WGS (whole genome sequencing) and WGBS called variants on the x-axis can be determined based on allele fraction estimates from direct counting of fragments containing a reliable set of single nucleotide variants (SNVs) called in WGBS biopsies. Figure 11 shows that the average regulated allele fraction from methyl variants is well calibrated to the allele fraction from WGS called variants. Thus, allele fraction can be directly estimated from tumor fraction.
[0205] In various embodiments, one or more biological samples described herein, including those discussed in Figures 6 and 9, can be enriched using a cancer assay panel that includes multiple probes or multiple probe pairs. Several targeted cancer assay panels are known in the art, for example, as described in International Publication No. WO 2019 / 195268, filed April 2, 2019, International Application No. PCT / US2019 / 053509, filed September 27, 2019, and International Application No. PCT / US2020 / 015082, filed January 24, 2020, which are incorporated herein by reference. For example, in some embodiments, a cancer assay panel can be designed to include multiple probes (or probe pairs) that can capture fragments that together can provide information relevant to a cancer diagnosis. In some embodiments, the panel comprises at least 50, 100, 500, 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, 25,000, or 50,000 pairs of probes, in other embodiments, the panel comprises at least 500, 1,000, 2,000, 5,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, or 100,000 probes. The plurality of probes together can comprise at least 100,000, 200,000, 400,000, 600,000, 800,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, or 10,000,000 nucleotides. The probe (or probe pair) is specifically designed to target one or more genomic regions that are differentially methylated in cancer and non-cancer samples. The target genomic regions can be selected to maximize classification accuracy according to a size budget (determined by the sequencing budget and the desired depth of sequencing).
[0206] Samples enriched using cancer assay panels can be subjected to targeted sequencing. Samples enriched using cancer assay panels can generally be used to detect the presence or absence of cancer, and / or provide cancer classification, such as cancer type, stage of cancer, such as I, II, III, or IV, or provide the primary tissue from which cancer is believed to originate. Depending on the purpose, the panel can include probes (or probe pairs) that target differentially methylated genomic regions between general cancerous (pan-cancer) samples and non-cancerous samples, or only in cancerous samples with specific cancer types (e.g., lung cancer-specific targets). Specifically, cancer assay panels are designed based on bisulfite sequencing data generated from cell-free DNA (cfDNA) or genomic DNA (gDNA) from cancer and / or non-cancerous individuals.
[0207] In some embodiments, the cancer assay panel designed by the methods provided herein comprises at least 1,000 pairs of probes, each pair comprising two probes configured to overlap each other by an overlapping sequence comprising a 30 nucleotide fragment. The 30 nucleotide fragment comprises at least 5 CpG sites, and at least 80% of the at least 5 CpG sites are either CpG or UpG. The 30 nucleotide fragment is configured to bind to one or more genomic regions in a cancerous sample, and the one or more genomic regions have at least 5 methylation sites with an aberrant methylation pattern. Another cancer assay panel comprises at least 2,000 probes, each of which is designed as a hybridization probe complementary to one or more genomic regions. Each of the genomic regions is selected based on the criteria that it comprises (i) at least 30 nucleotides, and (ii) at least 5 methylation sites, and the at least 5 methylation sites have an aberrant methylation pattern and are either hypomethylated or hypermethylated.
[0208] Each of the probes (or probe pairs) is designed to target one or more target genomic regions. The target genomic regions are selected based on several criteria designed to increase selective enrichment of relevant cfDNA fragments while reducing noise and non-specific binding. For example, the panel can include probes that can selectively bind and enrich cfDNA fragments that are differentially methylated in cancerous samples. In this case, sequencing of the enriched fragments can provide information relevant to the diagnosis of cancer. Furthermore, the probes can be designed to target genomic regions determined to have aberrant methylation patterns and / or hypermethylation or hypomethylation patterns to provide further selectivity and specificity of detection. For example, genomic regions can be selected when they have a methylation pattern with a low p-value according to a Markov model trained on a set of non-cancerous samples, which further covers at least 5 CpGs, 90% of which are either methylated or unmethylated. In other embodiments, genomic regions can be selected utilizing a mixture model as described herein.
[0209] Each of the probes (or probe pairs) can target a genomic region that includes at least 25bp, 30bp, 35bp, 40bp, 45bp, 50bp, 60bp, 70bp, 80bp, or 90bp. The genomic region can be selected by containing less than 20, 15, 10, 8, or 6 methylation sites. The genomic region can be selected when at least 80, 85, 90, 92, 95, or 98% of at least 5 methylation (e.g., CpG) sites are either methylated or unmethylated in the non-cancerous or cancerous sample.
[0210] Genomic regions can be further filtered to select only those that are potentially informative based on their methylation patterns, e.g., CpG sites that are differentially methylated between cancerous and non-cancerous samples (e.g., aberrantly methylated or unmethylated in cancer vs. non-cancer). In selecting, a calculation can be performed for each CpG site. In some embodiments, a first count is determined, which is the number of cancer-containing samples that contain fragments that overlap with the CpG (cancer count), and a second count is determined, which is the number of all samples that contain fragments that overlap with the CpG (total). Genomic regions can be selected based on a criterion that is positively correlated with the number of cancer-containing samples that contain fragments that overlap with the CpG (cancer count) and inversely correlated with the number of all samples that contain fragments that overlap with the CpG (total).
[0211] In one embodiment, the number of non-cancerous samples that have fragments overlapping the CpG site (n 非癌 ) and the number of cancer samples (n 癌 ) is counted. The probability that the sample is cancer is then calculated, for example, as (n 癌 +1) / (n 癌 +n 非癌 +2). We rank CpG sites by this metric and greedily add them to the panel until the panel size budget is exhausted.
[0212] Depending on whether the assay is intended to be a pan-cancer or single-cancer assay, or depending on what kind of flexibility is desired in choosing which CpG sites contribute to the panel, which samples are used for cancer counting can vary. A similar process can be used to design a panel for diagnosing a specific cancer type (e.g., TOO). In this embodiment, for each cancer type and each CpG site, the information gain is calculated to determine whether to include a probe targeting that CpG site. The information gain is calculated for a sample with a given cancer type compared to all other samples. For example, two random variables "AF" and "CT". "AF" is a binary variable indicating whether there is an abnormal fragment (yes or no) that overlaps with a specific CpG site in a particular sample. "CT" is a binary random variable indicating whether the cancer is of a particular type (e.g., lung cancer or non-lung cancer). Given "AF", mutual information can be calculated with respect to "CT". That is, how many information bits are gained for a cancer type (lung vs. non-lung in this example) if we know whether there are any anomalous fragments that overlap a particular CpG site? This can be used to rank CpGs based on how specific they are for a particular cancer type (e.g., TOO). This procedure is repeated for multiple cancer types. For example, if a particular region is typically differentially methylated only in lung cancers (and not so methylated in other cancer types or non-cancers), then CpGs in that region should tend to have high information gain for lung cancer. For each cancer type, CpG sites are ranked by this information gain metric and then greedily added to the panel until the size budget for that cancer type is exhausted.
[0213] Further filtering can be performed to select target genome regions that have less than a threshold number of off-target genome regions. For example, genome regions are selected only when there are less than 15, 10, or 8 off-target genome regions. In other cases, filtering is performed to remove genome regions when the sequence of the target genome region appears more than 5, 10, 15, 20, 25, or 30 times in the genome. Further filtering can be performed to select target genome regions when the sequence that is 90%, 95%, 98%, or 99% homologous to the target genome region appears less than 15, 10, or 8 times in the genome, or to remove target genome regions when the sequence that is 90%, 95%, 98%, or 99% homologous to the target genome region appears more than 5, 10, 15, 20, 25, or 30 times in the genome. This is to exclude repeat probes that may pull down off-target fragments, which are undesirable and may affect assay efficiency.
[0214] In some embodiments, it has been shown that at least 45 bp of fragment probe overlap is required to achieve a non-negligible amount of pull-down (although this number may vary depending on the assay details). Furthermore, it has been suggested that a mismatch rate of more than 10% between the probe and fragment sequence in the overlap region is sufficient to significantly disrupt binding and therefore pull-down efficiency. Thus, sequences that can be aligned to the probe along at least 45 bp with at least 90% match rate are candidates for off-target pull-down. Thus, in one embodiment, the number of such regions is scored. The best probes have a score of 1, meaning that they match only in one place (the intended target region). Probes with low scores (e.g., less than 5 or 10) are accepted, but any probes above this score are discarded. Other cut-off values can be used for specific samples.
[0215] In various embodiments, the selected target genomic region can be located at various locations in the genome, including, but not limited to, exons, introns, intergenic regions, and other portions. In some embodiments, probes can be added that target non-human genomic regions, such as those that target viral genomic regions.
[0216] VB Example 2 - Analysis of tumor fraction in urinary cfDNA Urological cancers, such as prostate, bladder and kidney cancers, have low detection sensitivity in the Circulating Cell-free Genome Atlas Study (CCGA, Clinical Trial.gov identifier NCT02889978), a prospective, multicenter, case-control observational study with long-term follow-up. This low detection rate may be due to the low tumor fraction of cfDNA in the blood of subjects with urothelial carcinoma. Analysis of cell-free DNA from urine may increase the detection sensitivity of urological cancers.
[0217] In an exemplary workflow, approximately 50 mL of urine is collected from a subject. A preservative is then added to the collected urine sample. Streck Urine Preserve (Streck, Nebraska, USA) or an equivalent preservative containing at least 0.5% weight / volume (w / v) nuclease inhibitor, at least 0.2-4.0% w / v preservative, and at least 0.01% w / v formaldehyde quencher can be used as the preservative. Other preservatives that can be used include Urine Collection Medium and UAS (Novosantis, Belgium). Alternatively, urine samples can be collected using a Urine Collection and Preservation Device (Norgen Biotek Corp., Canada) or an equivalent cup containing 10-30% w / v nuclease inhibitor, e.g., EDTA, and 0.1-1.0% w / v bacteriostatic preservative, e.g., sodium azide.
[0218] After adding the preservative, the urine sample is centrifuged at 4000×g for 20 minutes to sediment and remove any cellular debris. The resulting supernatant is then concentrated approximately 15-fold. Concentration of the stored urine sample can be achieved using a diafiltration column, for example, a filter unit with a regenerated cellulose membrane with a 3 kDa cutoff. For example, if the resulting supernatant has a volume of 15 mL or less, the sample can be concentrated by spinning at 4000×g for at least 30 minutes in an Amicon Ultra-15 filter unit (Thermo Fisher Scientific, Massachusetts, US). When working with volumes of supernatant ranging from 15 mL to 50 mL, the sample can instead be concentrated by spinning at 2500×g in a Centricon Plus-70 centrifugal filter unit (Thermo Fisher Scientific, Massachusetts, US) for at least 40 minutes or until the sample is concentrated to a volume of less than 4.2 mL. Concentrating the urine sample to a low volume results in a sample suitable for automated bead-based extraction methods (eg, MagMax extraction) and other benchtop available techniques.
[0219] After concentration of the urine sample, the sample can be used immediately or frozen at -80°C for later use or for batch processing of samples. The concentrated urine sample can then be subjected to cfDNA extraction and library preparation, which can then be sequenced for methylation analysis and urological cancer detection.
[0220] Adding preservative within 30 minutes of collecting urine samples preserved nucleosomes and produced cfDNA fragments up to 7000bp in size. In contrast, delaying the addition of preservative for one hour or more after collection resulted in the loss of the nucleosome peak at >700bp, a significant reduction in yield, and a narrower distribution of fragment lengths biased toward low molecular weight fragments, indicating lysis of cells in urine along with degradation and fragmentation of cfDNA. It was also found that samples frozen for later processing by concentrating urine exhibited a reduced cryoprecipitate formation.
[0221] Cancer-specific methylation signatures, such as those corresponding to samples from subjects with stage I, high-grade non-muscle invasive bladder cancer, were detected in urinary cfDNA processed from urine samples.
[0222] To evaluate how detection of cancer-specific methylation status from urine-derived and plasma-derived cfDNA differs, urine samples were collected and processed as described in Example 1. Further studies were performed to evaluate detection of cancer-specific methylation markers in urine-derived cfDNA compared to plasma tumor fraction estimates. Urine and blood were collected from patients with bladder, renal, and prostate cancer, as well as from age- and sex-matched non-cancer patients. To generate biopsy-free estimates in urine cfDNA, the plasma-based workflow was modified with the following steps (illustrated in Figure 7): (1) for non-cancer WGBS and TM data in the workflow, an external reference dataset of non-cancer urinary cfDNA (N = approx. 200) was used instead of plasma; (2) the noise threshold and false counts were adjusted to account for the smaller reference dataset; and (3) WGBS data from healthy urinary system tissues were used to further filter noisy methylation variants.
[0223] Sequencing libraries were then prepared from the resulting cfDNA for methylation analysis of a panel of urological cancer methylation markers. Tumor fraction estimation was also performed using the methylation marker panel, and samples with estimated tumor fractions above a threshold were identified as cancer detected based on urine-derived cfDNA. The estimated tumor fractions in urine-derived cfDNA samples were then compared with the estimated tumor fractions from the corresponding plasma-derived cfDNA fractions.
[0224] Scatter plots in Figures 13-15 show the tumor fraction distribution in each cancer type in matched urine and plasma cfDNA (each point is a single patient). Fills represent whether the multi-cancer classifier detected plasma cfDNA (at 99% specificity) for that patient. An increase in the tumor fraction in urine cfDNA compared to plasma was observed in all bladder cancer patients analyzed and in a subset of prostate cancer patients. Plasma tumor fraction estimates were consistent with classifier detection. In kidney cancer patients, the signal in urine was not increased above that in plasma, but a higher tumor fraction in urine cfDNA was observed compared to non-cancer patients. In a minority of patients, an increase in the signal in plasma relative to urine cfDNA was observed.
[0225] The estimated tumor fraction for subjects with any stage of bladder cancer was found to be higher when using urine-derived cfDNA compared to plasma-derived cfDNA (Figure 13). When evaluating samples from subjects with early stage non-muscle invasive bladder cancer, estimates based on urine tumor fraction stratified some bladder cancer samples that were not detected in plasma, and the majority of those samples that were also detected in plasma.
[0226] When analyzing a cohort of subjects with stage II or IV prostate cancer, estimated tumor fraction was also higher in urine-derived cfDNA than in plasma-derived cfDNA (Figure 14).In particular, 5 out of 10 samples analyzed had urine-derived cfDNA estimated tumor fraction above the minimum threshold, while only 1 out of 10 plasma-derived cfDNA estimated tumor fraction was above the threshold.The estimated tumor fraction in samples from subjects with renal cancer was found to be comparable between urine-derived cfDNA and plasma-derived cfDNA (Figure 15).
[0227] In conclusion, the determination of urological cancer status based on the analysis of methylation markers and tumor fractions was comparable or more accurate when analyzing urinary cfDNA than plasma-derived cfDNA. Concentration of urine samples improved the processing and analysis of urinary cfDNA for urological cancer detection.
[0228] VI. Cancer application examples In some embodiments, the disclosed methods, analysis systems and / or tumor fraction estimation models can be used to detect the presence (or absence) of cancer, monitor the progression or recurrence of cancer, monitor therapeutic response or effectiveness, determine the presence or minimal residual disease (MRD), or any combination thereof. In some embodiments, the analysis systems and / or classifiers can be used to identify the tissue or primary source of the cancer. For example, the systems and / or classifiers can be used to identify cancer, such as any of the following cancer types: head and neck cancer, liver / bile duct cancer, upper gastrointestinal cancer, pancreatic / gallbladder cancer, colorectal cancer, ovarian cancer, lung cancer, multiple myeloma, lymphoma, melanoma, sarcoma, breast cancer, and uterine cancer. For example, as described herein, the classifiers can be used to generate a likelihood or probability score (e.g., 0-100) that a sample feature vector is from a subject with cancer. In some embodiments, the probability score is compared to a threshold probability to determine whether the subject has cancer. In other embodiments, the likelihood or probability score can be determined at different time points (e.g., before or after treatment) to monitor disease progression or to monitor treatment effectiveness (e.g., therapeutic effectiveness). In still other embodiments, the likelihood or probability score can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, determining treatment effectiveness, etc.). For example, in one embodiment, if the likelihood or probability score exceeds a threshold, a physician can prescribe an appropriate treatment. In some embodiments, a test report can be generated to provide the patient with those test results, including, for example, a probability score that the patient has a disease state (e.g., cancer), a type of disease (e.g., a type of cancer), and / or a disease tissue of origin (e.g., a cancer tissue of origin).
[0229] VI.A. Early Detection of Cancer In some embodiments, the methods, tumor fraction estimation models, and / or classifiers of the present disclosure are used to detect the presence or absence of cancer in a subject suspected of having cancer. For example, a classifier (described herein) can be used to determine the likelihood or probability score that a sample feature vector is from a subject with cancer.
[0230] In one embodiment, a probability score of 60 or more can indicate that the subject has cancer. In yet other embodiments, a probability score of 65 or more, 70 or more, 75 or more, 80 or more, 85 or more, 90 or more, or 95 or more indicated that the subject has cancer. In other embodiments, the probability score can indicate the severity of the disease. For example, a probability score of 80 can indicate a more severe form or a later stage of cancer compared to a score below 80 (e.g., a score of 70). Similarly, an increase in the probability score over time (e.g., at a second later time point) can indicate disease progression, or a decrease in the probability score over time (e.g., at a second later time point) can indicate successful treatment.
[0231] In another embodiment, the cancer log odds ratio can be calculated by taking the logarithm of the ratio of the probability of being cancerous to the probability of being non-cancerous (i.e., 1 minus the probability of being cancerous) for a test subject, as described herein. According to this embodiment, a cancer log odds ratio greater than 1 can indicate that the subject has cancer. In yet other embodiments, a cancer log odds ratio greater than 1.2, greater than 1.3, greater than 1.4, greater than 1.5, greater than 1.7, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4 indicated that the subject has cancer. In other embodiments, the cancer log odds ratio can indicate the severity of the disease. For example, a cancer log odds ratio greater than 2 can indicate a more severe form or later stage of cancer compared to a score below 2 (e.g., a score of 1). Similarly, an increase in the log odds ratio of cancer over time (e.g., at a second later time point) can indicate disease progression, or a decrease in the log odds ratio of cancer over time (e.g., at a second later time point) can indicate successful treatment.
[0232] According to aspects of the present disclosure, the methods and systems of the present disclosure can be trained to detect or classify multiple cancer indicators. For example, the methods, systems, and classifiers of the present disclosure can be used to detect the presence of one or more, two or more, three or more, five or more, or ten or more different types of cancer.
[0233] In some embodiments, the cancer is one or more of head and neck cancer, liver / bile duct cancer, upper gastrointestinal cancer, pancreatic / gallbladder cancer, colorectal cancer, ovarian cancer, lung cancer, multiple myeloma, lymphoma, melanoma, sarcoma, breast cancer, and uterine cancer.
[0234] The method and system may further quantify the cancer signal present in the subject, for example, by determining a tumor fraction prediction based on the count of methylation variants. The analysis system may chart patients diagnosed with cancer with tumor fraction predictions over the course of the cancer. The analysis system may utilize this cancer cohort to generate a predictive model as a post-diagnosis tool. The quantified cancer signal may be used to predict the subject's prognosis, staging of the cancer, whether the cancer is benign, malignant, aggressive, or metastatic, prediction of the likelihood of recurrence after treatment, likelihood of treatment success, etc. For example, the predictive model may predict the subject's prognosis based on the tumor fraction prediction. The prognosis may indicate life expectancy, likelihood of remission (partial or complete), likelihood of recurrence, likelihood of metastasis, etc. Based on the prognosis, the analysis system may provide a list of recommended treatments. For example, a poor prognosis may recommend clinical trials, chemotherapy, surgery, etc., among other treatments. A good prognosis may recommend less intensive treatment options.
[0235] VI.B. Cancer and Treatment Monitoring In certain embodiments, the first time point is before the cancer treatment (e.g., before the resection surgery or therapeutic intervention) and the second time point is after the cancer treatment (e.g., after the resection surgery or therapeutic intervention), and the method is utilized to monitor the effectiveness of the treatment. For example, if the second likelihood or probability score decreases compared to the first likelihood or probability score, the treatment is deemed successful. However, if the second likelihood or probability score increases compared to the first likelihood or probability score (or the likelihood or probability score remains substantially the same), the treatment is deemed unsuccessful. In response to an ineffective treatment option, the analysis system may recommend an alternative treatment option that excludes the current ineffective treatment. In other embodiments, both the first and second time points are before the cancer treatment (e.g., before the resection surgery or therapeutic intervention). In yet other embodiments, both the first and second time points are after the cancer treatment (e.g., before the resection surgery or therapeutic intervention), and the method is used to monitor the effectiveness of the treatment or loss of effectiveness of the treatment. In yet other embodiments, cfDNA samples may be obtained from a cancer patient at a first and second time point and analyzed, e.g., to monitor cancer progression, to determine whether the cancer is in remission (e.g., following treatment), to monitor or detect residual disease or disease recurrence, or to monitor treatment (e.g., therapeutic) effectiveness.
[0236] One of skill in the art will readily appreciate that test samples can be obtained from a cancer patient over any desired set of time points and analyzed according to the disclosed methods to monitor the cancer status in the patient. In some embodiments, the first and second time points are, such as about 30 minutes, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or about 24 hours, about 1, 2, 3, 4, 5, 10, 15, 20, 25, or about 30 days, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months, or about 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 8, 9, 10, 11, or 12 months. , 20, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5, or about 30 years. In other embodiments, test samples may be obtained from a patient at least once every three months, at least once every six months, at least once every year, at least once every two years, at least once every three years, at least once every four years, or at least once every five years.
[0237] In some cancer patients, tumors can be detected early enough to be cured or controlled with long-term stable maintenance therapy (e.g., hormone replacement or ablation therapy for estrogen- and progesterone-induced breast cancer, or androgen- or testosterone-induced prostate cancer). Patients with previous cancers can have regular surveillance to catch recurrence or increase in minimal residual disease. Patients with previous cancers may be at above average risk of developing a second independent primary cancer and therefore may be advised to have regular screening and early detection for all other cancers. Surveillance for previous or actively treated cancers may be independent of screening or early detection of other cancers, which increases effort, cost, and reduces patient compliance with this long-term high-load testing regimen.
[0238] Using samples (e.g., non-cancerous samples, biopsy samples, or plasma from blood draws) and assays (e.g., targeted methylation), cancer surveillance or minimal residual disease detection can be performed along with cancer screening and early detection of second primary cancers. Surveillance can be performed more frequently (e.g., every 6 months) than cancer screening (e.g., annually). Regular, more frequent test intervals can be used to determine changes in the signal over time. Such longitudinal data can inform cancer prognosis and therapy decisions when either recurrence of a previously treated cancer or the development of a new cancer is detected.
[0239] Early detection of multiple cancers can be performed for possible secondary primary cancers, regardless of minimal molecular residual disease based on tumor information or not based on tumor information, or recurrence detection for previously known and previously treated cancer types. Minimal molecular residual disease or recurrence detection based on tumor information can use genetic or epigenetic variants determined from tissue samples, for example, biopsies of cancerous tissue. If cancer variants from tissue samples are known, the same plasma sample can be used with two different analysis and classification methods. One method can be a cancer screening method that uses a trained classifier that detects patterns indicating the presence of cancer and possible cancer types and cancer variants without prior knowledge. Another method can be a method for tissue-based residual disease or recurrence detection that focuses on variants that are known to occur in the cancer that the patient has been diagnosed with and is undergoing treatment or surveillance.
[0240] Testing for MRD or recurrence of previously treated cancers can estimate the presence and fraction of tumor-derived fragments in a sample. Measurements can be supported by information from tumor-derived genomic or methylation variants observed on tissue samples of that tumor. Recurrence or MRD can be reported if the estimated tumor fraction exceeds a pre-specified threshold that is established to reach the target specificity of the analysis in the presence of biological and technical noise. As a result, the tumor fraction estimate may be above zero even for negative results. However, if recurrence or MRD is detected in a subsequent blood draw, the previous tumor fraction estimate (even if below the detection threshold) can be used to evaluate the gradient of tumor fraction increase over time, aiding in patient prognosis and informing treatment decisions for recurrent cancers.
[0241] Similarly, a cancer early detection test can estimate the probability that any cancer is present. Cancer detection can be reported if this probability exceeds a pre-specified threshold established to reach the target specificity of the analysis. For a cancer early detection test, the probability of this cancer can also be above zero but below the detection threshold. If cancer is detected in a subsequent blood draw, the increase in the probability of the same type of cancer over time can be determined, again informing patient prognosis and treatment decisions. In this case, the time derivative value of the cancer signal intensity (e.g., how much the signal changes over time) can be input instead of the absolute value for cancer detection and / or patient management decisions. In one or more embodiments, the cancer signal over time can include an estimated tumor fraction of one or more primary tissues (e.g., as described above with reference to FIG. 7).
[0242] In one or more embodiments of MRD detection, plasma samples are collected for screening and early detection to provide a first reference for later treatment monitoring and MRD detection. Cancer diagnosis generally involves performing a biopsy of the cancerous tissue, and in many cases, treatment involves surgical removal of the tumor. Residual disease can be detected from plasma samples collected after treatment (e.g., after surgical resection), for example, several days after the start of therapy. Plasma samples can be shipped for processing and analysis immediately after blood collection, independent of tissue collection and processing, for example, according to method 100 of FIG. 1. Analysis of plasma samples can begin independent of tissue processing status. Non-tumor-informed analysis can determine whether a cancer signal has been detected and can provide a tumor burden measurement, such as VAF, and the confidence of this value (e.g., 95% confidence interval). For example, if a pre-treatment plasma sample from a screening or early detection test is available, tumor-derived fragments from this sample can be used for a first attempt with tumor-informed analysis.
[0243] The analysis system 200 may then evaluate whether to proceed with the analysis based on the detected cancer signal. A non-tumor-informed or tumor-informed analysis based on a previous plasma sample may or may not detect a cancer signal. If a cancer signal is detected, the use of training data to determine a measurement such as VAF may result in only a few informative ctDNA fragments in the plasma sample, and therefore low confidence in quantification. If the plasma sample contained a sufficient number of informative fragments that allow for providing clinically sufficient information on the VAF, the analysis system may return the VAF with high confidence as the test result, and avoid sequencing the tissue sample, biopsy slide review, biopsy slide markup, biopsy slide dissection, or any combination thereof. However, if a cancer signal is not detected and / or the confidence in the VAF estimate is low, the analysis system may continue with tissue processing and sequencing, i.e., the analysis system continues with the tumor-informed analysis. The analysis system 200 may return a VAF or tumor burden based on the tumor-informed analysis. Analysis system 200 may also return whether no cancer signal was detected in the non-tumor-informed analysis, or whether the VAF estimate was low confidence in the non-tumor-informed analysis.
[0244] Detection of inflammatory processes may depend on minimal residual disease detection from targeted methylation assays. The described method may allow for the detection and quantification of the presence of tumor, abnormal, abnormal, or disease-specific methylation patterns on a fragment-by-fragment basis. These cfDNA fragments with distinct methylation patterns (e.g., using one or more of the variant models) may then be used to quantify the presence of target types of tumors in blood collected from cancer patients.
[0245] In exemplary implementations, the presence of cfDNA fragments from apoptotic non-cancerous cells can be detected using the same assays and methodologies described throughout that are used to detect and quantify the presence of cancer-derived cfDNA fragments. Aberrant or cell type-specific methylation patterns from the following cells can be used to detect severe side effects from immune checkpoint inhibition, in addition to different tumor types: colonic epithelial cells that contribute to enteritis, gastric epithelial cells that contribute to enteritis, pancreatic ductal cells that contribute to pancreatitis, pancreatic acinar cells that contribute to pancreatitis, pancreatic endocrine cells that contribute to pancreatitis, hepatic cells that contribute to hepatitis, and bile duct cells that contribute to hepatitis.
[0246] Immune checkpoint inhibitors are used as treatments for, for example, lung cancer, bladder cancer, and melanoma, but severe side effects due to uncontrolled inflammation can occur in the colon, pancreas, or liver. As a result, apoptotic DNA of cells released by tumors and cells released due to inflammation originate from different cell origins and can be detected and quantified independently from the same blood draw, assay, and bioinformatics pipeline. The analysis system may identify such a tendency to uncontrolled inflammation in a patient based on the quantified presence of apoptotic non-cancerous cells. The analysis system and / or the physician (or other relevant health care provider) may recommend avoiding immune checkpoint inhibitors to minimize uncontrolled inflammation.
[0247] VI.C. Treatment In yet another embodiment, information obtained from any of the methods described herein (e.g., likelihood or probability scores) can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, determining treatment efficacy, etc.). For example, in one embodiment, if the likelihood or probability score exceeds a threshold, a physician can prescribe an appropriate treatment (e.g., resective surgery, radiation therapy, chemotherapy, and / or immunotherapy). In some embodiments, information such as the likelihood or probability score can be provided as a readout to a physician or to the subject.
[0248] A classifier or model (described herein) can be used to determine the likelihood or probability score that the sample feature vector is from a subject with cancer. In one embodiment, when the likelihood or probability exceeds a threshold, an appropriate treatment (e.g., resection surgery or therapeutic) is prescribed. For example, in one embodiment, if the likelihood or probability score is 60 or greater, one or more appropriate treatments are prescribed. In another embodiment, if the likelihood or probability score is 65 or greater, 70 or greater, 75 or greater, 80 or greater, 85 or greater, 90 or greater, or 95 or greater, one or more appropriate treatments are prescribed. In other embodiments, the cancer log odds ratio can indicate the effectiveness of the cancer treatment. For example, an increase in the cancer log odds ratio over time (e.g., in the second after treatment) can indicate that the treatment was not effective. Similarly, a decrease in the cancer log odds ratio over time (e.g., in the second after treatment) can indicate successful treatment. In another embodiment, if the cancer log odds ratio is greater than 1, greater than 1.5, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4, one or more appropriate treatments are prescribed.
[0249] In some embodiments, the treatment is one or more cancer therapy agents selected from the group including chemotherapeutic agents, targeted cancer therapy agents, differentiation therapy agents, hormonal therapy agents, and immunotherapy agents. For example, the treatment can be one or more chemotherapeutic agents selected from the group including alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeletal destaxans, topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based drugs, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapy agents selected from the group including signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteasome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapy agents including retinoids, e.g., tretinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormonal therapy agents selected from the group including antiestrogens, aromatase inhibitors, progestins, estrogens, antiandrogens, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapy agents selected from the group including monoclonal antibody therapy, such as rituximab (RITUXAN) and alemtuzumab (CAMPATH), non-specific immunotherapy and adjuvants, such as BCG, interleukin-2 (IL-2), and interferon-α, immunomodulatory agents, such as thalidomide and lenalidomide (REVLIMID). It is within the ability of a skilled physician or oncologist to select the appropriate cancer therapy agent based on characteristics such as tumor type, cancer stage, previous exposure to cancer treatments or therapeutic agents, and other characteristics of the cancer.
[0250] VII. Kit implementation Also disclosed herein are kits for carrying out the above methods, including those related to cancer classifiers. The kits may include one or more collection containers for collecting samples from individuals containing genetic material. The samples may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. Such kits may include reagents for isolating nucleic acids from the samples. The reagents may further include reagents for sequencing the nucleic acids, including buffers and detection agents. In one or more embodiments, the kits may include one or more sequencing panels that include probes for targeting specific genomic regions, specific mutations, specific genetic variants, or any combination thereof. In other embodiments, the samples collected via the kits are provided to a sequencing laboratory that may use the sequencing panels to sequence the nucleic acids in the samples.
[0251] The kit may further include instructions for the use of the reagents included in the kit. For example, the kit may include instructions for collecting a sample and extracting nucleic acid from a test sample. Exemplary instructions may be the order in which the reagents are added, the centrifugation speed used to isolate nucleic acid from the test sample, how to amplify nucleic acid, how to sequence nucleic acid, or any combination thereof. The instructions may also illuminate how to operate a computing device as the analysis system 200 to perform any of the steps of the described methods.
[0252] In addition to the above components, the kit may include a computer readable storage medium that stores computer software for carrying out the various methods described throughout this disclosure. One form in which these instructions may be present is as information printed on a suitable medium or substrate, such as a paper strip or strips on which the information is printed, in the kit packaging, in a package insert. Yet another means is a computer readable medium on which the instructions are stored in the form of computer code, such as a diskette, CD, hard drive, network data storage device. Yet another means may be a website address or QR code that can be used via the Internet to access information at a remote site.
[0253] VIII. Additional Considerations The foregoing description of embodiments of the present disclosure has been presented for purposes of illustration. It is not intended to be exhaustive or to limit the invention to the precise form disclosed. Those skilled in the art will recognize that numerous modifications and variations are possible in light of the above disclosure.
[0254] In some portions of this specification, embodiments of the present disclosure are described in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. While these operations are described functionally, computationally, or logically, it will be understood that they may be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Further, without loss of generality, it has proven convenient at times to refer to these configurations of operations as modules. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.
[0255] Any of the steps, operations, or processes described herein may be performed or implemented using one or more hardware or software modules, alone or in combination with other devices. In some embodiments, the software modules are implemented using a computer program product that includes a computer-readable non-transitory medium containing computer program code, which may be executed by a computer processor to perform any or all of the steps, operations, or processes described.
[0256] Embodiments may also relate to products produced by the computing processes described herein. Such products may include information resulting from the computing processes, the information being stored on a non-transitory tangible computer-readable storage medium and may include any embodiment of the computer program product or other data combination described herein.
[0257] Finally, the language used herein has been selected primarily for ease of reading and instructional purposes, and not to define or limit the subject matter of the present invention. Accordingly, the scope of the present invention is not intended to be limited by this detailed description, but rather by any claims issued on an application based thereon. Accordingly, the disclosure of the embodiments herein is intended to illustrate, rather than limit, the scope of the present invention, which is set forth in the following claims.
Claims
1. 1. A computer-implemented method for generating a tumor fraction estimate from a cell-free deoxyribonucleic acid (cfDNA) sample of a subject, the computer-implemented method comprising: receiving a dataset of methylated sequence reads from the cfDNA sample of the subject; partitioning the dataset into a plurality of variants, each variant comprising a methylation pattern across one or more CpG sites; filtering the plurality of variants based on a bank of reference sequence reads to generate a subset of filtered variants, the bank including reads generated from non-cancer cfDNA samples and biopsy samples of a plurality of tissues of a reference individual; for each variant in the filtered subset, determining a count of methylated sequence reads that contain the variant; inputting the counts of methylated sequence reads for the variants of the filtered subset into a model trained based on recurrence rates of the plurality of variants; and using the model to generate a tumor fraction prediction for the cfDNA sample.
2. 2. The computer-implemented method of claim 1, wherein the recurrence rates of the plurality of variants are determined based on the reference sequence reads in the bank.
3. 2. The computer-implemented method of claim 1, wherein filtering the plurality of variants based on a reference sequence read to generate the filtered subset of variants comprises filtering out one or more variants whose abundance in the non-cancer sample exceeds a threshold.
4. 2. The computer-implemented method of claim 1, wherein a particular recurrence rate of a particular variant corresponds to an observation rate of the particular variant among the reference sequence reads in the bank.
5. 2. The computer-implemented method of claim 1, wherein the tumor fraction prediction is a distribution of probabilities of the fraction of fragments in the cfDNA sample that are tumor-derived.
6. 2. The computer-implemented method of claim 1, wherein the tumor fraction prediction is the fraction of fragments in the cfDNA sample that are tumor-derived.
7. 2. The computer-implemented method of claim 1, wherein the model comprises at least one probability model, the probability model comprising a Poisson distribution for a particular variant, the Poisson distribution weighted by the recurrence rate of the particular variant.
8. 2. The computer-implemented method of claim 1, wherein the model comprises a plurality of probability distributions, each probability distribution corresponding to a particular variant and parameterized based on a site-specific noise rate for the particular variant and a sequencing depth per site for the particular variant.
9. 10. The computer-implemented method of claim 8, wherein each probability distribution corresponding to a particular variant is further parameterized based on at least one of the depth of the cfDNA sample, the targeted panel pull-down efficiency of the cfDNA sample, and the estimated tumor fraction of the cfDNA sample.
10. 2. The computer-implemented method of claim 1, wherein the count for each variant of the filtered subset comprises a count of methylated sequence reads of the cfDNA sample that include the methylation pattern across the one or more CpG sites of the variant.
11. 2. The computer-implemented method of claim 1, wherein a particular variant comprising multiple adjacent CpG sites is encoded by a series of binary values, the series corresponding to the adjacent CpG sites, a first binary value at a particular CpG site representing observed methylation and a second binary value at the particular CpG site representing observed unmethylation.
12. The computer-implemented method of claim 1 , wherein the tumor fraction prediction includes multiple fractions for subsets of tissues.
13. 13. The computer-implemented method of claim 12, wherein each fraction represents a percentage of fragments of the cfDNA sample derived from each tissue of the subset of tissues.
14. The computer-implemented method of claim 12 , wherein the model is a binomial mixed model that assumes independence between the variants in the filtered set.
15. 15. The computer-implemented method of claim 14, wherein the model comprises a plurality of methylation sub-models, each methylation sub-model associated with a variant in the filtered set and parameterized by the recurrence rate and estimated tumor fraction of the variant across the subset of tissues, and each methylation sub-model configured to calculate the likelihood of observing the methylated sequence read count based on the methylated sequence read count.
16. 16. The computer-implemented method of claim 15, wherein the model comprises a weighted sum of methylation submodels enumerating all possibilities for tissues in the subset to be activated or inactivated for the variant, each methylation submodel representing the likelihood of the possibility and comprising a weight parameterized by the recurrence rate of the activated tissue.
17. 16. The computer-implemented method of claim 15, wherein each methylation sub-model is Poisson distributed.
18. 16. The computer-implemented method of claim 15, wherein the model generates the tumor fraction prediction by identifying the estimated tumor fraction having the greatest likelihood as calculated by the methylation submodel.
19. 20. The computer-implemented method of claim 18, wherein the maximum likelihood is determined via a grid search.
20. The computer-implemented method of claim 15 , wherein the subset of organizations is selected for a plurality of organizations.
21. 21. The computer-implemented method of claim 20, wherein the model generates the tumor fraction prediction by identifying an estimated tumor fraction having a maximum likelihood as calculated by the methylation sub-model across two or more subsets of tissues, each subset of tissue being a different combination of tissues selected from the plurality of tissues from other subsets of tissues.
22. 22. The computer-implemented method of claim 21, wherein the tissue subset is of one size.
23. 23. The computer-implemented method of claim 22, wherein the size of the tissue subset is selected from two, three, or four.
24. 21. The computer-implemented method of claim 20, wherein the model includes a cancer classifier that generates predictions for one or more tissues from which the fragments originate, and the subset of tissues includes the one or more tissues predicted by the cancer classifier.
25. The computer-implemented method of claim 1 , wherein the methylated sequence reads in the dataset are generated by a targeted methylation assay.
26. 2. The computer-implemented method of claim 1, wherein the plurality of tissues of the reference individual used to generate the reference sequence reads are selected from the group comprising breast tissue, thyroid tissue, lung tissue, bladder tissue, cervical tissue, small intestine tissue, colorectal tissue, esophageal tissue, stomach tissue, tonsil tissue, liver tissue, ovarian tissue, fallopian tube tissue, pancreatic tissue, prostate tissue, kidney tissue, and uterine tissue.
27. The computer-implemented method of claim 1 , wherein the tumor fraction estimate is generated without a matched biopsy sample.
28. 10. The computer-implemented method of claim 1, wherein the cfDNA sample is obtained from a bodily fluid without an invasive biopsy.
29. 2. The computer-implemented method of claim 1, wherein one or more reference individuals in the bank are determined to have cancer or a type of cancer that is breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis and ureter, non-urothelial renal cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal squamous cell carcinoma, non-squamous esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from liver cells, hepatobiliary carcinoma arising from cells other than liver cells, pancreatic cancer, human papillomavirus-associated head and neck cancer, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung cancer, lung cancer other than squamous cell carcinoma, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, or leukemia.
30. The computer-implemented method of claim 1 , wherein the model is a machine learning model.
31. 31. The computer-implemented method of claim 30, wherein the machine learning model is one or more of a constant model, a binomial model, an independent sites model, a neural network model, or a Markov model.
32. The machine learning model: For each reference sample in the bank, including the non-cancer cfDNA sample and the biopsy sample, for each variant of the filtered variants, identifying a count of reads containing the variant; For each variant of the filtered variants, determining a non-cancer recurrence rate based on the read count for the variant in the non-cancer sample; For each variant of the filtered variants, determining a cancer recurrence rate based on the count of reads of the variant in the biopsy sample; and training the model using the non-cancer recurrence rate and the cancer recurrence rate, wherein the model is configured to predict a tumor fraction prediction based on read counts of the filtered variants in a given sample.
33. the cfDNA sample Cancer surveillance for previously diagnosed cancers, and The computer-implemented method of claim 1 , wherein the method is used to perform one or more of early cancer screening for a plurality of cancer types.
34. wherein the cfDNA sample of the subject is a liquid biopsy collected after initiation of treatment for cancer, and the computer-implemented method comprises: determining a confidence score for the tumor fraction prediction based on the count of the methylated sequence reads that comprise the filtered subset of variants; in response to determining that the confidence score is below a confidence threshold, sequencing a tissue sample collected after the initiation of the treatment for the cancer; receiving a second dataset of methylation sequence reads from the tissue sample of the subject; dividing the second data set into a second plurality of variants; filtering the second plurality of variants based on the bank of reference sequence reads to generate a second subset of filtered variants; for each variant in the filtered subset, determining a second count of methylated sequence reads that contain the variant; inputting the second counts of methylated sequence reads for the variants of the second filtered subset into the model; and 10. The computer-implemented method of claim 1, further comprising: using the model to generate a second tumor fraction prediction for the tissue sample.
35. 35. The computer-implemented method of claim 34, further comprising returning the tumor fraction prediction with the second tumor fraction prediction and the confidence score.
36. wherein the cfDNA sample of the subject is a liquid biopsy sample collected after initiation of treatment for cancer, and the computer-implemented method comprises: determining that the tumor fraction prediction of the liquid biopsy sample is below a threshold signal; In response to determining that the tumor fraction prediction of the liquid biopsy sample is below a threshold signal, sequencing a tissue sample collected after the initiation of the treatment for the cancer; receiving a second dataset of methylation sequence reads from the tissue sample of the subject; dividing the second data set into a second plurality of variants; filtering the second plurality of variants based on the bank of reference sequence reads to generate a second subset of filtered variants; for each variant in the filtered subset, determining a second count of methylated sequence reads that contain the variant; inputting the second counts of methylated sequence reads for the variants of the second filtered subset into the model; and 10. The computer-implemented method of claim 1, further comprising: using the model to generate a second tumor fraction prediction for the tissue sample.
37. 37. The computer-implemented method of claim 36, further comprising returning the second tumor fraction prediction and the tumor fraction prediction below the threshold signal.
38. determining a first fraction of the first tissue that is apoptotic non-cancerous tissue; determining whether the first fraction of the first tissue is above a threshold; 10. The computer-implemented method of claim 1, further comprising: in response to determining that the first fraction of the first tissue is above the threshold, returning an indication that the test subject has a higher propensity for uncontrolled inflammation as a side effect to an immune checkpoint inhibitor.
39. 10. The computer-implemented method of claim 1, wherein the cfDNA sample is derived from urine of the subject.
40. 40. The computer-implemented method of claim 39, wherein the tumor fraction prediction is for bladder cancer, prostate cancer, kidney cancer, or any combination thereof.
41. 10. The computer-implemented method of claim 1, further comprising determining a poor prognosis for the subject in response to determining that the tumor fraction prediction for the cfDNA sample is above a threshold.
42. 42. The computer-implemented method of claim 41, wherein the poor prognosis indicates that the subject has progressive disease.
43. 42. The computer-implemented method of claim 41, wherein the poor prognosis indicates the subject is at higher risk of disease recurrence.
44. 42. The computer-implemented method of claim 41, further comprising providing a list of recommended treatments based on the poor prognosis for the subject.
45. wherein the cfDNA sample is collected from the subject after initiation of treatment, and the method comprises:
10. The computer-implemented method of claim 1, further comprising determining the presence of residual disease in the subject in response to determining that the tumor fraction prediction for the cfDNA sample is above a threshold.
46. 46. The computer-implemented method of claim 45, further comprising providing a list of recommended treatments based on the presence of residual disease in the subject.
47. wherein the cfDNA sample is collected from the subject after initiation of treatment for a disease in the subject, and the method comprises: The computer-implemented method of claim 1 , further comprising evaluating the treatment based on the tumor fraction prediction.
48. evaluating the treatment, 48. The computer-implemented method of claim 47, comprising determining that the treatment is efficacious in response to determining that the tumor fraction prediction for the cfDNA sample collected after initiation of the treatment is less than an initial tumor fraction prediction for an initial cfDNA sample collected before initiation of the treatment.
49. evaluating the treatment, 48. The computer-implemented method of claim 47, comprising determining that the treatment is ineffective in response to determining that the tumor fraction prediction for the cfDNA sample collected after initiation of the treatment is substantially equal to or greater than an initial tumor fraction prediction for an initial cfDNA sample collected before initiation of the treatment.
50. 50. The computer-implemented method of claim 49, further comprising, in response to determining that the treatment is ineffective, providing a list of alternative treatments excluding the treatment.
51. 51. A non-transitory computer readable medium configured to store computer code including instructions for generating a tumor fraction estimate from a cell-free deoxyribonucleic acid (cfDNA) sample of a subject, the instructions, when executed by one or more processors, causing the one or more processors to perform the method of any one of claims 1-50.
52. 1. A system comprising: one or more processors; 51. A system comprising: a memory configured to store computer code including instructions for generating a tumor fraction estimate from a cell-free deoxyribonucleic acid (cfDNA) sample of a subject, the instructions, when executed by the one or more processors, causing the one or more processors to perform the method of any one of claims 1-50.