Methylation-based age prediction as a feature for cancer classification

JP2025529015A5Pending Publication Date: 2026-08-05GRAIL INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GRAIL INC
Filing Date
2023-07-28
Publication Date
2026-08-05

AI Technical Summary

Technical Problem

Existing cancer detection methods struggle to accurately distinguish between cancer-related methylation patterns and other biological or clinical variables, such as age or gender, due to the complexity and volume of data generated by next-generation sequencing (NGS), leading to errors in cancer identification.

Method used

A method involving the use of machine learning models trained on methylation patterns of nucleic acid fragments to predict chronological age and identify cancer, utilizing a feature set of genomic regions with strong indicative scores, and applying linear regression and p-value filtering to enhance accuracy.

Benefits of technology

This approach enables early and accurate cancer detection by reducing computational burden and improving the specificity of cancer identification, allowing for faster and more efficient processing of vast genetic data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method and system for covariate prediction from methylation features is disclosed. The system identifies feature sets for genomic regions by training one or more regressions to evaluate covariate scores for the genomic regions. The system may select the feature set with the highest predictive score and may consider other selection criteria. The system trains an age prediction model using training samples with reported chronological age labels. The system further uses the chronological age prediction to predict the likelihood of cancer in a test sample. To do this, the system may compare predicted covariate values ​​and / or labels with reported values ​​and / or labels. In one embodiment, the system may use an age residual threshold to identify whether there is a strong likelihood of the presence of cancer. In another embodiment, the system may use the predicted chronological age as a feature for a cancer classifier.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 392,980, filed July 28, 2022, which is incorporated by reference in its entirety. [Background technology]

[0002] Deoxyribonucleic acid (DNA) methylation plays an important role in the regulation of gene expression. Aberrant DNA methylation is observed in many disease processes, including cancer. DNA methylation profiling using methylation sequencing (e.g., whole genome bisulfite sequencing (WGBS) or targeted methylation sequencing) is increasingly recognized as a valuable diagnostic tool for cancer detection, diagnosis, and / or monitoring. For example, specific patterns of differentially methylated regions and / or allele-specific methylation patterns can be useful molecular markers for minimally invasive diagnosis using circulating cell-free (cf) DNA. As part of cancer classification, there remains a need to understand the impact that covariate variables (or, more generally, variables indicative of cancer or non-cancer) may have on the human genome. Furthermore, there remains a need to be able to distinguish between variables that may be indicative of cancer and / or any other biological attributes, such as age or gender. Summary of the Invention [Problem to be solved by the invention]

[0003] The present disclosure is directed to addressing the above-mentioned problems. The background discussion provided herein is intended to provide a general overview of the contents of the present disclosure. Unless otherwise indicated herein, the statements in this section are not prior art to the claims of this application, and no admission is made that they are prior art or suggest prior art by their inclusion in this section. [Means for solving the problem]

[0004] In some embodiments, the technology described herein relates to a method for detecting the presence or absence of cancer in a test sample. The method includes obtaining a plurality of training samples, each training sample comprising a plurality of nucleic acid fragments, each of the plurality of nucleic acid fragments having a genomic location overlapping with at least one genomic region of a plurality of genomic regions and labeled with the chronological age of the individual from whom the training sample was collected. The method includes sequencing the plurality of nucleic acid fragments of each training sample and identifying a methylation pattern for each nucleic acid fragment. For each genomic region of the plurality of genomic regions, the method identifies a nucleic acid fragment from the plurality of nucleic acid fragments having a genomic location overlapping with the genomic region, and calculates, for that genomic region, an indicatory score representing a correlation between chronological age and methylation pattern, the indicatory score being calculated based on the chronological age of the individual from whom the identified nucleic acid fragment was collected and the methylation pattern of the identified nucleic acid fragment. The method includes generating a feature set comprising one or more genomic regions of the plurality of genomic regions, wherein one or more genomic regions in the feature set have an indicatory score higher than a threshold. The method includes training a machine learning age prediction model to identify a predicted chronological age of a subject from whom a test sample was taken, the training being based on methylation patterns of nucleic acid fragments of a plurality of training samples that overlap with one or more genomic regions in a feature set.

[0005] In some embodiments, the method further includes training a linear regression for each genomic region of the feature set based on methylation patterns of nucleic acid fragments that overlap with each genomic region from a plurality of training samples labeled as non-cancer. The method includes obtaining a plurality of additional training samples, each additional training sample including a plurality of additional nucleic acid fragments having additional genomic locations that overlap with at least one genomic region of the plurality of genomic regions, labeled with the chronological age of the individual from whom the additional training sample was taken, and labeled as non-cancer or cancer based on a prior identification of the presence of cancer in the additional training sample. The method includes sequencing the plurality of additional nucleic acid fragments to identify a methylation pattern for each additional nucleic acid fragment. For each genomic region of the plurality of genomic regions, the method applies linear regression to the methylation patterns of nucleic acid fragments of a plurality of additional training samples to determine a predicted chronological age of the individual from whom the additional training sample was taken, calculates an age residual for each additional training sample as the difference between its predicted chronological age and its labeled chronological age, and compares the age residuals of the additional training samples labeled as cancer with the age residuals of the additional training samples labeled as non-cancer. The method includes generating a reduced feature set from the feature set based on the comparison of the age residuals, the reduced feature set including a smaller number of genomic regions than the feature set, and the reduced feature set is used to train a machine learning age prediction model.

[0006] In some embodiments, the method further comprises obtaining a test sample, the test sample comprising a plurality of additional nucleic acid fragments, labeled with the chronological age of the subject from whom the test sample was collected. The method comprises sequencing the plurality of additional nucleic acid fragments of the test sample to identify methylation patterns of the plurality of additional nucleic acid fragments. The method comprises applying the trained age prediction model to determine a predicted chronological age of the subject from whom the test sample was collected based on the methylation patterns of the additional nucleic acid fragments that overlap with one or more genomic regions of the feature set, and calculating an age residual as the difference between the labeled chronological age and the predicted chronological age of the subject. The method comprises identifying the test sample as having a strong likelihood of the presence of cancer in response to identifying the age residual as being above a residual threshold.

[0007] In some embodiments, the method includes applying the trained age prediction model to a second plurality of training samples identified as non-cancer to determine a predicted age for each of the second plurality of training samples, and calculating an age residual for each of the second plurality of training samples by comparing the predicted age to the labeled chronological age of the second plurality of training samples, wherein the determining includes identifying a residual threshold based on the calculated age residuals of the second plurality of training samples, wherein at least a majority of the calculated age residuals of the second plurality of training samples satisfy the residual threshold.

[0008] In some embodiments, the method further includes, in response to identifying the test sample as having a strong likelihood of the presence of cancer, filtering the methylation patterns of the plurality of additional nucleic acid fragments using p-value filtering to identify a set of aberrant methylation patterns; generating a feature vector for the test sample based on the age residuals and the set of aberrant methylation patterns; and identifying a cancer prediction for the test sample by inputting the feature vector into a trained cancer classifier.

[0009] In some embodiments, the cancer prediction used in the method is a binary prediction of the presence or absence of cancer or other disease state.

[0010] In some embodiments, the cancer prediction used in the method is a multi-class prediction across multiple cancer types.

[0011] In some embodiments, the cancer prediction used in the method is a multi-class prediction between multiple disease states.

[0012] In some embodiments, the method further includes identifying the presence of cancer in the test sample using a second machine learning cancer classifier configured to receive as input the predicted chronological age of the subject and the methylation patterns of the plurality of additional nucleic acid fragments, and to output a prediction of the presence of cancer in the test sample.

[0013] In some embodiments, the second machine learning cancer classifier is further configured to receive as input the clinical information and genetic background of the subject and to output a prediction of the presence of cancer in the test sample.

[0014] In some embodiments, the indicator score is a Pearson correlation or a covariate score.

[0015] In some embodiments, the indicator score is determined by training a linear regression to predict chronological age from methylation density of non-cancer training samples, where methylation density is calculated as the percentage of nucleic acid fragments having genomic locations that overlap with a particular genomic region that have a methylation state at that particular genomic region.

[0016] In some embodiments, the machine learning age prediction model includes a multivariate regression. The multivariate regression can be penalized based on the number of one or more genomic regions in the feature set. The machine learning age prediction model can receive as input the methylation density corresponding to each of the genomic regions in the feature set.

[0017] In some embodiments, the number of one or more genomic regions in the feature set is selected from the range of 5 to 10,000. Further, sequencing the nucleic acid fragments comprises whole genome bisulfite sequencing (WGBS), and / or sequencing the nucleic acid fragments comprises targeted sequencing.

[0018] In some embodiments, each training sample of the plurality of training samples is pre-identified as not having a cancer presence, or each training sample of the plurality of training samples is pre-identified as having a cancer presence. Further, each training sample of the plurality of training samples is pre-identified as having a cancer presence or an absence of cancer. When each training sample of the plurality of training samples is labeled as having a cancer presence or not having a cancer presence, the label is based on the pre-identification of the cancer status for the training sample.

[0019] In some aspects, the technology described herein relates to a method for training a classifier, the method comprising: obtaining a plurality of training samples, each training sample comprising a plurality of nucleic acid fragments, each of the plurality of nucleic acid fragments having a genomic location overlapping with at least one genomic region of a plurality of genomic regions, and labeling the training sample with a characteristic of the individual from whom the training sample was taken; determining the sequence of the plurality of nucleic acid fragments of each training sample to identify a methylation pattern for each nucleic acid fragment; for each genomic region of the plurality of genomic regions, the method identifies from the plurality of nucleic acid fragments a nucleic acid fragment having a genomic location overlapping with the genomic region; and, for each genomic region, calculating an indicatory score representing a correlation between the characteristic and the methylation pattern, the indicatory score being calculated based on the characteristic of the individual from whom the identified nucleic acid fragment was taken and the methylation pattern of the identified nucleic acid fragment; the method comprises generating a feature set comprising one or more genomic regions of the plurality of genomic regions, wherein one or more genomic regions in the feature set have an indicatory score higher than a threshold. The method includes training a machine learning trait prediction model to identify a predictive trait for a subject from whom a test sample was taken, the training being based on methylation patterns of nucleic acid fragments in a plurality of training samples that overlap with one or more genomic regions in the feature set.

[0020] In some embodiments, the characteristic is the individual's biological sex, where the characteristic is biological male or biological female. Alternatively or additionally, the characteristic is the individual's smoking status, where the characteristic is either smoking or non-smoking.

[0021] In some embodiments, the machine learning trait prediction model comprises a logistic regression analysis implementing a sigmoid function.

[0022] In some embodiments, the method further comprises obtaining a test sample, the test sample comprising a plurality of additional nucleic acid fragments, labeled with a label indicating a characteristic of the test sample. The method further comprises sequencing the additional plurality of nucleic acid fragments of the test sample to identify a test methylation pattern for each of the additional nucleic acid fragments, and applying the trained machine learning characteristic prediction model to predict a characteristic of the test sample based on the methylation patterns of the additional nucleic acid fragments that overlap with the feature set of the genomic region. The method further comprises flagging the test sample as contaminated and excluding the test sample from further analysis if the predicted label differs from the label of the test sample.

[0023] In some embodiments, if the predicted signature matches the labeled signature, the method further includes filtering the methylation patterns of additional nucleic acid fragments of the test sample by p-value filtering to identify a set of aberrant methylation patterns; generating a feature vector for the test sample based on the set of aberrant methylation patterns; and identifying a cancer prediction for the test sample by inputting the feature vector into a trained cancer classifier.

[0024] In some embodiments, the cancer prediction is a binary prediction of the presence or absence of cancer or other disease state. The cancer prediction can be a multi-class prediction between multiple cancer types or multiple disease states.

[0025] In another aspect, a system includes a hardware processor and a non-transitory computer-readable storage medium storing instructions that, when executed by the hardware processor, cause the hardware processor to perform the methods disclosed herein, as well as a non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the processors to perform the methods disclosed herein. [Brief explanation of the drawings]

[0026] [Figure 1] 1 is an exemplary flowchart illustrating the overall workflow of cancer classification of a sample, according to one or more embodiments. [Figure 2A] 1 is an exemplary flowchart illustrating the process of sequencing fragments of cell-free (cf) DNA to obtain a methylation status vector, according to one or more embodiments. [Figure 2B] FIG. 2B is an illustrative diagram of the sequencing process of FIG. 2A of fragments of cell-free (cf) DNA to obtain a methylation state vector, according to one or more embodiments. [Figure 3A] 1 shows methylation features that can be derived from a single CpG site as a genomic region, according to one or more embodiments. [Figure 3B] 1 shows methylation features that can be derived from multiple CpG sites as a genomic region, according to one or more embodiments. [Figure 4A] 1 is an exemplary flowchart illustrating a process for training a chronological age prediction model, according to one or more embodiments. [Figure 4B] 1 illustrates the development of a chronological age prediction model, according to one or more embodiments. [Figure 5A] 1 is an exemplary flowchart illustrating a process for generating a control group data structure for identifying aberrantly methylated fragments, according to one or more embodiments. [Figure 5B] 1 is an exemplary flowchart illustrating a process for identifying fragments that will be aberrantly methylated based on a symmetry group data structure, according to one or more embodiments. [Figure 6A] 1 is an exemplary flowchart illustrating a process for training a cancer classifier, according to one or more embodiments. [Figure 6B] 1 illustrates an exemplary generation of feature vectors used to train a cancer classifier, according to one or more embodiments. [Figure 7A]1 shows an exemplary flow chart of an apparatus for sequencing a nucleic acid sample, according to one or more embodiments. [Figure 7B] FIG. 1 is an exemplary block diagram of an analysis system according to one or more embodiments. [Figure 8] 1 illustrates genomic regions associated with age, according to one or more embodiments. [Figure 9A] 1 illustrates one process for identifying a feature set of genomic regions that exhibit covariance with age, according to one or more embodiments. [Figure 9B] 1 shows a graph of age prediction results for a non-cancer holdout cohort, according to an example. [Figure 9C] 1 shows a graph of age prediction results for a cancer cohort, according to an example. [Figure 10A] 10 illustrates another process for identifying a feature set of genomic regions indicative of age covariates, according to one or more embodiments. [Figure 10B] 1 shows a graph of age prediction results for a non-cancer holdout cohort, according to an example. [Figure 10C] 1 shows a graph of age prediction results for a cancer cohort, according to an example. [Figure 11] 1 shows the spread of the test cohort across cancer stages, according to an example. [Figure 12A] 10 shows a top series of graphs representing test samples predicted to result in a negative outcome by a cancer classifier, according to some embodiments. [Figure 12B] 10 shows a sub-series of graphs representing test samples predicted to result positive by a cancer classifier, according to some embodiments. [Figure 13] 1 shows one genomic region showing age deceleration across cancer types, according to an example. [Figure 14A] By way of example, we present results for the first genomic region where there appears to be consistent age acceleration in hematological cancers, with less consistent age acceleration seen in other non-hematological cancers. [Figure 14B]By way of example, we present results for a second genomic region where there appears to be consistent age deceleration in hematological cancers, with little or no significant age deceleration in other non-hematological cancers. [Figure 15A] 1 illustrates identification of a feature set of genomic regions for predicting biological sex, according to one or more embodiments. [Figure 15B] 1 shows results from a trained biological sex prediction model, according to an example. [Figure 16A] 1 illustrates identification of a feature set of genomic regions for predicting smoking status, according to one or more embodiments. [Figure 16B] 1 shows results from a trained smoking status gender prediction model, according to an example. DETAILED DESCRIPTION OF THE INVENTION

[0027] The drawings depict various embodiments for purposes of illustration only, and those skilled in the art will readily appreciate from the following description that alternative embodiments of the structures and methods illustrated herein may be utilized without departing from the principles described herein.

[0028] I. Overview Early detection and classification of cancer is an important technology. Being able to detect cancer before symptoms appear benefits all parties involved, including patients, physicians, and loved ones. For patients, early detection of cancer increases the patient's chances of a favorable outcome; for physicians, early detection of cancer increases the number of treatment avenues that can lead to favorable outcomes; and for loved ones, early detection of cancer increases the chance of not losing a friend or family member to the disease.

[0029] In recent years, early cancer detection techniques have evolved to analyze genetic fragments (e.g., DNA) from a patient's blood or other sources to identify whether any of these genetic fragments originate from cancer cells. These new technologies allow physicians to identify the presence of cancer in patients who might otherwise have gone undetected. For example, consider a patient at high risk for breast cancer. Traditionally, this person would regularly visit the doctor for mammograms, which generate images of their breast tissue (e.g., X-rays) that physicians use to identify cancerous tissue. Unfortunately, even with the highest resolution mammograms, physicians cannot identify tumors until they are approximately one millimeter in size. This means that the cancer may have been present in the person's body for some time, remaining undiagnosed and untreated. This visual identification is typical of most cancers—that is, they cannot be identified until they are large enough to be identified with some type of imaging technology.

[0030] Cancer detection using analysis of gene fragments, such as those in a patient's blood, alleviates this problem. For example, cancer cells begin shedding DNA fragments into a person's bloodstream as soon as they form. This occurs when cancer cells are very scarce and before they become visible using imaging techniques. Thus, with the right methodology, a system that analyzes DNA fragments in the bloodstream can identify the presence of cancer in a person's body based on the shed cancer DNA fragments, and more importantly, the system can detect cancer before it becomes identifiable using more traditional cancer detection techniques.

[0031] Cancer detection based on the analysis of DNA fragments is made possible by next-generation sequencing ("NGS") technology. Broadly speaking, NGS is a collection of technologies that enable high-throughput sequencing of genetic material. As described in more detail herein, NGS broadly consists of (1) sample preparation, (2) DNA sequencing, and (3) data analysis. Sample preparation refers to the laboratory methods required to prepare DNA fragments for sequencing, sequencing is the process of reading the nucleotides arranged in a specific order in the sample, and data analysis involves processing and analyzing the genetic information in the sequence data to identify the presence of cancer.

[0032] While these steps of NGS can help enable early cancer detection, they also pose their own complex and detrimental challenges to cancer detection; therefore, any improvement to sample preparation, DNA sequencing, and / or data analysis, including pre-processing, algorithmic processing, and summarizing or presenting predictions or conclusions, will improve cancer detection technology and, more generally, early cancer detection.

[0033] To illustrate, by way of example, (1) sample preparation problems include DNA sample quality, sample contamination, fragmentation bias, and accurate indexing. Correcting these problems will yield better genetic data for cancer detection. Similarly, (2) sequencing problems include, for example, errors in accurate fragment transcription (e.g., reading "A" instead of "C"), inaccurate or difficult fragment synthesis and duplication, different coverage uniformity, sequencing depth vs. cost vs. specificity, and insufficient sequencing length. Again, correcting any of these problems will yield improved genetic data for cancer detection.

[0034] (3) The problem of data analysis is the most daunting and challenging. The challenges stem from the sheer volume of data generated by NGS sequencing technology. Genetic datasets are typically on the order of terabytes, and effectively and efficiently analyzing such volumes of data is procedurally and computationally demanding. For example, analysis of NGS sequencing involves several fundamental processing steps, such as co-alignment of reads, alignment and mapping of reads to a reference genome, identification and calling of mutated genes, identification and calling of aberrantly methylated genes, and generation of functional annotations. Performing any of these processes on terabytes of genetic data is computationally expensive even for the most powerful computer architectures and completely infeasible for the average human brain. Additionally, the error-prone process of sample preparation and sequencing results in genetic sequencing data that is largely of low quality and unusable for cancer identification. For example, large amounts of genetic data may contain contaminated samples, transcription errors, mismatched regions, duplicated regions, etc., making them unsuitable for accurate cancer detection. Identifying and accounting for low-quality genetic data from the vast amounts of genetic data obtained from NGS sequencing is also procedurally and computationally demanding to accomplish and is not practically feasible for the human brain. Overall, any process that leads to more efficient processing of vast amounts of sequence data could improve cancer detection using NGS sequencing.

[0035] Finally, and perhaps most importantly, accurately identifying abnormal DNA from NGS data to identify the presence of cancer is also challenging. To be effective, algorithms are needed to compensate for errors arising from, for example, sample preparation and sequencing, and to overcome the challenges of analyzing the vast amounts of data associated with NGS technology. That is, the design of one or more machine learning models or other computational algorithms that enable early cancer detection based on next-generation sequencing technologies must be configured to address the challenges arising from these technologies. Some of these techniques and models are described below, followed by specific improvements over the state of the art.

[0036] One of the challenges in creating and properly applying a cancer detection model, as mentioned above, is the sheer volume of sequence data to which the model may be applied. Within that sequence data, there are numerous “cancer signals” and “other signals.” Cancer signals refer to sequence data that indicate the presence or absence of cancer in a subject's body, while other signals refer to sequence data that indicate additional biological or clinical characteristics of the subject (e.g., age, sex, smoking, etc.). Further complicating this issue, some other signals may appear to share similar characteristics with the cancer signals, and some cancer signals may appear to share similar characteristics with other signals. Furthermore, some of the sequence data may be both cancer signals and other signals. For example, age-related methylation patterns may appear to be similar to methylation patterns for aging, or, in a similar context, a particular methylation pattern may indicate both cancer and aging.

[0037] This relationship between signals in sequencing data is problematic because it can lead to errors in cancer identification. For example, an improperly trained cancer classifier may identify a person as having a signal of interest based on analysis of their cfDNA, because some methylation in the cfDNA may appear to indicate "older chronological age and cancer" when in fact the subject is simply older chronologically. Therefore, care should be taken when selecting sequencing data for training a cancer classifier when the sequencing data may indicate several closely related biological processes or clinical characteristics. Any method to improve cancer identification in such situations will improve the field of early cancer detection.

[0038] Additionally, carefully collecting a set of data that accurately indicates cancer and / or another biological and / or clinical characteristic reduces the complexity and corresponding computational cost of running a machine learning model to indicate or predict the presence of a cancer call in a sample (i.e., a cancer "call"). Indeed, carefully selecting genomic sites that are "strongly" associated with cancer and / or biological or clinical characteristics to be processed by the model reduces the computational burden of the model, making it faster and more efficient, which improves the field of cancer identification. For example, consider the exemplary machine learning model trained to identify cancer based on methylation described above. However, in this case, each indicative feature is also associated with the "strength" of the cancer indication for that site or the "strength" of the chronological age indication for that site. That is, aberrant methylation at a first site may strongly indicate cancer, aberrant methylation at a second site may strongly or weakly indicate chronological age, etc. In this case, having the machine learning model process genomic sites that are only "weakly" indicative of cancer would incur computational costs without providing a corresponding benefit to accuracy or specificity in cancer identification.

[0039] To provide a quantitative example, applying a machine learning model to 100 weakly instructive sites may not yield as much benefit as applying the model to one strongly instructive site. Therefore, by selecting appropriate sites for a feature set for processing by a machine learning model to identify the presence and / or biological or clinical characteristics of cancer, processing costs can be significantly reduced without significantly reducing the accuracy or specificity of the model. Indeed, the analytical load can be reduced, for example, from millions of genomic sites to tens of thousands of genomic sites without sacrificing model accuracy. More succinctly, reducing the analytical load from millions of genomic sites to tens of thousands of genomic sites reduces processing load and time by several orders of magnitude. This reduction enables faster cancer detection and, more importantly, earlier cancer detection than was possible with conventional methods, freeing up computational resources for other models and classifications (e.g., processing additional samples), improving the performance of computers implementing the models, reducing the monetary cost of such systems, and improving fields such as public health, medicine, diagnosis, and treatment.

[0040] Another challenge with the vast amount of data generated by NGS sequencing is appropriately training a machine learning model to identify cancer from the vast amount of data. For example, a machine learning model can be trained to identify cancer by comparing a feature vector with genomic data. A "feature" in this feature vector can be any genomic site with sufficient depth of aberrantly methylated genomic locations that correspond to the presence of cancer, as described below. By constructing a genome-wide feature vector, this can typically lead to tens of thousands of features, and as described above, some of these features may be more indicative of the presence of cancer than others. In this regard, selecting which features and corresponding genomic data to use to train the machine learning model can be challenging. The machine learning model should be trained and configured to accurately identify the presence of cancer, yet the resulting model should not be computationally expensive. In other words, appropriate selection of data and features for training the machine learning model improves early cancer detection.

[0041] Overview of the IA Cancer Classification Workflow FIG. 1 is an exemplary flowchart illustrating an overall workflow 100 for cancer classification of a sample, according to one or more embodiments. Workflow 100 may involve one or more entities, including, for example, a healthcare provider, a sequencing device, an analysis system, etc. The purpose of the workflow includes detecting and / or monitoring cancer in an individual. From a medical perspective, workflow 100 may serve to complement other existing cancer diagnostic tools. Workflow 100 may provide early cancer detection and / or routine cancer monitoring to better inform treatment plans for individuals diagnosed with cancer. Overall workflow 100 may include additional / fewer steps than those shown in FIG. 1 .

[0042] A healthcare provider performs sample collection 110. An individual whose cancer is to be classified visits their healthcare provider. The healthcare provider collects a sample for cancer classification. Examples of biological samples include, but are not limited to, a subject's tissue biopsy, blood, whole blood, plasma, serous fluid, urine, cerebrospinal fluid, stool, saliva, sweat, tears, pleural fluid, pericardial fluid, and peritoneal fluid. The sample contains genetic material belonging to the individual that can be extracted and sequenced for cancer classification. Once the sample is collected, the sample is provided to a sequencing device. Along with the sample, the healthcare provider may also collect other information about the individual, such as biological sex, age, ethnicity, smoking status, medical history, etc.

[0043] The sequencing device sequences 120 the sample. A clinician at the lab may perform one or more processing steps on the sample in preparation for sequencing. Once prepared, the clinician loads the sample into the sequencing device. An example of a device used in sequencing is further described with respect to Figures 7A and 7B. Sequencing devices generally extract and separate nucleic acid fragments, which are sequenced to identify the sequences of nucleic acid bases corresponding to the fragments. Sequencing may also involve amplification of the nucleic acid material. Different sequencing processes include Sanger sequencing, fragment analysis, and next-generation sequencing. Sequencing can be whole-genome sequencing or targeted sequencing using a targeted panel. With respect to DNA methylation, bisulfite sequencing (e.g., as further described with respect to Figures 2A and 2B) can identify the methylation status through bisulfite conversion of unmethylated cytosines at CpG sites. Sample sequencing 120 obtains the sequences of multiple nucleic acid fragments in the sample. In one or more embodiments, the sequence may include methylation state vectors, each of which describes the methylation state for a CpG site on a fragment.

[0044] The analysis system performs pre-analysis processing 130. An exemplary analysis system is illustrated in Figure 7B. Pre-analysis processing 130 may include, but is not limited to, de-duplication of sequence reads, determining coverage metrics, determining whether a sample is contaminated, removing contaminated fragments, calling sequencing errors, etc.

[0045] The analysis system performs one or more analyses 140. The analyses may be statistical analyses or application of one or more trained models to predict at least the cancer status of the individual from whom the sample was collected. Different genetic features, such as methylation at CpG sites, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), and other types of genetic mutations, may be evaluated and considered. With respect to methylation, the analysis 140 may include identifying aberrant methylation 142 (as further described, e.g., in Figures 5A and 5B), feature extraction 144 (as further described, e.g., in Figures 3A, 3B, 4A, 4B, 5A, and 5B), and applying a cancer classifier 146 to identify a cancer prognosis (as further described, e.g., in Figures 6A and 6B). In one or more embodiments of feature extraction, the analysis system may utilize one or more age prediction models to generate one or more age-covariate residuals as features for cancer classification. The cancer classifier 146 inputs the extracted features for cancer prognosis identification. The cancer prognosis may be a label or a numerical value. The label may indicate a particular cancer state, e.g., a binary label may indicate the presence or absence of cancer, or a multi-class label may indicate one or more cancer types among multiple cancer types being screened for. The numerical value may indicate the likelihood of a particular cancer state, e.g., the likelihood of cancer and / or the likelihood of a particular cancer type.

[0046] The analytics system returns the prediction 150 to the healthcare provider, who may establish or adjust a treatment plan based on the cancer prediction. Treatment optimization is further discussed in Section "IV.C. Treatment."

[0047] Overview of IB methylation As described herein, cfDNA fragments from an individual are processed, e.g., by converting unmethylated cytosines to uracils, sequenced, and the sequence reads are compared to a reference genome to identify the methylation status at specific CpG sites within the DNA alignment. Each CpG site can be methylated or unmethylated. Identifying aberrantly methylated fragments compared to healthy individuals can provide insight into the subject's cancer status. As is well known in the art, aberrant DNA methylation (compared to healthy controls) can produce differential effects that may contribute to cancer. Identifying aberrantly methylated cfDNA fragments presents several challenges. First, identifying a DNA fragment as having an aberrantly methylated status can carry weight relative to an individual's control group, and therefore, when the control group is small, the identification becomes less reliable due to statistical variability within the smaller control group. Additionally, methylation status may vary within an individual's control group, which can be difficult to take into account when identifying a subject's DNA fragments as having an aberrantly methylated status. On the other hand, cytosine methylation at a CpG site can necessarily influence methylation at subsequent CpG sites, and encapsulating this dependency can be another challenge in itself.

[0048] Methylation typically occurs in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. In particular, methylation can occur at the dinucleotide of cytosine and guanine, referred to herein as a "CpG site." In other instances, methylation can occur at cytosines that are not part of a CpG site or at other nucleotides that are not cytosines, although these occur more rarely. In this disclosure, methylation is discussed in terms of CpG sites for clarity. Aberrant DNA methylation can be identified as hypermethylation or hypomethylation, both of which can indicate a cancerous state. Throughout this disclosure, hypermethylation and hypomethylation can be characterized for a DNA fragment when the DNA fragment contains more than a threshold number of CpG sites and more than a threshold percentage of these CpG sites are methylated or unmethylated.

[0049] The principles described herein can be applied to the detection of non-CpG methylation, including non-cytosine methylation. In such embodiments, the wet-lab assay used to detect methylation may differ from that described herein. Furthermore, the methylation state vectors described herein may include elements that are generally methylated or unmethylated sites (even if these sites are not specifically CpG sites). With such substitutions, the remainder of the processes described herein may remain the same, and as a result, the inventive concepts described herein apply to other forms of methylation.

[0050] IC definition The term "cell-free nucleic acid" or "cfNA" refers to nucleic acid fragments circulating in an individual's body (e.g., blood) and derived from one or more healthy cells and / or one or more unhealthy cells (e.g., cancer cells). The term "cell-free DNA" or "cfDNA" refers to deoxyribonucleic acid fragments circulating in an individual's body (e.g., blood). In addition, cfNA or cfDNA in an individual's body may be from other non-human sources.

[0051] The terms "genomic nucleic acid," "genomic DNA," or "gDNA" refer to nucleic acid or deoxyribonucleic acid molecules obtained from one or more cells. In various embodiments, gDNA can be extracted from healthy cells (e.g., non-tumor cells) or from tumor cells (e.g., biopsy samples). In some embodiments, gDNA can be extracted from cells obtained from the blood cell lineage, such as white blood cells.

[0052] The term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments derived from tumor cells or other types of cancer cells that may be released into an individual's bodily fluids (e.g., blood, sweat, urine, or saliva) as a result of biological processes such as apoptosis or necrosis of dying cells, or that may be actively released by viable tumor cells.

[0053] The terms "DNA fragment," "fragment," or "DNA molecule" may generally refer to any deoxyribonucleic acid fragment, i.e., cfDNA, gDNA, ctDNA, etc.

[0054] The terms "aberrant fragment," "aberrantly methylated fragment," or "fragment with an aberrant methylation pattern" refer to a fragment with aberrant methylation of a CpG site. The aberrant methylation of a fragment can be identified using a probabilistic model to identify the surprise of observing the methylation pattern of the fragment in a control group.

[0055] The term "abnormal fragments with extreme methylation" or "UFXM" refers to hypomethylated or hypermethylated fragments. Hypomethylated and hypermethylated fragments refer to fragments that have at least some CpG sites (e.g., 5) that are above a threshold percentage (e.g., 90%) of either methylated or unmethylated, respectively.

[0056] The term "anomaly score" refers to a score for a CpG site based on the number of aberrant fragments (or UFXMs, in some embodiments) from a sample that overlap that CpG site. The anomaly score is used in conjunction with feature identification of the sample for classification.

[0057] As used herein, the terms "about" or "approximately" can mean within an acceptable error for a particular numerical value as determined by one of ordinary skill in the art, which may depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" can mean within 1 or more standard deviations, as is customary in the industry. "About" can mean a range of ±20%, ±10%, ±5%, or ±1% of a numerical value. The terms "about" or "nearly" can mean within one order of magnitude, i.e., within 5-fold or 2-fold, of a numerical value. When a specific numerical value is described in the specification and claims, unless otherwise specified, it should be assumed that the term "about" means within an acceptable error range for that particular numerical value. The term "about" can have the meaning commonly understood by one of ordinary skill in the art. The term "about" can mean ±10%. The term "about" can mean ±5%.

[0058] As used herein, "biological sample," "patient sample," or "sample" refers to any sample obtained from a subject, which may reflect a biological state associated with the subject and includes cell-free DNA. Examples of biological samples include, but are not limited to, a subject's blood, white blood cells, plasma, serous fluid, urine, cerebrospinal fluid, stool, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid. A biological sample may include any tissue or substance obtained from a living or dead subject. A biological sample may be a cell-free sample. A biological sample may contain nucleic acid (e.g., DNA or RNA) or a fragment thereof. The term "nucleic acid" may refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a hybrid or fragment of either. Nucleic acid in a sample may be cell-free nucleic acid. A sample may be a liquid sample or a solid sample (e.g., a cell or tissue sample). The biological sample can be blood, plasma, serous fluid, urine, saliva, fluid from a hydrocele (e.g., testicle), pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from other parts of the body (e.g., thyroid, breast), etc. The biological sample can be a stool sample. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained through a centrifugation protocol) can be cell-free (e.g., greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free). The biological sample can be treated to physically disrupt tissue or cellular structures (e.g., centrifugation and / or cell lysis), thereby releasing intracellular components into a solution that can further include enzymes, buffers, salts, detergents, etc., that can be used to prepare the sample for analysis.

[0059] As used herein, the terms "control," "control sample," "reference," "reference sample," "normal," and "normal sample" refer to a sample from a subject who does not have a specific disease or is otherwise healthy. In some examples, the methods disclosed herein can be performed on a subject with a tumor, in which case the reference sample is a sample obtained from the subject's healthy tissue. The reference sample can be obtained from the subject or from a database. The reference can be, for example, a reference genome used to map nucleic acid fragment sequences obtained from sequencing a sample from the subject. The reference genome can refer to a haploid or diploid genome, to which nucleic acid fragment sequences from a biological sample and a structural sample can be aligned and compared. An example of a structural sample can be the DNA of white blood cells obtained from the subject. In the case of a haploid genome, there can be only one nucleotide at each locus. In the case of a diploid genome, heterozygous loci can be identified, and each heterozygous locus can have two alleles, with either allele being available for alignment with the locus.

[0060] As used herein, the term "cancer" or "tumor" refers to an abnormal mass of tissue in which the growth of the mass is more rapid and uncoordinated than the growth of normal tissue.

[0061] As used herein, the term "healthy" refers to a subject who has good health. A healthy subject can be proven to be free of any malignant or non-malignant disease. A "healthy individual" may have other diseases or illnesses unrelated to the disease being assayed and that would not normally be considered "healthy."

[0062] As used herein, the term "methylation" refers to a modification of deoxyribonucleic acid (DNA) in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. In particular, methylation tends to occur at the dinucleotides of cytosine and guanine, referred to herein as "CpG sites." In other instances, methylation can occur at cytosines that are not part of CpG sites or at other nucleotides that are not cytosines, although these occurrences are rare. Aberrant cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which can indicate a cancerous state. Aberrant DNA methylation (compared to healthy controls) can produce differential effects, which may contribute to cancer. The principles herein apply equally to the detection of methylation at CpGs and non-CpGs, including non-cytosine methylation. Furthermore, a methylation status vector can generally include elements that are vectors of sites with or without methylation (even if these sites are not CpG sites).

[0063] As used interchangeably herein, the terms "methylation fragment" or "nucleic acid methylation fragment" refer to a sequence of methylation states for each CpG site in a plurality of CpG sites identified by methylation sequencing of a nucleic acid (e.g., a nucleic acid molecule and / or a nucleic acid fragment). In a methylation fragment, the position and methylation state of each CpG site in a nucleic acid fragment are identified based on alignment of sequence reads (e.g., obtained from sequencing the nucleic acid) with a reference genome. A nucleic acid methylation fragment includes a methylation state (e.g., a methylation state vector) for each CpG site in a plurality of CpG sites, which indicates the position of the nucleic acid fragment in the reference genome (e.g., indicated by the position of the first CpG site in the nucleic acid fragment using a CpG index or other similar metric) and the number of CpG sites in the nucleic acid fragment. Alignment of sequence reads based on methylation sequencing of a nucleic acid molecule with a reference genome can be performed using a CpG index. As used herein, the term "CpG index" refers to a list of each CpG site in a plurality of CpG sites (e.g., CpG1, CpG2, CpG3, etc.) in a reference genome, such as the human reference genome, which may be in an electronic format. The CpG index further includes, for each respective CpG site in the CpG index, a corresponding genomic location in the corresponding reference genome. Each CpG site in each nucleic acid methylated fragment is therefore indexed to a specific location in each reference genome, which can be identified using the CpG index.

[0064] As used herein, the term "true positive" (TP) refers to a subject who has a disease. A "true positive" can refer to a subject who has a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), localized or metastatic cancer, or a non-malignant disease. A "true positive" can refer to a subject who has a disease and is identified as having the disease by an assay or method of the present disclosure. As used herein, the term "true negative" (TN) refers to a subject who does not have a disease or who does not have detectable disease. A true negative can refer to a subject who does not have a disease or detectable disease, such as a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), localized or metastatic cancer, a non-malignant disease, or who is otherwise healthy. A true negative can refer to a subject who does not have a disease or who does not have detectable disease or who is identified as not having the disease by an assay or method of the present disclosure.

[0065] As used herein, the term "reference genome" refers to any specific, known, sequenced, or characterized partial or complete genome of any organism or virus that can be used to reference identified sequences from a subject. Exemplary reference genomes for human subjects, as well as many other organisms, are provided in online genome browsers hosted by the National Center for Biotechnology Information (NCBI) or the University of California, Santa Cruz (UCSC). "Genome" refers to the complete genetic information of an organism or virus expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome is often a synthesized or partially synthesized genome sequence from one or more individuals. In some embodiments, a reference genome is a synthesized or partially synthesized genome sequence from one or more human individuals. A reference genome can be viewed as a genetic representative of a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).

[0066] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence generated by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read") or from both ends of a nucleic acid (e.g., a paired-end read, a double-end read). In some embodiments, a sequence read (e.g., a single-end or paired-end read) can be generated from one or both strands of a target nucleic acid fragment. The length of a sequence read is often associated with a particular sequencing technology. High-throughput methods provide sequence reads that can vary in size, for example, from tens to hundreds of base pairs (bp). In some embodiments, sequence reads have a mean, median, or average length of about 15 bp to 900 bp in length (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 450 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, sequence reads have a mean, median, or average length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. The length of a sequence read is typically a page length. For example, nanopore sequencing can provide sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads that are less variable, e.g., most of the sequence reads can be less than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, or a string of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of an entire nucleic acid fragment.Sequence reads can be obtained in a variety of ways, for example using sequencing techniques, probes, e.g. in hybridization arrays or capture probes, or amplification techniques such as polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.

[0067] As used herein, terms like "sequencing" generally refer to any biochemical process that can be used to determine the order of biological macromolecules, such as nucleic acids and proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as a DNA fragment.

[0068] As used herein, the term "sequencing depth" is used interchangeably with the term "coverage" and refers to the number of times a locus is covered by consensus sequence reads corresponding to unique nucleic acid target molecules aligned with that locus. For example, sequencing depth is equal to the number of unique nucleic acid target molecules covering the locus. A locus can be as small as a nucleotide, as large as a chromosome arm, or as large as an entire genome. Sequencing depth can be expressed as "Yx," e.g., 50x, 100x, etc., where "Y" refers to the number of times a locus is covered by sequences corresponding to nucleic acid targets, e.g., the number of independent sequence information covering a particular locus is obtained. In some embodiments, sequencing depth corresponds to the number of genomes sequenced. Sequencing depth can also apply to multiple loci or the entire genome, in which case Y refers to the average number of times a locus, a haploid genome, or the entire genome is sequenced, respectively. When an average depth is cited, the actual depth of different loci included in a dataset may range over a certain range of values. Ultra-deep sequencing can refer to a sequencing depth of at least 100x at a given locus.

[0069] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the number of true positives plus false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify the proportion of a population that actually has the disease. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects in a population who have cancer. In another example, sensitivity can characterize the ability of a method to correctly identify one or more markers that are indicative of cancer.

[0070] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the number of true negatives plus false positives. Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that is actually disease-free. For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population that do not have cancer. In another example, specificity characterizes the ability of a method to correctly identify one or more markers of cancer.

[0071] As used herein, the term "subject" refers to any living or non-living organism, including, but not limited to, humans (e.g., males, females, fetuses, pregnant females, children, etc.), non-human animals, plants, bacteria, fungi, or protozoans. The subject can be either a human or a non-human animal, including, but not limited to, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cattle), equines (e.g., horses), caprines (e.g., sheep, goats), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursinians (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, the subject is a male or female of any stage (e.g., adult male, adult female, child). The subject from whom the sample is taken or who is treated by any of the methods or compositions described herein may be of any age, and may be an adult, an infant, or a child.

[0072] As used herein, the term "tissue" can refer to a collection of cells organized as a functional unit. Multiple cell types can be found within a single tissue. Different types of tissue can consist of different cell types (e.g., liver cells, lung cells, or blood cells), as well as tissues from different organisms (maternal and fetal) and healthy and tumor cells. The term "tissue" can generally refer to any collection of cells found in the human body (e.g., cardiac tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some embodiments, the term "tissue" or "tissue type" can refer to the tissue from which cell-free nucleic acid is derived. In one example, viral nucleic acid fragments can be obtained from blood tissue. In another example, viral nucleic acid fragments can be obtained from tumor tissue.

[0073] As used herein, the term "genomic" refers to a characteristic of an organism's genome. Examples of genomic characteristics include, but are not limited to, those related to the primary nucleic acid sequence of all or a portion of the genome (e.g., the presence or absence of single nucleotide polymorphisms, indels, sequence rearrangements, mutation frequencies, etc.), the copy number of one or more specific nucleotide sequences within the genome (e.g., copy number, allele frequency fraction, single chromosome, or whole genome ploidy, etc.), the epigenetic state of all or a portion of the genome (e.g., nucleic acid covalent modifications such as methylation, histone modifications, nucleosome topology, etc.), and the expression profile of the organism's genome (e.g., gene expression levels, isotype expression levels, gene expression ratios, etc.).

[0074] The terms used herein are for the purpose of describing particular cases only and are not intended to be limiting. As used herein, the singular articles (a, an, the) are intended to include the plural forms as well, unless the context clearly dictates otherwise. Furthermore, when the terms "including," "includes," "having," "has," "with," or variations thereof are used in the detailed description and / or claims, these terms are intended to be as inclusive as "comprising."

[0075] Example of an ID analysis system 7A is an exemplary flowchart of an apparatus for sequencing a nucleic acid sample, according to one or more embodiments. This exemplary flowchart includes apparatus such as a sequencer 720 and an analysis system 700. The sequencer 720 and analysis system 700 may cooperate to perform one or more steps of the processes of 300 in FIG. 3A, 400 in FIG. 4A, 420 in FIG. 4B, and other processes described herein.

[0076] In various embodiments, a sequencer 720 receives the enriched nucleic acid sample 710. As shown in FIG. 7A , the sequencer 720 can include a graphical user interface 725 that allows a user to interact with specific tasks (e.g., start sequencing or stop sequencing) as well as another loading station 730 for loading a sequencing cartridge with an enriched fragment sample and / or loading buffers necessary to run a sequencing assay. Thus, once a user of the sequencer 720 provides the necessary reagents and sequencing cartridge to the loading station 730 of the sequencer 720, the user can initiate sequencing by manipulating the graphical user interface 725 of the sequencer 720. Once initiated, the sequencer 720 performs sequencing and outputs sequence reads of the enriched fragments from the nucleic acid sample 710.

[0077] In some embodiments, sequencer 720 is communicatively coupled to analysis system 700. Analysis system 700 includes several computing devices used to process sequence reads for various applications, such as assessing methylation status at one or more CpG sites, calling variants, or quality control. Sequencer 720 may provide sequence reads to analysis system 700 in BAM file format. Analysis system 700 can be communicatively coupled to sequencer 720 via wireless, wired, or a combination of wireless and wired communication technologies. Generally, analysis system 700 comprises a processor and a non-transitory computer-readable storage medium storing computer instructions that, when executed by the processor, cause the processor to process sequence reads or perform one or more steps of any of the methods or processes disclosed herein.

[0078] In some embodiments, sequence reads can be aligned to a reference genome using methods known in the art to determine alignment position information. Alignment positions generally describe the start and end positions of a region within a reference genome, corresponding to the first and last nucleotide bases of a sequence read. In support of methylation sequencing, alignment position information can be generalized to indicate the first and last CpG sites contained within a sequence read according to alignment with the reference genome. Alignment position information can further indicate the methylation status and positions of all CpG sites within a sequence read. Regions within a reference genome can be associated with genes or gene segments, and thus, the analysis system 700 can label a sequence read with one or more genes aligned to the sequence read. In one embodiment, the length (or size) of a fragment is determined from the start and end positions.

[0079] In various embodiments, for example, when a paired-end sequencing process is used, a sequence read consists of a read pair denoted as R_1 and R_2. For example, a first read R_1 can be sequenced from a first end of a double-stranded DNA (dsDNA) molecule, and a second read R_2 can be sequenced from a second end of the double-stranded DNA (dsDNA). Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 can always be aligned (e.g., in opposite orientations) with the nucleotide bases of the reference genome. The alignment position information obtained from the read pair R_1 and R_2 can include the start position of the reference genome corresponding to the end of the first read (e.g., R_1) and the end position in the reference genome corresponding to the end of the second read (e.g., R_2). In other words, the start and end positions of the reference genome can represent possible positions in the reference genome to which nucleic acid fragments correspond. An output file having a SAM (sequence alignment map) format or a BAM (binary) format can be generated and output for subsequent analysis.

[0080] 7B, which is a block diagram of an analysis system 700 for processing DNA samples according to one embodiment. The analysis system is implemented on one or more computing devices used to analyze DNA samples. The analysis system includes a sequence processor 740, a sequence database 745, a model database 755, a model 750, a parameter database 765, and a score engine 760. In some embodiments, the analysis system 700 performs some or all of the processes described in this disclosure.

[0081] The sequence processor 740 generates methylation state vectors for fragments from the sample. For each CpG site on a fragment, the sequence processor 740 generates a methylation state vector for each fragment via process 200 of FIG. 2A, which specifies the location of the fragment in the reference genome, the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment: methylated, unmethylated, or uncertain. The sequence processor 740 may store the methylation state vectors for the fragments in a sequence database 745. The data in the sequence database 745 may be organized so that the methylation state vectors from the samples are related to each other.

[0082] Additionally, multiple different models 750 may be stored in a model database 755 or retrieved for use with test samples. In one example, the model is a trained cancer classifier for identifying a cancer prediction for a test sample using feature vectors obtained from abnormal fragments. The training and use of cancer classifiers is further discussed in Section "III. Cancer Classifiers for Cancer Identification." The analysis system 700 may train one or more models 750 and store various trained parameters in a parameter database 765. The analysis system 700 stores the models 750, along with their functions, in the model database 755.

[0083] During inference, the score engine 760 uses one or more models 750 to return an output. The score engine 760 accesses the models 750 in the model database 755 and trained parameters from the parameter database 765. According to each model, the score engine receives appropriate inputs for the model and calculates an output based on the received inputs, parameters, and a function of each model on the inputs and outputs. In some use cases, the score engine 760 also calculates metrics that correlate with confidence in the calculated output from the model. In other use cases, the score engine 760 calculates other intermediate values ​​for use within the model.

[0084] II. Sample Sequencing and Processing II.A. Generation of DNA fragment methylation status vectors 2A is an exemplary flowchart illustrating a process 200 for sequencing fragments of cfDNA to obtain a methylation state vector, according to one or more embodiments. To analyze DNA methylation, an analysis system first obtains 210 a sample from an individual, the sample containing a plurality of cfDNA molecules. In additional embodiments, process 200 may also be applied to sequencing other types of DNA molecules. Process 200 is one embodiment of sample sequencing 120 of FIG. 1.

[0085] From the sample, the analysis system can separate individual cfDNA molecules 210. The cfDNA molecules can be treated to convert unmethylated cytosines to uracil 220. In one embodiment, the method utilizes bisulfite treatment of DNA, which converts unmethylated cytosines to uracil without converting methylated cytosines. For example, commercially available kits such as EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, or EZ DNA Methylation™-Lightning kits (available from Zymo Research, Inc., Irvine, CA) are used for bisulfite conversion. In other embodiments, the conversion of unmethylated cytosines to uracil is performed using an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for converting unmethylated cytosines to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).

[0086] Sequencing libraries can be created from the converted cfDNA molecules. 230 During library creation, molecular barcodes (UMIs, unique molecular identifiers) can be added to nucleic acid molecules (e.g., DNA molecules) through adapter ligation. UMIs can be short nucleic acid sequences (e.g., 4–10 base pairs) added to the ends of DNA fragments (DNA molecules fragmented by physical shearing, enzymatic digestion, and / or chemical fragmentation) during adapter ligation. UMIs can be degenerate base pairs that serve as unique tags that can be used to identify sequence reads derived from specific DNA fragments. During PCR amplification after adapter ligation, UMIs can replicate along with the attached DNA fragments, allowing sequence reads from the same original fragment to be identified during downstream analysis.

[0087] Optionally, sequencing libraries can be enriched for cfDNA molecules or genomic regions235, which can signal cancer status using multiple hybridization probes. Hybridization probes are short oligonucleotides that can hybridize to specifically designated cfDNA molecules or target regions and enrich these fragments or regions for subsequent sequencing and analysis. Hybridization probes can be used to perform targeted, high-depth analysis of designated CpG sites of interest to researchers. Hybridization probes can be aligned with one or more target sequences at 1X, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, or 11X or greater coverage. For example, hybridization probes aligned with 2X coverage include overlapping probes, whereby each portion of the target sequence hybridizes to two independent probes. Hybridization probes can be aligned with one or more target sequences at less than 1X coverage.

[0088] In one embodiment, hybridization probes are designed to enrich for DNA molecules that have been treated to convert unmethylated cytosines to uracils (e.g., with bisulfite). During enrichment, hybridization probes (also referred to herein as "probes") can be used to target and extract nucleic acid fragments that indicate the presence or absence of cancer (or disease), the state of the cancer, or the classification of the cancer (e.g., cancer class or tissue of origin). Probes can be designed to anneal (i.e., hybridize) to a target (complementary) strand of DNA. The target strand can be the "positive" strand (e.g., the strand transcribed into mRNA and then translated into protein) or the complementary "negative" strand. Probes can range in length from 10s, 100s, or 1000s of base pairs. Probes can be designed based on a methylation site panel. Probes can be designed based on a panel of target genes to analyze target regions of the genome (human or other organism) suspected of corresponding to specific mutations or specific cancers or other types of diseases. Additionally, probes can cover overlapping portions of the target regions.

[0089] Once prepared, the sequencing library or a portion thereof can be sequenced to obtain multiple sequence reads 240. The sequence reads can be in a computer-readable digital format for processing and interpretation by computer software. The sequence reads can be aligned with the reference genome to determine alignment position information. The alignment position information can indicate the start and end positions of a region in the reference genome, corresponding to the starting and ending nucleotide bases of a sequence read. The alignment position information can also include the length of the sequence read, which can be determined from the start and end positions. The region in the reference genome can be associated with a gene or gene segment. The sequence reads can consist of a read pair, denoted as R1 and R2. For example, the first read R1 can be sequenced from a first end of a nucleic acid fragment, and the second read R2 can be sequenced from a second end of the nucleic acid fragment. Thus, the nucleotide base pairs of the first read R1 and the second read R2 can always be aligned (e.g., in opposite orientations) with the nucleotide bases of the reference genome. The alignment position information obtained from the read pair R1 and R2 may include a start position in the reference genome corresponding to the end of the first read (e.g., R1) and an end position in the reference genome corresponding to the end of the second read (e.g., R2). In other words, the start and end positions in the reference genome may represent positions in the reference genome to which the nucleic acid fragment may correspond. An output file having a sequence alignment map (SAM) format or a binary array map (BAM) format may be generated and output for subsequent analysis, such as methylation status identification.

[0090] From the sequence reads, the analysis system identifies the location and methylation state of each CpG site based on alignment with the reference genome 250. The analysis system generates a methylation state vector for each fragment 260, which specifies the location of the fragment within the reference genome (e.g., specified by the location of the first CpG site in each fragment or other similar metric), the number of CpG sites within the fragment, and the methylation state of each CpG site within the fragment: methylated (e.g., indicated by M), unmethylated (e.g., indicated by U), or indeterminate (e.g., indicated by I). Observed states can be methylated and unmethylated states, while unobserved states are indeterminate. Indeterminate methylation states can result from sequencing errors and / or mismatches between the methylation states of complementary strands of DNA fragments. The methylation state vector can be persisted in temporary or persistent computer memory for later use or processing. Additionally, the analysis system can remove duplicate reads or duplicate methylation state vectors from a single sample. The analysis system may determine that the uncertain methylation status of particular fragments with one or more CpG sites exceeds a threshold number or percentage and may exclude such fragments or selectively include such fragments but build a model that takes into account such uncertain methylation status.

[0091] Figure 2B is an illustrative diagram of the process 200 of Figure 2A for sequencing a cfDNA molecule to obtain a methylation state vector, according to one or more embodiments. As an example, an analysis system receives a cfDNA molecule 212, which in this example includes three CpG sites. As shown, the first and third CpG sites of the cfDNA molecule 212 are methylated 214. During processing step 220, the cfDNA molecule 212 is converted to generate a converted cfDNA molecule 222. During processing 220, the second CpG site, which was unmethylated, has its cytosine converted to uracil. However, the first and third CpG sites are not converted.

[0092] After conversion, a sequencing library 230 is created and sequenced 240 to generate sequence reads 242. The analysis system aligns 250 the sequence reads 242 to a reference genome 244, which provides context for where in the human genome the fragmented cfDNA originated. In this simplified example, the analysis system aligns 250 the sequence reads 242 so that three CpG sites correlate with CpG sites 23, 24, and 25 (arbitrary reference identifiers used for convenience of illustration). The analysis system can therefore generate information about both the methylation status of every CpG site on the cfDNA molecule 212 and the location to which the CpG site maps within the human genome. As shown, methylated CpG sites on the sequence read 242 are read as cytosines. In this example, cytosines are found only in the first and third CpG sites in the sequence read 242, which infers that the first and third CpG sites of the original cfDNA molecule are methylated. In contrast, the second CpG site can be read as thymine (U is converted to T during the sequencing process), and therefore, it can be inferred that the second CpG site is unmethylated in the original cfDNA molecule. From these two pieces of information, i.e., methylation status and location, the analysis system generates a methylation status vector 252 for the fragment cfDNA 212 260. In this example, the resulting methylation status vector 252 is <M 23 ,U 23 ,M 25 >, where M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript numbers correspond to the position of each CpG site in the reference genome.

[0093] One or more alternative sequencing methods can also be used to obtain sequence reads from nucleic acids in biological samples. The one or more sequencing methods can include any form of sequencing that can be used to obtain multiple sequence reads measured from nucleic acids (e.g., cell-free nucleic acids), including, but not limited to, high-throughput sequencing systems such as the Roche 454 platform, Applied Biosystems' SOLID platform, Helicos True Single Molecule DNA sequencing, Affymetrix's sequencing-by-hybridization platform, Pacific Biosciences' single molecule real-time (SMRT) method, sequencing-by-synthesis platforms from 454 Life Science, Illumina / Solexa, and Helicos Biosciences, and Applied Biosystems' sequencing-by-ligation platform. Life technologies' ION TORRENT technology and nanopore sequencing can also be used to obtain sequence reads from nucleic acids (e.g., cell-free nucleic acids) in biological samples. Sequencing-by-synthesis and reversible terminator-based sequencing (e.g., Illumina's Genome Analyzer, Genome Analyzer II, HISEQ 2000; HISEQ 4500 (Illumina, San Diego, Calif.)) can be used to obtain sequence reads from cell-free nucleic acids obtained from biological samples of training subjects to form genotype datasets. Millions of cell-free nucleic acid (e.g., DNA) fragments can be sequenced in parallel. One example of this type of sequencing technology uses a flow cell containing an optically transparent slide with eight individual lanes, on whose surface oligonucleotide anchors (e.g., adapter plasma) are immobilized.The cell-free nucleic acid sample can include a signal or tag to facilitate detection. Obtaining sequence reads from cell-free nucleic acids obtained from a biological sample can include obtaining quantitative information from the signal or tag through various techniques, such as flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene chip analysis, microarray, mass spectrometry, cytofluorimetric analysis, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electroprecipitation, sequencing, and combinations thereof.

[0094] The one or more sequencing methods can include a whole genome sequencing assay. A whole genome sequencing assay can include a physical assay that generates sequence reads of the entire genome or a substantial portion of the entire genome, which can be used to identify large-scale abnormalities such as copy number variations and copy number aberrations. Such a physical assay can utilize whole genome sequencing or whole exome sequencing technology. The average sequencing depth of the whole genome sequencing assay can be at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, at least 20x, at least 30x, or at least 40x across the subject's genome. In some embodiments, the sequencing depth is about 30,000x. The one or more sequencing methods can include a targeted panel sequencing assay. The average sequencing depth of the targeted panel sequencing assay can be at least 50,000x, at least 55,000x, at least 60,000x, or at least 70,000x for a targeted gene panel. The targeted gene panel can include 450 to 500 genes. The targeted gene panel can include a range of 500±5 genes, a range of 500±10 genes, or a range of 500±25 genes.

[0095] The one or more sequencing methods can include paired-end sequencing. The one or more sequencing methods can generate multiple sequence reads. The average length of the multiple sequence reads can be 10 to 700, 50 to 400, or 100 to 300. The one or more sequencing methods can include a methylation sequencing assay. The methylation sequencing can be i) whole-genome methylation sequencing or ii) targeted DNA methylation sequencing using multiple nucleic acid probes. For example, the methylation sequencing can be whole-genome bisulfite sequencing (e.g., WGBS). The methylation sequencing can be targeted DNA methylation sequencing using multiple nucleic acid probes targeting the most informative regions of the methylome, a unique methylation database, and previous prototype whole-genome and targeted sequencing assays.

[0096] Methylation sequencing can detect one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in each nucleic acid methylated fragment. Methylation sequencing can include converting one or more unmethylated cytosines or one or more methylated cytosines in each nucleic acid methylated fragment to one or more corresponding uracils. One or more uracils can be detected as one or more corresponding thymines during methylation sequencing. The conversion of one or more unmethylated cytosines or one or more methylated cytosines can include chemical conversion, enzymatic conversion, or a combination thereof.

[0097] For example, bisulfite conversion involves converting cytosine to uracil while leaving methylated cytosine (e.g., 5-methylcytosine or 5-mC) intact. In some DNA, approximately 95% of the cytosines may be unmethylated, and the resulting DNA fragments may contain many uracils represented by thymine. Enzymatic conversion processes can be used to treat nucleic acids prior to sequencing, which can be performed in a variety of ways. One example of bisulfite-free conversion is TET-assisted pyridine borane sequencing (TAPS), a bisulfite-free or single-base resolution sequencing method for nondestructively detecting 5-methylcytosine and 5-hydroxymethylcytosine directly without affecting unmethylated cytosine. The methylation state of a CpG site among the corresponding plurality of CpG sites in each nucleic acid methylation fragment can be methylated when the CpG site is identified as methylated by methylation sequencing, or can be unmethylated when the CpG site is identified as unmethylated by methylation sequencing.

[0098] The average sequencing depth of a methylation sequencing assay (e.g., WGBS and / or targeted methylation sequencing) can include, but is not limited to, up to about 1,000x, 2,000x, 3,000x, 5,000x, 10,000x, 15,000x, 20,000x, or 30,000x. The sequencing depth of methylation sequencing can be greater than 3,000x, for example, at least 40,000x or 50,000x. The average sequencing depth of a whole-genome bisulfite sequencing method can be 20x to 50x, and the average effective depth of a targeted methylation sequencing method can be 100x to 1000x, where the effective depth is the equivalent whole-genome bisulfite sequencing coverage required to obtain the same number of sequence reads as obtained by targeted methylation sequencing.

[0099] For further details on methylation sequencing (e.g., WGBS and / or targeted methylation sequencing), see, e.g., U.S. Patent Application No. 16 / 352,602, entitled "Methylation Fragment Anomaly Detection," filed March 13, 2019, and U.S. Patent Application No. 16 / 719,902, entitled "Systems and Methods for Estimating Cell Source Fractions Using Methylation Information," filed December 18, 2019, each of which is incorporated herein by reference. Other methods of methylation sequencing, including those disclosed herein and / or modifications, alternatives, or combinations thereof, can be used to obtain fragment methylation patterns. Methylation sequencing can be used to identify one or more methylation state vectors, for example, as described in U.S. patent application Ser. No. 16 / 352,602, filed March 13, 2019, entitled "Anomalous Fragment Detection and Classification," or according to any of the techniques disclosed in U.S. patent application Ser. No. 15 / 931,022, filed May 13, 2020, entitled "Model-Based Featurization and Classification," each of which is incorporated herein by reference.

[0100] Methylation sequencing of nucleic acids and the resulting one or more methylation state vectors can be used to obtain a plurality of nucleic acid methylation fragments. Each corresponding plurality of nucleic acid methylation fragments (e.g., for each respective genotype data set) can include more than 100 nucleic acid methylation fragments. The average number of nucleic acid methylation fragments among each of the corresponding plurality of nucleic acid methylation fragments can include 1,000 or more nucleic acid methylation fragments, 500 or more nucleic acid methylation fragments, 10,000 or more nucleic acid methylation fragments, 20,000 or more nucleic acid methylation fragments, or 30,000 or more nucleic acid methylation fragments. The average number of nucleic acid methylation fragments among each of the corresponding plurality of nucleic acid methylation fragments can be between 10,000 and 50,000 nucleic acid methylation fragments. The corresponding plurality of nucleic acid methylation fragments can include 1,000 or more, 10,000 or more, 100,000 or more, 1 million or more, 10 million or more, 100 million or more, 500 million or more, 1 billion or more, 2 billion or more, 3 billion or more, 4 billion or more, 5 billion or more, 6 billion or more, 7 billion or more, 8 billion or more, 9 billion or more, or 10 billion or more nucleic acid methylation fragments. The average length of the corresponding plurality of nucleic acid methylation fragments can be 140-480 nucleotides.

[0101] Further details regarding methods for sequencing nucleic acid and methylation sequencing data are disclosed in U.S. Patent Application No. 17 / 191,914, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," filed March 4, 2021, the entire contents of which are incorporated herein by reference.

[0102] III. Cancer Classifier for Identifying Cancer Cancer classification can include extracting genetic features and applying one or more models to the extracted features to identify a cancer prediction. The extracted features can include a feature vector generated for the test sample. Cancer classification of the test sample can include identifying a cancer prediction based on the feature vector. The cancer prediction can include a label and / or a numerical value. The label can be binary, indicating the presence or absence of cancer in the subject, and / or multi-class, including one or more specific cancer types among a plurality of screened cancer types. In particular, the cancer classifier can be a machine learning model including multiple classification parameters and a function representing the relationship between the feature vector as input and the cancer prediction as output. Inputting the feature vector along with the cancer parameters into the function can result in a cancer prediction. In one or more embodiments, an age prediction model is used to predict the age of an individual associated with the test sample based on methylation features. The residual between the subject's predicted age and reported age can be utilized as a feature within the cancer classifier. In one or more embodiments, the feature vector input to the cancer classifier is based on a set of aberrant fragments (also called "aberrant methylation" or "unusual fragments of extreme methylation" (UFXM)) identified from the test sample. The aberrant fragments may be identified through process 520 of FIG. 5B, or more specifically, hypermethylated or hypomethylated fragments identified through step 570 of process 520, or aberrant fragments identified by some other process. Prior to utilizing the cancer classifier, the analysis system may train the cancer classifier.

[0103] Cancer classification can include extracting genetic features and applying one or more models to the extracted features to identify a cancer prediction. The extracted features can include a feature vector generated for the test sample, and a cancer prediction can be identified based on the input feature vector. The cancer prediction can include a label and / or a numerical value. The label can be binary, indicating the presence or absence of cancer in the subject, and / or multi-class, including one or more specific cancer types among a plurality of screened cancer types. In particular, the cancer classifier can be a machine learning model including multiple classification parameters and a function representing the relationship between the feature vector as input and the cancer prediction as output. Inputting the feature vector along with the cancer parameters into the function can result in a cancer prediction. In one or more embodiments, an age prediction model is used to predict the age of an individual associated with the test sample based on methylation features. The residual between the subject's predicted age and reported age can be utilized as a feature within the cancer classifier. In one or more embodiments, the feature vector input to the cancer classifier is based on a set of aberrant fragments (also called "aberrant methylation" or "unusual fragments of extreme methylation" (UFXM)) identified from the test sample. The aberrant fragments may be identified through process 520 of FIG. 5B, or more specifically, hypermethylated or hypomethylated fragments identified through step 570 of process 520, or aberrant fragments identified by some other process. Prior to utilizing the cancer classifier, the analysis system may train the cancer classifier.

[0104] III.A. Age Prediction Model The age prediction model can predict the age of a sample based on methylation features extracted from the methylation pattern in the sample. The age prediction model can evaluate methylation features for multiple age-information or age-indicating genomic regions. The genomic region can be a region covering one CpG site or multiple CpG sites. The methylation features can be obtained from the methylation patterns of sequence reads in the sample. The number of age-indicating regions, and therefore age-indicating features, can be 1, 5, 10, 25, 50, 100, 1,000, 10,000, 100,000, or more genomic regions.

[0105] FIG. 3A illustrates methylation features that can be obtained from a single CpG site as a genomic region, according to one or more embodiments. For example, there are six fragments 310 that overlap a single CpG site 305. At each CpG site, indicated by a diamond on the fragment, the fragment has a methylation state. Methylation states can include methylation, indicated as filled, unmethylation, indicated as unfilled, and variants, indicated by diagonal hatches. Methylation variants can include indeterminate states resulting from mutations or sequencing errors. One methylation feature is the methylation density at CpG site 305. In this example, four of the six fragments (fragment 1 310A, fragment 2 310B, fragment 5 310E, and fragment 6 310F) have a methylation state at CpG site 305. The methylation density is 4 / 6, 0.66, or 66%. Another methylation feature indicates the percentage of highly methylated fragments that overlap with CpG site 305. The highly methylated fragments can have the above-mentioned percentages of overlapping CpG sites. Exemplary threshold percentages include 75%, 80%, 85%, 90%, 95%, etc. In this example, using a threshold percentage of 80%, three of the six fragments (fragment 1 310A, fragment 2 310B, and fragment 3 310C) are highly methylated and at least four of the five CpG sites are methylated. The percentage of overlapping highly methylated fragments is 3 / 6, 0.50, or 50%. Another methylation feature indicates the percentage of highly unmethylated fragments that overlap with CpG site 305. Exemplary threshold percentages include 75%, 80%, 85%, 90%, 95%, etc. In this example, using 80% as the threshold percentage, one of the six fragments (fragment 310B) is highly unmethylated and at least four of the five CpG sites are unmethylated. The percentage of overlapping highly unmethylated fragments is 1 / 6, or 0.16, or 16%.Fragment 5 310E and fragment 6 310F are in a mixed methylation state, i.e., neither hypermethylated nor highly unmethylated. The hypermethylated or unmethylated methylation feature is used to identify more significant fragments that overlap with CpG site 305. Take fragment 3 310C as an example, although fragment 3 310C is unmethylated at CpG site 305, as a highly methylated fragment, fragment 3 310C contributes to the count of hypermethylated overlapping fragments. Other methylation features can be obtained based on the aforementioned methylation features. For example, other methylation features can include a comparison between the number of hypermethylated overlapping fragments and the number of highly unmethylated overlapping fragments, e.g., the ratio of the two numbers.

[0106] Figure 3B illustrates methylation features that can be derived from multiple CpG sites as a genomic region 315, according to one or more embodiments. CpG site 317 includes CpG sites 1, 2, 3, 4, and 5. As in Figure 3A, filled diamonds indicate methylation, open diamonds indicate unmethylation, and diagonally hatched diamonds indicate variants. As in Figure 3A, methylation features include, but are not limited to, the methylation density within the genomic region 315, the percentage of highly methylated fragments that overlap with the genomic region 315, and the percentage of highly unmethylated fragments that overlap with the genomic region 315. For methylation density, the methylation state of all fragments 420 on the CpG site 317 is divided by the total number of CpG sites on the fragments 320. In this example, the methylation density is 0.63, or 63%. Fragment 1 320A, fragment 2 320B, and fragment 3 320C are hypermethylated, with over 80% of the fragments being hypermethylated (at least as shown in the figure). Fragment 4 320D is hyperunmethylated, with over 80% of the fragments being unmethylated (at least as shown in the figure). Fragment 5 320E and fragment 6 320F are mixed methylated, i.e., neither hypermethylated nor hyperunmethylated. Thus, one methylation feature, the percentage of hypermethylated fragments that overlap with genomic region 315, is 0.50, or 50%. Another methylation feature, the percentage of hyperunmethylated fragments that overlap with genomic region 315, is 0.17, or 17%. The fragments that overlap with genomic region 315 may be fragments that overlap with at least one of CpG sites 317. In some embodiments, the fragment that overlaps the genomic region 315 overlaps with at least some percentage of the CpG sites 317, for example, at least 20% of the CpG sites 317 of the genomic region 315.

[0107] FIG. 4A illustrates training 400 of an age prediction model according to one or more embodiments. The analysis system may perform some or all of training 400. In other embodiments, other components in FIGS. 6A and 6B may perform some or all of training 400. Training 400 results in a trained age prediction model that may input methylation features for a set of age-informative genomic regions and output a predicted age. The age prediction model in the training 400 process is similarly applicable to training other covariate prediction models. In this covariate embodiment, the analysis system utilizes training samples with reported values ​​for the covariate prediction model being trained.

[0108] The analysis system acquires 405 a plurality of training samples. The plurality of training samples may include (1) only cancer training samples, (2) only non-cancer training samples, or (3) a combination of cancer and non-cancer training samples. Each sensor training sample is collected from an individual with a confirmed cancer diagnosis. The cancer diagnosis can be confirmed before or after sample collection. In some examples, the type of cancer may be known. Each non-cancer training sample is collected from an individual who has not been diagnosed with any cancer and may be considered a generally healthy individual. Each training sample contains genetic material that can be sequenced and analyzed. In some embodiments, the sample is a blood sample containing nucleic acid fragments, such as cfDNA fragments. Additionally, each training sample includes the individual's reported chronological age. That is, each training sample is labeled with the chronological age of the subject from whom the sample was obtained. For example, if the cancer subject was obtained from a 39-year-old male, the cancer training sample would be labeled as being obtained from a 39-year-old male. In some cases, the labels may contain errors (e.g., due to swapping samples between subjects). For example, a non-cancer training sample may be labeled as coming from a 62-year-old woman when in fact it came from a 76-year-old man.

[0109] The analysis system sequences the nucleic acid fragments in each training sample to identify a methylation pattern for each nucleic acid fragment 410. The methylation pattern, as described above, represents the methylation status of CpG sites within DNA fragments in a genomic region in a sample. Methylation patterns may be determined for other samples or other sample populations. Sequencing may include bisulfite sequencing to convert unmethylated CpG sites. In other embodiments, a sequencer performs sequencing of the nucleic acid fragments, and the analysis system processes the sequence reads to identify the methylation pattern. The analysis system may further perform one or more processing steps on the sequence reads, such as removing duplicate copies of the same original fragment, identifying contaminating fragments 36, identifying sequencing errors, etc. The process of sequencing nucleic acid fragments and identifying methylation patterns has already been described with reference to Figures 2A and 2B.

[0110] For each genomic region, the analysis system calculates a predictability score between chronological age and the methylation patterns of nucleic acid fragments overlapping the genomic region 415. More generally, the predictability score represents the correlation between the subject's chronological age and the methylation pattern. The analysis system may identify one or more methylation features for each genomic region of the initial set of genomic regions. The initial set of genomic regions may be a broad set covering a majority of CpG sites in the human genome. For example, the analysis system may identify any combination of methylation features described in FIG. 3A for a single-CpG site genomic region, such as methylation density, percentage of overlapping hypermethylated fragments, and percentage of overlapping hypermethylated fragments. The analysis system can calculate a predictability score for a genomic region by training a regression of chronological age based on the methylation features in that genomic region. The analysis system trains the regression using the methylation features extracted for each training sample and the reported age of the training sample (e.g., the chronological age labeled on the sample). Types of regression include linear regression, logarithmic regression, exponential regression, multivariate regression, logistic regression, polynomial regression, Lasso regression, etc. From the trained regression, the analytics system may measure various metrics for use as indication scores. Exemplary metrics may include, but are not limited to, covariates, Pearson's correlation coefficient (or simply "Pearson's correlation"), R2, the sum of squares of residuals (RSS), the total sum of squares (TSS), the t-statistic for the slope, the two-sided p-value of the t-statistic (which may be adjusted for, e.g., multiple hypothesis testing), other statistical metrics for regression, etc.

[0111] By way of example, the analysis system may calculate a multivariate chronological age regression for a genomic region based on methylation density, percentage of overlapping hypermethylated fragments, percentage of overlapping hypermethylated fragments, etc. Based on the trained multivariate regression, the analysis system may identify a Pearson correlation, which may serve as an indicator score for that genomic region. In another example, the analysis system may calculate a linear chronological age regression for a genomic region based solely on methylation density and identify a Pearson correlation as the indicator score. In some embodiments, the analysis system may train multiple regressions for each genomic region based on different sets of methylation features under consideration, for example, a first regression based solely on methylation density, a second regression based solely on percentage of overlapping hypermethylated fragments, a third regression based solely on percentage of overlapping hypermethylated fragments, and a fourth regression based on multiple combinations of the above three methylation features.

[0112] The analysis system generates 420 feature sets of genomic regions based on the covariate scores for use as features in the age prediction model. To do so, the analysis system may calculate one or more indicator scores for each of the initial genomic regions, e.g., one indicator score for each methylation feature set. The methylation feature set achieving the highest absolute indicator score for a genomic region is identified as the most informative methylation feature set for that genomic region. In one embodiment, the analysis system uses a threshold absolute indicator score to identify genomic regions to use as part of the feature set for the genomic region. For example, the threshold absolute indicator score is 0.5, and thus indicator scores greater than 0.5 or less than -0.5 (e.g., indicator scores with an absolute value greater than 0.5) exceed the threshold absolute indicator score. The analysis system then generates a feature vector including the identified genomic regions with indicator scores greater than the threshold. For example, the feature vector may include all genomic regions with absolute indicator scores greater than 0.5 or less than -0.5. In other embodiments, the analysis system ranks genomic regions based on their highest indicator scores. From this ranking, the analysis system can select enough genomic regions that use the full budget of genomic regions, e.g., that are targeted by the assay panel, and generate feature vectors accordingly.

[0113] In additional embodiments, the analysis system may further consider one or more other factors of the genomic region to identify a feature set for the genomic region, which is then included in the generated feature vector. For example, the analysis system may further consider a balance of genomic regions showing negative and positive correlations. Negatively correlated genomic regions reflect increasing methylation feature values ​​that correlate with decreasing age values. Positively correlated genomic regions reflect increasing methylation feature values ​​that correlate with increasing age values. The analysis system may further consider the rate of change between age and methylation features. For example, some genomic regions may show a flat correlation with age (small differences in methylation features relative to an individual's age), some genomic regions may show a steady correlation with age (moderate differences in methylation features relative to an individual's age), and some genomic regions may show a clear correlation with age (large differences in methylation features relative to an individual's age). The analysis system may further consider the location of the genomic region in the human genome, for example, to ensure an appropriate spread of the genomic region across the human genome. This prevents situations where all locations of age-informative genomic regions are identified in one section of the genome, which may, for any number of reasons, have sparse signal in the test sample to which age prediction is applied, causing age prediction to fail if all age-informative genomic regions were located in that section.

[0114] In some embodiments, the analysis system utilizes penalized regression to identify feature sets for genomic regions and generate feature vectors accordingly. The penalization process aims to optimize the feature set utilized to the smallest feature set that still provides optimal predictive power. In other embodiments, relaxed lasso regression is utilized to achieve similar results.

[0115] In some embodiments, the analysis system reduces the feature set of genomic regions to those genomic regions that have a high correlation with the cancer signal 425. That is, the analysis system may remove or exclude genomic regions from the feature vector or feature set that do not show a high correlation with the cancer signal. For example, the analysis system may separately identify genomic regions that correlate with the cancer signal (or other disease signals). The analysis system may then identify genomic regions that intersect with the correlation with age and the correlation with the cancer signal. One method for identifying genomic regions that correlate with the cancer signal is disclosed below (e.g., with respect to Figures 6A and 6B). Using genomic regions that intersect with the correlation with age and the correlation with the cancer signal allows for efficient use of the budget for target regions on an assay panel. That is, the panel is created with fewer probes that target regions that are more indicative of the presence of age and / or cancer. Essentially, probes in a panel that target regions that are strongly correlated with cancer and / or age are "worthier" to include in the panel than probes that show a weak correlation with cancer and / or age. Furthermore, utilizing intersecting genomic regions may prove advantageous in utilizing predicted age as a feature in cancer classification.

[0116] At this point, the analysis system has identified a feature set of genomic regions to use in training the age prediction model and generated a corresponding feature vector. Each genomic region in the feature set (included in the feature vector) may include one CpG site, multiple CpG sites, or multiple groups of CpG sites. Each genomic region in the feature set has a specific methylation feature set, which may differ from other genomic regions. For example, a first genomic region covers one CpG site and considers the methylation density at that CpG site, while a second genomic region covers multiple CpG sites and considers both the percentage of overlapping hypermethylated fragments and the percentage of overlapping hypermethylated fragments. In another embodiment, the analysis system penalizes to minimize the number of genomic regions that maintain the accuracy of age prediction. Penalization is a factor that negatively impacts age prediction based on the number of genomic regions. Penalization forces the analysis system to identify the minimum number of genomic regions that maintain optimal performance of the age prediction model.

[0117] The analysis system trains an age prediction model based on the feature set of genomic regions and the methylation patterns of nucleic acid fragments from the training samples 430. For each training sample, the analysis system identifies the numerical values ​​of methylation features of the feature set of genomic regions. In some embodiments, the analysis system trains the age prediction model as a machine learning model. Exemplary machine learning models include linear regression, logarithmic regression, exponential regression, multivariate regression, logistic regression, polynomial regression, Lasso regression, etc. The analysis system may train multiple age prediction models with different feature sets of genomic regions and evaluate performance between the various models. For example, a first model is trained with a small feature set of genomic regions, and a second model is trained with a large feature set of genomic regions that includes the small feature set. The analysis system evaluates the performance of the two age prediction models using a validation set of training samples.

[0118] 4B illustrates deployment 440 of a chronological age prediction model, according to one or more embodiments. The analysis system may perform some or all of deployment 440. In other embodiments, other components of FIGS. 6A and 6B may perform some or all of deployment 440. Utilizing 440 the age prediction model includes determining an age prediction for an individual associated with a test sample based on methylation features for a feature set of genomic regions of the test sample.

[0119] The analysis system obtains 445 a test sample having a plurality of nucleic acid fragments and a reported age (a label representing the chronological age of the subject from whom the sample was taken). A physician or other healthcare provider collects the test sample and may also obtain the reported age of the individual providing the test sample. In some embodiments, the age may be a single number or an age range. For example, an individual may report their age as 47 years old or an age range of 40-50 years old. The sample may be any type of biological sample containing an individual's nucleic acid material. In blood sample embodiments, the blood sample includes at least cfDNA fragments sheared from cells.

[0120] The analysis system sequences the nucleic acid fragments of the test sample and identifies the methylation pattern of each nucleic acid fragment 450. Sequencing can include bisulfite sequencing, which converts unmethylated CpG sites. In other embodiments, a sequencer performs the sequencing of the nucleic acid fragments, and the analysis system processes the sequence reads to identify the methylation pattern. The analysis system can further perform one or more processing steps on the sequence reads, such as removing duplicate copies of the same original fragment, identifying contaminating fragments, identifying sequencing errors, etc. The process of sequencing nucleic acid fragments and identifying methylation patterns has already been described in Figures 2A and 2B.

[0121] The analysis system applies the trained age prediction model to predict chronological age of the test sample based on the methylation patterns of the nucleic acid fragments of the test sample 455. The analysis system identifies methylation features for the age prediction model trained, for example, in process 400 of FIG. 4A. The age prediction model is configured to receive as input the methylation features of the feature set of genomic regions and output a predicted chronological age based on the methylation features. Like the reported age, the predicted age can be a single number or an age range.

[0122] The analysis system compares the predicted age to the reported age 460. The comparison can be a determination of whether the predicted age matches the reported age. For example, if the reported age is in the age range of 20-30 years and the predicted age is 26 years or in the range of 20-30 years, the predicted age matches the reported age range. In some embodiments, the comparison can be a residual, which is the difference between the predicted age and the reported age. For example, if the reported age is 63 years and the predicted age is 72 years, the residual is 9 years relative to the reported age. The residual can also be an absolute value, such as a 9-year difference from the reported age.

[0123] The analysis system uses the predicted age to proceed with the analysis. In some embodiments, the analysis system may perform sample swap validation 465. Sample swap validation, in this context, can refer to identifying whether a test sample was correctly labeled. For example, a "sample swap" may occur when a sample obtained from a 54-year-old woman is labeled as having been obtained from a 24-year-old man. However, more generally, identifying a sample swap is more akin to identifying that the reported age of a test sample is incorrect.

[0124] In some configurations, the analysis system may use the residual to identify whether the sample has been swapped and thereby identify that the test sample does not actually originate from the individual expected to be associated with the sample. The analysis system may call or identify a sample swap if the predicted age differs from the reported age. In other embodiments, the analysis system may call a sample swap if the residual difference is greater than a threshold difference. For example, the residual threshold may be set at 10 years, and the analysis system may call a sample swap if the residual difference between the predicted age and the reported age is greater than the 10-year residual threshold. In yet other embodiments, the analysis system may call a sample swap based on a comparison of the predicted age and the reported age and other analyses. For example, the analysis system may train a separate model for ethnicity identification to identify whether the predicted ethnicity matches the individual's reported ethnicity, or a separate model for gender identification to identify whether the predicted gender matches the individual's reported gender. When a sample swap is called, the analysis system may prevent the sample from proceeding further into downstream analysis. Samples for which a sample swap is not called may be proceeded further into downstream analysis. For example, when a sample swap is called for a training sample, the analytical system may avoid using that training sample in training one or more models or constructing one or more distributions. As another example, when a sample swap is called for a test sample, the analytical system may withhold a cancer prediction for that test sample.

[0125] The analysis system uses the comparison of predicted age to reported age as part of the cancer classification 470. In one or more embodiments, the analysis system uses the residual of predicted age relative to reported age as a feature for cancer classification, e.g., along with other features extracted from the sequencing data. For example, the residual may indicate an abnormally high or low amount of methylation in the test sample (leading to a high residual), which indicates the presence or absence of cancer, as discussed above. As such, a high residual may be used as an indicative feature in identifying whether the test sample has cancer.

[0126] In some embodiments, the analysis system may compare the residual to a residual threshold to identify whether a sample has a strong likelihood of cancer. The analysis system may set the residual threshold using a set of training samples. The analysis system identifies methylation features for a feature set of genomic regions for each training sample. The analysis system inputs the methylation features for each training sample into an age prediction model to identify a predicted age for each training sample. The analysis system may calculate the residual for each training sample by calculating the difference between the predicted age and the reported age. The analysis system may set a residual threshold that includes a majority of the training samples. For example, the analysis system may want to use a residual threshold that captures 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, 99.5%, or 99.9% of the training samples. In effect, setting the threshold residual trains the model to recognize samples with methylation patterns that allow errors in age prediction to be associated with the presence of cancer or non-cancer. The residual threshold is therefore used as a measure to determine whether the amount of difference in methylation patterns is due to age or whether it is due to cancer. With that in mind, if a test sample has residuals that fall outside the residual threshold, the analytical system can make an initial determination that the test sample has a strong likelihood of cancer. Depending on the configuration, a strong likelihood can indicate that the sample is more likely to have at least a 60%, 65%, 70%, 80%, 90%, etc. probability of containing cancer than not indicating cancer, or may have an indicative score above the threshold. The analytical system can proceed to cancer classification to confirm the initial determination.

[0127] III.B. Identification of Aberrant Fragments The analysis system can identify aberrant fragments of a sample using the methylation state vector of the sample. For each fragment in a sample, the analysis system can identify whether the fragment is an aberrant fragment using the methylation state vector corresponding to the fragment. In some embodiments, the analysis system calculates a p-value for each methylation state vector, which indicates the likelihood of observing that methylation state vector or other methylation state vectors that are much less likely to be seen in healthy controls. In some examples, the p-value score can be adjusted for multiple hypothesis testing, for example, by controlling the false positive rate, family-wise error rate, false discovery rate, etc. The process of calculating the p-value score is further described in Section "II. Cp Value Filtering" below. The analysis system can identify fragments with methylation state vectors that have p-value scores below a threshold as aberrant fragments. In some embodiments, the analysis system further labels fragments with at least some CpG sites with a percentage of methylation or unmethylation above either threshold as hypermethylated and hypomethylated fragments, respectively. Hypermethylated and hypomethylated fragments may also be referred to as extremely methylated aberrant fragments (UFXM). In other embodiments, the analysis system may implement various other probabilistic models for identifying aberrant fragments. Examples of other probabilistic models include mixture models, deep probabilistic models, etc. In some embodiments, the analysis system may use any combination of the processes described below for identifying aberrant fragments. Using the identified aberrant fragments, the analysis system may filter methylation state vectors for the sample for use in other processes, such as in training and developing a cancer classifier.

[0128] III.C. I .P-value filtering In some embodiments, the analysis system calculates a p-value score for each methylation state vector compared to methyl state vectors from fragments in a healthy control group. The p-value score can indicate the likelihood of observing a methylation state that matches that methylation state vector or other methylation states that are much less likely in the healthy control group. To identify a DNA fragment as abnormally methylated, the analysis system can use a healthy control group in which the majority of the fragments are normally methylated. When performing this probabilistic analysis to identify abnormal fragments, the identification can be weighted relative to a group of control subjects that constitute the healthy control group. To ensure the robustness of the healthy control group, the analysis system can select a threshold number of healthy individuals from which to draw samples containing the DNA fragment. Figure 5A below describes how the analysis system can generate a healthy control group data structure from which the analysis system can calculate a p-value score. Figure 5B describes how to calculate a p-value score using the generated data structure.

[0129] 5A is a flowchart illustrating a process 500 for generating a healthy control data structure, according to one embodiment. To create the healthy control data structure, an analysis system can receive multiple DNA fragments (e.g., cfDNA) from multiple healthy individuals. The analysis system can generate 505 a methylation state vector for each fragment, for example, via process 200.

[0130] Using the methylation state vector of each fragment, the analysis system can subdivide the methylation state vector into strings of CpG sites 510. In some embodiments, the analysis system subdivides the methylation state vector so that all resulting strings are shorter than a certain length 510. For example, subdividing a methylation state vector of length 11 into strings of lengths 3 or less results in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, subdividing a methylation state vector of length 7 into strings of lengths 4 or less results in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If the methylation state vector is shorter than or equal to a specified string length, the methylation state vector can be converted into a single string containing all of the CpG sites in the vector.

[0131] For each possible CpG site and possible methylation state in the vector, the analysis system tallies the strings by counting the number of strings present in the control set that have the CpG site designated as the first CpG site in the string and have that possible methylation state 515. For example, at a CpG site, given a string length of 3, there are 2^3 or 8 possible string configurations. For each of the 8 possible string configurations at that CpG site, the analysis system tallies 510 how many occurrences of each possible methylation state vector are seen in the control set. Continuing with this example, this may include tallying the following quantities: <M x ,M x+1 ,M x+2 >, <M x ,M x+1 ,U x+2 >,..., x ,U x+1 ,U x+2 The analysis system creates a data structure 515 that stores the aggregated counts for each starting CpG site and string probability. ​

[0132] Setting an upper limit on string length has several advantages. First, depending on the maximum length of a string, the size of the data structure created by the analysis system can be significantly increased. For example, a maximum string length of 4 means that all CpG sites have at least 2^4 counts to be tallied for strings of length 4. Increasing the maximum string length to 5 means that all CpG sites have an additional 2^4, or 16, counts to be tallied, doubling the number of counts (and the required computer memory) compared to the previous string length. Reducing the string size can help keep the creation and performance of the data structure (e.g., for subsequent access, as described below) at a reasonable level in terms of computation and storage. Second, a statistical consideration for limiting the maximum string length may be to avoid overfitting of downstream models that use string counts. If long strings of CpG sites do not biologically have a strong impact on the outcome (e.g., predicting an abnormality that predicts the presence of cancer), calculating probabilities based on large strings of CpG sites can be problematic because it uses a large amount of data that may not be available and therefore may be too sparse for the model to function properly. For example, a calculation of the probability of abnormality / cancer conditioned on the previous 100 CpG sites can use counts of strings of length 100 in a data structure, ideally some that exactly match the methylation status of the previous 100. If only sparse counts of strings of length 100 are available, there may be insufficient data to identify whether any of the 100 string lengths in a test sample are abnormal or not.

[0133] 5B is a flowchart illustrating a process 530 for identifying aberrantly methylated fragments from an individual, according to one embodiment. In process 530, the analysis system generates 540 methylation state vectors from the subject's cfDNA fragments, e.g., via process 200. The analysis system can treat each methylation state vector as follows:

[0134] For a given methylation state vector, the analysis system enumerates all possible methylation state vectors that have the same starting CpG site and the same length within the methylation state vector (i.e., a group of CpG sites). Since each methylation state is generally either methylated or unmethylated, there are actually two possible states at each CpG site, and therefore the count of different possibilities for a methylation state vector may depend on a power of 2, so a methylation state vector of length n is 2 of the methylation state vector. n A methylation state vector that includes intermediate states for one or more CpG sites may allow the analysis system to enumerate the possibilities of the methylation state vector, considering only CpG sites with the observed state 530.

[0135] The analysis system calculates the probability of observing each possible methylation state vector for the identified starting CpG site and methylation state vector length by accessing a healthy control data structure 550. In some embodiments, the calculation of the probability of observing a certain possibility uses Markov chain probability to model joint probability calculations. The Markov model can be trained, at least in part, based on an evaluation of the methylation state of each CpG site at a plurality of corresponding CpG sites of each fragment (e.g., nucleic acid methylation fragment) across nucleic acid methylation fragments in a healthy non-cancer cohort dataset having a corresponding plurality of CpG sites. For example, a Markov model (e.g., a hidden Markov model, or HMM) is used to consider a probability set that specifies, for each state in a sequence, the likelihood of observing the next state in the sequence, and to determine the probability that a sequence of methylation states (e.g., including "M" or "U") can be observed for a nucleic acid methylation fragment among a plurality of nucleic acid methylation fragments. The probability set can be obtained by training an HMM. Such training may include calculating statistical parameters (e.g., the probability that a first state can transition to a second state (transition probability) and / or the probability that a methylation state can be observed for each CpG site (output probability)) given an initial training dataset of observed methylation state sequences (e.g., methylation patterns). HMMs may be trained using supervised training (e.g., using samples where the base sequences and observed states are known) and / or unsupervised training (e.g., Viterbi learning, maximum likelihood estimation, expectation-maximization training, and / or Baum-Welch training). In other embodiments, computational methods other than Markov chain probabilities are used to identify the probability of observing each possible methylation state vector. For example, such computational methods may include learned representations. The p-value threshold may be between 0.01 and 0.10, or between 0.03 and 0.06. The p-value threshold may be 0.05. The p-value threshold may be less than 0.01, less than 0.001, or less than 0.0001.

[0136] The analysis system uses the calculated probability for each possibility to calculate a p-value score for the methylation state vector 555. In some embodiments, this involves identifying a calculated probability that corresponds to the likelihood of matching the methylation state vector in question. Specifically, this can be the likelihood of having the same CpG group, or equivalently, the same starting CpG site and length as the methylation state vector. The analysis system can sum the calculated probabilities of any possibilities that have a probability lower than or equal to the identified probabilities to generate a p-value score.

[0137] This p-value can represent the probability of observing the methylation state vector of that fragment, or other methylation state vectors that are much less likely in healthy controls. A low p-value score can therefore generally correspond to a methylation state vector that is rare in healthy individuals, resulting in the fragment being labeled as abnormally methylated relative to the healthy control group. A high p-value score can generally relate to a methylation state vector that is expected to be present in healthy individuals in a relative sense. For example, if the healthy control group is a non-cancer group, a low p-value can indicate a fragment that is abnormally methylated relative to the non-cancer group and thus may indicate the presence of cancer in the subject.

[0138] As described above, the analysis system can calculate a p-value score for each of a plurality of methylation state vectors, each representing a cfDNA fragment in the test sample. To identify which of the fragments are abnormally methylated, the analysis system can filter the methylation state vectors based on their p-value scores 565. In some embodiments, filtering is performed by comparing the p-value scores to a threshold and retaining only fragments below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001, 0.0001, or even smaller.

[0139] Exemplary results from process 500 indicate that the analysis system can obtain a median (range) of 2,800 (1,500-12,000) fragments with aberrant methylation patterns for non-cancer training participants and a median (range) of 3,000 (1,200-420,000) fragments with aberrant methylation patterns for cancer training participants. These filtered fragments with aberrant methylation patterns can be used for downstream analysis, as described below in Section III.

[0140] In some embodiments, the analysis system uses a sliding window to identify possibilities for the methylation state vector and calculate p-values ​​560. Rather than enumerating possibilities and calculating p-values ​​for the entire methylation state vector, the analysis system can count possibilities and calculate p-values ​​only for a window of contiguous CpG sites, where the window is shorter in length (of CpG sites) than at least some of the fragments (otherwise the window would serve no purpose). The length of the window may be fixed, user-determined, dynamic, or otherwise selected.

[0141] When calculating p-values ​​for a methylation state vector longer than a window, the window can identify contiguous CpG sites from the vector within the window starting from the first CpG site in the vector. The analysis system can calculate a p-value score for the window that includes the first CpG site. The analysis system can then "slide" the window to a second CpG site in the vector and calculate another p-value score for the second window. Thus, for a window size of l and a methylation vector length of m, each methylation state vector can generate m-1+l p-value scores. After completing the p-value calculations for each portion of the vector, the lowest p-value score from all sliding windows can be used as the overall p-value score for that methylation state vector. In another embodiment, the analysis system adds up the p-value scores of the methylation state vectors to generate an overall p-value score.

[0142] The use of a sliding window can help reduce the number of possible methylation state vectors enumerated and their corresponding probability calculations that would otherwise need to be performed. To take a realistic example, a fragment may have upwards of 54 CpG sites. Instead of calculating probabilities for 2^54 (approximately 1.8 x 10^16) possibilities to generate a single p-score, the analysis system can instead use a window of size 5 (for example), resulting in 50 p-value calculations for each of the 50 windows of the fragment's methylation state vector. Each of the 50 calculations can enumerate 2^5 (32) possible methylation state vectors, resulting in an overall resulting probability calculation of 50 x 2^5 (1.6 x 10^3). This significantly reduces the number of calculations that must be performed, without significantly impacting the accurate identification of aberrant fragments.

[0143] In embodiments with intermediate states, the analysis system may calculate a p-value score summing out CpG sites with intermediate states in the methylation state vector of a fragment. The analysis system may identify all possibilities that are consistent with all methylation states in the methylation state vector, excluding the intermediate states. The analysis system may assign that possibility to the methylation state vector as the sum of the probabilities of the identified possibilities. In one example, the analysis system may:<M1,I2,U3> The probability of the methylation state vector of<M1,M2,U3> and<M1,U2,U3> , because the methylation states for CpG sites 1 and 3 are observed and match the methylation states of the fragment at CpG sites 1 and 3. This method of summing out CpG sites with intermediate states can use a calculation of probabilities for up to 2^i possibilities, where i denotes the number of intermediate states in the methylation state vector. In additional embodiments, a dynamic programming algorithm can be used to calculate the probability of a methylation state vector having one or more intermediate states. Advantageously, the dynamic programming algorithm operates in linear computation time.

[0144] In some embodiments, the computational burden of calculating the probabilities and / or p-value scores may be further reduced by temporarily storing at least some of the calculations. For example, the analysis system may temporarily store probability calculations for the likelihood of a methylation state vector (or a window thereof) in temporary or persistent memory. When other fragments have the same CpG sites, temporarily storing the likelihood probabilities allows for efficient calculation of p-score values ​​without recalculating the underlying likelihood probabilities. Equivalently, the analysis system may calculate a p-value score for each of the methylation state vector possibilities associated with a group of CpG sites from the vector (or its window). The analysis system may temporarily store the p-value scores for use in determining p-value scores for other fragments containing the same CpG sites. In general, the p-value score of a methylation state vector possibility with the same CpG site may be used to determine a p-value score for another possibility from the same group of CpG sites.

[0145] In some instances, p-value scores may be adjusted for multiple hypothesis testing according to various suitable techniques. By way of example and without limitation, such techniques may include controlling for false positive rate, family-wise error rate, per-experiment error rate, false discovery rate, etc. Known applicable techniques for p-value adjustment include, but are not limited to, Bonferroni, Holm, Hochberg, harmonic mean p-value, Benjamini-Hochberg, Benjamini-Yekutieli, Storey-Tibshirani, etc. Adjusting p-value scores for multiple hypothesis testing can be used to improve the accuracy associated with the underlying positive detection call based on the p-value and reduce the incidence of false positives.

[0146] The one or more nucleic acid methylation fragments can be filtered before training a region model or a cancer classifier. Filtering the nucleic acid methylation fragments can include removing each of the respective nucleic acid methylation fragments that does not satisfy one or more selection criteria (e.g., below or above one selection criterion) from the corresponding plurality of nucleic acid methylation fragments. The one or more selection criteria can include a p-value threshold. The output p-value for each nucleic acid methylation fragment can be determined, at least in part, based on comparing the corresponding methylation pattern of each nucleic acid methylation fragment with a corresponding distribution of methylation patterns of nucleic acid methylation fragments having the corresponding plurality of CpG sites of each nucleic acid methylation fragment in a healthy non-cancer cohort dataset.

[0147] Filtering the plurality of nucleic acid methylation fragments can include removing each of the respective nucleic acid methylation fragments that does not meet a p-value threshold. A filter can be applied to each of the methylation patterns of each of the respective nucleic acid methylation fragments using the methylation patterns observed across the first plurality of nucleic acid methylation fragments. Each of the respective methylation patterns of each of the respective nucleic acid methylation fragments (e.g., fragment 1, ..., fragment N) can include one or more corresponding methylation sites (e.g., CpG sites) identified by a methylation site identifier and a corresponding methylation pattern represented as a sequence of 1's and 0's, where each "1" represents a methylated CpG site in the one or more CpG sites and each "0" represents an unmethylated CpG site in the one or more CpG sites. The methylation patterns observed across the first plurality of nucleic acid methylation fragments can be used to construct a methylation state distribution for the CpG site states (e.g., CpG site A, CpG site B, ..., CpG site ZZZ) collectively represented by the first plurality of nucleic acid methylation fragments. Further details regarding the processing of nucleic acid methylated fragments are disclosed in U.S. Patent Application No. 17 / 191,914, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," filed March 4, 2021, the entire contents of which are incorporated herein by reference.

[0148] Each nucleic acid methylation fragment may not satisfy one of the selection criteria among the one or more selection criteria if the nucleic acid methylation fragment has an abnormal methylation score lower than the abnormal methylation score threshold. In this situation, the abnormal methylation score can be determined using a mixture model. For example, the mixture model can detect abnormal methylation patterns in nucleic acid methylation fragments by determining the likelihood of a methylation state vector (e.g., a methylation pattern) for each nucleic acid methylation fragment based on the number of possible methylation state vectors of the same length at the same corresponding genomic location. This can be done by generating multiple possible methylation states for a vector of a particular length at each genomic location in the reference genome. Using the multiple possible methylation states, the total number of possible methylation states and therefore the probability of each predicted methylation state at that genomic location can be determined. The likelihood of a sample nucleic acid methylation fragment corresponding to a genomic location in the reference genome can then be determined by matching the sample nucleic acid methylation fragment with the predicted (e.g., possible) methylation state and deriving the calculated probability of the predicted methylation state. An abnormal methylation score can then be calculated based on the probability of the same nucleic acid methylation fragment.

[0149] Each nucleic acid methylated fragment may not satisfy a selection criterion among the one or more selection criteria if the respective nucleic acid methylated fragment has a residual count lower than a threshold. The threshold residual count can be 10 to 50, 50 to 100, 100 to 150, or more than 150. The threshold residual count can be a fixed value between 20 and 90. Each nucleic acid methylated fragment may not satisfy a selection criterion among the one or more selection criteria if the respective nucleic acid methylated fragment has a number of CpG sites lower than a threshold. The threshold number of CpG sites can be 4, 5, 6, 7, 8, 9, or 10. Each nucleic acid methylated fragment may not satisfy a selection criterion among the one or more selection criteria if the genomic start and end positions of the respective nucleic acid methylated fragment indicate that the respective nucleic acid methylated fragment represents fewer than a threshold number of nucleotides in the human genome reference sequence.

[0150] Filtering can remove nucleic acid methylated fragments from the corresponding plurality of nucleic acid methylated fragments that have the same corresponding methylation pattern and the same corresponding genomic start and end positions as other nucleic acid methylated fragments from the corresponding plurality of nucleic acid methylated fragments. This filtering step can remove redundant fragments that are exact duplicates, including PCR duplicates in some examples. Filtering can remove nucleic acid methylated fragments that have the same corresponding genomic start and end positions as other nucleic acid methylated fragments from the corresponding plurality of nucleic acid methylated fragments and fewer than a threshold number of different methylation states. The threshold number of different methylation states used to retain nucleic acid methylated fragments can be 1, 2, 3, 4, 5, 6, or more. For example, a first nucleic acid methylated fragment that has the same corresponding genomic start and end positions as a second nucleic acid methylated fragment but has at least 1, at least 2, at least 3, at least 4, or at least 5 different methylation states at each CpG site (e.g., aligned to a reference genome) is retained. As another example, a first nucleic acid methylated fragment having the same methylation state vector (e.g., methylation pattern) as a second nucleic acid methylated fragment, but having different corresponding genomic start and end positions, is also retained.

[0151] Filtering can remove assay artifacts in multiple nucleic acid methylation fragments. Removing assay artifacts can include removing sequence reads obtained from sequenced hybridization probes and / or sequence reads obtained from sequences where conversion did not occur during bisulfite conversion. Filtering can remove contaminants (e.g., from sequencing, nucleic acid isolation, and / or sample preparation).

[0152] The filtering can remove a subset of methylation fragments from the plurality of methylation fragments based on mutual information filtering of each methylation fragment relative to the cancer state across a plurality of untrained individuals. For example, mutual information can provide a measure of interdependence between two simultaneously sampled states of interest. Mutual information can be determined by selecting independent CpG sites (e.g., among all or a portion of the nucleic acid methylation fragments) from one or more datasets and comparing the probability of the methylation state for that CpG site group between two sample groups (e.g., subsets and / or groups of genotype datasets, biological samples, and / or subjects). The mutual information score can indicate the probability of a methylation pattern for a first state versus a second state in each region within each frame of the sliding window, and thus indicates the discriminatory power of each region. The mutual information score can be similarly calculated for each region in each frame of the sliding window as it progresses across the selected CpG sites and / or selected genomic regions. Further details regarding mutual information filtering are disclosed in U.S. Patent Application No. 17 / 119,606, entitled "Cancer Classification using Patch Convolutional Neural Networks," filed December 11, 2020, the entire contents of which are incorporated herein by reference.

[0153] III.C. II Hypermethylated and hypomethylated fragments In some embodiments, the analysis system identifies hypomethylated or hypermethylated fragments as aberrant fragments from the filtered set. The analysis system identifies hypermethylated fragments with more than a threshold number of CpG sites and a threshold percentage of methylated CpG sites. The analysis system identifies hypomethylated fragments with more than a threshold number of CpG sites and a threshold percentage of unmethylated CpG sites. Exemplary threshold fragment (or CpG site) length include more than 3, 4, 5, 6, 7, 8, 9, 10, etc. Exemplary methylated or unmethylated threshold percentages include more than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50-100%.

[0154] III.C. Training a Cancer Classifier 6A is a flowchart illustrating a process 600 for training a cancer classifier, according to one embodiment. The analysis system obtains 610 a plurality of training samples, each having a set of abnormal fragments and a cancer type label. The plurality of training samples can include any combination of samples from healthy individuals with a general label of "non-cancer," or samples from subjects with a general or specific label of "cancer" (e.g., "breast cancer," "lung cancer," etc.). Training samples from subjects with a single cancer type can be referred to as a cohort for that cancer type or cancer type cohort.

[0155] The analysis system identifies a feature vector for each training sample based on the set of aberrant fragments for that training sample 620. The analysis system can calculate an aberration score for each CpG site in the initial set of CpG sites. The initial set of CpG sites can be all CpG sites in the human genome or a portion thereof, which can be 10 4 , 10 5 , 10 6 , 10 7 , 10 8and so on. In one embodiment, the analysis system determines an anomaly score for the feature vector using binary scoring based on whether there is an anomaly fragment encompassing the CpG site within the set of anomaly fragments. In other embodiments, the analysis system determines an anomaly score based on a count of anomaly fragments that overlap the CpG site. In one example, the analysis system may use a ternary scoring system that assigns a first score for the absence of an anomaly fragment, a second score for the presence of a small number of anomaly fragments, and a third score for the presence of more than a small number of anomaly fragments. For example, the analysis system may count five anomaly fragments in a sample that overlap a CpG site and calculate an anomaly score based on the count of five. In one or more embodiments, the feature vector further includes one or more features based on chronological age prediction (e.g., covariate prediction) described with respect to FIGS. 4A and 4B . For example, the feature vector may include an age residual, which is the difference between predicted chronological age (e.g., by applying a trained age prediction model) and reported chronological age. In other examples, the feature vector may include other features based on predictive covariates, such as one or more of the predictive covariates. In some embodiments, the feature vector further includes one or more methylation features from the feature set evaluated in chronological age prediction (e.g., the feature set of the chronological age prediction model identified in step 420 of FIG. 4A).

[0156] Once all anomaly scores have been identified for the training sample, the analysis system can identify a feature vector as a vector of elements, each of which contains one of the anomaly scores associated with one of the CpG sites in the initial set. The analysis system can normalize the anomaly scores of the feature vector based on the coverage of the sample. Here, coverage can refer to the mean or average sequencing depth of all CpG sites covered by the initial set of CpG sites used in the classifier, or can be based on the set of anomaly fragments of a training sample.

[0157] By way of example, refer now to FIG. 6B, which illustrates a training feature vector matrix 622. In this example, the analysis system has identified a CpG site [K] 626 to consider when generating a feature vector for a cancer classifier. The analysis system selects a training sample [N] 624. The analysis system determines a first anomaly score 628 for a first arbitrary CpG site [k1] to be used in the feature vector of the training sample [n1]. The analysis system checks each aberrant fragment within the set of aberrant fragments. If the analysis system identifies at least one aberrant fragment that includes the first CpG site, the analysis system determines the first aberration score 628 for the first CpG site as 1, as shown in FIG. 6B. Considering a second arbitrary CpG site [k2], the analysis system similarly checks the set of aberrant fragments for at least one that includes the second CpG site [2]. If the analysis system does not find such an abnormal fragment containing the second CpG site, the analysis system identifies the second anomaly score 629 of the second CpG site [k2] as 0, as shown in Figure 6B. Once the analysis system has identified all of the anomaly scores for the initial group of CpG sites, the analysis system identifies a feature vector for the first training sample [n1] that includes an anomaly score, with the feature vector including the first anomaly score 628 of the first CpG site [k1] as 1 and the second anomaly score 629 of the second CpG site [k2] as 0, and the subsequent anomaly scores, thus forming the feature vector [1, 0, ...].

[0158] Other sample feature identification methods can be found in U.S. patent application Ser. No. 15 / 931,022, entitled "Model-Based Featurization and Classification," U.S. patent application Ser. No. 16 / 579,805, entitled "Mixture Model for Targeted Sequencing," U.S. patent application Ser. No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification," and U.S. patent application Ser. No. 16 / 723,716, entitled "Source of Origin Deconvolution Based on Methylation Fragments in Cell-Free DNA Samples," all of which are incorporated by reference in their entireties.

[0159] The analysis system may further limit the CpG sites considered for use in the cancer classifier. For each CpG site in the initial set of CpG sites, the analysis system calculates 630 an information gain based on the feature vectors of the training samples. From step 620, each training sample has a feature vector that may include anomaly scores for all CpG sites in the initial set of CpG sites, which may include up to all CpG sites in the human genome. However, some CpG sites in the initial set of CpG sites may be less informative than others in distinguishing between cancer types or may overlap with other CpG sites.

[0160] In one embodiment, the analysis system calculates an information gain for each CpG site in each cancer type and initial group to determine whether to include the CpG site in the classifier 630. The information gain is calculated for a training sample of a cancer type compared to all other samples. For example, two random variables are used: "abnormal fragment" ("AF") and "cancer type" ("CT"). In one embodiment, AF is a binary variable indicating whether there is an abnormal fragment that overlaps with a CpG site in a sample, as identified for the anomaly score / feature vector above. CT is a random variable indicating whether the cancer is of a particular type. The analysis system calculates the mutual information of an AF with respect to the CT, i.e., how many bits of information about the cancer type are obtained when knowing whether there is an abnormal fragment that overlaps with a particular CpG site. In practice, for a first cancer type, the analysis system calculates the pairwise mutual information gain for each other cancer type and sums the mutual information gains across all other cancer types.

[0161] For a given cancer type, the analysis system can use this information to rank CpG sites based on how cancer-specific they are. This procedure can be repeated for all cancer types under consideration. If a particular region is commonly aberrantly methylated in training samples of one cancer but not in training samples of other cancer types or healthy training samples, the CpG sites where these aberrant fragments overlap may have high information gain for that cancer type. The ranked CpG sites for each cancer type can be added (selected) en masse to a set of CpG sites selected based on their rank for use in a cancer classifier.

[0162] In additional embodiments, the analysis system may consider other selection criteria for selecting informative CpG sites to be used in the cancer classifier. One selection criterion may be that the selected CpG sites exceed a threshold separation from other selected CpG sites. For example, the selected CpG sites may be more than a threshold number of base pairs (e.g., 100 base pairs) away from any other selected CpG sites, such that neither CpG site within the threshold separation is selected for consideration in the cancer classifier.

[0163] In one embodiment, the set of selected CpG sites from the initial set may cause the analysis system to modify the feature vector of the training sample as needed 650. For example, the analysis system may truncate the feature vector to remove anomaly scores corresponding to CpG sites that are not within the set of selected CpG sites.

[0164] Using the feature vectors of the training samples, the analysis system may train a cancer classifier in any of a variety of ways. The feature vectors may correspond to the initial set of CpG sites from step 620 or to the set of selected CpG sites from step 650. In one embodiment, the analysis system trains 660 a binary cancer classifier to distinguish between cancer and non-cancer based on the feature vectors of the training samples. In this manner, the analysis system uses training samples that include both non-cancer samples from healthy individuals and cancer samples from subjects. Each training sample can have one of two labels: "cancer" or "non-cancer." In this embodiment, the classifier outputs a cancer prediction indicating the likelihood of the presence or absence of cancer.

[0165] In another embodiment, the analysis system trains 670 a multi-class cancer classifier that distinguishes between many cancer types (also called tissue of origin (TOO) labels). The cancer types can include one or more cancers and can include non-cancer types (including additional other diseases or genetic disorders, etc.). To do so, the analysis system can use a cancer type cohort, which may or may not include a non-cancer type cohort. In this multi-cancer embodiment, the cancer classifier is trained to identify cancer predictions (or, more specifically, TOO predictions) that include a predictive value for each of the classified cancer types. The predictive value may correspond to the likelihood that a given training sample (and, during estimation, a test sample) has each of the cancer types. In one example, the predictive values ​​are scored from 0 to 100, with the cumulative sum of the predictive values ​​being 100. For example, the cancer classifier returns a cancer prediction that includes predictive values ​​for breast cancer, lung cancer, and non-cancer. For example, the classifier may return a cancer prediction that the test sample has a 65% likelihood of breast cancer, a 25% likelihood of lung cancer, and a 10% likelihood of non-cancer. The analysis system may further evaluate the predictive values ​​to generate one or more predictions of the presence of cancer in the sample, which may also be referred to as TOO predictions, indicating one or more TOO labels, e.g., a first TOO label with the highest prediction, a second TOO label with the next highest prediction, etc. Continuing with the above example, given these percentages, in this example, the system may identify the sample as having breast cancer because the likelihood of breast cancer is highest.

[0166] In either embodiment, the analysis system trains the cancer classifier by inputting a set of training samples along with their feature vectors into the cancer classifier and adjusting the classification parameters so that the classifier's function accurately relates the training feature vectors to their corresponding labels. The analysis system may group the training samples into one or more training sample sets for iterative batch training of the cancer classifier. After inputting all of the training sample sets, including their training feature vectors, and adjusting the classification parameters, the cancer classifier can be sufficiently trained to label test samples according to their feature vectors within any acceptable error range. The analysis system may train the cancer classifier according to any one of a number of methods. As an example, a binary cancer classifier can be an L2 regularized logistic regression classifier trained using a log-loss function. As another example, a multi-cancer classifier can be polynomial logistic regression. In fact, either type of cancer classifier can be trained using other techniques. These techniques are numerous and include machine learning algorithms such as kernel methods, random forest classifiers, mixture models, autoencoder models, multi-layer neural networks, and other possibilities.

[0167] The classifier may include a logistic regression algorithm, a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a neighborhood algorithm, a boosted decision tree algorithm, a random forest algorithm, a decision tree algorithm, a polynomial logistic regression algorithm, a linear model, and a linear regression algorithm.

[0168] III.D. Deployment of Cancer Classifiers During use of the cancer classifier, the analysis system can obtain a test sample from a subject with an unknown cancer type. The analysis system can process the test sample, consisting of DNA molecules, using any combination of processes 200 and 530 to achieve a set of aberrant fragments. The analysis system can identify a test feature vector for use by the cancer classifier according to principles similar to those discussed in process 600. The analysis system can calculate an aberration score for each CpG site within a plurality of CpG sites used by the cancer classifier. For example, the cancer classifier receives as input a feature vector containing an aberration score for 1,000 selected CpG sites. The analysis system can therefore identify a test feature vector containing an aberration score for the 1,000 selected CpG sites based on the set of aberrant fragments. The analysis system can calculate the aberration score in the same manner as for the training sample. In some embodiments, the analysis system defines the aberration score as a binary score based on whether there is a hypermethylated or hypomethylated fragment encompassing that CpG site within the set of aberrant fragments. In some embodiments, the analysis system performs covariate prediction (e.g., process 440 in FIG. 4) to predict covariate values ​​and / or labels. The analysis system may generate a test feature vector including one or more features based on the covariate prediction.

[0169] The analysis system can then input the test feature vector into a cancer classifier. The cancer classifier function can then generate a cancer prediction based on the classification parameters trained in process 600 and the test feature vector. In a first approach, the cancer prediction is binary and can be selected from the group consisting of "cancer" or "non-cancer." In a second approach, the cancer prediction is selected from the group of a number of cancer types and "non-cancer." In additional embodiments, the cancer prediction has a predictive value for each of a number of cancer types. Furthermore, the analysis system can determine that the test sample is most likely to have one of the cancer types. Following the above example of a cancer prediction for a test sample with a 65% likelihood of breast cancer, a 25% likelihood of lung cancer, and a 10% likelihood of non-cancer, the analysis system can determine that the test sample is most likely to have breast cancer. In another example, if the cancer prediction is binary, with a 60% likelihood of non-cancer and a 40% likelihood of cancer, the analysis system can determine that the test sample is most likely not to have cancer. In additional embodiments, the most likely cancer prediction may be further compared to a threshold (e.g., 40%, 50%, 60%, 70%) to call the subject as having that cancer type. If the most likely cancer prediction does not exceed the threshold, the analysis system may return an inconclusive result.

[0170] In additional embodiments, the analysis system may link the cancer classifier trained in step 660 of process 600 to other cancer classifiers trained in step 670 or process 600. The analysis system may input the test feature vector into the cancer classifier trained as a binary classifier in step 660 of process 600. The analysis system may receive a cancer prediction output. The cancer prediction may be binary, indicating whether the subject is likely to have cancer or not. In other examples, the cancer prediction includes a predictive value that accounts for the likelihood of cancer and the likelihood of non-cancer. For example, the cancer prediction may have a cancer predictive value of 85% and a non-cancer predictive value of 15%. The analysis system may identify the subject as likely to have cancer. If the analysis system identifies the subject as likely to have cancer, the analysis system may input the test feature vector into a multi-class cancer classifier trained to distinguish between different cancer types. The multi-class cancer classifier may receive the test feature vector and return a cancer prediction for one of multiple cancer types. For example, the multi-class cancer classifier may provide a cancer prediction indicating that the subject is most likely to have ovarian cancer. In another embodiment, the multi-class cancer classifier provides a predictive value for each cancer type of a plurality of cancer types, for example, the cancer predictions may include a breast cancer predictive value of 40%, a colon cancer predictive value of 15%, and a liver cancer predictive value of 45%.

[0171] According to a general embodiment of a binary cancer classifier, an analysis system can identify a cancer score for a test sample based on sequencing data (e.g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.) for the test sample. The analysis system compares the cancer score for the test sample to a binary threshold cutoff to predict whether the test sample is likely to have cancer. The binary threshold cutoff can be adjusted using a TOO threshold based on one or more TOO subtype classes. The analysis system can further generate a feature vector for the test sample that is used in a multi-class cancer classifier to identify a cancer prediction indicative of one or more possible cancer types.

[0172] The classifier can be used to identify a disease status of a subject, for example, a subject whose disease status is unknown. The method can include obtaining a test genomic data configuration (e.g., point-in-time test data) in electronic form, which includes a value for each genomic feature among a plurality of genomic features of a corresponding plurality of nucleic acid fragments in a biological sample obtained from the subject. The method can then include applying the test genomic data configuration to the test classifier, thereby identifying the disease status of the subject. The subject may not have been previously diagnosed with the disease.

[0173] The classifier can be a temporary classifier that uses at least (i) a first test genomic data configuration generated from a first biological sample obtained from the subject at a first time point and (ii) a second test genomic data configuration generated from a second biological sample obtained from the subject at a second time point.

[0174] The trained classifier can be used to identify a disease status of a subject, for example, a subject whose disease status is unknown. In this case, the method can include obtaining a test time series dataset in electronic form for the subject, the test time series dataset including, for each respective time point among a plurality of time points, a corresponding test genotype data configuration including values ​​for a plurality of genotype characteristics of a corresponding plurality of nucleic acid fragments in a corresponding biological sample obtained from the subject at each time point, and, for each respective pair of consecutive time points among the plurality of time points, an indication of the length of time between each pair of consecutive time points. The method can then include applying the test genotype data configuration to the test classifier, thereby identifying the disease status of the subject. The subject may not have been previously diagnosed with the disease.

[0175] IV. Application In some embodiments, the methods, analysis systems, and / or classifiers of the present invention can be used to detect the presence of cancer, monitor cancer progression or recurrence, monitor treatment response or effectiveness, identify or monitor the presence of minimal residual disease (MRD), or any combination thereof. For example, as described herein, a classifier can be used to generate a probability score (e.g., 0 to 100) describing the likelihood that a test feature vector was obtained from a subject with cancer. In some embodiments, the probability score is compared to a threshold probability to identify whether the subject has cancer. In other embodiments, the likelihood or probability score can be assessed at multiple different time points (e.g., before or after treatment) to monitor disease progression or the effectiveness of treatment (e.g., treatment response). In yet other embodiments, the likelihood or probability score can be used to make or influence clinical decisions (e.g., diagnosing cancer, selecting a treatment, assessing treatment response, etc.). For example, in one embodiment, if the probability score exceeds a threshold, a physician can prescribe an appropriate treatment.

[0176] IV.A. Early Cancer Detection In some embodiments, the methods and / or classifiers of the present invention are used to detect the presence or absence of cancer in a subject suspected of having cancer. For example, a classifier (e.g., as described above in Section III and exemplified in Section V) can be used to identify a cancer predictor that describes the likelihood that a test feature vector is from a subject with cancer.

[0177] In one embodiment, the cancer prediction is a likelihood (e.g., scored from 0 to 100) of whether a test sample has cancer (i.e., a binary classification). Thus, the analysis system may identify a threshold for identifying whether a subject has cancer. For example, a cancer prediction of 60 or greater may indicate that the subject has cancer. In yet another embodiment, a cancer prediction of 65 or greater, 70 or greater, 75 or greater, 80 or greater, 85 or greater, 90 or greater, or 95 or greater may indicate that the subject has cancer. In other embodiments, the cancer prediction may indicate disease severity. For example, a cancer prediction of 80 may indicate a more severe form or more advanced stage of cancer compared to a cancer prediction below 80 (e.g., a probability prediction of 70). Similarly, an increase in cancer prediction over time (e.g., determined by classifying test feature vectors of multiple samples from the same subject taken at two or more time points) may indicate disease progression, or a decrease in cancer prediction over time may indicate successful treatment.

[0178] In other embodiments, the cancer prediction includes multiple predictive values, with each of the multiple cancer types classified (i.e., multi-class classification) having a predictive value (e.g., scored from 0 to 100). The predictive values ​​may correspond to the likelihood that a training sample (and the training sample during estimation) has each of the cancer types. The analysis system may identify the cancer type with the highest predictive value and indicate that the subject is likely to have that cancer type. In other embodiments, the analysis system may further compare the highest predictive value to a threshold value (e.g., 50, 55, 60, 65, 70, 75, 80, 85, etc.) to identify that the subject is likely to have that cancer type. In other embodiments, the predictive value may also indicate the severity of the disease. For example, a predictive value higher than 80 may indicate a more severe form or more advanced stage of cancer compared to a predictive value of 60. Similarly, an increase in cancer prediction over time (e.g., identified by classifying test feature vectors of multiple samples from the same subject taken at two or more time points) may indicate disease progression, or a decrease in cancer prediction over time may indicate treatment success.

[0179] According to aspects of the invention, the methods and systems of the invention can be trained to detect or classify multiple cancer indications. For example, the methods, systems, and classifiers of the invention can be used to detect the presence of one or more, two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different types of cancer.

[0180] Examples of cancers that can be detected using the methods, systems, and classifiers of the present invention include carcinoma, lymphoma, blastoma, sarcoma, leukemia, or lymphoid tumors. More specific examples of such cancers include squamous cell cancer (e.g., epithelial squamous cell cancer), skin cancer, melanoma, lung cancer, including small cell lung cancer, non-small cell lung cancer (NSCL) and adenocarcinoma of the lung, cancer of the peritoneum, gastric cancer, including gastrointestinal cancer, pancreatic cancer (e.g., pancreatic ductal adenocarcinoma), cervical cancer, ovarian cancer (e.g., ovarian high-grade serous carcinoma), liver cancer (e.g., hepatocellular carcinoma (HCC)), and ovarian cancer (e.g., ovarian high-grade serous carcinoma). carcinoma), primary hepatocellular carcinoma, bladder cancer (e.g., urothelial carcinoma), testicular (germ cell tumor) cancer, breast cancer (e.g., HER2-positive, HER2-negative, and triple-negative breast cancer), brain tumors (e.g., astrocytoma, glioma (e.g., glioblastoma)), colon cancer, rectal cancer, colorectal cancer, endometrial or uterine cancer, salivary gland cancer, kidney cancer (e.g., renal cell carcinoma, nephroblastoma, or Wilms' tumor), prostate cancer, vulvar cancer, thyroid cancer, anal cancer, penile cancer, head and neck cancer, esophageal cancer, nasopharyngeal cancer (NPC) Other examples of cancer include, but are not limited to, retinoblastoma, thecoma, virilizing germinoma, non-Hodgkin's lymphoma (NHL), blood cancers including multiple myeloma and acute myeloid leukemia, endometriosis, fibrosarcoma, choriocarcinoma, laryngeal cancer, Kaposi's sarcoma, schwannoma, oligodendroglioma, neuroblastoma, rhabdomyosarcoma, osteosarcoma, leiomyosarcoma, and urinary tract cancer.

[0181] In some embodiments, the cancer is anorectal cancer, bladder cancer, breast cancer, cervical cancer, colon cancer, esophageal cancer, gastric cancer, head and neck cancer, biliary tract cancer, leukemia, lung cancer, lymphoma, melanoma, multiple myeloma, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer, thyroid cancer, endometrial cancer, or any combination thereof.

[0182] In some embodiments, one or more cancers may be "high signal" cancers (defined as cancers with a 5-year cancer survival rate greater than 50%), such as anorectal, colon, esophageal, head and neck, biliary tract, lung, ovarian, and pancreatic cancers, as well as lymphoma and multiple osteosarcoma. High signal cancers tend to be more aggressive and typically have higher-than-average cell-free nucleic acid concentrations in a test sample obtained from a patient.

[0183] IV.B. Cancer and Treatment Monitors In some embodiments, the cancer prognosis can be assessed at multiple different time points (e.g., before or after treatment) to monitor disease progression or to monitor the effectiveness of a treatment (e.g., therapeutic response). For example, the invention includes a method comprising obtaining a first sample (e.g., a first plasma cfDNA sample) from a cancer patient at a first time point and identifying a first cancer prognosis therefrom (as described herein), and obtaining a second test sample (e.g., a second plasma cfDNA sample) from the cancer patient at a second time point and identifying a second cancer prognosis therefrom (as described herein).

[0184] In certain embodiments, the first time point is before cancer treatment (e.g., before resection surgery or intervention) and the second time point is after cancer treatment (e.g., after resection surgery or intervention), and the classifier is utilized to monitor the effectiveness of the treatment. For example, if the second cancer prediction is decreased compared to the first cancer prediction, the treatment is considered successful. However, if the second cancer prediction is increased compared to the first cancer prediction, the treatment is considered unsuccessful. In other embodiments, the first and second time points are both before cancer treatment (e.g., before resection surgery or intervention). In yet another embodiment, the first and second time points are both after cancer treatment (e.g., after resection surgery or intervention). In yet another embodiment, cfDNA samples may be obtained from a cancer patient at the first and second time points and analyzed, for example, to monitor ocular progression, determine whether cancer is in remission (e.g., post-treatment), monitor or detect residual disease or disease recurrence, or monitor treatment (e.g., therapy) efficacy.

[0185] One of skill in the art will readily appreciate that test samples can be obtained from a cancer patient at any desired time point and analyzed according to the methods of the present invention to monitor the patient's cancer status. In some embodiments, the first and second time points are between about 15 minutes and about 30 years, e.g., about 30 minutes, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or about 24 hours, e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, or about 50 days, or e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months, or e.g., about 1, 1.5, 2, 2.5, 3, 3.5, or 4.5 years. , 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5, or separated by approximately 30 years. In other embodiments, test samples can be obtained from a patient at least once every five months, at least once every six months, at least once a year, at least once every two years, at least once every three years, at least once every four years, or at least once every five years.

[0186] IV.C. Treatment In yet another embodiment, the cancer prediction can be used to make or influence clinical decisions (e.g., diagnosing cancer, selecting treatment, assessing the effectiveness of treatment, etc.) For example, in one embodiment, if the cancer prediction (e.g., for cancer, or for a particular cancer type) exceeds a threshold, a physician can prescribe an appropriate treatment (e.g., resective surgery, radiation therapy, chemotherapy, and / or immunotherapy).

[0187] A classifier (described herein) can be used to identify a cancer prediction that the sample feature vector is from a subject with cancer. In one embodiment, an appropriate treatment (e.g., resective surgery or therapy) is prescribed if the cancer prediction exceeds a threshold. For example, in one embodiment, if the cancer prediction is 60 or greater, one or more appropriate treatments are prescribed. In other embodiments, if the cancer prediction is 65 or greater, 70 or greater, 75 or greater, 80 or greater, 85 or greater, 90 or greater, or 95 or greater, one or more appropriate treatments are prescribed. In other embodiments, the cancer prediction can indicate the severity of the disease. An appropriate treatment matching the severity of the disease can then be prescribed.

[0188] In some embodiments, the treatment is one or more cancer therapy agents selected from the group consisting of chemotherapy agents, targeted cancer therapy agents, differentiation inducers, hormonal therapy agents, and immunotherapy agents. For example, the treatment can be one or more chemotherapeutic agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeletal disrupting agents (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based drugs, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapy agents selected from the group consisting of signal transduction inhibitors (e.g., tyrosine kinases, growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapy agents, including retinoids, e.g., tretinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormone therapy agents selected from the group consisting of antiestrogens, aromatase inhibitors, progestins, antiandrogens, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapeutic agents selected from the group consisting of monoclonal antibody therapy, such as rituximab (RITUXAN) and alemtuzumab (CAMPATH), nonspecific immunotherapy and adjuvants, such as BCG, interleukin-2 (IL-2), and interferon-alpha, and immunomodulatory agents, e.g., thalidomide and lenalidomide (REVLIMID). A skilled physician or oncologist can select the appropriate cancer treatment based on characteristics such as tumor type, stage of cancer, previous cancer treatment or exposure to therapeutic agents, and other characteristics of the cancer.

[0189] V. Example Results VA sample collection and processing Study Design and Samples: CCGA (NCT02889978) is a prospective, multicenter, case-control, observational study with a fixed-term follow-up. De-identified biospecimens were collected from approximately 15,000 participants at 342 sites. Samples were divided into a training group (1,785) and a test group (1,015). Samples were selected to ensure a predetermined distribution of cancer and non-cancer types across each cohort site. Cancer and non-cancer samples were frequency- and age-matched by sex.

[0190] Whole-genome bisulfite sequencing: cfDNA was isolated from plasma and analyzed using whole-genome bisulfite sequencing (WGBS; 30x depth). cfDNA was extracted from two plasma containers per patient (maximum total volume of 10 ml) using a modified QIAamp Circulating Nucleic Acid Kit (Qiagen; Germantown, MD). Up to 75 ng of plasma cfDNA was bisulfite converted using the EZ-96 DNA Methylation Kit (Zymo Research, D5003). The converted cfDNA was used to generate dual-index sequencing libraries using the Accel-NGS Methyl-Seq DNA Library Construction Kit (Swift BioSciences, Ann Arbor, MI), and the constructed libraries were quantified using the KAPA Library Quantification Kit for Illumina Platforms (Kapa Biosystems; Wilmington, MA). The four libraries were pooled together with a 10% PhiX x3 library (Illumina, FC-110-3001) on an Illumina NovaSeq 7000 S2 flow cell, clustered, and then subjected to 150-bp paired-end sequencing (30x).

[0191] For each sample, the WGBS fragment set was reduced to a smaller subset of fragments with aberrant methylation patterns. Additionally, hyper- or hypomethylated cfDNA fragments were selected. cfDNA fragments with aberrant methylation patterns and selected as hyper- or hypomethylated, i.e., UFXM, were used. Fragments occurring frequently in individuals without cancer or with unstable methylation are unlikely to generate discriminatory features for cancer status classification. Therefore, we created a statistical model and data structure for representative fragments using an independent reference group of 108 cancer-free, non-smoking participants from the CCGA study (age: 58 ± 14 years, 79 [73%] women). These samples were used to train a Markov chain model (third order) that estimates the likelihood of sequences with a certain CpG methylation status within a fragment, as previously described in Section II.C. This model was demonstrated to be calibrated within the normal fragment range (p-value > 0.001) and was used to reject fragments with a p-value ≥ 0.001 from the Markov model as insufficiently abnormal.

[0192] As previously described, a separate data reduction step selected only fragments with at least five CpGs covered and an average methylation of either >0.9 (hypermethylated) or <0.1 (hypomethylated). This procedure resulted in a median (range) of 2,800 (1,500–12,000) UFXM fragments for the training cancer-free participants. Because this data reduction procedure only used the reference set data, this step only needed to be applied once for each sample.

[0193] VB covariate prediction results Figure 8 illustrates genomic regions associated with age, according to one or more embodiments. The analysis system can train a regression using methylation features from training samples. These examples use only non-cancer training samples for illustrative purposes. In this example, the analysis system calculates a two-sided p-value for the regression slope from a t-statistic for each genomic region based on the trained linear regression. A lower p-value indicates a higher likelihood of observing that slope. In the graph shown, the x-axis plots the chromosomes of the human body, and the y-axis plots the position within each chromosome. Each mark represents a region containing a genomic cluster (CAE) within 500 CpG sites with a support score higher than any threshold support score. Each CAR indicates the lowest p-value of the genomic regions clustered within the CAR. The legend to the right of the graph is the negative logarithm of the p-value; the higher the negative logarithm, the lower the p-value.

[0194] Figure 9A illustrates one process for identifying a feature set of age-indicative genomic regions to be used in the generated feature vector, according to one or more embodiments. The training samples used were obtained from the CCGA study. The analog system performed GLMNET relaxed lasso regression on a lambda grid to identify the optimal range of genomic regions to use as the feature set. The optimal set of genomic regions, ranging from 58 to 83 genomic regions, yielded the lowest mean squared error (plotted on the y-axis). The feature set of 83 genomic regions was used to train the age prediction model for the test set.

[0195] Figures 9B and 9C show age residuals using an age prediction model trained with the feature set identified in Figure 9A, according to an example. The x-axis plots reported chronological age ("Actual") against predicted chronological age ("Predicted") by the age prediction model on the y-axis. Figure 9B shows a graph of age prediction results for a non-cancer holdout cohort, according to an example. The holdout cohort was not used in training the regression. The total number of samples in the holdout cohort was 369. Residuals are calculated as reported age minus predicted age. Notably, the highest residuals in the non-cancer cohort were near approximately -35 and +25, with the majority of the non-cancer cohort having residuals between -10 and +10. Figure 9C shows a graph of age prediction results for a cancer cohort, according to an example. Because the age prediction model was fit to non-cancer samples, no cancer samples were used in training. The total number of samples in the cancer cohort was 1561. Here, the highest residual is approximately -155. The spread of the residuals is also more scattered compared to the non-cancer cohort in Figure 9B.

[0196] FIG. 10A illustrates another process for identifying a feature set of genomic regions indicative of chronological age information, according to one or more embodiments. The training samples used were obtained from a follow-up observation of the CCGA study, called CCGA2. Similar to FIG. 9A, the analysis system performs GLMNET relaxed lasso regression on a lambda grid to identify an optimal range of genomic regions to use as a feature set. The optimal set of genomic regions, ranging from 31 to 57, yielded the lowest mean squared error (plotted on the y-axis). The feature set of the 57 genomic regions specifically was used to train other age prediction models on the dataset.

[0197] Figures 10B and 10C show age residuals using an age prediction model trained with the feature set identified in Figure 10A, according to an example. The x-axis plots reported age ("Actual") against the y-axis of predicted age ("Predicted") by the age prediction model. Figure 10B shows a graph of age prediction results for a non-cancer holdout cohort, according to an example. The holdout cohort was not used in training the regression. The total number of samples in the holdout cohort was 466. Residuals are calculated as reported age minus predicted age. Notably, the highest residuals for the non-cancer cohort were approximately -10 and +10, with the majority of the non-cancer cohort having residuals within +5 to +5. Figure 10C shows a graph of age prediction results by cancer code, according to an example. Because the age prediction model was fitted to the non-cancer samples, none of the cancer samples were used in training the regression. The total number of samples in the cancer cohort was 967. Here, the highest residuals are −193 and +135 (either side of underestimation or overestimation). The spread of residuals is also more scattered compared to the non-cancer cohort in FIG. 10B.

[0198] Figure 11 shows the spread of the test cohort against cancer stage, according to an example. The x-axis of the graph separates the known cancer status of the samples: non-cancer, cancer stages 1-4, and other statuses ("not_expected" and "missing"). The left graph contains samples for which the classifier predicted no cancer, i.e., negative results, including both true negatives and false negatives. The right graph contains samples for which the classifier predicted cancer, i.e., positive results, including both true positives and false positives. The residual threshold (shown as a "z-score" less than 4) was four standard deviations from the mean. All samples whose chronological age residuals were greater than the threshold were colored red, while the rest were colored yellow. Two important things to note in the left graph: First, the chronological age residuals of the true non-cancer samples did not exceed the chronological age residual threshold (no samples are marked red). Second, although there were many cancer samples that were false negatives, the chronological age residuals of some of these false negatives exceeded the residual threshold. Using a chronological age residual threshold would have identified these false negatives as having cancer. The graph on the right also has two important findings. First, a significant number of cancer samples were higher than the chronological age residual threshold compared to the non-cancer samples in either graph. Second, the later stages of cancer had a larger number of samples whose chronological age residuals were higher than the chronological age residual threshold. This indicates that accelerating cancer status leads to greater degradation of chronological age predictions from sample methylation signatures.

[0199] Figures 12A and 12B show the spread of the test cohort for cancer type, according to an example. Figure 12A shows the top series of graphs representing test samples predicted by the cancer classifier as negative results, i.e., predicted not to have cancer. Figure 12B shows the bottom series of graphs representing test samples predicted by the cancer classifier as positive results, i.e., predicted to have cancer. A similar chronological age residual threshold is used, e.g., a z-score greater than or less than 4. In the top series of graphs, at least 9 false-negative samples (samples known to have cancer but predicted not to have cancer by the cancer classifier) ​​have chronological age residuals (calculated by the chronological age prediction model) that exceed the chronological age residual prediction. These are samples identified as having a strong likelihood of cancer.

[0200] Figure 13 shows one genomic region showing chronological age deceleration across cancer types, according to an embodiment. All samples in the figure are cancer samples from different cancer types. The x-axis of each graph shows the actual chronological age of the sample, and the y-axis shows the predicted chronological age as a percentage of the actual age. For many cancer types, the predicted chronological age is decelerated. This genomic region generally does not discriminate between cancer types, but it does decelerate chronological age across almost all cancer types (some of which is hindered by low sampling).

[0201] Figures 14A and 14B show two genomic regions that distinguish between hematological and non-hematological cancers, according to an embodiment. All samples in the figures are from different cancer types. The x-axis of each graph shows the actual chronological age of the sample, and the y-axis shows the predicted chronological age as a percentage of the actual age. Figure 14A shows results for a first genomic region where there appears to be consistent chronological age acceleration in hematological cancers and less consistent chronological age acceleration in other non-hematological cancers. Figure 14B shows results for a second genomic region where there appears to be consistent chronological age deceleration in hematological cancers and little to no chronological age deceleration in other non-hematological cancers.

[0202] FIG. 15A illustrates the identification of a genomic region feature set for predicting biological sex, according to one or more embodiments. The training sample used is the CCGA2 study. The analysis system runs GLMNET relaxed lasso regression on a lambda grid to identify optimal genomic region ranges to use as feature sets. The optimal set of two or more ranges of genomic regions yielded the smallest mean squared error (plotted on the y-axis). Three genomic region feature sets were used to train a biological sex prediction model on the test set.

[0203] Figure 15B shows results obtained from a trained gender prediction model, according to an embodiment. The gender prediction model was trained on the three genomic regions identified in Figure 15A. The test cohort included non-cancer samples. In summary, the gender prediction model had a specificity of 100% and an accuracy of 99.8%.

[0204] FIG. 16A illustrates the identification of a feature set of genomic regions for predicting smoking status, according to one or more embodiments. The training sample used is the CCGA2 study. The analysis system runs GLMNET relaxed lasso regression on a lambda grid to identify the optimal range of genomic regions to use as feature sets. The optimal set of genomic regions, ranging from 1 to 4, yielded the lowest mean squared error (plotted on the y-axis). Two genomic region feature sets were used to train a smoking status prediction model on the test set.

[0205] Figure 16B shows results obtained from a trained smoking status gender prediction model, according to an embodiment. The smoking status prediction model was trained on the two genomic regions identified in Figure 16A. The test cohort included non-cancer samples. In summary, the smoking status prediction model had a specificity of 99.6% and an accuracy of 96.2%.

[0206] VI. Other Considerations The detailed description of certain embodiments refers to the accompanying drawings, which illustrate specific embodiments of the present disclosure. Other embodiments having different structure and operation do not depart from the scope of the present disclosure. Terms such as "the present invention" are used with reference to specific examples of the many alternative aspects or embodiments of Applicant's invention described herein, and neither their use nor their absence is intended to limit the scope of Applicant's invention or the claims.

[0207] Embodiments of the present invention may also relate to apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes and / or it may include a general-purpose computing device selectively activated or reconfigured by a computer program stored within the computer. Such a computer program may be stored on a non-transitory tangible computer-readable storage medium or any type of medium suitable for storing electronic instructions that may be coupled to a computer system bus. Furthermore, any computer system referred to herein may include a single processor or may be an architecture utilizing multiple processors designed to increase computing power.

[0208] Any of the steps, operations, or processes described herein as being performed by the analysis system may be performed or implemented in one or more hardware or software modules of the apparatus, alone or in combination with other computing devices. In one embodiment, the software modules are implemented in a computer program product that includes a computer-readable medium having stored thereon computer program code executable by a computer processor to perform any or all of the steps, operations, or processes described above.

Claims

1. In the method, A step in which multiple training samples are obtained, each training sample is: It comprises multiple nucleic acid fragments, each of which has a genomic position that overlaps with at least one of the multiple genomic regions. The training sample is labeled by the chronological age of the individual from whom it was collected. The steps include sequencing the plurality of nucleic acid fragments of each training sample to identify the methylation pattern of each nucleic acid fragment, For each of the multiple genomic regions, The step of identifying a nucleic acid fragment from among the plurality of nucleic acid fragments that has a genomic position overlapping with the genomic region, The steps include: expressing the correlation between chronological age and methylation pattern for the aforementioned genomic region, and calculating an indication score based on the chronological age of the individual from whom the identified nucleic acid fragment was collected and the methylation pattern of the identified nucleic acid fragment; A step of generating a feature set that includes one or more genomic regions from among the plurality of genomic regions, wherein the one or more genomic regions in the feature set have an indication score higher than a threshold, A step of training a machine learning age prediction model to identify the predicted chronological age of a subject from whom the test sample was collected, wherein the training is based on the methylation patterns of nucleic acid fragments of the plurality of training samples that overlap with one or more genomic regions in the feature set. A method that includes this.

2. A step of training a linear regression for each genomic region of the feature set based on the methylation patterns of the nucleic acid fragments overlapping with each genomic region from the plurality of training samples labeled as non-cancer, A step of obtaining multiple additional training samples, where each additional training sample is: It includes a plurality of additional nucleic acid fragments having additional genomic locations that overlap with at least one of the plurality of genomic regions, The aforementioned additional training samples were labeled by the chronological age of the individuals from whom they were collected. A step of labeling as non-cancerous or cancerous based on past identification of the presence of cancer in the aforementioned additional training samples, The steps include sequencing the plurality of additional nucleic acid fragments and identifying the methylation pattern of each additional nucleic acid fragment, For each of the aforementioned multiple genomic regions, The steps include applying the linear regression to the methylation patterns of nucleic acid fragments of the plurality of additional training samples to identify the predicted chronological age of the individual from whom the additional training samples were collected, The steps include: calculating the age residual for each additional training sample as the difference between the predicted chronological age and the labeled chronological age; A step of comparing the age residual of the additional training sample labeled as cancerous with the age residual of the additional training sample labeled as non-cancerous, A step of generating a reduced feature set from the feature set based on a comparison of age residuals, wherein the reduced feature set includes fewer genomic regions than the feature set, and the reduced feature set is used to train the machine learning age prediction model. The method according to claim 1, further comprising:

3. A step of obtaining a test sample, wherein the test sample comprises a plurality of additional nucleic acid fragments, and the test sample is labeled with the chronological age of the subject from whom it was collected. The steps include sequencing the plurality of additional nucleic acid fragments of the test sample to identify the methylation patterns of the plurality of additional nucleic acid fragments, The steps include applying the trained age prediction model to identify the predicted chronological age of the subject from which the test sample was taken, based on the methylation patterns of the additional nucleic acid fragments that overlap with one or more genomic regions in the feature set, The steps include: calculating the age residual as the difference between the labeled chronological age and the predicted chronological age of the subject; In response to the identification that the age residual exceeds the residual threshold, the step of identifying that the test sample has a strong likelihood of cancer presence, The method according to claim 1, further comprising:

4. The residual threshold is, The trained age prediction model is applied to a second set of training samples identified as non-cancer to determine the predicted age of each of the second set of training samples. By comparing the predicted age with the labeled chronological age of the second set of training samples, the age residual of each of the second set of training samples is calculated. Identifying the residual threshold based on the calculated age residuals of the second set of training samples. The method according to claim 3, wherein the calculated age residuals of the second plurality of training samples are identified and at least a large portion of them satisfy the residual threshold.

5. In response to the identification of the aforementioned test sample as having a strong likelihood of cancer presence, The steps include filtering the methylation patterns of the plurality of additional nucleic acid fragments by p-value filtering to identify groups of abnormal methylation patterns, The steps include generating a feature vector of the test sample based on the age residual and the abnormal methylation pattern group, The steps include identifying the cancer prediction for the test sample by inputting the feature vector into a trained cancer classifier, The method according to claim 3, further comprising:

6. The method according to claim 5, wherein the cancer prediction is a binary prediction of the presence or absence of cancer or other disease state, or a multi-class prediction among multiple cancer types, or a multi-class prediction among multiple disease states.

7. A step of identifying the presence of cancer in the test sample using a second machine learning cancer classifier, wherein the second cancer classifier is configured to receive the subject's predicted chronological age and the methylation patterns of a plurality of additional nucleic acid fragments as input, and to output a prediction of the presence of cancer in the test sample. The method according to claim 3, further comprising:

8. The method according to claim 7, wherein the second machine learning cancer classifier is further configured to receive clinical information and genetic background of the subject as input and to output the prediction of the presence of the cancer in the test sample.

9. The aforementioned indicator score is either a Pearson correlation or a covariate score, or The method according to claim 1, wherein the indication score is determined by training a linear regression to predict chronological age by regression from the methylation density of a non-cancer training sample, the methylation density is calculated as the percentage of nucleic acid fragments having a methylation state in a particular genomic region, where the methylation density has a genomic location overlapping with a particular genomic region.

10. The aforementioned machine learning age prediction model includes multivariate regression, Optionally, The multivariate regression is penalized based on the number of one or more genomic regions in the feature set, or The method according to claim 1, wherein the machine learning age prediction model receives as input the methylation density corresponding to each of the genomic regions in the feature set.

11. The method according to claim 1, wherein the number of one or more genomic regions in the feature set is selected from the range of 5 to 10,000.

12. The method according to claim 1, wherein the sequencing of the nucleic acid fragments includes whole-genome bisulfite sequencing (WGBS) or targeted sequencing.

13. Each of the aforementioned training samples is identified in advance as containing or not containing cancer, or The method according to claim 1, wherein each of the plurality of training samples is labeled as having cancer or not having cancer, and the label is based on prior identification of the cancer status of the training sample.

14. When executed by one or more processors, the one or more processors will Multiple training samples are obtained, and each training sample is: It comprises multiple nucleic acid fragments, each of which has a genomic position that overlaps with at least one of the multiple genomic regions. The aforementioned training samples are labeled by the chronological age of the individuals from whom they were collected. The plurality of nucleic acid fragments of each training sample are sequenced to identify the methylation pattern of each nucleic acid fragment. For each of the multiple genomic regions, From among the plurality of nucleic acid fragments, a nucleic acid fragment having a genomic position that overlaps with the genomic region is identified. For the aforementioned genomic region, the correlation between chronological age and methylation pattern is expressed, and an indication score is calculated based on the chronological age of the individual from whom the identified nucleic acid fragment was collected and the methylation pattern of the identified nucleic acid fragment. A feature set is generated that includes one or more genomic regions from the plurality of genomic regions, and the one or more genomic regions in the feature set have an indication score higher than a threshold. A machine-learned age prediction model is trained to identify the predicted chronological age of the subject from whom the test sample was taken, and the training is based on the methylation patterns of nucleic acid fragments that overlap with one or more genomic regions in the feature set of the multiple training samples. A non-temporary computer-readable storage medium containing computer programs.

15. In the system, One or more processors, When executed by the one or more processors, the one or more processors Multiple training samples are obtained, and each training sample is: It comprises multiple nucleic acid fragments, each of which has a genomic position that overlaps with at least one of the multiple genomic regions. The aforementioned training samples are labeled by the chronological age of the individuals from whom they were collected. The plurality of nucleic acid fragments of each training sample are sequenced to identify the methylation pattern of each nucleic acid fragment. For each of the multiple genomic regions, From among the plurality of nucleic acid fragments, a nucleic acid fragment having a genomic position that overlaps with the genomic region is identified. For the aforementioned genomic region, the correlation between chronological age and methylation pattern is expressed, and an indication score is calculated based on the chronological age of the individual from whom the identified nucleic acid fragment was collected and the methylation pattern of the identified nucleic acid fragment. A feature set is generated that includes one or more genomic regions from the aforementioned plurality of genomic regions, and the one or more genomic regions in the feature set are above a threshold. A machine learning age prediction model is trained to identify the predicted chronological age of the subject from whom the test sample was taken, and the training is based on the methylation patterns of nucleic acid fragments that overlap with one or more genomic regions in the feature set of the multiple training samples. A non-temporary computer-readable storage medium for storing computer program instructions, A system that includes this.