Component mixture models for tissue identification in DNA specimens

A machine-learned mixture model addresses noise and heterogeneity in DNA methylation profiling by deconvolving tissue contributions, improving cancer classification accuracy and subtype prediction.

JP2025536913APending Publication Date: 2025-11-12GRAIL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025521419
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-10-18
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Existing DNA methylation profiling for cancer detection is limited by noise in samples, misdiagnosis due to confounding tissue signals, and heterogeneity in tumors, leading to inaccurate cancer classification.

Method used

A machine-learned mixture model is trained to deconvolve component contributions in DNA fragments using maximum likelihood estimation, generating methylation signatures to improve cancer classification by determining tissue purity and predicting cancer type.

Benefits of technology

Enhances cancer classification accuracy by distinguishing between different tissue types and predicting cancer subtypes, reducing contamination and heterogeneity-related errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536913000001_ABST
    Figure 2025536913000001_ABST
Patent Text Reader

Abstract

A method and system for methylation-informed component deconvolution using a mixture model is disclosed. The mixture model can be trained independently of labels or known component contributions. The system generates a methylation signature for each of a plurality of training samples. The methylation signature can be based on the count or proportion of one or more methylation variants represented in the methylated sequence reads of the training sample in each genomic region of a plurality of genomic regions. The system can deconvolve the component contributions by training the mixture model using maximum likelihood estimation. The mixture model can include component submodels and a deconvolution submodel. The component submodel predicts component likelihoods based on the methylation signatures. The deconvolution submodel predicts component contributions based on the component likelihoods.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 417,616, filed October 19, 2022, which is incorporated by reference in its entirety. [Background technology]

[0002] Technical Field Deoxyribonucleic acid (DNA) methylation plays an important role in regulating gene expression. Aberrant DNA methylation has been implicated in many disease processes, including cancer. DNA methylation profiling using methylation sequencing (e.g., whole-genome bisulfite sequencing (WGBS)) is increasingly recognized as a useful diagnostic tool for cancer detection, diagnosis, and / or monitoring. For example, specific patterns of differentially methylated regions and / or allele-specific methylation patterns can be useful as molecular markers for noninvasive diagnostics using circulating cell-free (cf) DNA. When training cancer classifiers, the classifier is limited by the amount of noise present in the sample. For example, samples may be misdiagnosed, contain confounding tissue signals, or contain multiple gene signatures from different clonal populations in heterogeneous tumors.

[0003] The present disclosure is directed to addressing the problems mentioned above. The background discussion provided herein is intended to provide a general context for the present disclosure. Unless otherwise indicated herein, the material described in this section is not prior art to the claims of this application, and no admission is made that it is prior art or that it is implied by its inclusion in this section. Summary of the Invention [Means for solving the problem]

[0004] Early detection of a subject's disease state (such as cancer) is important because it allows for early treatment and thus increases the chances of survival. Sequencing DNA fragments in cell-free (cf) DNA samples can be used to identify features that can be used for disease classification. For example, in cancer diagnosis, features based on cell-free DNA from blood samples (such as the presence or absence of somatic mutations, methylation status, or other genetic abnormalities) can provide insight into whether a subject may have cancer, and even what type of cancer the subject may have. To this end, the present description includes systems and methods for analyzing cell-free DNA (cfDNA) sequencing data to determine the likelihood that a subject has disease.

[0005] The present disclosure addresses the above-identified problems by providing an improved cancer classification system and method. The system trains a mixture model to deconvolve component contributions (also referred to as "proportions") of DNA fragments to training samples. The mixture model may be trained independently of labels or known component proportions. A system generates a methylation signature for each of multiple training samples. The methylation signature is based on the count or proportion of alternative methylation variants represented in the methylated sequence reads of the training sample in each of multiple genomic regions. The system may deconvolve the component contributions by training the mixture model using maximum likelihood estimation. The mixture model may include component submodels and deconvolution submodels. The component submodel predicts component likelihoods based on the methylation signatures. The component likelihoods are the predicted likelihoods that a nucleic acid molecule represented by the methylation signature originates from a certain component. The number of component submodels may be tuned during the training of the mixture model. The deconvolution submodel predicts component proportions based on component likelihoods.

[0006] This mixture model may have a variety of applications within large-scale cancer classification schemes.

[0007] In one or more embodiments, the mixture model can be used for tissue purity determination. Tissue purity determination can involve verifying whether a sample labeled as having known component contributions actually has such component contributions. Tissue purity determination can also include determining whether a cancer sample has significant tissue contributions for use in training a cancer classifier. Tissue purity determination can also be incorporated into contamination detection. For example, if a sample is said to be derived from a particular tissue, but the expected tissue proportions are different, then the sample can be considered contaminated.

[0008] In one or more embodiments, the mixture model may be used to predict the cancer type, e.g., the organ type of primary disease, of a test subject. The mixture model may be used to output a predominant component ratio corresponding to the predicted cancer type. In one or more embodiments, the mixture model may be used to quantify the cancer signal of a test subject.

[0009] In one or more embodiments, the analytics system may utilize component submodels to determine the origin of methylated sequence reads. In one or more embodiments, the analytics system may extract learned information from mixture models. In one embodiment, the analytics system may learn correlations between methylation variations in genomic regions that contribute to component proportions. For example, the analytics system may identify methylation variations that provide information about specific components. In another embodiment, the analytics system may identify methylation variations in genomic regions that correlate with non-cancer impurities.

[0010] Clause 1. A method for training a machine-learned mixture model to identify tissue types, the method comprising: obtaining a set of training samples comprising at least 1,000 methylated sequence reads obtained from sequencing deoxyribonucleic acid (DNA) fragments; modifying each training sample to generate a corresponding sample methylation signature by determining, for each genomic region among a plurality of genomic regions, a first set of methylated sequence reads that overlap the genomic region and a second set of methylated sequence reads that include methylation mutations in the genomic region, wherein the sample methylation signature is generated based, at least in part, on the first set of methylated sequence reads and the second set of methylated sequence reads; generating a training dataset comprising the sample methylation signature; and training a machine-learned mixture model using the training dataset, wherein the machine-learned mixture model is configured to identify the contribution of each of a plurality of tissue-of-origin types to the DNA fragments in the sample.

[0011] Clause 2. The method of clause 1, wherein at least one training sample is known to contain a first tissue type of origin among a plurality of tissue types of origin, and training the mixture model includes training the mixture model to identify a contribution of the first tissue type of origin to DNA fragments in the one training sample.

[0012] Clause 3. The method of clause 1 or 2, wherein at least one training sample is known to have a contribution rate of DNA fragments from each of a plurality of tissue-of-origin types, and training the mixture model includes training the mixture model to identify the contribution rate of each of the plurality of tissue-of-origin types to the DNA fragments in the one training sample.

[0013] Clause 4. The method of any one of clauses 1 to 3, wherein at least one training specimen is one of a liquid biopsy specimen, a tissue biopsy specimen, and a purified specimen.

[0014] Clause 5. A method according to any one of clauses 1 to 4, wherein at least one genomic region consists of one CpG site.

[0015] Clause 6. A method according to any one of clauses 1 to 5, wherein at least one genomic region comprises multiple CpG sites.

[0016] Clause 7. The method of any one of clauses 1 to 6, further comprising: determining an average sequencing depth for each of an initial set of genomic regions based on the methylated sequence reads of the training sample; and selecting a plurality of genomic regions by filtering out genomic regions whose average sequencing depth is below a threshold depth.

[0017] Clause 8. A method according to any one of clauses 1 to 7, wherein the methylation variation in the genomic region is one of two methylation patterns in the genomic region.

[0018] Clause 9. The method of clause 8, wherein the two methylation patterns in one genomic region include methylation and unmethylation.

[0019] Clause 10. The method of any one of clauses 1 to 9, wherein the methylation variation in the genomic region is one of more than three methylation patterns in the genomic region.

[0020] Clause 11. The method of any one of clauses 1 to 10, wherein modifying each training sample to generate a corresponding sample methylation signature further comprises determining, for each genomic region among the plurality of genomic regions, a third set of methylated sequence reads having a reference state for that genomic region, wherein the reference state is any methylation pattern that does not belong to a methylation variant, and wherein the sample methylation signature is further generated based on the third set of methylated sequence reads.

[0021] Clause 12. The method of any one of clauses 1 to 11, wherein the tissue types include a combination of non-cancerous impurities; squamous cell carcinoma tissue; skin cancer tissue; melanoma tissue; lung cancer tissue; lung adenocarcinoma tissue; lung squamous cell carcinoma tissue; peritoneal cancer tissue; gastrointestinal cancer tissue; pancreatic cancer tissue; cervical cancer tissue; ovarian cancer tissue; liver cancer tissue; hepatocellular carcinoma tissue; liver cancer tissue; bladder cancer tissue; testicular cancer tissue; breast cancer tissue; brain cancer tissue; colon cancer tissue; rectal cancer tissue; colorectal cancer tissue; endometrial or uterine cancer tissue; salivary gland cancer tissue; kidney or renal cancer tissue; prostate cancer tissue; vulvar cancer tissue; thyroid cancer tissue; anal cancer tissue; penile cancer tissue; head and neck cancer tissue; esophageal cancer tissue; and nasopharyngeal carcinoma (NPC) tissue.

[0022] Clause 13. The method of clause 12, wherein the non-cancerous impurities include one or more of lymphocytes, macrophages, fibroblasts, vascular endothelial cells, or non-cancerous tissue.

[0023] Clause 14. The method of clause 12 or 13, wherein the methylation signature of the non-cancer impurity is obtained from a reference database containing a plurality of methylation signatures for non-cancer impurities.

[0024] Clause 15. A method according to any one of clauses 1 to 14, wherein training the machine-learned mixture model is by maximum likelihood estimation.

[0025] Clause 16. The method of any one of clauses 1 to 15, wherein training the machine-learned mixture model includes tuning the number of tissue types as one hyperparameter of the machine-learned mixture model.

[0026] Clause 17. The method of clause 16, wherein tuning the number of tissue types as one hyperparameter of the machine-learned mixture model includes, for each number of origin tissue types within a range: training a machine-learned mixture model having that number as a hyperparameter, determining maximum likelihood by cross-validating the trained machine-learned mixture model with a holdout sample set, and implementing a penalty on the maximum likelihood based on the number; and selecting an optimal number from the range as the hyperparameter based on the penalized maximum likelihood.

[0027] Clause 18. The method of any one of clauses 1 to 17, wherein training the machine-learned mixture model includes training with one or more machine learning algorithms.

[0028] Clause 19. The method of any one of clauses 1-18, wherein the machine-learned mixture model includes a first set of tissue type models, each tissue type modeling a methylation signature of DNA fragments of a tissue type of origin, and training the machine-learned mixture model includes training the first set of tissue type models.

[0029] Clause 20. The method of clause 19, wherein training the first sub-model includes training each tissue type model according to a beta distribution.

[0030] Clause 21. A method according to any one of clauses 1 to 20, wherein the machine-learned mixture model includes a deconvolution model for deconvolving the contribution of tissue type of origin for each training sample, and training the machine-learned mixture model includes training the deconvolution model.

[0031] Clause 22. The method of clause 21, wherein training the deconvolution model includes training the deconvolution model according to a binomial distribution.

[0032] Clause 23. A method according to any one of clauses 1 to 22, wherein the machine-learned mixture model includes a first tier of one or more submodels for predicting the contribution rate of a macro-tissue type of origin and a second tier of one or more submodels for predicting the contribution rate of tissue types of origin below that macro-tissue type, wherein one first tier submodel predicts the contribution rate of one macro-tissue type and a set of one or more second tier submodels predicts the contribution rate of a set of tissue types below that one macro-tissue type that is equal to the contribution rate of the one macro-tissue type.

[0033] Clause 24. A method for identifying the contribution of tissue-of-origin type to DNA fragments in a test specimen, the method comprising: obtaining a test specimen comprising at least 1,000 methylated sequence reads for DNA fragments in the test specimen; for each of a plurality of genomic regions: determining a first set of methylated sequence reads that overlap the genomic region, and determining a second set of methylated sequence reads having alternative methylation signatures for the genomic region; generating a sample methylation signature for the training specimen based on the first set of methylated sequence reads and the second set of methylated sequence reads across the plurality of genomic regions; and identifying the contribution of each tissue-of-origin type to the DNA fragments in the test specimen by applying a machine learning mixture model to the sample methylation signature, wherein optionally the machine learning mixture model is trained by a method described in any one of clauses 1 to 23.

[0034] Clause 25. The method of clause 24, further comprising reporting the identified contribution of the tissue-of-origin type to a client device.

[0035] Clause 26. The method of clause 24 or 25, further comprising identifying the tissue type of origin with the highest contribution rate and reporting a treatment recommendation for treating cancer originating primarily from the identified tissue type of origin.

[0036] Clause 27. A method according to any one of clauses 24 to 26, further comprising: comparing the contribution rate of the tissue type of origin with a prior contribution rate of the tissue type of origin predicted before treatment; determining treatment efficacy based on the comparison; and reporting a treatment recommendation based on treatment efficacy.

[0037] Clause 28. The method of any one of clauses 24 to 27, further comprising determining the treatment as successful if the contribution rate of the first tissue type of origin is lower than the predicted prior contribution rate of the first tissue type of origin before the treatment; and reporting the success of the treatment.

[0038] Clause 29. The method of any one of clauses 24 to 28, further comprising: determining the treatment as a failure if the contribution rate of the first tissue type of origin is higher than the prior contribution rate of the first tissue type of origin predicted before the treatment; and reporting the failure of the treatment.

[0039] Clause 30. The method of clause 29, further comprising reporting a treatment modification based on treatment failure.

[0040] Clause 31. A method according to any one of clauses 24 to 30, wherein the test specimen is predicted to have cancer.

[0041] Clause 32. The method of clause 31, wherein the test specimen is predicted to have cancer by a cancer classifier.

[0042] Clause 33. The method of clause 32, wherein the cancer classifier is trained on a cancer cohort of training samples and a non-cancer cohort of training samples, and each training sample from the cancer cohort and the non-cancer cohort contains at least 1,000 methylated sequence reads for DNA fragments in the training sample.

[0043] Clause 34. The method of clause 24, further comprising: determining a subset of methylated sequence reads originating from non-cancer impurity types by applying a machine learning mixture model to the sample methylation signature; excluding the subset of methylated sequence reads originating from the non-cancer impurity types to result in a feature set of methylated sequence reads; and predicting cancer in the test sample by applying a cancer classifier to the feature set of methylated sequence reads.

[0044] Clause 35. A method of training a cancer classifier, comprising obtaining a cancer cohort of training samples and a non-cancer cohort of training samples, wherein each training sample from the cancer cohort and the non-cancer cohort comprises at least 1000 methylated sequence reads for DNA fragments in the training sample; for each of a plurality of genomic regions: determining a first set of methylated sequence reads that overlap the genomic region; determining a second set of methylated sequence reads having an alternative methylation signature for the genomic region; and optionally generating a sample methylation signature by: A method comprising: applying a machine learning mixed model trained by any one of the methods described in clauses 1 to 23 to a sample methylation signature of each training sample in a cancer cohort to identify a subset of methylated sequence reads originating from non-cancer impurity types; excluding, for each training sample in the cancer cohort, the subset of methylated sequence reads originating from non-cancer impurity types, thereby obtaining a feature set of methylated sequence reads; generating, for each training sample in the non-cancer cohort, a feature set of methylated sequence reads; and training a cancer classifier with the feature set of methylated sequence reads for the training sample from the cancer cohort and the feature set of methylated sequence reads for the training sample from the non-cancer cohort.

[0045] Clause 36. The method of clause 35, wherein the cancer classifier is trained as a machine learning model.

[0046] Clause 37. A method for training a cancer classifier, comprising: obtaining, for each training sample of a plurality of training samples, a set of methylated sequence reads obtained from sequencing DNA fragments of the training sample, the plurality of training samples including training samples obtained from a first cohort of subjects diagnosed with cancer and training samples obtained from a second cohort of subjects not diagnosed with cancer; for each training sample in the first cohort and for each methylated sequence read of the training sample: applying a tissue type model trained to predict the likelihood that the methylated sequence read is obtained from non-cancerous tissue, and excluding the methylated sequence read if the predicted likelihood exceeds a threshold; generating, for each training sample in the first cohort, a methylation signature based on the methylated sequence reads of the training sample that were not excluded; generating, for each training sample in the second cohort, a methylation signature based on the methylated sequence reads of the training sample; and training a cancer classifier to detect the presence of a cancer signal having the methylation signature of the training sample.

[0047] Clause 38. The method of clause 37, wherein the tissue type model is trained according to a beta distribution using methylation signatures of non-cancer training specimens generated from methylated sequence reads obtained by sequencing DNA fragments of the non-cancer training specimens.

[0048] Clause 39. The method of clause 38, wherein the methylation signatures of non-cancer training samples are obtained from a reference database.

[0049] Clause 40. The method of any one of clauses 37 to 39, wherein the cancer classifier is trained as a machine learning model.

[0050] Clause 41. The method of any one of clauses 37 to 40, wherein the methylation signature is based on identification of one or more methylation mutations in each genomic region of the plurality of genomic regions.

[0051] Clause 42. A method for identifying a tissue type of origin for a DNA fragment, the method comprising: obtaining methylated sequence reads for the DNA fragment, the methylated sequence reads comprising a methylation signature across one or more genomic regions; predicting the likelihood that the methylation signature of the methylated sequence read is from each of a plurality of tissue types of origin by applying each of a plurality of tissue type models to the methylation signature, the plurality of tissue type models optionally being part of a mixture model trained by a method described in any one of clauses 1 to 23; determining that the DNA fragment is from the most likely tissue type of origin; and returning a report of the determination.

[0052] Clause 43. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the computer processor to perform the method described in any one of clauses 1 to 42.

[0053] Clause 44. A system including a computer processor; and the non-transitory computer-readable storage medium described in Clause 43.

[0054] Clause 45. A computer program product including a non-transitory computer-readable storage medium storing a machine learning mixture model for predicting tissue type proportions from a test specimen, the product being made by a method described in any one of clauses 1-23.

[0055] Clause 46. A computer program product including a non-transitory computer-readable storage medium storing a machine learning cancer classifier for predicting cancer in a test specimen, the product being made by a method described in any one of clauses 33-41.

[0056] Clause 47. A treatment kit comprising: a collection container for collecting a DNA sample from a subject; optionally, one or more reagents for isolating DNA fragments in the DNA sample; optionally, one or more probes targeting one or more genomic loci determined to be indicative of a cancerous condition; and a non-transitory computer-readable storage medium described in Clause 43 or a computer program product described in claim 45 or 46. [Brief explanation of the drawings]

[0057] [Figure 1] 1 is an exemplary flowchart describing the overall workflow of cancer classification of a specimen, according to one or more embodiments. [Figure 2A] 1 is an illustrative flow chart of a nucleic acid sample sequencing device according to one or more embodiments. [Figure 2B] FIG. 1 is a block diagram of an analytics system for processing a DNA specimen, according to one or more embodiments. [Figure 3A] 1 is an illustrative flowchart describing the process of sequencing cell-free (cf) DNA fragments to obtain a methylation status vector, according to one or more embodiments. [Figure 3B] FIG. 2B is an illustrative illustration of the cell-free (cf) DNA fragment sequencing process to obtain the methylation state vector of FIG. 2A according to one or more embodiments. [Figure 4A] 1 is an exemplary flowchart describing the process of training a mixture model to predict component proportions in a sample, according to one or more embodiments. [Figure 4B] 1 is an exemplary flowchart describing a deployment process of a mixture model for predicting component proportions in a specimen, according to one or more embodiments. [Figure 5A] 1 illustrates the methylation signature that can be obtained from a single CpG site as a genomic region, according to one or more embodiments. [Figure 5B] 1 illustrates a methylation signature that can be obtained from multiple CpG sites as a genomic region, according to one or more embodiments. [Figure 6] 1 illustrates an exemplary methylation signature matrix used in training a mixture model, according to an exemplary embodiment. [Figure 7] 1 illustrates an example architecture of a mixed model according to one or more embodiments. [Figure 8A] 1 is an illustrative flowchart describing a process for generating a control group data structure for determining aberrant methylation fragments according to one or more embodiments. [Figure 8B] 1 is an illustrative flowchart describing a process for determining a fragment as aberrantly methylated based on a control group data structure according to one or more embodiments. [Figure 9A] 1 is an exemplary flowchart describing a process for training a cancer classifier according to one or more embodiments. [Figure 9B] 1 illustrates an exemplary generation of feature vectors used to train a cancer classifier, according to one or more embodiments. [Figure 10] 10 is an exemplary result of a two-component mixture model according to an exemplary embodiment. [Figure 11] 10 is an exemplary result of a five-component mixture model according to an exemplary embodiment. [Figure 12] 10A-10C are exemplary results of deconvolution of hormone receptor status positive breast tissue, hormone receptor status negative breast tissue, prostate tissue, and uterine tissue with a mixture model, according to an exemplary embodiment. [Figure 13A] 10 is an exemplary result of deconvolving non-cancer impurities from colorectal tissue in a colorectal specimen with a mixture model, according to an exemplary embodiment. [Figure 13B] 10 is an exemplary result of deconvolving non-cancerous impurities from bladder tissue in a bladder specimen with a mixture model, according to an exemplary embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0058] These figures depict various embodiments for illustrative purposes only, and those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.

[0059] Detailed Description I. Overview Early detection and classification of cancer is an important technology. Detecting cancer before it becomes symptomatic benefits everyone involved, including the patient, physician, and loved ones. For the patient, early cancer detection allows for a greater chance of a beneficial outcome; for the physician, early cancer detection allows for an increased number of treatment pathways that may lead to beneficial outcomes; for loved ones, early cancer detection increases the chance of avoiding losing friends and family to the disease.

[0060] In recent years, early cancer detection technology has advanced toward analyzing genetic fragments (e.g., DNA) in a person's blood, for example, to determine whether those genetic fragments originate from cancer cells. These novel techniques enable physicians to identify the presence of cancer in patients that may otherwise be undetectable by traditional screening processes. For example, consider a person at high risk for breast cancer. Traditionally, this person would undergo regular mammograms with their physician, and the physician would use the images (e.g., X-rays) of their breast tissue produced by the mammogram to identify cancerous tissue. Unfortunately, even the highest-resolution mammograms only allow physicians to identify tumors once they are approximately millimeters in size. This means that the cancer has been present in the person for some time and has remained undiagnosed and untreated. This visual determination is typical for most cancers—that is, they can only be identified once they have grown large enough to be detected by some type of imaging technology.

[0061] Detecting cancer using analysis of gene fragments in a patient's blood, for example, alleviates this problem. Specifically, as soon as cancer cells form, they begin shedding DNA fragments into a person's bloodstream. This occurs when cancer cells are very scarce and would not yet be visible using imaging techniques. Thus, in the right way, a system that analyzes DNA fragments in the bloodstream could potentially use the shed cancer DNA fragments to identify the presence of cancer in a person, and more importantly, it could do so before the cancer is identifiable using conventional cancer detection techniques.

[0062] Cancer detection based on the analysis of DNA fragments is made possible by next-generation sequencing ("NGS") techniques. NGS, broadly defined, is a group of technologies that allow high-throughput sequencing of genetic material. As discussed in more detail herein, NGS broadly consists of (1) specimen preparation, (2) DNA sequencing, and (3) data analysis. Specimen preparation is the laboratory method required to prepare DNA fragments for sequencing, sequencing is the process of reading the ordered nucleotides in the specimen, and data analysis is the processing and analysis of the genetic information of the sequencing data to identify the presence of cancer.

[0063] While these NGS steps can help achieve early cancer detection, they also introduce their own complex and detrimental problems to cancer detection; therefore, any improvement in specimen preparation, DNA sequencing, and / or data analysis, including preprocessing, algorithmic processing, and summarizing or presenting predications or conclusions, will lead to improvements in cancer detection technology and, more generally, in early cancer detection.

[0064] To illustrate, for example, (1) issues introduced in specimen preparation include DNA specimen quality, specimen contamination, fragmentation bias, and accurate indexing. Addressing these issues may result in better genetic data for cancer detection.

[0065] Similarly, (2) problems introduced into sequencing include, for example, errors in the precise transcription of fragments (e.g., reading "A" instead of "C"), inaccurate or difficult fragment assembly and overlap, inconsistent coverage uniformity, the conflict between sequencing depth, cost, and specificity, and insufficient sequencing length. Again, addressing any of these problems could result in improved genetic data for cancer detection.

[0066] (3) The data analysis problem is the most daunting and complex. The challenge posed is due to the sheer volume of data generated by NGS sequencing techniques. Sequencing data for a single sample can amount to several terabytes of data, with sequence reads on the order of hundreds of thousands (up to millions). This is multiplied by the thousands (up to tens of thousands) of samples collected for use in training analytical models. Effectively and efficiently analyzing this amount of data is both procedurally and computationally demanding. For example, analysis of NGS sequencing involves several baseline processing steps, such as aligning reads to each other, aligning and mapping reads to a reference genome, deduplication of duplicated reads, detecting sample contamination, identifying and calling mutated genes, identifying and calling aberrantly methylated genes, and generating functional annotations. Performing any of these processes on terabytes of genetic data is computationally expensive, even with the most powerful computer architectures, and simply impossible for the average human brain. Additionally, with genetic sequencing data obtained from error-prone specimen preparation and sequencing processes, a large portion of the resulting genetic data may be of low quality or unusable for cancer identification. For example, large amounts of genetic data may contain contaminated samples, transcription errors, mismatched regions, overrepresented regions, etc., making them unsuitable for high-precision cancer detection. Identifying and accounting for low-quality genetic data across the vast amount of genetic data obtained from NGS sequencing is also procedurally and computationally demanding to achieve and is practically impossible for the human brain. Overall, creating any process that leads to more efficient processing of large-scale array sequencing data could improve cancer detection using NGS sequencing. Moreover, such processes are conceived as solutions to various obstacles encountered in NGS sequencing and are therefore atypical and unconventional tasks in the field of technical endeavors.

[0067] In particular, under (3) data analysis, accurately identifying abnormal DNA from NGS data to identify the presence of cancer is also an immediate challenge. To be effective, algorithms are required that compensate for errors generated by, for example, sample preparation and sequencing, and overcome the large-scale data analysis problems associated with NGS techniques. That is, the design of one or more machine learning models or other computational algorithms that achieve early cancer detection based on next-generation sequencing techniques must be configured to take into account the challenges these techniques create. Some of these techniques and models are discussed below, and detailed improvements to current state-of-the-art techniques and models are further discussed. Furthermore, such techniques are atypical and unconventional tasks in the field of technical endeavor.

[0068] One particular challenge arises with tissue specimens. Generally, the process of obtaining a tissue specimen requires removal of tissue from a subject via a tissue biopsy. The tissue is thinly sliced, placed on a slide, stained (e.g., with hematoxylin and eosin (H&E)), annotated by a pathologist, and dissociated based on the pathologist's annotation. The dissociated portion is then used to isolate nucleic acid fragments associated with the tissue, e.g., by lysing the cells and isolating the nucleic acid fragments. This process relies on accurate annotation by the pathologist and accurate dissociation of the target portion to serve as the source of nucleic acid fragments. Impurities can be introduced into the tissue specimen at various points during this process. For example, inaccurate annotation and / or dissociation can result in the dissociated portion being overly inclusive, including tissues other than the target portion. In cancer tissue specimens, this can result in the tissue specimen containing both tumorous and healthy tissue, thereby obscuring the purity of the tissue specimen. Or in other cases, the tissue specimen may contain excess non-target tissue types, thereby also obscuring the purity of the specimen.

[0069] One or more present inventions describe mixture models that can determine the tissue purity of a specimen, thus overcoming the above-mentioned challenges. Specimens that are deemed impure, i.e., specimens with contamination higher than an acceptable tolerance, can be removed from the training dataset to prevent distortion of downstream analyses. For example, when training a cancer classifier to classify tumor specimens among different cancer types, the mixture model can determine the purity of the cancer specimens used to train the cancer classifier.

[0070] Another particular challenge arises in the context of intratumor heterogeneity, where cancer specimens originating from subjects diagnosed with cancer may contain heterogeneous tumors. Such tumors are a mixture of two or more cancer subtypes, each potentially with its own tumor biology and / or genetic mutations. Training on such specimens in the context of a single tumor biology can blur classification lines and distort the predictive accuracy of one or more analytical models.

[0071] One or more present inventions describe a mixture model that can parse different cancer subtypes with various methylation signatures in heterogeneous tumor samples. The mixture model can be trained to deconvolve various subtypes in heterogeneous tumor samples. Samples with a significant proportion of two or more cancer subtypes can be left out for training the analytical model. In another exemplary application, the mixture model can be trained to predict cancer subtype, for example, as a downstream analysis upon detecting the likelihood of cancer in a test sample.

[0072] IA Cancer Classification Workflow FIG. 1 is an exemplary flowchart describing an overall workflow 100 for cancer classification of a specimen, according to one or more embodiments. Workflow 100 may be performed by one or more entities, including, for example, a healthcare provider, a sequencing device, an analytics system, etc. The purpose of the workflow may include detecting and / or monitoring cancer in an individual. From a medical perspective, workflow 100 may serve to complement other existing cancer diagnostic tools. Workflow 100 may serve to provide early cancer detection and / or routine cancer monitoring to better inform treatment plans for individuals diagnosed with cancer. This overall workflow 100 may include additional / fewer steps than those shown in FIG. 1.

[0073] A healthcare provider performs specimen collection 110. An individual undergoing cancer classification visits their healthcare provider. The healthcare provider collects a specimen for cancer classification. Examples of biological specimens include, but are not limited to, a subject's tissue biopsy, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid. The specimen contains genetic material belonging to the individual that may be extracted and sequenced for cancer classification. Once specimen collection is complete, the specimen is provided to a sequencing device. The healthcare provider may collect other information about the individual along with the specimen, such as biological sex, age, race, smoking status, any prior diagnoses, etc.

[0074] A sequencing device performs specimen sequencing 120. A laboratory clinician may perform one or more processing steps on the specimen to prepare it for sequencing. Once prepared, the clinician loads the specimen into the sequencing device. Examples of devices utilized for sequencing are further described in conjunction with Figures 2A and 2B. The sequencing device generally extracts and isolates fragments of the nucleic acid to be sequenced and determines the sequence of nucleic acid bases corresponding to the fragments. Sequencing may also include amplification of the nucleic acid material. Various sequencing processes include Sanger sequencing, fragment analysis, and next-generation sequencing. Sequencing may be whole-genome sequencing or targeted sequencing using a targeted panel. In the context of DNA methylation, bisulfite sequencing (e.g., as further described in Figures 3A and 3B) can determine methylation status through the conversion of unmethylated cytosines at CpG sites with bisulfite. Sample sequencing 120 generates sequences for multiple nucleic acid fragments in the sample. In one or more embodiments, these sequences may include methylation state vectors, where each methylation state vector describes the methylation state for CpG sites on the fragment.

[0075] The analytics system performs pre-analysis processing 130. An exemplary analytics system is depicted in Figure 2B. Pre-analysis processing 130 may include, but is not limited to, de-duplication of sequence reads, determining coverage metrics, determining whether a sample is contaminated, removing contaminating fragments, calling sequencing errors, etc.

[0076] The analytics system performs one or more analyses 140, which involve predicting at least the cancer status of the individual from whom the sample was obtained, through statistical analysis or application of one or more trained models. Various genetic features may be evaluated and considered, such as methylation of CpG sites, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), and other types of genetic mutations. In the context of methylation, analysis 140 may include tissue purity determination 142 (further described, e.g., in Figures 4A, 4B, 5A, 5B, 6, and 7), feature extraction 144, and application of a cancer classifier 146 to determine a cancer prediction (further described, e.g., in Figures 9A and 9B). Tissue purity determination 142 involves applying a mixture model to deconvolute the proportions of tissue components that contribute DNA fragments to the sample. In general, tissue purity determination 142 can be used to determine, for a given sample, the proportion of methylated sequence reads that are omitted from cancerous tissue compared to the proportion of methylated sequence reads from non-cancerous cells. Tissue purity determination 142 can be particularly useful in deconvolving heterogeneous tumors containing multiple clonal populations that may have different gene signatures. Application of a mixture model can also quantify the cancer signal or grade the cancer state based on the determined proportions. Cancer classifier 146 inputs the extracted features to determine a cancer prediction. The cancer prediction can be a label or a value. The label may indicate a particular cancer state; for example, a binary label can indicate the presence or absence of cancer, while a multi-class label can indicate one or more cancer types, cancer stages, etc., from multiple cancer types being screened for. The value may indicate the likelihood of a particular cancer state, e.g., the likelihood of cancer and / or the likelihood of a particular cancer type.

[0077] The analytics system returns a prediction 150 to the healthcare provider. The prediction 150 may include a binary prediction of cancer presence or absence, detailed cancer type, cancer stage, tissue proportions, etc. The healthcare provider may develop or adjust a treatment plan based on the returned prediction 150. Treatment optimization is further described in Section VC, Treatment.

[0078] Overview of IB methylation As described herein, cfDNA fragments from an individual are processed, for example, by converting unmethylated cytosines to uracils, and sequenced. The sequence reads are compared with a reference genome to identify the methylation status of specific CpG sites within the DNA fragments. Each CpG site can be methylated or unmethylated. Identifying abnormally methylated fragments compared with healthy individuals can provide insight into the cancer status of the subject. As is well known in the art, abnormalities in DNA methylation (compared to healthy controls) can cause various effects, which can contribute to cancer. Identifying abnormally methylated cfDNA fragments presents various challenges. First, determining that a DNA fragment is abnormally methylated can carry weight in comparison with a group of control individuals, and when the number of controls is small, the determination becomes unreliable due to statistical variability within the small size of the control group. In addition, there may be variation in the methylation status among a group of control individuals, which can be difficult to consider when determining that a subject's DNA fragment is abnormally methylated. Alternatively, methylation of cytosines at CpG sites can causally influence the methylation of subsequent CpG sites, and including this dependency can pose another challenge in itself.

[0079] Methylation typically occurs in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. Specifically, methylation can occur at the dinucleotide of cytosine and guanine, referred to herein as a "CpG site." In other cases, methylation can occur at cytosine that is not part of a CpG site or at other nucleotides that are not cytosine; however, these occurrences are rare. In this disclosure, methylation is discussed in terms of CpG sites for clarity. DNA methylation abnormalities can be identified as hypermethylation or hypomethylation, both of which can be indicative of a cancerous state. Throughout this disclosure, a DNA fragment can be characterized as hypermethylated or hypomethylated if the DNA fragment contains a number of CpG sites greater than a threshold, and a percentage of those CpG sites is methylated or unmethylated above a threshold.

[0080] The principles described herein may be equally applicable to detecting methylation in non-CpG contexts, including non-cytosine methylation. In such embodiments, the wet-lab assays used to detect methylation may differ from those described herein. Furthermore, the methylation state vectors discussed herein may include elements that are generally methylated or unmethylated sites (even if such sites are not specifically CpG sites). With this substitution, the remainder of the process described herein may remain the same, and consequently, the inventive concepts described herein may be applicable to such other forms of methylation.

[0081] IC definition The term "cell-free nucleic acid" or "cfNA" refers to nucleic acid fragments originating from one or more healthy cells and / or one or more unhealthy cells (e.g., cancer cells) circulating in an individual's body (e.g., blood). The term "cell-free DNA" or "cfDNA" refers to deoxyribonucleic acid fragments circulating in an individual's body (e.g., blood). In addition, cfNA or cfDNA in an individual's body can be derived from other non-human sources.

[0082] The terms "genomic nucleic acid," "genomic DNA," or "gDNA" refer to nucleic acid or deoxyribonucleic acid molecules obtained from one or more cells. In various embodiments, gDNA can be extracted from healthy cells (e.g., non-tumor cells) or from tumor cells (e.g., biopsy specimens). In some embodiments, gDNA can be extracted from cells derived from the blood cell lineage, such as white blood cells.

[0083] The term "circulating tumor DNA" or "ctDNA" refers to nucleic acid fragments that originate from tumor cells or other types of cancer cells and that may be released into an individual's body fluids (e.g., blood, sweat, urine, or saliva) as a result of biological processes such as apoptosis or necrosis of dying cells, or that may be actively released by viable tumor cells.

[0084] The terms "DNA fragment," "fragment," or "DNA molecule" may generally refer to any deoxyribonucleic acid fragment, i.e., cfDNA, gDNA, ctDNA, etc.

[0085] The terms "aberrant fragment," "aberrant methylation fragment," or "fragment with an abnormal methylation pattern" refer to a fragment that has abnormal methylation at CpG sites. The methylation abnormality of a fragment can be determined by using a probability model to identify how unexpected it is to observe the methylation pattern of the fragment in a control group.

[0086] The term "unusual fragment with extreme methylation" or "UFXM" refers to a hypomethylated or hypermethylated fragment. Hypomethylated and hypermethylated fragments refer to fragments in which at least a certain number (e.g., 5) of CpG sites are methylated or unmethylated, respectively, above a certain threshold percentage (e.g., 90%).

[0087] The term "abnormality score" refers to a score for a CpG site based on the number of abnormal fragments (or UFXMs, in some embodiments) from a sample that overlap that CpG site. Anomaly scores are used in the context of featurizing samples for classification.

[0088] As used herein, the term "about" or "approximately" may mean within an acceptable range of error for a particular value as determined by one of ordinary skill in the art; such error may depend, in part, on how the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" may mean within one standard deviation or more, according to practice in the art. "About" may mean within ±20%, ±10%, ±5%, or ±1% of a given value. The term "about" or "approximately" may mean within a certain magnitude, within 5-fold, or within 2-fold of a value. When a particular value is described in the present application and claims, unless otherwise specified, the term "about" should be assumed to mean within an acceptable range of error for that particular value. The term "about" may have the meaning commonly understood by one of ordinary skill in the art. The term "about" may refer to ±10%. The term "about" may refer to ±5%.

[0089] As used herein, the terms "biological specimen," "patient specimen," or "specimen" refer to any specimen taken from a subject, which may reflect a biological state associated with the subject, including cell-free DNA. Examples of biological specimens include, but are not limited to, a subject's blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid. A biological specimen may include any tissue or material derived from a living or deceased subject. A biological specimen may be a cell-free specimen. A biological specimen may contain nucleic acid (e.g., DNA or RNA) or a fragment thereof. The term "nucleic acid" may refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or any hybrid or fragment thereof. Nucleic acid in a specimen may be cell-free nucleic acid. A specimen may be a liquid specimen or a solid specimen (e.g., a cell or tissue specimen). The biological specimen may be a bodily fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., testicular), vaginal washings, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple secretions, aspirates from various body sites (e.g., thyroid, breast), etc. The biological specimen may also be a stool specimen. In various embodiments, the majority of the DNA in a biological specimen enriched for cell-free DNA (e.g., a plasma specimen obtained by a centrifugation protocol) may be cell-free (e.g., greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free). The biological specimen may be subjected to a process that physically disrupts tissue or cellular structure (e.g., centrifugation and / or cell lysis), which may result in the release of intracellular components into a solution that may further contain enzymes, buffers, salts, detergents, etc., which may be used to prepare the specimen for analysis.

[0090] As used herein, the terms "control," "control specimen," "reference," "reference specimen," "normal," and "normal specimen" refer to specimens from subjects who do not have a particular condition or who are otherwise healthy. As an example, a method as disclosed herein can be performed on a subject with a tumor, where the reference specimen is a specimen taken from the subject's healthy tissue. The reference specimen can be obtained from the subject or from a database. The reference can be, for example, a reference genome used to map nucleic acid fragment sequences obtained by sequencing a specimen from the subject. A reference genome can refer to a haploid or diploid genome to which nucleic acid fragment sequences from a biological specimen and a constitutional specimen can be aligned and compared. An example of a constitutional specimen can be the DNA of white blood cells obtained from a subject. For a haploid genome, there can only be one nucleotide at each locus. For a diploid genome, heterozygous loci can be identified; each heterozygous locus has two alleles, where either allele can match when aligned with the locus.

[0091] As used herein, the term "cancer" or "tumor" refers to an abnormal mass of tissue in which the growth of the mass exceeds and is uncoordinated with the growth of normal tissue.

[0092] As used herein, the phrase "healthy" refers to a subject in good health. A healthy subject may demonstrate the absence of any malignant or non-malignant disease. A "healthy individual" may have other diseases or conditions unrelated to the condition under assay and which may not normally be considered "healthy."

[0093] As used herein, the term "methylation" refers to a modification of deoxyribonucleic acid (DNA) in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. Specifically, methylation tends to occur at the dinucleotide of cytosine and guanine, referred to herein as "CpG sites." In other cases, methylation may occur at cytosine that is not part of a CpG site or at other nucleotides that are not cytosine; however, these occurrences are rare. cfDNA methylation abnormalities can be identified as hypermethylation or hypomethylation, both of which may indicate a cancerous state. Abnormalities in DNA methylation (compared to healthy controls) can cause various effects, which may contribute to cancer. The principles described herein are equally applicable to detecting methylation in CpG contexts and non-CpG contexts, including non-cytosine methylation. Furthermore, a methylation state vector may contain elements that are vectors of generally methylated or unmethylated sites (even if such sites are not specifically CpG sites).

[0094] As used interchangeably herein, the terms "methylation fragment" or "nucleic acid methylation fragment" refer to a sequence of methylation states for each CpG site among a plurality of CpG sites, as determined by methylation sequencing of a nucleic acid (e.g., a nucleic acid molecule and / or a nucleic acid fragment). In a methylation fragment, the position and methylation state for each CpG site in a nucleic acid fragment are determined based on alignment of sequence reads (e.g., obtained from sequencing the nucleic acid) to a reference genome. A nucleic acid methylation fragment includes a methylation state (e.g., a methylation state vector) for each CpG site among a plurality of CpG sites, which specifies the position of the nucleic acid fragment in the reference genome (e.g., as specified by the position of the first CpG site in the nucleic acid fragment using a CpG index or another similar metric) and the number of CpG sites in the nucleic acid fragment. Alignment of sequence reads to a reference genome based on methylation sequencing of a nucleic acid molecule can be performed using a CpG index. As used herein, the term "CpG index" refers to a list of each CpG site (e.g., CpG1, CpG2, CpG3, etc.) in a plurality of CpG sites in a reference genome, such as a human reference genome, which may be in electronic format. The CpG index further includes, for each CpG site in the CpG index, a corresponding genomic location in the corresponding reference genome. Thus, for each nucleic acid methylation fragment, each CpG site is indexed to a specific location in the respective reference genome, which can be determined using the CpG index.

[0095] As used herein, the term "methylation mutation" refers to a characteristic methylation pattern in a set of k (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 25, etc.) consecutive CpG sites. Methylation mutations can be used to distinguish DNA from non-cancerous plasma from specific cancer types. A given methylation mutation has a single, defined "mutation state," which is a single, specific, detectable pattern of methylation status. A "reference state" is any methylation state pattern that is not a mutation state. As one or more examples, a set of five consecutive CpGs can have a mutation state of five methylated CpGs (e.g., MMMMM, where "M" represents methylated), another mutation state of five unmethylated CpGs (e.g., UUUUU, where "U" represents unmethylated), or a third mutation state that includes a mixture of methylation states (e.g., UMMMU). The reference state may include all other patterns that are not considered to be one of the mutant states.

[0096] As used herein, the term "true positive" (TP) refers to a subject who has a condition. A "true positive" may refer to a subject who has a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), localized or metastatic cancer, or a non-malignant disease. A "true positive" may refer to a subject who has a condition and is identified as having the condition by an assay or method of the present disclosure. As used herein, the term "true negative" (TN) refers to a subject who does not have a condition or who does not have a detectable condition. A true negative may refer to a subject who does not have a disease, such as a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), localized or metastatic cancer, or a non-malignant disease, or who does not have a detectable disease, or who is otherwise healthy. A true negative may refer to a subject who does not have a condition, or who does not have a detectable condition, or who is identified as not having the condition by an assay or method of the present disclosure.

[0097] As used herein, the term "reference genome" refers to any particular known, sequenced, or characterized genome, whether partial or complete, of any organism or virus that can be used as a reference for sequences identified from a subject. Exemplary reference genomes for human subjects, as well as many other organisms, are provided in online genome browsers hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz (UCSC). "Genome" refers to the complete genetic information of an organism or virus expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from one or more individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. A reference genome can be considered a representative example of the gene set of a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).

[0098] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence generated by any sequencing process described herein or known in the art. A read may be generated from one end of a nucleic acid fragment (a "single-end read"), or sometimes from both ends of the nucleic acid (e.g., a paired-end read, a double-end read). In some embodiments, a sequence read (e.g., a single-end or paired-end read) may be generated from one or both strands of a targeted nucleic acid fragment. The length of the sequence read is often related to the detailed sequencing technique. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, sequence reads are between about 15 bp and 900 bp in length (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 450 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp in average, median, or average length). In some embodiments, sequence reads are between about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more in average, median, or average length. or average length. For example, nanopore sequencing may provide sequence reads that can vary in size from tens to hundreds or thousands of base pairs. Illumina parallel sequencing may provide sequence reads that do not vary significantly, e.g., most may be smaller than 200 bp. A sequence read (or sequencing read) may refer to sequence information (e.g., a string of nucleotides) corresponding to a nucleic acid molecule. For example, a sequence read may correspond to a string of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, a string of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of an entire nucleic acid fragment.Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes, for example, in hybridization arrays or capture probes, or using amplification techniques such as polymerase chain reaction (PCR) or linear or isothermal amplification using a single primer.

[0099] As used herein, the term "sequencing" or the like generally refers to any biochemical process that can be used to determine the order of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as a DNA fragment.

[0100] As used herein, the term "sequencing depth" is used synonymously with the term "coverage" and refers to the number of times a locus is covered by consensus sequence reads corresponding to unique nucleic acid target molecules aligned with the locus; for example, sequencing depth is equal to the number of unique nucleic acid target molecules covering the locus. A locus can be as small as a nucleotide, as large as a chromosome arm, or as large as the entire genome. Sequencing depth can be expressed as "Yx", for example, 50x, 100x, etc., where "Y" refers to the number of times a locus is covered by sequences corresponding to nucleic acid targets; for example, the number of independent sequence information covering a particular locus is obtained. In some embodiments, sequencing depth corresponds to the number of genomes sequenced. Sequencing depth can also apply to multiple loci or the entire genome, in which case Y can refer to the average number or average number of times a locus, a haploid genome, or the entire genome is sequenced, respectively. When depth averages are quoted, the actual depths for various loci included in a dataset can span a range of values. Ultra-deep sequencing can refer to a sequencing depth of at least 100x at a locus.

[0101] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity may characterize the ability of an assay or method to correctly identify the proportion of a population that truly has a condition. For example, sensitivity may characterize the ability of a method to correctly identify the number of subjects in a population that have cancer. In another example, sensitivity may characterize the ability of a method to correctly identify one or more markers that are indicative of cancer.

[0102] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly does not have a condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population that do not have cancer. In another example, specificity characterizes the ability of a method to correctly identify one or more markers that indicate cancer.

[0103] As used herein, the term "subject" refers to any living or non-living organism, including, but not limited to, humans (e.g., male humans, female humans, fetuses, pregnant women, children, etc.), non-human animals, plants, bacteria, fungi, or protists. Any human or non-human animal can serve as a subject, including, but not limited to, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cows), equines (e.g., horses), caprines and ovines (e.g., sheep, goats), porcines (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursinians (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, the subject is male or female (e.g., male, female, or child) at any stage of life. The subject from whom the sample is taken or who is treated by any of the methods or compositions described herein can be of any age, and can be an adult, infant, or child.

[0104] As used herein, the term "tissue" may correspond to a group of cells held together as a functional unit. More than one type of cell may be found in a single tissue. Different types of tissue may consist of different types of cells (e.g., liver cells, lung cells, or blood cells), but may also correspond to tissues from different organisms (maternal or fetal), or to healthy or tumor cells. The term "tissue" may generally refer to any group of cells found in the human body (e.g., cardiac tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some embodiments, the terms "tissue" or "tissue type" or "tissue type of origin" can be used to refer to the tissue from which a cell-free nucleic acid molecule or fragment originates. In one example, a viral nucleic acid fragment may be derived from blood tissue. In another example, a viral nucleic acid fragment may be derived from tumor tissue.

[0105] As used herein, the term "genomic" refers to characteristics of an organism's genome. Examples of genomic characteristics include, but are not limited to, the primary nucleic acid sequence of all or a portion of the genome (e.g., the presence or absence of nucleotide polymorphisms, indels, sequence rearrangements, mutation frequencies, etc.), the copy number of one or more specific nucleotide sequences within the genome (e.g., copy number, allele frequency rate, monochromosomal or whole genome ploidy, etc.), the epigenetic state of all or a portion of the genome (e.g., covalent nucleic acid modifications such as methylation, histone modifications, nucleosome positioning, etc.), and the expression profile of an organism's genome (e.g., gene expression levels, isotype expression levels, gene expression rates, etc.).

[0106] The terminology used herein is for the purpose of describing particular cases only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof are used in either the detailed description and / or claims, such terms are intended to be inclusive, similar to the term "comprising."

[0107] ID exemplary analytics system 2A is an exemplary flowchart of an apparatus for sequencing a nucleic acid sample, according to one or more embodiments. This schematic flowchart includes apparatus such as a sequencer 220 and an analytics system 200. The sequencer 220 and the analytics system 200 may cooperate to perform one or more steps of the process.

[0108] In various embodiments, sequencer 220 receives enriched nucleic acid sample 210. As shown in FIG. 2A , sequencer 220 may include a graphical user interface 225 that allows a user to interact with a particular task (e.g., start sequencing or stop sequencing), as well as another loading station 230 for loading a sequencing cartridge containing the enriched fragment sample and / or loading buffers necessary to perform a sequencing assay. Thus, once a user of sequencer 220 has provided the necessary reagents and sequencing cartridge to loading station 230 of sequencer 220, the user can initiate sequencing by interacting with graphical user interface 225 of sequencer 220. Once initiated, sequencer 220 performs sequencing and outputs sequence reads of the enriched fragments from nucleic acid sample 210.

[0109] In some embodiments, the sequencer 220 is communicatively coupled to the analytics system 200. The analytics system 200 comprises several computational devices used to process sequence reads for various applications, such as determining the methylation status at one or more CpG sites, calling variants, or quality control. The sequencer 220 may provide the sequence reads in a BAM file format to the analytics system 200. The analytics system 200 may be communicatively coupled to the sequencer 220 via wireless, wired, or a combination of wireless and wired communication techniques. Generally, the analytics system 200 comprises a processor and a non-transitory computer-readable storage medium storing computer instructions that, when executed by the processor, cause the processor to process the sequence reads or perform one or more steps of any of the methods or processes disclosed herein.

[0110] In some embodiments, alignment position information may be determined by aligning sequence reads with a reference genome using methods known in the art. The alignment position may generally describe the start and end positions of a region in the reference genome corresponding to the starting and ending nucleotide bases of a given sequence read. Generalizing alignment position information to accommodate methylation sequencing may indicate the first and last CpG sites included in the sequence read according to alignment with the reference genome. The alignment position information may further indicate the methylation status and positions of all CpG sites in a given sequence read. Regions in the reference genome may be associated with genes or gene segments; in this way, the analytics system 200 may label the sequence read with one or more genes that align with the sequence read. In one embodiment, fragment length (or size) can be determined from the start and end positions.

[0111] In various embodiments, for example, when a paired-end sequencing process is used, the sequence read is composed of a read pair denoted as R_1 and R_2. For example, the first read R_1 may be sequenced from a first end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 may be sequenced from a second end of the double-stranded DNA (dsDNA). Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 may be aligned (e.g., in reverse orientation) with the nucleotide bases of the reference genome. The alignment position information obtained from the read pair R_1 and R_2 may include a start position in the reference genome corresponding to the end of the first read (e.g., R_1) and an end position in the reference genome corresponding to the end of the second read (e.g., R_2). In other words, the start position and end position in the reference genome may represent the expected position of a nucleic acid fragment in the corresponding reference genome. An output file having a SAM (Sequence Alignment Map) format or a BAM (binary) format can be generated and output for further analysis.

[0112] 2B is a block diagram of an analytics system 200 for processing DNA samples, according to one or more embodiments. The analytics system implements one or more computing devices used to analyze DNA samples. Analytics system 200 includes a sequence processor 240, a sequence database 245, a model database 255, a model 250, a parameter database 265, and a score engine 260. In some embodiments, analytics system 200 performs some or all of the processes described throughout this disclosure.

[0113] The sequence processor 240 generates methylation state vectors for fragments from the specimen. For each CpG site on a fragment, the sequence processor 240 generates a methylation state vector for each fragment, per process 200 of Figure 2A, that specifies the location of the fragment in the reference genome, the number of CpG sites in the fragment, and whether the methylation state of each CpG site in the fragment is methylated, unmethylated, or uncertain. The sequence processor 240 may store the methylation state vectors for the fragments in a sequence database 245. The data in the sequence database 245 may be organized so that the methylation state vectors from the specimen are related to each other.

[0114] Additionally, multiple different models 250 may be stored in the model database 255 or retrieved for use with test samples. Generally, a model receives input and generates an output based on a function that operates on the input according to one or more parameters. In one example, the model is a trained mixture model that deconvolves component proportions based on the methylation signature of the sample. The mixture model may include various submodels, each with its own parameters and one or more functions. In another example, the model is a trained cancer classifier that uses feature vectors derived from abnormal fragments to determine a cancer prediction for a test sample. The training and use of cancer classifiers will be discussed further in connection with Section IV. Cancer Classifiers for Determining Cancer. The analytics system 200 trains one or more models 250 and stores the various trained parameters in the parameter database 265. The analytics system 200 stores the models 250, along with their functions, in the model database 255.

[0115] During inference, the score engine 260 uses one or more models 250 to return an output. The score engine 260 accesses the models 250 in the model database 255 along with the trained parameters from the parameter database 265. For each model, the score engine receives appropriate inputs for that model and calculates an output based on the received inputs, parameters, and a function of each model on the inputs and outputs. In some use cases, the score engine 260 also calculates metrics that correlate to confidence in the output calculated from the model. In other use cases, the score engine 260 calculates other intermediate values ​​for use in the model.

[0116] II. Methylation sequencing of DNA fragments 3A is an exemplary flowchart describing a process 300 for sequencing fragments of cfDNA to obtain a methylation state vector, according to one or more embodiments. To analyze DNA methylation, an analytics system first obtains a specimen from an individual that contains a plurality of cfDNA molecules 310. In further embodiments, process 300 may be applied to sequencing other types of DNA molecules. Process 300 is an embodiment of specimen sequencing 120 of FIG. 1.

[0117] The analytics system may isolate each cfDNA molecule from the specimen 310. The cfDNA molecules may be treated to convert unmethylated cytosines to uracil 320. In one embodiment, the method uses treatment of DNA with bisulfite, which converts unmethylated cytosines to uracil but not methylated cytosines. For example, bisulfite conversion is performed using a commercially available kit, such as the EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, or EZ DNA Methylation™-Lightning kit (available from Zymo Research Corp, Irvine, CA). In another embodiment, the conversion of unmethylated cytosines to uracil is achieved using an enzymatic reaction. For example, this conversion can be performed using a commercially available kit for converting unmethylated cytosines to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).

[0118] Sequencing libraries can be prepared from the converted cfDNA molecules. 330 During library preparation, unique molecular identifiers (UMIs) can be added to nucleic acid molecules (e.g., DNA molecules) by adapter ligation. UMIs can be short nucleic acid sequences (e.g., 4-10 base pairs) that are added to the ends of DNA fragments (e.g., DNA molecules fragmented by physical shearing, enzymatic digestion, and / or chemical fragmentation) during adapter ligation. UMIs can also be degenerate base pairs that act as unique tags that can be used to identify sequence reads originating from a specific DNA fragment. During PCR amplification after adapter ligation, the UMIs can be replicated along with the attached DNA fragments. This can provide a method for identifying sequence reads originating from the same fragment of origin in downstream analysis.

[0119] Optionally, the sequencing library may be enriched for informative cfDNA molecules or genomic regions regarding cancer status using multiple hybridization probes. 335 Hybridization probes are short oligonucleotides capable of hybridizing to specifically designated cfDNA molecules or targeted regions, enriching those fragments or regions for subsequent sequencing and analysis. Hybridization probes can be used to perform targeted, deep analysis of a specified set of CpG sites of interest to researchers. Hybridization probes can be tiled across one or more target sequences at 1X, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, or greater than 10X coverage. For example, hybridization probes tiled at 2X coverage include overlapping probes such that each portion of the target sequence hybridizes to two independent probes. Hybridization probes can be tiled across one or more target sequences at less than 1X coverage.

[0120] In one embodiment, hybridization probes are designed to enrich for DNA molecules that have undergone a treatment to convert unmethylated cytosines to uracils (e.g., using bisulfite). During enrichment, hybridization probes (also referred to herein as "probes") can be used to target and pull down nucleic acid fragments that are informative about the presence or absence of cancer (or disease), cancer status, or cancer classification (e.g., cancer class or tissue of origin). Probes may be designed to anneal (or hybridize) to a target (complementary) strand of DNA. The target strand may be the "plus" strand (e.g., the strand transcribed into mRNA and subsequently translated into protein) or the complementary "minus" strand. Probes can range in length from tens, hundreds, or thousands of base pairs. Probes can be designed based on a methylation site panel. Probes can be designed based on a panel of target genes to analyze detailed mutations or target regions of the genome (e.g., of a human or another organism) suspected to correspond to a particular cancer or other type of disease. Moreover, the probes may cover overlapping portions of the target region.

[0121] Once prepared, the sequencing library or a portion thereof can be sequenced 340 to obtain a plurality of sequence reads. The sequence reads can be in a computer-readable digital format for processing and interpretation by computer software. These sequence reads can be aligned with a reference genome to determine alignment position information. The alignment position information can indicate the start and end positions of a region in the reference genome corresponding to the starting and ending nucleotide bases of a given sequence read. The alignment position information can also include the length of the sequence read, which can be determined from the start and end positions. The region in the reference genome can be associated with a gene or gene segment. The sequence reads can consist of a read pair, denoted R1 and R2. For example, the first read R1 can be sequenced from a first end of a nucleic acid fragment, while the second read R2 can be sequenced from a second end of the nucleic acid fragment. Thus, the nucleotide base pairs of the first read R1 and the second read R2 can be aligned (e.g., in reverse orientation) with the nucleotide bases of the reference genome. The alignment position information obtained from the read pair R1 and R2 may include a start position in the reference genome corresponding to the end of the first read (e.g., R1) and an end position in the reference genome corresponding to the end of the second read (e.g., R2). In other words, the start and end positions in the reference genome may represent the expected positions of the nucleic acid fragments in the corresponding reference genome. An output file having a SAM (Sequence Alignment Map) format or a BAM (Binary Alignment Map) format may be generated and output for further analysis, such as determining methylation status.

[0122] From the sequence reads, the analytics system determines the location and methylation state for each CpG site based on alignment with the reference genome 350. The analytics system generates a methylation state vector 360 that specifies, for each fragment, the location of the fragment in the reference genome (e.g., as specified by the location of the first CpG site in each fragment, or another similar metric), the number of CpG sites in the fragment, and whether the methylation state of each CpG site in the fragment is methylated (e.g., designated M), unmethylated (e.g., designated U), or indeterminate (e.g., designated I). Observed states can be methylated and unmethylated states; while unobserved states are indeterminate. Indeterminate methylation states can originate from sequencing errors and / or mismatches between the methylation states of complementary strands of the DNA fragments. The methylation state vector may be stored in temporary or persistent computer memory for later use and processing. Additionally, the analytics system may remove duplicate reads or duplicate methylation state vectors from a single sample. The analytics system may determine that certain fragments with one or more CpG sites have an uncertain methylation state at a certain threshold number or percentage, and may exclude such fragments, or may selectively include such fragments, but may build a model taking into account such uncertain methylation states.

[0123] 3B is an exemplary illustration of the process 300 of FIG. 3A for sequencing a cfDNA molecule to obtain a methylation state vector, according to one or more embodiments. As an example, an analytics system receives a cfDNA molecule 312, which in this example has three CpG sites. As shown, the first and third CpG sites of the cfDNA molecule 312 are methylated 314. In processing step 320, the cfDNA molecule 312 is converted to generate a converted cfDNA molecule 322. In processing 320, the second CpG site, which was previously unmethylated, has its cytosine converted to uracil. However, the first and third CpG sites were not converted.

[0124] After conversion, a sequencing library 330 is prepared and sequenced 340 to generate sequence reads 342. An analytics system aligns 350 the sequence reads 342 to a reference genome 344, which provides context for where in the human genome the fragment cfDNA originated. In this simplified example, the analytics system aligns 350 the sequence reads 342 to correlate three CpG sites with CpG sites 33, 24, and 25 (arbitrary reference identification numbers are used for ease of explanation). Thus, the analytics system can generate information about both the methylation status of all CpG sites in the cfDNA molecule 312 and where the CpG sites map in the human genome. As shown, methylated CpG sites in the sequence read 342 are read as cytosines. In this example, cytosine appears only at the first and third CpG sites in sequence read 342, allowing for the inference that the first and third CpG sites are methylated in the original cfDNA molecule. However, the second CpG site is read as thymine (U is converted to T during sequencing), and therefore, it can be inferred that the second CpG site is unmethylated in the original cfDNA molecule. With these two pieces of information, methylation state and location, the analytics system generates 360 a methylation state vector 352 for fragment cfDNA 312. In this example, the resulting methylation state vector 352 is <M 23 , U 24 , M 25 >, where M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript numbers correspond to the position of each CpG site in the reference genome.

[0125] One or more alternative sequencing methods can be used to obtain sequence reads from nucleic acids in a biological specimen. Such one or more sequencing methods can include any form of sequencing that can be used to obtain measured sequence reads from nucleic acids (e.g., cell-free nucleic acids), including high-throughput sequencing systems such as, but not limited to, the Roche 454 platform, the Applied Biosystems SOLID platform, Helicos True single-molecule DNA sequencing technology, sequencing-by-hybridization platforms from Affymetrix Inc., single-molecule real-time (SMRT) technology from Pacific Biosciences, sequencing-by-synthesis platforms from 454 Life Sciences, Illumina / Solexa, and Helicos Biosciences, and sequencing-by-ligation platforms from Applied Biosystems. ION TORRENT technology and nanopore sequencing from Life technologies can also be used to obtain sequence reads from nucleic acids (e.g., cell-free nucleic acids) in a biological specimen. Sequencing-by-synthesis and reversible terminator-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 3000; HISEQ 4500 (Illumina, San Diego, Calif.)) can be used to obtain sequence reads from cell-free nucleic acids obtained from biological specimens of training subjects to form genotype datasets. Millions of cell-free nucleic acid (e.g., DNA) fragments can be sequenced in parallel. One example of this type of sequencing technology uses a flow cell containing an optically transparent slide with eight individual lanes, to which oligonucleotide anchors (e.g., adapter primers) are attached. Cell-free nucleic acid specimens can contain signals or tags to facilitate detection.Obtaining sequence reads from cell-free nucleic acids obtained from biological specimens can include obtaining quantification information of signals or tags by various techniques, such as flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, genechip analysis, microarrays, mass spectrometry, quantitative cytofluorimetry, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electrolytic suspension, sequencing, and combinations thereof.

[0126] One or more sequencing methods may include a whole genome sequencing assay. A whole genome sequencing assay may include a physical assay that generates sequence reads for the entire genome or a substantial portion of the entire genome, which can be used to determine large variations, such as copy number variations or copy number abnormalities. Such a physical assay may utilize a whole genome sequencing technique or a whole exome sequencing technique. A whole genome sequencing assay may have an average sequencing depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, at least 20x, at least 30x, or at least 40x across the genome of the test subject. In some embodiments, the sequencing depth is about 30,000x. One or more sequencing methods may include a targeted panel sequencing assay. The targeted panel sequencing assay may have an average sequencing depth of at least 50,000x, at least 55,000x, at least 60,000x, or at least 70,000x for the targeted gene panel. The targeted gene panel may include 450 to 500 genes. The targeted gene panel may include a range of 500±5 genes, a range of 500±10 genes, or a range of 500±25 genes.

[0127] The one or more sequencing methods may include paired-end sequencing. The one or more sequencing methods may generate multiple sequence reads. The multiple sequence reads may have an average length ranging from 10 to 700, 50 to 400, or 100 to 300. The one or more sequencing methods may include a methylation sequencing assay. The methylation sequencing may be i) whole-genome methylation sequencing or ii) targeted DNA methylation sequencing using multiple nucleic acid probes. For example, the methylation sequencing is whole-genome bisulfite sequencing (e.g., WGBS). The methylation sequencing may be targeted DNA methylation sequencing using multiple nucleic acid probes targeting the most informative regions of the methylome, a unique methylation database, and a pre-prototype whole-genome and targeted sequencing assay.

[0128] Methylation sequencing can detect one or more 5-methylcytosines (5mC) and / or 5-hydroxymethylcytosines (5hmC) in each nucleic acid methylated fragment. Methylation sequencing can include converting one or more unmethylated cytosines or one or more methylated cytosines in each nucleic acid methylated fragment to one or more corresponding uracils. One or more uracils can be detected as one or more corresponding thymines in methylation sequencing. The conversion of one or more unmethylated cytosines or one or more methylated cytosines can include chemical conversion, enzymatic conversion, or a combination thereof.

[0129] For example, bisulfite conversion involves converting cytosine to uracil while leaving methylated cytosines (e.g., 5-methylcytosine or 5-mC) intact. In some DNA, approximately 95% of the cytosines in the DNA may be unmethylated, and the resulting DNA fragments may contain an abundance of uracil, represented by thymine. Nucleic acids may be treated prior to sequencing using an enzymatic conversion process, which can be performed in a variety of ways. One example of bisulfite-free conversion includes TET-assisted pyridine borane sequencing (TAPS), a bisulfite-free, single-base resolution sequencing method for the non-destructive, direct detection of 5-methylcytosine and 5-hydroxymethylcytosine without affecting unmodified cytosines. The methylation status of a CpG site in a corresponding plurality of CpG sites in each nucleic acid methylated fragment can be methylated when the CpG site is determined to be methylated by methylation sequencing, and can be unmethylated when the CpG site is determined to be unmethylated by methylation sequencing.

[0130] Methylation sequencing assays (e.g., WGBS and / or targeted methylation sequencing) can have an average sequencing depth, including, but not limited to, up to about 1,000x, 2,000x, 3,000x, 5,000x, 10,000x, 15,000x, 20,000x, or 30,000x. Methylation sequencing can have a sequencing depth of more than 30,000x, for example, at least 40,000x or 50,000x. Whole-genome bisulfite sequencing methods can have an average sequencing depth of 20x to 50x, and targeted methylation sequencing methods have an average effective depth of 100x to 1000x, where effective depth can be the equivalent whole-genome bisulfite sequencing coverage to obtain the same number of sequence reads as obtained by targeted methylation sequencing.

[0131] For further details regarding methylation sequencing (e.g., WGBS and / or targeted methylation sequencing), see, e.g., U.S. patent application Ser. No. 16 / 352,602, entitled "Methylation Fragment Anomaly Detection," filed March 13, 2019; U.S. patent application Ser. No. 16 / 719,902, entitled "Systems and Methods for Estimating Cell Source Fractions Using Methylation Information," filed December 18, 2019; and U.S. patent application Ser. No. 17 / 191,914, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," filed March 4, 2021. Other methylation sequencing methods can also be used to obtain methylation patterns of fragments, including those disclosed herein and / or any variations, alternatives, or combinations thereof. For example, methylation sequencing can be used to identify one or more methylation state vectors as described in U.S. patent application Ser. No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification," filed March 13, 2019, or according to any of the techniques disclosed in U.S. patent application Ser. No. 15 / 931,022, entitled "Model-Based Featurization and Classification," filed May 13, 2020. Each reference is incorporated by reference in its entirety.

[0132] Methylation sequencing of nucleic acids and the resulting one or more methylation state vectors can be used to obtain a plurality of nucleic acid methylation fragments. Each corresponding plurality of nucleic acid methylation fragments (e.g., for each genotype dataset) can include more than 100 nucleic acid methylation fragments. The average number of nucleic acid methylation fragments across each corresponding plurality of nucleic acid methylation fragments can include 1,000 or more nucleic acid methylation fragments, 5,000 or more nucleic acid methylation fragments, 10,000 or more nucleic acid methylation fragments, 20,000 or more nucleic acid methylation fragments, or 30,000 or more nucleic acid methylation fragments. The average number of nucleic acid methylation fragments across each corresponding plurality of nucleic acid methylation fragments can be between 10,000 nucleic acid methylation fragments and 50,000 nucleic acid methylation fragments. The corresponding plurality of nucleic acid methylated fragments can include 1,000 or more, 10,000 or more, 100,000 or more, 1 million or more, 10 million or more, 100 million or more, 500 million or more, 1 billion or more, 2 billion or more, 3 billion or more, 4 billion or more, 5 billion or more, 6 billion or more, 7 billion or more, 8 billion or more, 9 billion or more, or 10 billion or more nucleic acid methylated fragments. The average length of the corresponding plurality of nucleic acid methylated fragments can be 140 to 480 nucleotides.

[0133] III. Deconvolution of component ratios Deconvolution of component proportions involves determining the breakdown of component proportions that are the source of nucleic acid fragments in a sample. Component proportions correspond to the breakdown of components that contribute nucleic acid fragments to a sample. In one or more embodiments, component proportions specify the proportion of a sample attributable to each component. For example, the deconvolution process determines a first proportion of nucleic acid fragments in a sample that originate from a first component, a second proportion of nucleic acid fragments in a sample that originate from a second component, and so on for the remaining components. In other embodiments, component proportions may specify a ranking of components according to their predicted proportions. For example, (in the example of a mixture model predicting proportions for three components) component 3 has the greatest contribution of nucleic acid fragments to the sample, followed by component 1, then component 2.

[0134] The deconvolution process models methylation signatures of various components. The methylation signatures correspond to the methylated sequence reads of the components. In some embodiments, the methylation signatures include counts of methylation mutations across multiple genomic regions. In other embodiments, the methylation signatures may further include counts of reference states across multiple genomic regions. The reference states in a genomic region may include methylation patterns that remain unassigned to a mutation state. With respect to these components, each component may be a tissue type, a cell type, or a finer subtype thereof.

[0135] The deconvolution process may be applied at various stages throughout the cancer classification workflow 100 depicted in Figure 1. In one exemplary embodiment, the deconvolution process may be utilized to determine specimen purity, for example, at step 142 of workflow 100. In another exemplary embodiment, the deconvolution process may be utilized in cancer classification 146 to determine detailed cancer type. In yet another exemplary embodiment, the deconvolution process may be utilized to quantify cancer signals, for example, for use in monitoring cancer status and / or progression. Other exemplary embodiments may utilize component ratios in other ways.

[0136] III.A. Mixed Models The analytics system may train a mixture model for deconvolution of component proportions. The mixture model generally predicts component proportions based on methylated sequence reads of a sample. The mixture model generally includes multiple component submodels. The component submodels model the methylation signatures of the components. The number of components to deconvolve in a set of training samples may be tuned as a hyperparameter of the mixture model. After training, the component submodels may input the methylation signatures of the samples and output component likelihoods indicating the likelihood that the methylation signature of the sample is derived from the components.

[0137] 4A and 4B provide an overview of the use of mixture models. An analytics system (or various components thereof) may perform the method of FIGS. 4A and 4B. In other embodiments, other general-purpose computing devices may perform any of the steps of the method. In other embodiments, the method may include additional steps, fewer steps, different steps, or some combination thereof.

[0138] FIG. 4A is an exemplary flowchart describing a method 400 of training a mixture model to predict component proportions in a sample, according to one or more embodiments.

[0139] The analytics system obtains 410 a plurality of training samples. Each training sample may include methylation sequence reads, for example, as generated by sample sequencing 120 in FIG. 1 or, more specifically, sample processing 300 in FIG. 3A. Each methylation sequence read may include methylation information for nucleic acid fragments in the sample. The sample may be a liquid biopsy sample, a tissue-derived sample, an isolated cell sample, another biological sample, or some combination thereof. The sample may include circulating tumor cells, disseminated tumor cells, other types of somatic cells, etc. The component ratios of the training samples may be known or unknown. For example, a first sample may be a tissue-derived sample containing nucleic acid fragments originating from an isolated tissue. This first sample may have a component ratio dominated by the isolated tissue. Another sample may be a purified plasma sample from a healthy subject containing predominantly cell-free nucleic acid fragments. This sample may have a component ratio dominated by non-cancer cell-free nucleic acid fragments. In some exemplary embodiments, the analytics system obtains 1,000 training samples, 2,000 training samples, 3,000 training samples, 4,000 training samples, 5,000 training samples, 6,000 training samples, 7,000 training samples, 8,000 training samples, 9,000 training samples, 10,000 training samples, 20,000 training samples, 30,000 training samples, 40,000 training samples, 50,000 training samples, or more than 50,000 training samples.

[0140] The analytics system generates a methylation signature for each training sample 420, for example, by counting methylation variations across multiple genomic regions. The analytics system may segment the genome into various genomic regions. Each genomic region may span multiple methylation sites, e.g., CpG sites. The analytics system may be genome-spanning, i.e., use whole-genome sequencing of the entire human genome for the human training sample. In other embodiments, the methylated sequence reads are generated from a targeted methylation assay that targets a subset of genomic regions in the genome. The analytics system may also filter out genomic regions whose average sequencing depth across the training sample is below a threshold depth. The methylation signature may be based on other types of methylation features (e.g., methylation density at each of multiple methylation sites, methylation density across the genomic region, counts of highly methylated fragments overlapping the genomic region, counts of highly unmethylated fragments overlapping the genomic region, etc.).

[0141] In each genomic region, the analytics system may count methylated sequence reads, each of which may be characterized by one of multiple methylation mutations. For example, a set of CpGs may be associated with multiple methylation mutations, each of which is indicative of a different cancer. Each methylation mutation may have a defined mutation state. If there are two mutation states, one mutation state may be considered the primary mutation state, while the other mutation state may be considered a secondary mutation state. The analytics system may count a first count of methylated sequence reads (or nucleic acid fragments) having the primary mutation state and a second count of methylated sequence reads (or nucleic acid fragments) having the secondary mutation state. If there are three or more mutation states, there may be a primary mutation state, a secondary mutation state, and multiple other mutation states. The analytics system may count methylated sequence reads having each of those mutation states, i.e., a first count for the primary mutation state and a count for each of the other mutation states. In other embodiments, the analytics system may perform a first count for the primary mutational state and a second count that aggregates one or more other mutational states. In some embodiments, the analytics system may count only alternative mutational states. Figures 5A and 5B illustrate mutational state counts.

[0142] In one or more embodiments, the analytics system may perform genomic region selection. The analytics system may exclude genomic regions according to one or more selection criteria. One exemplary selection criterion sets a limit on the number of methylation variants that can be included in a genomic region. For example, if three or more methylation variants are present in a genomic region at a significantly high rate in a population, the analytics system may exclude such a genomic region. A significantly high rate may also be referred to as prevalence, and may be 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 20%, 25%, or 30%. Another selection criterion may be rank, which may exclude genomic regions according to their discriminatory power between components. For example, when training a mixture model, the analytics system may generate an information gain score for each genomic region based on how informative the genomic region is in deconvolution. Genomic regions with scores below a threshold may be excluded.

[0143] 5A illustrates counting methylation variations from a single CpG site 505 as a genomic region, according to one or more embodiments. A single CpG site genomic region generally has two methylation states (with respect to methylation variations): a reference state and a mutated state. The analytics system may count a first count of methylated sequence reads having the reference state and a second count of methylated sequence reads having the mutated state.

[0144] For example, in FIG. 5A, there are six fragments 510 overlapping a single CpG site 505. At each CpG site on the fragment, represented by a diamond, the fragment has a methylation state. Methylation states may include methylated, represented as solid, unmethylated, represented as unfilled, and unknown, represented as diagonal shading. Unknown methylation may include states that are undetermined due to mutation or sequencing error. As a single site, there are two possible methylation states (M or U) for this genomic region. One methylation state is considered the reference state, while the other is considered the variant state. For example, for CpG site 505 as a genomic region, the reference genome provides that this CpG site is predominantly methylated, thus the reference state is methylated. In terms of counting, the analytics system may count four fragments (or methylated sequence reads) as having a reference state of methylation and two fragments (or methylated sequence reads) as having a variant state of unmethylated.

[0145] FIG. 5B illustrates counting methylation variations from a genomic region 515 spanning multiple CpG sites 517, according to one or more embodiments. CpG sites 517 include CpG sites 1, 2, and 3. As in FIG. 5A, filled diamonds indicate methylation, open diamonds indicate unmethylation, and diagonally shaded diamonds indicate unknown. In the illustrated embodiment, fragment 1 520A may have a methylation pattern that includes methylation at CpG site 1, unmethylation at CpG site 2, and methylation at CpG site 3. By way of example, this may be a reference state. The analytics system may count the number of fragments in genomic region 515 that have the reference states of methylation, unmethylation, and methylation. In this example, four fragments have this reference state. Fragment 2 520B has a methylation variation that differs from the reference state. Fragment 2 520B has a methylation pattern in which all three CpG sites 517 are methylated. The analytics system may count a total of one fragment (or methylated sequence read) as having this first methylation mutation. Fragment 4 520D has another methylation mutation that differs from the reference state and the first methylation mutation. Fragment 4 520D has a methylation pattern in which all three CpG sites 517 are unmethylated. The analytics system may also count a total of one fragment (or methylated sequence read) as having this second methylation mutation. The analytics system may also add up the counts of all methylation mutations. In such an example, the count of all methylation mutations is 2, including Fragment 2 520A with the first methylation mutation and Fragment 4 520D with the second methylation mutation.

[0146] Returning to FIG. 4A , the analytics system generates 420 a methylation signature based on the counts of methylation variations. In one embodiment, the methylation signature is a count of one or more methylation variations for each of a plurality of genomic regions. In another embodiment, the methylation signature is a ratio of one or more methylation variations for each of a plurality of genomic regions. To calculate the ratio of one or more methylation variations for a genomic region, the analytics system counts the total count of methylated sequence reads present in the genomic region and the count of methylated sequence reads with one or more alternative methylation variations. The ratio is equal to the count of methylated sequence reads with one or more methylation variations relative to the total count of methylated sequence reads present in the genomic region. For example, if there are 10,000 genomic regions, the methylation signature is a vector of length 10,000. The first value of the vector corresponds to the count or percentage of one or more methylation variations in the first genomic region, the second value of the vector corresponds to the count or percentage of one or more methylation variations in the second genomic region, and so on up to 10,000 genomic regions.

[0147] In some embodiments, the methylation signature may further include counts or percentages of reference states for each genomic region. For the example of 10,000 genomic regions, the methylation signature may include 20,000 values, including 10,000 counts or percentages of reference states and 10,000 counts or percentages of methylation variations across the genomic regions.

[0148] In some embodiments, the methylation signature may further include one or more counts or percentages for each methylation variation in a genomic region. For example, a population expresses two or more alternative methylation variations in a genomic region at a statistically significant proportion of the population (e.g., greater than 5%, 10%, 15%, etc.). The methylation signature may also include counts or percentages for each methylation variation in the genomic region, e.g., a first count for a first methylation variation, a second count for a second methylation variation with a different methylation pattern than the first methylation variation, and so on.

[0149] In one or more embodiments, the methylation signature may be normalized. Normalization may be based on the sequencing depth of the specimen. Normalization based on sequencing depth normalizes the distribution of genetic material across all specimens.

[0150] The analytics system trains a mixture model with the methylation signatures of the training samples 430. The analytics system may train the mixture model to predict component proportions based on the input methylation signatures (also referred to as deconvolving the component proportions). The analytics system may train the model to fit the methylation signatures of the training samples by performing maximum likelihood estimation.

[0151] An analytics system may train the mixture model as a machine learning model. Training a machine learning model may include supervised training, unsupervised training, or semi-supervised training. Supervised training generally involves using known labels or values ​​that can be used to calculate an error or loss in the model's prediction. The analytics system that trains the model may adjust the model's parameters to minimize the error. In a mixture model, the analytics system may use training examples with known component proportions to provide a supervised loss in training the mixture model. Unsupervised training generally involves learning patterns in training data without known labels or values ​​to supervise the learning. Exemplary unsupervised techniques include clustering, anomaly detection, latent variable learning, and other types of machine learning techniques that do not rely on known labels or values. In the context of a mixture model, an analytics system may use training examples without known component proportions and have the mixture model learn patterns within the training data. Semi-supervised learning generally involves using training data where some of the training data have known labels or values. Types of machine learning models that may be implemented include decision trees, neural networks, multi-layer perceptrons, support vector machines, models relying on other types of machine learning derivatives of these, and combinations thereof. An embodiment of a mixed model architecture is described in Figures 6 and 7.

[0152] 6 illustrates an exemplary methylation signature matrix 610 used in training a mixture model, according to an exemplary embodiment. The methylation signature matrix 610 represents the methylation signatures of training samples used to train the mixture model. The methylation signature matrix 610 includes multiple samples 620 for multiple genomic regions 630. The methylation signature of each training sample 620 is represented in a row of the methylation signature matrix 610.

[0153] FIG. 7 illustrates an exemplary architecture 700 of a mixture model according to one or more embodiments. The mixture model 700 includes multiple component submodels and a deconvolution model 740. Each component submodel is configured to input a methylation signature of a sample and output a predicted likelihood that the methylation signature originated from that component. The predicted likelihood may be expressed as a percentage, e.g., the likelihood that the methylation signature originated from component 1 is 65%. For example, component 1 submodel 710 outputs component 1 likelihood 715, component 2 submodel 720 outputs component 2 likelihood 725, and so on, until component K submodel 730 outputs component K likelihood 735. These component likelihoods are input to a deconvolution, which outputs the component proportions of the methylation signature.

[0154] The analytics system may train the component sub-models and the deconvolution sub-model 740 simultaneously. Simultaneous training involves the analytics system simultaneously inputting the methylation signature matrix 610 through all sub-models while simultaneously adjusting the parameters of the various sub-models. In training with a supervised learning approach, the analytics system may predict the component proportions of a training sample by feeding one or more training samples through the mixture model 700. The analytics system may calculate the loss of a training sample by comparing the known component proportions of the sample with the predicted component proportions of the sample. In training with an unsupervised learning approach, the analytics system may cluster similar methylation signatures as predominantly originating from the same component by feeding the training samples through the mixture model 700. In such an unsupervised approach, the analytics system may be independent of the true component proportions of the training sample.

[0155] In other embodiments, the analytics system may train one or more of the sub-models individually. Individual training involves training a sub-model using a subset of the training data. After training, the analytics system may fix the parameters of that sub-model while training other sub-models. For example, the analytics system may train each of the component sub-models individually, and then train the deconvolution sub-model 740 while holding the parameters of the component sub-models. The analytics system may use training examples that predominantly originate from one component to train that component's corresponding component sub-model. The analytics system may then use training examples of the mixture proportions to train the deconvolution model 740.

[0156] In one or more embodiments, the mixture model 700 may further divide the component submodels into additional sub-component submodels. For example, component 1 submodel 710 may be connected to one or more sub-component submodels (not shown). The sub-component submodels may further determine the sub-component likelihood that the methylation signature belongs to that subcomponent. For example, component 1 may represent breast tissue shed from potential breast tumor cells. Component 1 may include human epidermal growth factor receptor 2 (HER2) positive and negative states as further divisions of breast tissue. Component 1 submodel 710 may provide a component 1 likelihood 715 of 65% breast tissue. The first sub-component submodel may output that, based on the methylation signature of the specimen, if the methylation signature originates from breast tissue (component 1), then the methylation signature has a 35% likelihood of being a HER2-positive state. The deconvolution submodel 740 inputs the component likelihoods and any sub-component likelihoods to determine component and sub-component proportions. Continuing with the above example, the deconvolution submodel 740 of the mixture model 700 may output the proportion of breast tissue as component 1, along with the proportion of HER2-positive status. In some embodiments with multiple subcomponents, the deconvolution submodel 740 may provide a subcomponent proportion for each subcomponent, with the sum of the subcomponent proportions resulting in a component proportion. The mixture model 700 may include a tree of components and subcomponents that spans multiple levels of depth. For example, a three-level depth may include a first component that includes a set of subcomponents and one of the subcomponents that includes a second set of subcomponents. The analytics system may train the subcomponent submodels individually, or may train the subcomponent submodels simultaneously with other submodels of the mixture model 700.

[0157] Returning to FIG. 4A , in training 430 the mixture model, the analytics system may tune 435 the number of components K as a hyperparameter. K represents the number of components from which the mixture model predicts component proportions (e.g., as shown in FIG. 7 ). The analytics system may iteratively train the mixture model while adjusting the number of components K, while cross-validating the trained mixture model. Cross-validation may include using a validation set of training samples with known component proportions. The analytics system may input methylation signatures of the training samples to predict component proportions. The analytics system may determine error or loss as the difference between the known component proportions and the predicted component proportions. The analytics system may score each trained mixture model with a different number of components K based on the error or loss from the validation set. In one or more embodiments, the analytics system may apply a penalty as a factor of the number of components to guide the mixture model to optimize the number of components. Too few components and models may result in a poor fit to the training data. Too many components and models may result in overfitting to the training data. The analytics system may also score how well the mixture model fits the training data.

[0158] In one or more embodiments, the analytics system can provide information to the mixture model about what each component corresponds to. In such embodiments, the analytics system can train the mixture model without knowing the labels of the training data. The analytics system can later use known labels or proportions of the training data to provide information about what each component corresponds to. The analytics system can use a training sample with known component proportions, for example, a training sample containing 65% nucleic acid fragments originating from breast tissue and 35% nucleic acid fragments corresponding to non-cancer cell-free nucleic acid fragments. The analytics system can input this training sample into the mixture model, which can output predicted component proportions of 70% component 1, 25% component 2, and 5% component 3 (in an exemplary mixture model trained across at least three components). The analytics system can match the largest known component proportion with the largest predicted component proportion to correspond to the same component. The analytics system can do the same for other known and predicted component proportions in the training sample. The analytics system can use multiple labeled training samples to support the component rankings in the mixture model. For example, the majority of training samples suggest that component 1 of the mixture model corresponds to lung tissue, and so on for the remaining components of the mixture model.

[0159] 4B is an exemplary flowchart describing a method 440 of deploying a mixture model to predict component proportions in a sample, according to one or more embodiments. This mixture model may be trained according to the method 400 for training a mixture model.

[0160] The analytics system obtains 450 a test sample for predicting component ratios. The test sample includes genetic material containing methylation information. The disease or cancer state of the subject providing the test sample may be known or unknown to the analytics system. The methylation information may include, for example, methylated sequence reads as sequenced by the process described in Figures 3A and 3B.

[0161] The analytics system generates a methylation signature for the test sample by counting methylation variations across multiple genomic regions 460. The multiple genomic regions correspond to the multiple genomic regions on which the mixture model is trained. The analytics system may normalize the methylation signature of the test sample, as well as the training sample, based on, for example, sequencing depth.

[0162] The analytics system may predict the component proportions of the test sample by applying a mixture model to the methylation signature 470. The analytics system inputs the methylation signature of the test sample into the mixture model and outputs the component proportions therefrom.

[0163] III.B. Mixed Model Application Depending on how the mixture model is trained, analytics systems can use the mixture model for a variety of applications.

[0164] In one or more embodiments, a mixture model may be used for tissue purity determination (e.g., step 142 of FIG. 1 ). An analytics system may use the mixture model to assess sample purity before using the sample, such as for training a cancer classifier. For example, an analytics system may obtain multiple training samples diagnosed with a particular cancer type. The analytics system may use the mixture model to predict the component proportions of the training sample. In one embodiment, the analytics system may exclude training samples that fall below a threshold proportion of tissue for that particular cancer type. For example, a training sample may originate from a subject diagnosed with lung cancer, but the mixture model predicts that the training sample will have 10% nucleic acid fragments originating from lung tissue. If the analytics system retains a training sample with more than 30% nucleic acid fragments originating from tissue corresponding to the diagnosed cancer type, then the analytics system will withhold use of the below-threshold training sample when training the cancer classifier. This advantageously focuses the training of the cancer classifier on training samples with substantial cancer signals, thereby improving the sensitivity of the cancer classifier. Using a mixture model for tissue purity can also identify misdiagnosed training samples. For example, a training sample originating from a subject diagnosed with prostate cancer may have a high contribution of liver tissue in a liquid biopsy sample, as determined by a mixture model. The analytics system can withhold such training samples in training a cancer classifier to prevent distortion of the cancer classifier.

[0165] In one or more embodiments, a mixture model may be used to predict a cancer type in a test subject. An analytics system may use the mixture model following or as part of a cancer classifier. The analytics system may use the cancer classifier to generate a cancer prediction. The cancer prediction may include a binary state (between cancer and non-cancer), a cancer type (between multiple cancer types), or some combination thereof. In one embodiment, in response to predicting the likelihood of cancer presence (a binary state), the analytics system may use the mixture model to determine component proportions as a way to determine cancer type. The largest component proportion that is not non-cancerous or other healthy tissue may be determined to be the cancer type of the test specimen. In other embodiments, the analytics system may use the component proportions of the mixture model to support the cancer type predicted by the cancer classifier. For example, a cancer classifier may predict a test specimen as likely to have head and neck cancer. The analytics system may use the mixture model to identify whether head and neck tissue is present in a substantial proportion in the test specimen. If not, then the analytics system may note the discrepancy, e.g., a decrease in confidence in the prediction. The analytics system may also utilize mixture models to determine whether multiple cancer types are likely to be present. For example, if the mixture model provides a substantial proportion of two distinct components (not non-cancerous or other healthy tissue), then the analytics system may determine that the test specimen likely has two cancer types.

[0166] In one or more embodiments, the mixture model may be used to quantify a cancer signal in a test subject. The analytics system may use the mixture model to predict component ratios for a test subject diagnosed with cancer or predicted to potentially have cancer. The analytics system may quantify the cancer signal based on the component ratios. In some embodiments, the analytics system may use the component likelihoods output by the component submodels to quantify the cancer signal. Quantifying the cancer signal may provide insight into what stage the cancer may be at or how aggressive the cancer may be. For example, the analytics system may analyze test specimens from an individual at various time points to determine changes in the cancer signal, e.g., based on the component ratios predicted by the mixture model. If the analytics system identifies an increase in the cancer signal, then the analytics system may determine that the cancer is progressing. The analytics system may further determine the aggressiveness of the cancer based on the rate of increase in the cancer signal. The analytics system may also use the quantification of the cancer signal to determine various treatment regimens. If there is a slowing of growth or a decrease in quantification, then the analytics system may determine that the ongoing treatment plan is successful. In other cases, the analytics system may determine that the treatment plan is unsuccessful and may even recommend adjustments to the treatment plan. Recommendations for adjustments to the treatment plan may include changing the therapy, changing the dosage or regimen, increasing the length or frequency of the treatment plan, etc.

[0167] In one or more embodiments, the analytics system may utilize component submodels separately from the mixture model. Generally, a component submodel represents the distribution of methylated variant allele proportions for genetic material from a component. The analytics system may then utilize the component submodel to predict the likelihood that a methylation signature originates from that component. The analytics system may also generate a distribution of methylation signatures for that component. Based on the component submodel, the analytics system can predict which component a methylated sequence read originates from. The analytics system may use such predictions to filter cancer training samples to remove non-cancer impurity sequence reads and retain methylated sequence reads that are more likely to be associated with cancer signals when training a cancer classifier. The non-cancer impurities may include one or more of lymphocytes, macrophages, fibroblasts, vascular endothelial cells, or non-cancerous tissues. In one or more embodiments, the analytics system may obtain methylation signatures of non-cancer impurities from a reference database. The reference database may include a distribution of methylation signatures observed from non-cancerous tissues or cells, for example, from healthy subjects who have not been diagnosed with cancer (or any other genetically based disease). In some embodiments, the reference database may be populated with methylation signatures observed during pre-characterization of genetic material. In some embodiments, the reference database may be populated with values ​​obtained from a third party, such as from a public methylation signature collection.

[0168] In one or more embodiments, the analytics system can extract information learned from the mixture model. In one embodiment, the analytics system can learn correlations between methylation variations in genomic regions that contribute to component proportions. For example, the analytics system can identify methylation variations that provide information about a particular component. In another embodiment, the analytics system can identify methylation variations in genomic regions that correlate with non-cancer impurities. The analytics system can filter out the use of such methylation variations in those genomic regions because those genomic regions may be confounded by non-cancer impurities.

[0169] IV. Cancer Classification Cancer classification involves extracting genetic features and applying one or more models to the extracted features to determine a cancer prediction. An analytics system may aggregate the extracted features into a feature vector, which may then be input into a trained cancer classifier to determine a cancer prediction based on the input feature vector. The cancer prediction may include a label and / or a value. The label may be binary, indicating the presence or absence of cancer in the test subject, and / or multi-class, indicating one or more specific cancer types from multiple screened cancer types. The value may be a grading of the cancer status, e.g., a cancer stage, or a quantification of the cancer signal. In particular, the cancer classifier may be a machine-learned model that includes multiple classification parameters and a function that represents the relationship between the feature vector as input and the cancer prediction as output. The feature vector, along with the classification parameters, is input into the function to produce a cancer prediction. In one or more embodiments, the cancer classification utilizes multiple models. Some models may be used to determine one or more features from the methylated sequence reads used for cancer classification. Prior to deployment of the cancer classifier, the analytics system trains the cancer classifier.

[0170] IV.A. Identification of the Abnormal Fragment For a given sample, the analytics system can determine abnormal fragments using the methylation state vector for that sample. For each fragment in the sample, the analytics system can determine whether the fragment is an abnormal fragment using the methylation state vector corresponding to that fragment. In some embodiments, the analytics system calculates, for each methylation state vector, a p-value score describing the probability of observing that methylation state vector or another methylation state vector, which may be even less likely in a healthy control group. The analytics system may determine as abnormal fragments those fragments whose methylation state vectors have a p-value score below a threshold. In some embodiments, the analytics system further labels fragments in which at least a portion of the CpG sites have a percentage of methylation or unmethylation above some threshold as hypermethylated and hypomethylated fragments, respectively. Hypermethylated or hypomethylated fragments may also be referred to as unusual fragments with extreme methylation (UFXM). In other embodiments, the analytics system may implement various other probability models for determining abnormal fragments. Examples of other probabilistic models include mixture models, deep probabilistic models, etc. In some embodiments, the analytics system may use any combination of the processes described below to identify abnormal fragments. According to the identified abnormal fragments, the analytics system may filter the set of methylation state vectors for samples used in other processes, for example, samples used in training and deployment of a cancer classifier.

[0171] IV. AIP Value Filtering In some embodiments, the analytics system calculates a p-value score for each methylation state vector compared to methylation state vectors from fragments in a healthy control group. The p-value score can describe the probability of observing a methylation state consistent with that methylation state vector or other methylation state vectors, which are even less likely in a healthy control group. To determine whether a DNA fragment is abnormally methylated, the analytics system can use a healthy control group in which the majority of fragments are normally methylated. When performing this probabilistic analysis to determine abnormal fragments, the determination can be weighted in comparison to a group of control subjects that constitute the healthy control group. To ensure robustness in the healthy control group, the analytics system can select some threshold number of healthy individuals to serve as the source of samples containing DNA fragments. Figure 8A below describes how the analytics system can generate a data structure for a healthy control group that can be used to calculate a p-value score. Figure 8B describes how the generated data structure can be used to calculate a p-value score.

[0172] 8A is a flowchart describing a process 800 for generating a data structure for a healthy control group, according to one embodiment. To create the healthy control group data structure, an analytics system can receive multiple DNA fragments (e.g., cfDNA) from multiple healthy individuals. The analytics system can generate 805 a methylation state vector for each fragment, for example, by process 300.

[0173] Using the methylation state vector for each fragment, the analytics system can subdivide the methylation state vector into strings of CpG sites 810. In some embodiments, the analytics system subdivides the methylation state vector 810 so that the resulting strings are all less than a given length. For example, if a methylation state vector of length 11 can be subdivided into strings of lengths 3 or less, there could be 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, if a methylation state vector of length 7 is subdivided into strings of lengths 4 or less, there could be 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If the methylation state vector is shorter than or equal to a specified string length, then the methylation state vector may be converted into a single string containing all of the CpG sites in the vector.

[0174] The analytics system tabulates the strings by, for each possible CpG site and methylation state possibility in the vector, counting the number of strings present in the control group that have the specified CpG site as the first CpG site in the string and have that possibility of methylation state 815. For example, at a given CpG site, given a string length of 3, 3 There are eight possible string configurations. For that given CpG site, the analytics system tallies 810 how many occurrences of each possible methylation state vector appear in the control population for each of the eight possible string configurations. Continuing with this example, this involves counting the following quantities for each starting CpG site x in the reference genome: <M x ,M x+1 ,M x+2 >, <M x ,M x+1 ,U x+2 >,..., x ,U x+1 ,U x+2 The analytics system creates a data structure 815 that stores the counts for each possible starting CpG site and string. ​

[0175] Setting an upper limit on string length has several benefits. First, the size of the data structures created by analytics systems can increase dramatically depending on the maximum length of the string. For example, a maximum string length of 4 means that every CpG site must be at least 2 for a string of length 4. 4 Increasing the maximum string length to 5 means that every CpG site counts an additional 2 4 This means that the number of counts (and required computer memory) will double compared to the previous string length, since 16 numbers will be counted as shown. Reducing the string size can help keep the creation and performance of the data structure reasonable in terms of computation and storage (e.g., for later access, as described below). Second, a statistical consideration limiting the maximum string length can be to avoid overfitting downstream models that use string counts. If long strings of CpG sites do not biologically have a strong effect on the outcome (e.g., predicting abnormalities that are predictive of the presence of cancer), calculating probabilities based on large strings of CpG sites can be problematic, as the amount of data used may not be available, and thus the model may be too sparse for adequate performance. For example, calculating the probability of abnormality / cancer conditional on the previous 100 CpG sites could ideally use string counts in a data structure of length 100, some of which exactly match the methylation states of the previous 100. If only a sparse count of strings of length 100 is available, there may be insufficient data to determine whether a given string of length 100 in a test sample is anomalous or not.

[0176] 8B is a flowchart describing a process 830 for identifying aberrant methylation fragments from an individual, according to one embodiment. In process 830, the analytics system generates 840 methylation state vectors from the subject's cfDNA fragments, for example, by process 300. The analytics system can handle each methylation state vector as follows:

[0177] For a given methylation state vector, the analytics system enumerates all possible methylation state vectors that have the same starting CpG site and the same length (i.e., set of CpG sites) in that methylation state vector 845. Because each methylation state is generally either methylated or unmethylated, there can effectively be two possible states for each CpG site, and therefore the count of different possibilities for a methylation state vector can depend on a power of 2, i.e., for a methylation state vector of length n, there can be 2 possible states for the methylation state vector. n Since the methylation state vector includes an uncertain state for one or more CpG sites, the analytics system can enumerate the possibilities of the methylation state vector by considering only CpG sites with an observed state 830.

[0178] The analytics system accesses the healthy control data structure to calculate 850 the probability of observing each possible methylation state vector for the identified starting CpG site and methylation state vector length. In some embodiments, the calculation of the probability of observing a given possibility uses Markov chain probability, which models joint probability calculations. The Markov model can be trained, at least in part, based on evaluating the methylation state of each CpG site in a plurality of corresponding CpG sites of each fragment (e.g., nucleic acid methylation fragment) spanning the nucleic acid methylation fragment in a healthy non-cancer cohort dataset having a corresponding plurality of CpG sites. For example, a Markov model (e.g., a hidden Markov model, or HMM) can be used to determine the probability of observing a sequence of methylation states (e.g., including "M" or "U") for a nucleic acid methylation fragment in a plurality of nucleic acid methylation fragments, given a set of probabilities that, for each state in the sequence, determine the likelihood of observing the next state in the sequence. The set of probabilities can be obtained by training an HMM. Such training may involve calculating statistical parameters (e.g., the probability that a first state can transition to a second state (transition probability) and / or the probability that a given methylation state can be observed for each CpG site (output probability)) given an initial training dataset of observed methylation state sequences (e.g., methylation patterns). HMMs can be trained using supervised training (e.g., using exemplars where the underlying sequences and observed states are known) and / or unsupervised training (e.g., Viterbi learning, maximum likelihood estimation, expectation-maximization training, and / or Baum-Welch training). In other embodiments, computational methods other than Markov chain probabilities are used to determine the probability of observing each possible methylation state vector. For example, such computational methods may include trained representations. The p-value threshold may be between 0.01 and 0.10, or between 0.03 and 0.06. The p-value threshold may be 0.05. The p-value threshold may be less than 0.01, less than 0.001, or less than 0.0001.

[0179] The analytics system calculates a p-value score for the methylation state vector 855 using the calculated probability for each possibility. In some embodiments, this involves identifying the calculated probability that corresponds to the possibility of matching the methylation state vector in question. Specifically, this may be the possibility of having the same set of CpG sites as the methylation state vector, or similarly the same starting CpG site and length. The analytics system may sum the calculated probabilities of any possibilities that have a probability less than or equal to the identified probability to generate a p-value score.

[0180] This p-value may correspond to the probability of observing the methylation state vector of the fragment or other methylation state vectors that are even less likely in healthy controls. Thus, a low p-value score may generally correspond to a methylation state vector that is rare in healthy individuals and that results in fragments labeled as aberrantly methylated compared to healthy controls. A high p-value score may generally relate to a methylation state vector that is expected to be present in healthy individuals in a relative sense. If the healthy control group is, for example, a non-cancer group, a low p-value may indicate that the fragment is aberrantly methylated compared to the non-cancer group, and therefore may potentially be indicative of the presence of cancer in the test subject.

[0181] As described above, the analytics system can calculate a p-value score for each of a plurality of methylation state vectors, each corresponding to a cfDNA fragment in the test specimen. To identify which fragments are aberrantly methylated, the analytics system can filter the set of methylation state vectors based on their p-value scores 865. In some embodiments, filtering is performed by comparing the p-value scores to a threshold and retaining only fragments that are below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001, 0.0001, or a similar order of magnitude.

[0182] According to exemplary results from process 800, the analytics system may produce a median (range) of 2,800 (1,500-12,000) fragments with aberrant methylation patterns for participants without cancer at training, and a median (range) of 3,000 (1,200-420,000) fragments with aberrant methylation patterns for participants with cancer at training. These filtered sets of fragments with aberrant methylation patterns may be used for downstream analysis, as described in Section IV.B below.

[0183] In some embodiments, the analytics system uses a sliding window 860 to determine the probabilities of the methylation state vector and calculate p-values. Rather than enumerating the probabilities and calculating p-values ​​for the entire methylation state vector, the analytics system can enumerate the probabilities and calculate p-values ​​only for a window of contiguous CpG sites, where the window is shorter in length than at least some fragment (of CpG sites) (otherwise the window may serve no purpose). The window length can be static, user-determined, dynamic, or some other selected length.

[0184] In calculating p-values ​​for methylation state vectors larger than a window, the window can identify a contiguous set of CpG sites from the vector that fall within the window, starting with the first CpG site in the vector. The analytics system can calculate a p-value score for the window that includes the first CpG site. The analytics system can then "slide" the window to a second CpG site in the vector and calculate another p-value score for the second window. In this way, for a window size l and a methylation vector length m, each methylation state vector can generate m-l+1 p-value scores. After completing the p-value calculations for each portion of the vector, the lowest p-value score from all sliding windows can be taken as the overall p-value score for that methylation state vector. In other embodiments, the analytics system aggregates the p-value scores for a methylation state vector to generate an overall p-value score.

[0185] The use of a sliding window can help reduce the number of possible methylation state vectors to be enumerated and the corresponding probability calculations that might otherwise need to be performed. To take a realistic example, a fragment may have 54 or more CpG sites.2 54 Street (approx. 1.8 x 10 16 Instead of calculating probabilities for 2 possibilities (number of possibilities) to generate a single p-score, the analytics system can instead use a window of size 5 (for example), resulting in 50 p-value calculations for each of the 50 windows of the methylation state vector for this fragment. 5 We can enumerate the possible paths (32), which gives us a total of 50 x 2 5 Street (1.6 x 10 3 This significantly reduces the computation that must be performed without significantly impacting the accurate identification of anomalous fragments.

[0186] In embodiments involving uncertain states, the analytics system may calculate a p-value score by summing and eliminating the uncertain CpG sites in the methylation state vector of the fragment. The analytics system may identify all possibilities that are in consensus with all methylation states in the methylation state vector excluding the uncertain states. The analytics system may assign a probability to the methylation state vector as the sum of the probabilities of the identified possibilities. As an example, since the methylation states of CpG sites 1 and 3 are observed and are in consensus with the methylation states at CpG sites 1 and 3 of the fragment, the analytics system may<M1,I2,U3> The probability of the methylation state vector of<M1,M2,U3> and<M1,U2,U3> This method of summing and eliminating uncertain CpG sites can be performed with a maximum of 2 i Calculation of the probability of one or more possible uncertain states in the methylation state vector may be used. In a further embodiment, a dynamic programming algorithm may be implemented to calculate the probability of one or more uncertain states in the methylation state vector. Advantageously, the dynamic programming algorithm operates in linear computation time.

[0187] In some embodiments, the computational burden of calculating probabilities and / or p-value scores may be further reduced by caching at least some of the calculations. For example, the analytics system may cache probability calculations for possibilities of a methylation state vector (or a window thereof) in temporary or persistent memory. If other fragments have the same CpG sites, caching the possibility probabilities may allow p-score values ​​to be calculated efficiently without recalculating the underlying possibility probabilities. In other words, the analytics system may calculate a p-value score for each of the possibilities of a methylation state vector associated with a set of CpG sites from the vector (or a window thereof). The analytics system may cache the p-value scores for use in determining p-value scores for other fragments containing the same CpG sites. In general, the p-value scores of possibilities of methylation state vectors with the same CpG sites may be used to determine p-value scores for different possibilities from the same set of CpG sites.

[0188] Prior to training the region model or cancer classifier, one or more nucleic acid methylation fragments can be filtered. Filtering the nucleic acid methylation fragments can include removing each respective nucleic acid methylation fragment that does not satisfy one or more selection criteria (e.g., falls below or exceeds one selection criterion) from the corresponding plurality of nucleic acid methylation fragments. The one or more selection criteria can include a p-value threshold. The output p-value of each nucleic acid methylation fragment can be determined, at least in part, based on a comparison of the corresponding methylation pattern of each respective nucleic acid methylation fragment with a corresponding distribution of methylation patterns of the respective nucleic acid methylation fragments in a healthy non-cancer cohort dataset having the corresponding plurality of CpG sites of each respective nucleic acid methylation fragment.

[0189] Filtering the plurality of nucleic acid methylation fragments can include removing each respective nucleic acid methylation fragment that does not meet a p-value threshold. A filter can be applied to the methylation pattern of each respective nucleic acid methylation fragment using the methylation patterns observed in the first plurality of nucleic acid methylation fragments. Each respective methylation pattern of each respective nucleic acid methylation fragment (e.g., fragment 1, ..., fragment N) can include one or more corresponding methylation sites (e.g., CpG sites) identified with a methylation site identifier, and a corresponding methylation pattern represented as a sequence of 1s and 0s (where each "1" represents a methylated CpG site in the one or more CpG sites and each "0" represents an unmethylated CpG site in the one or more CpG sites). The methylation patterns observed in the first plurality of nucleic acid methylation fragments can be used to construct a methylation state distribution for the CpG site states collectively represented by the first plurality of nucleic acid methylation fragments (e.g., CpG site A, CpG site B, ..., CpG site ZZZ). Further details regarding processing of nucleic acid methylated fragments are disclosed in U.S. Patent Application No. 17 / 191,914, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," filed March 4, 2021, which is hereby incorporated by reference in its entirety.

[0190] Each nucleic acid methylated fragment may not satisfy one of the one or more selection criteria when the methylation anomaly score of the respective nucleic acid methylated fragment is less than a methylation anomaly score threshold. In this situation, the methylation anomaly score may be determined by a mixture model. For example, the mixture model can detect methylation pattern anomalies in nucleic acid methylated fragments by determining the likelihood of a methylation state vector (e.g., a methylation pattern) of each nucleic acid methylated fragment based on the number of possible methylation state vectors of the same length at the same corresponding genomic location. This may be performed by generating multiple possible methylation states for a vector of a specified length at each genomic location in the reference genome. The multiple possible methylation states can be used to determine the total number of possible methylation states and subsequently the probability of each predicted methylation state at that genomic location. The likelihood of a sample nucleic acid methylated fragment corresponding to a genomic location in the reference genome can be determined by matching the sample nucleic acid methylated fragment with the predicted (e.g., possible) methylation states to obtain the calculated probability of the predicted methylation state. A methylation anomaly score can then be calculated based on the probabilities of the sample nucleic acid methylated fragments.

[0191] Each nucleic acid methylated fragment may not satisfy a certain selection criterion among the one or more selection criteria if the number of residues in the respective nucleic acid methylated fragment is less than a threshold. The threshold number of residues may be 10 to 50, 50 to 100, 100 to 150, or more than 150. The threshold number of residues may be a fixed value of 20 to 90. Each nucleic acid methylated fragment may not satisfy a certain selection criterion among the one or more selection criteria if the number of CpG sites in the respective nucleic acid methylated fragment is less than a threshold. The threshold number of CpG sites may be 4, 5, 6, 7, 8, 9, or 10. Each nucleic acid methylated fragment may not satisfy a certain selection criterion among the one or more selection criteria if the genomic start and end positions of the respective nucleic acid methylated fragment indicate that the respective nucleic acid methylated fragment corresponds to less than a threshold number of nucleotides in the human genome reference sequence.

[0192] Filtering can remove nucleic acid methylated fragments from the corresponding plurality of nucleic acid methylated fragments that have the same corresponding methylation pattern and the same corresponding genomic start and end positions as another nucleic acid methylated fragment from the corresponding plurality of nucleic acid methylated fragments. This filtering step can remove overlapping fragments that are exact duplicates, including PCR duplicates in some examples. Filtering can remove nucleic acid methylated fragments that have the same corresponding genomic start and end positions as another nucleic acid methylated fragment from the corresponding plurality of nucleic acid methylated fragments and less than a threshold number of different methylation states. The threshold number of different methylation states used to retain nucleic acid methylated fragments can be 1, 2, 3, 4, 5, or more than 5. For example, a first nucleic acid methylated fragment that has the same corresponding genomic start and end positions as a second nucleic acid methylated fragment but has at least 1, at least 2, at least 3, at least 4, or at least 5 different methylation states at each CpG site (e.g., aligned to a reference genome) is retained. As another example, a first nucleic acid methylated fragment having the same methylation state vector (e.g., methylation pattern) but a corresponding genomic start and end position that differs from the second nucleic acid methylated fragment is also retained.

[0193] Filtering can remove assay artifacts in multiple nucleic acid methylation fragments. Removing assay artifacts can include removing sequence reads obtained from hybridization probes after sequencing and / or sequence reads obtained from sequences that did not undergo bisulfite conversion. Filtering can remove contaminants (e.g., resulting from sequencing, nucleic acid isolation, and / or sample preparation).

[0194] The filtering can remove a subset of methylation fragments from the plurality of methylation fragments based on mutual information filtering of each methylation fragment for the cancer conditions of the plurality of training subjects. For example, mutual information can provide a measure of interdependence between two simultaneously sampled conditions of interest. Mutual information can be determined by selecting an independent set of CpG sites (e.g., within all or a portion of the nucleic acid methylation fragments) from one or more datasets and comparing the methylation state probabilities for that set of CpG sites between two sample groups (e.g., genotype datasets, biological specimens, and / or subject subsets and / or groups). The mutual information score can indicate the probability of a methylation pattern for a first condition versus a second condition in each region within each sliding window, and thus indicates the discriminatory power of each region. Mutual information scores can similarly be calculated for each region within each sliding window as the window progresses across the selected set of CpG sites and / or selected genomic regions. Further details regarding mutual information filtering are disclosed in U.S. Patent Application No. 17 / 119,606, entitled "Cancer Classification using Patch Convolutional Neural Networks," filed December 11, 2020, which is hereby incorporated by reference in its entirety.

[0195] IV.A.II. Hypermethylated and Hypomethylated Fragments In some embodiments, the analytics system identifies 870 and determines hypomethylated or hypermethylated fragments as aberrant fragments from the filtered set. The analytics system identifies hypermethylated fragments, in which more than a threshold number and more than a threshold percentage of CpG sites are methylated. The analytics system identifies hypomethylated fragments, in which more than a threshold number and more than a threshold percentage of CpG sites are unmethylated. Exemplary thresholds for fragment (or CpG site) length include lengths of more than 3, 4, 5, 6, 7, 8, 9, 10, etc. Exemplary percentage thresholds for methylation or unmethylation include percentages of more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%.

[0196] IV.B. Training a Cancer Classifier 9A is a flowchart describing a process 900 for training a cancer classifier, according to one embodiment. An analytics system obtains 910 a plurality of training samples, each having a set of abnormal fragments and a cancer type label. The plurality of training samples may include any combination of samples from healthy individuals with a generic label of "non-cancer," samples from subjects with a generic or specific label of "cancer" (e.g., "breast cancer," "lung cancer," etc.). Training samples from subjects for one cancer type may be referred to as a cohort for that cancer type or a cancer type cohort.

[0197] The analytics system determines 920, for each training sample, a feature vector based on the set of abnormal fragments of the training sample. The analytics system can calculate an anomaly score for each CpG site in the initial set of CpG sites. The initial set of CpG sites can be all CpG sites in the human genome or a subset thereof—this can be 10 4 , 10 5 , 10 6 , 10 7 , 10 8In one embodiment, the analytics system defines an anomaly score for the feature vector using binary scoring based on whether an aberrant fragment is present among a set of aberrant fragments encompassing the CpG site. In another embodiment, the analytics system defines an anomaly score based on a count of aberrant fragments overlapping the CpG site. In one example, the analytics system may use ternary scoring, assigning a first score to the absence of aberrant fragments, a second score to the presence of a few aberrant fragments, and a third score to the presence of many aberrant fragments. For example, the analytics system may count five aberrant fragments overlapping the CpG site in a given sample and calculate the anomaly score based on the count of five. In one or more embodiments, the feature vector further includes one or more features based on methylation variations, e.g., as identified and counted for use in the mixture model in FIGS. 4A and 4B.

[0198] After determining all the anomaly scores for the training samples, the analytics system can determine a feature vector as a vector of elements, each element containing one of the anomaly scores associated with one of the CpG sites in the initial set. The analytics system can normalize the anomaly scores of the feature vector based on the coverage of the sample, where coverage can refer to all CpG sites covered by the initial set of CpG sites used for the classifier, or the median or average sequencing depth based on the set of anomaly fragments for a given training sample.

[0199] By way of example, reference is now made to FIG. 9B, which illustrates a matrix of training feature vectors 922. In this example, the analytics system has identified CpG sites [S] 926 for consideration in generating feature vectors for a cancer classifier. The analytics system selects training samples [N] 924. The analytics system determines a first anomaly score 928 for a first arbitrary CpG site [s1] to be used in the feature vector of the training sample [n1]. The analytics system examines each abnormal fragment in the set of abnormal fragments. If the analytics system identifies at least one abnormal fragment that includes the first CpG site, then the analytics system determines a first anomaly score 928 for the first CpG site as 1, as illustrated in FIG. 9B. Considering a second arbitrary CpG site [s2], the analytics system similarly examines the set of abnormal fragments for at least one that includes the second CpG site [s2]. If the analytics system does not find any such abnormal fragments containing the second CpG site, the analytics system determines a second anomaly score 929 for the second CpG site [s2] as 0, as shown in Figure 9B. Once the analytics system has determined all of the anomaly scores for the initial set of CpG sites, the analytics system determines a feature vector for the first training sample [n1] containing the anomaly scores, the feature vector including a first anomaly score 928 of 1 for the first CpG site [s1] and a second anomaly score 929 of 0 for the second CpG site [s2], and subsequent anomaly scores, thus forming a feature vector [1, 0, ...].

[0200] For additional techniques for featurizing samples, see U.S. patent application Ser. No. 15 / 931,022, entitled "Model-Based Featurization and Classification"; U.S. patent application Ser. No. 16 / 579,805, entitled "Mixture Model for Targeted Sequencing"; U.S. patent application Ser. No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification"; and U.S. patent application Ser. No. 16 / 723,716, entitled "Source of Origin Deconvolution Based on Methylation Fragments in Cell-Free DNA Samples," all of which are incorporated by reference in their entireties.

[0201] The analytics system may further limit the CpG sites considered for use in the cancer classifier. For each CpG site in the initial set of CpG sites, the analytics system calculates 930 an information gain based on the feature vectors of the training samples. From step 920, each training sample has a feature vector that may have an anomaly score for all CpG sites in the initial set of CpG sites, which may include up to all CpG sites in the human genome. However, some CpG sites in the initial set of CpG sites may not be as informative in distinguishing between cancer types as others, or may overlap with other CpG sites.

[0202] In one embodiment, the analytics system calculates 930 an information gain for each cancer type and for each CpG site in the initial set to determine whether to include that CpG site in the classifier. The information gain is calculated for a training sample of a given cancer type compared to all other samples. For example, two random variables are used: "abnormal fragment" ("AF") and "cancer type" ("CT"). In one embodiment, AF is a binary variable indicating whether a given sample has an abnormal fragment that overlaps a given CpG site, as determined for the anomaly score / feature vector above. CT is a random variable indicating whether the cancer is of a particular type. The analytics system calculates the mutual information for CT given AF, i.e., how many bits of information are gained about the cancer type when it is known whether there is an abnormal fragment overlapping a particular CpG site. In effect, for a first cancer type, the analytics system calculates pairwise mutual information gains for each of the other cancer types and sums these mutual information gains across all other cancer types.

[0203] For a given cancer type, the analytics system can use this information content to rank CpG sites based on how cancer-specific they are. This procedure can be repeated for all cancer types under consideration. If a particular region is commonly aberrantly methylated in training samples of a given cancer, but is unmethylated in training samples of other cancer types or healthy training samples, then the CpG sites where these aberrant segments overlap may have high information gain for the given cancer type. The ranked CpG sites for each cancer type can be greedily added (selected) 940 to a set of selected CpG sites based on their ranking for use in the cancer classifier.

[0204] In further embodiments, the analytics system may consider other selection criteria for selecting informative CpG sites to use in the cancer classifier. One selection criterion may be that the selected CpG site is more than a threshold distance from other selected CpG sites. For example, the selected CpG site may be more than a threshold number of base pairs (e.g., 100 base pairs) away from any other selected CpG site, such that both CpG sites within the threshold distance are not selected for consideration in the cancer classifier.

[0205] In one embodiment, according to the set of selected CpG sites from the initial set, the analytics system may modify the feature vectors of the training samples as needed 950. For example, the analytics system may truncate the feature vectors such that anomaly scores corresponding to CpG sites that are not in the set of selected CpG sites are removed.

[0206] Using the feature vectors of the training samples, the analytics system may train a cancer classifier in any of a number of ways. The feature vector may correspond to the initial set of CpG sites from step 920, or may correspond to the set of selected CpG sites from step 950. In one embodiment, the analytics system trains 960 a binary cancer classifier to distinguish between cancer and non-cancer based on the feature vectors of the training samples. In this manner, the analytics system uses training samples that include both non-cancer samples from healthy individuals and cancer samples from the subject. Each training sample can have one of two labels: "cancer" or "non-cancer." In this embodiment, the classifier outputs a cancer prediction indicating the likelihood of the presence or absence of cancer.

[0207] In another embodiment, the analytics system trains a multi-class cancer classifier 970 to distinguish between many cancer types (also referred to as tissue of origin (TOO) labels). The cancer types may include one or more cancers and may also include non-cancer types (and any additional other diseases or genetic disorders, etc.). To do so, the analytics system may use a cancer type cohort, which may or may not include a non-cancer type cohort. In this multi-cancer embodiment, the cancer classifier is trained to determine cancer predictions (or, more specifically, TOO predictions) that include a predictive value for each of the cancer types under classification. The predictive value may correspond to the likelihood that a given training sample (and, in inference, a test sample) has each of the cancer types. In one embodiment, the predictive value is scored between 0 and 100, where the sum of the predictive values ​​equals 100. For example, the cancer classifier returns a cancer prediction that includes predictive values ​​for breast cancer, lung cancer, and non-cancer. For example, the classifier may return a cancer prediction that the test sample has a 65% likelihood of breast cancer, a 25% likelihood of lung cancer, and a 10% likelihood of non-cancer. The analytics system may further evaluate the predictions to generate a prediction of the presence of one or more cancers in the sample, indicating one or more TOO labels, which may also be referred to as TOO predictions, e.g., a first TOO label with the highest predictive value, a second TOO label with the second highest predictive value, etc. Continuing with the example above, given these proportions, the system may determine that the sample has breast cancer, given that the likelihood of breast cancer is highest. The analytics system may further support the TOO prediction by predicting the component proportions of the sample using the mixture models of Figures 4A and 4B. See Figure 4B above for further explanation.

[0208] In both embodiments, the analytics system trains the cancer classifier by inputting a set of training samples along with their feature vectors into the cancer classifier and adjusting classification parameters so that the classifier's function accurately relates the training feature vectors to their corresponding labels. The analytics system may group the training samples into one or more sets of training samples for repeated batch training of the cancer classifier. After inputting the entire set of training samples, including their training feature vectors, and adjusting the classification parameters, the cancer classifier may be sufficiently trained to label test samples according to their feature vectors within some margin of error. The analytics system may train the cancer classifier according to any one of several methods. As an example, a binary cancer classifier may be an L2 regularized logistic regression classifier trained using a logarithmic loss function. As another example, a multi-cancer classifier may be multinomial logistic regression. In fact, any type of cancer classifier may be trained using other techniques. There are many such techniques, including the possibility of using machine learning algorithms such as kernel methods, random forest classifiers, mixture models, autoencoding models, and multi-layer neural networks.

[0209] The classifier may include a logistic regression algorithm, a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a decision tree algorithm, a multinomial logistic regression algorithm, a linear model, or a linear regression algorithm.

[0210] IV.C. Cancer Classifier Deployment When using the cancer classifier, the analytics system may obtain a test sample from a subject with an unknown cancer type. The analytics system may process the test sample, which is composed of DNA molecules, through any combination of processes 300 and 830 to obtain a set of aberrant fragments. The analytics system may determine the test feature vector used by the cancer classifier according to principles similar to those discussed in process 900. The analytics system may calculate an aberration score for each CpG site among the plurality of CpG sites being used by the cancer classifier. For example, the cancer classifier may receive as input a feature vector containing an aberration score for 1,000 selected CpG sites. The analytics system may then determine a test feature vector containing an aberration score for the 1,000 selected CpG sites based on the set of aberrant fragments. The analytics system may calculate the aberration score in the same manner as for the training sample. In some embodiments, the analytics system defines the aberration score as a binary score based on whether the set of aberrant fragments encompassing the CpG sites includes a hypermethylated fragment or a hypomethylated fragment. The analytics system can generate a test feature vector that includes one or more features based on the methylation variations (as described in Figures 4A and 4B).

[0211] The analytics system can then input the test feature vector into a cancer classifier. The cancer classifier function can then generate a cancer prediction based on the classification parameters trained in process 900 and the test feature vector. In a first method, the cancer prediction may be binary and selected from the group consisting of "cancer" or "non-cancer"; in a second method, the cancer prediction is selected from a group of many cancer types and "non-cancer." In a further embodiment, the cancer prediction has a predictive value for each of many cancer types. Additionally, the analytics system may determine that the test specimen is most likely to be one of the cancer types. Following the example above, where the cancer prediction for the test specimen is 65% likely to be breast cancer, 25% likely to be lung cancer, and 10% likely to be non-cancer, the analytics system may determine that the test specimen is most likely to have breast cancer. In another example, if the cancer prediction is binary, 60% likely to be non-cancer and 40% likely to be cancer, the analytics system may determine that the test specimen is most likely not to have cancer. In further embodiments, the most likely cancer prediction may still be compared to a threshold (e.g., 40%, 50%, 60%, 70%) to call the test subject as having that cancer type. If the most likely cancer prediction does not exceed the threshold, the analytics system may return an indeterminate result.

[0212] In a further embodiment, the analytics system chains the cancer classifier trained in step 960 of process 900 with another cancer classifier trained in step 970 or process 900. The analytics system can input a test feature vector to a cancer classifier trained as a binary classifier in step 960 of process 900. The analytics system can receive an output of a cancer prediction. The cancer prediction can be a binary value related to whether the test subject is likely to have cancer or likely not to have cancer. In other embodiments, the cancer prediction includes a predictive value describing the likelihood of cancer and the likelihood of non-cancer. For example, the cancer prediction has a cancer predictive value of 85% and a non-cancer predictive value of 15%. The analytics system can determine that the test subject is likely to have cancer. Once the analytics system determines that the test subject is likely to have cancer, the analytics system can input the test feature vector to a multi-class cancer classifier trained to distinguish between different cancer types. The multi-class cancer classifier can receive the test feature vector and return a cancer prediction for a cancer type among multiple cancer types. For example, the multi-class cancer classifier provides a cancer prediction that designates the test subject as most likely to have ovarian cancer. In another embodiment, the multi-class cancer classifier provides a predictive value for each cancer type among a plurality of cancer types. For example, the cancer predictions may include a breast cancer type predictive value of 40%, a colorectal cancer type predictive value of 15%, and a liver cancer predictive value of 45%.

[0213] According to a generalized embodiment of binary cancer classification, an analytics system can determine a cancer score for a test specimen based on sequencing data (e.g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.) for the test specimen. The analytics system can compare the cancer score for the test specimen to a binary threshold cutoff for predicting whether the test specimen is likely to have cancer. The binary threshold cutoff can be tuned using TOO thresholding based on one or more TOO subtype classes. The analytics system can further generate a feature vector for the test specimen for use in a multi-class cancer classifier to determine a cancer prediction indicating one or more likely cancer types.

[0214] The classifier can be used to determine the disease state of a test subject, for example, a subject whose disease state is unknown. The method can include obtaining a test genomic data construct (e.g., test data at a single time point) in electronic format, the test genomic data construct including a value for each genomic feature in a plurality of genomic features of a corresponding plurality of nucleic acid fragments in a biological specimen obtained from the test subject. The method can then include applying the test genomic data construct to a test classifier, thereby determining the state of the disease condition in the test subject. The test subject need not have previously been diagnosed with the disease condition.

[0215] The classifier may be a temporary classifier that uses at least (i) a first test genomic data construct generated from a first biological specimen collected from the test subject at a first time point, and (ii) a second test genomic data construct generated from a second biological specimen collected from the test subject at a second time point.

[0216] The trained classifier can be used to determine the disease status of a test subject, e.g., a subject whose disease status is unknown. In this case, the method may include obtaining a test time series dataset in electronic format for a test subject, where the test time series dataset includes, for each time point among a plurality of time points, a corresponding test genotype data construct including values ​​for a plurality of genotype characteristics of a plurality of corresponding nucleic acid fragments in a corresponding biological specimen obtained from the test subject at each time point, and an indication of the length of time between each pair of consecutive time points among the plurality of time points. The method may then include applying the test genotype data construct to the test classifier, thereby determining the status of the disease condition in the test subject. The test subject need not have been previously diagnosed with the disease condition.

[0217] V. Application In some embodiments, the methods, analytics systems, and / or classifiers of the present invention can be used to detect the presence of cancer, monitor cancer progression or recurrence, monitor treatment response or effectiveness, determine or monitor the presence of minimal residual disease (MRD), or any combination thereof. For example, as described herein, classifiers can be used to generate a probability score (e.g., 0-100) describing the likelihood that a test feature vector is from a subject with cancer. In some embodiments, the probability score is compared to a threshold probability to determine whether a subject has cancer. In other embodiments, disease progression or treatment effectiveness (e.g., therapeutic efficacy) can be monitored by determining the likelihood or probability score at multiple different time points (e.g., before or after treatment). In still other embodiments, the likelihood or probability score can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, determining treatment effectiveness, etc.). For example, in one embodiment, if the probability score exceeds a threshold, a physician can prescribe an appropriate treatment.

[0218] Early detection of VA cancer In some embodiments, the methods and / or classifiers of the present invention can be used to detect the presence or absence of cancer in a subject suspected of having cancer. For example, a classifier (e.g., as described in Section IV above) can be used to determine a cancer prediction that describes the likelihood that a test feature vector is from a subject with cancer.

[0219] In one embodiment, the cancer prediction is a likelihood (e.g., scored from 0 to 100) of whether a test specimen has cancer (i.e., a binary classification). Accordingly, the analytics system may determine a threshold for determining whether a test subject has cancer. For example, a cancer prediction of 60 or greater may indicate that the subject has cancer. In yet other embodiments, a cancer prediction of 65 or greater, 70 or greater, 75 or greater, 80 or greater, 85 or greater, 90 or greater, or 95 or greater indicates that the subject has cancer. In other embodiments, the cancer prediction may indicate disease severity. For example, a cancer prediction of 80 may indicate a more severe, severe form of cancer or a later stage of cancer compared to a cancer prediction below 80 (e.g., a probability score of 70). Similarly, an increase in cancer prediction over time (e.g., determined by classifying test feature vectors from multiple specimens from the same subject taken at two or more time points) may indicate disease progression, or a decrease in cancer prediction over time may indicate successful treatment.

[0220] In another embodiment, the cancer prediction includes multiple predictive values, where each of multiple cancer types under classification (i.e., multi-class classification) has a predictive value (e.g., scored between 0 and 100). The predictive values ​​may correspond to the likelihood that a given training sample (and, upon inference, the training sample) has each of the cancer types. The analytics system may identify the cancer type with the highest predictive value and indicate that the test subject is likely to have that cancer type. In other embodiments, the analytics system further determines that the test subject is likely to have that cancer type by comparing the highest predictive value to a threshold value (e.g., 50, 55, 60, 65, 70, 75, 80, 85, etc.). In other embodiments, the predictive value may also indicate disease severity. For example, a predictive value higher than 80 may indicate a more severe, severe, or later stage cancer compared to a predictive value of 60. Similarly, an increase in predictive value over time (e.g., as determined by classifying test feature vectors from multiple samples from the same subject taken at two or more time points) may indicate disease progression, or a decrease in predictive value over time may indicate successful treatment.

[0221] According to aspects of the invention, the methods and systems of the invention can be trained to detect or classify multiple adaptive cancers. For example, the methods, systems and classifiers of the invention can be used to detect the presence of one or more, two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different types of cancer.

[0222] Examples of cancers that can be detected using the methods, systems, and classifiers of the present invention include carcinoma, lymphoma, blastoma, sarcoma, and leukemia or lymphoid malignancies. More specific examples of such cancers include, but are not limited to, squamous cell carcinoma (e.g., squamous cell carcinoma), skin cancer, melanoma, lung cancer, including small cell lung cancer, non-small cell lung cancer ("NSCLC"), lung adenocarcinoma, and lung squamous cell carcinoma, peritoneal cancer, gastric cancer, including gastrointestinal cancer or stomach cancer. cancer), pancreatic cancer (e.g., pancreatic ductal adenocarcinoma), cervical cancer, ovarian cancer (e.g., high-grade serous ovarian cancer), liver cancer (e.g., hepatocellular carcinoma (HCC)), hepatoma, liver cancer, bladder cancer (e.g., urothelial bladder cancer), testicular (germ cell tumor) cancer, breast cancer (e.g., HER2-positive, HER2-negative, and triple-negative breast cancer), brain cancer (e.g., astrocytoma, glioma (e.g., glioblastoma)), colon cancer, rectal cancer, colorectal cancer, endometrial or uterine cancer, salivary gland cancer, kidney or renal cancer (e.g., renal cell carcinoma, nephroblastoma, or Wilms' tumor), prostate cancer, vulvar cancer, thyroid cancer, anal cancer, penile cancer, head and neck cancer, esophageal cancer, and nasopharyngeal carcinoma (NPC). Further examples of cancer include, without limitation, retinoblastoma, thecoma, androgenic tumor, non-Hodgkin's lymphoma (NHL), multiple myeloma and hematological malignancies, including but not limited to acute hematological malignancies, endometriosis, fibrosarcoma, choriocarcinoma, laryngeal carcinoma, Kaposi's sarcoma, Schwannoma, oligodendroglioma, neuroblastoma, rhabdomyosarcoma, osteosarcoma, leiomyosarcoma, and urinary tract cancer.

[0223] In some embodiments, the cancer is one or more of anorectal cancer, bladder cancer, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastric cancer, head and neck cancer, hepatobiliary cancer, leukemia, lung cancer, lymphoma, melanoma, multiple myeloma, ovarian cancer, pancreatic cancer, prostate cancer, renal cancer, thyroid cancer, uterine cancer, or any combination thereof.

[0224] In some embodiments, the one or more cancers may be "high signal" cancers (defined as cancers with a 5-year cancer-specific mortality rate greater than 50%), such as anorectal cancer, colorectal cancer, esophageal cancer, head and neck cancer, hepatobiliary cancer, lung cancer, ovarian cancer, and pancreatic cancer, as well as lymphoma and multiple myeloma. High signal cancers tend to be more aggressive and typically exhibit above-average cell-free nucleic acid concentrations in test specimens obtained from patients.

[0225] VB Cancer and Treatment Monitoring In some embodiments, the cancer prognosis can be determined at multiple different time points (e.g., before or after treatment) to monitor disease progression or to monitor treatment effectiveness (e.g., therapeutic efficacy). For example, the invention includes methods involving obtaining a first specimen (e.g., a first plasma cfDNA specimen) from a cancer patient at a first time point and determining a first cancer prognosis therefrom (as described herein), and obtaining a second test specimen (e.g., a second plasma cfDNA specimen) from the cancer patient at a second time point and determining a second cancer prognosis therefrom (as described herein).

[0226] In certain embodiments, the first time point is before cancer treatment (e.g., before resection surgery or therapeutic intervention), and the second time point is after cancer treatment (e.g., after resection surgery or therapeutic intervention), and the classifier is used to monitor the effectiveness of the treatment. For example, if the second cancer prediction decreases compared to the first cancer prediction, then the treatment is considered successful. However, if the second cancer prediction increases compared to the first cancer prediction, then the treatment is considered unsuccessful. In other embodiments, both the first and second time points are before cancer treatment (e.g., before resection surgery or therapeutic intervention). In yet other embodiments, both the first and second time points are after cancer treatment (e.g., after resection surgery or therapeutic intervention). In still other embodiments, cfDNA samples may be obtained and analyzed from a cancer patient at the first and second time points to, for example, monitor cancer progression, determine whether the cancer is in remission (e.g., after treatment), monitor or detect residual disease or disease recurrence, or monitor the (e.g., therapeutic) effectiveness of the treatment.

[0227] Those skilled in the art will readily recognize that test specimens can be obtained from a cancer patient over any desired set of time points and analyzed by the methods of the present invention to monitor the patient's cancer status. In some embodiments, the first and second time points are between about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or about 24 hours, about 1, 2, 3, 4, 5, 10, 15, 20, 25, or about 50 days, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months, or about 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7. 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5, or about 30 years, or about 30 years, or about 15 minutes up to about 30 years apart. In other embodiments, test specimens can be obtained from a patient at least once every 5 months, at least once every 6 months, at least once every year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years.

[0228] VC treatment In yet another embodiment, the cancer predictions can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, determination of treatment efficacy, etc.). For example, in one embodiment, if the cancer prediction (e.g., for the cancer or for a particular cancer type) exceeds a threshold, a physician can prescribe an appropriate treatment (e.g., resective surgery, radiation therapy, chemotherapy, and / or immunotherapy). The physician can prescribe the appropriate treatment based on an analysis performed by the analytics system, e.g., analysis 140 of FIG. 1.

[0229] A classifier (as described herein) can be used to determine a cancer prognosis that a sample feature vector is from a subject with cancer. In one embodiment, when the cancer prognosis exceeds a threshold, an appropriate treatment (e.g., resection surgery or therapy) is prescribed. For example, in one embodiment, if the cancer prognosis is 60 or greater, one or more appropriate treatments are prescribed. In another embodiment, if the cancer prognosis is 65 or greater, 70 or greater, 75 or greater, 80 or greater, 85 or greater, 90 or greater, or 95 or greater, one or more appropriate treatments are prescribed. In other embodiments, the cancer prognosis may indicate the severity of the disease. An appropriate treatment may then be prescribed that matches the severity of the disease.

[0230] In some embodiments, the treatment is one or more cancer therapeutic agents selected from the group consisting of chemotherapeutic agents, targeted cancer therapeutic agents, differentiation therapeutic agents, hormonal therapeutic agents, and immunotherapeutic agents. For example, the treatment can be one or more chemotherapeutic agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeleton disrupting agents (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapeutic agents selected from the group consisting of signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapeutic agents, including retinoids such as tretinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormonal therapy agents selected from the group consisting of antiestrogens, aromatase inhibitors, progestins, estrogens, antiandrogens, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapeutic agents selected from the group including monoclonal antibody therapy, such as rituximab (RITUXAN) and alemtuzumab (CAMPATH), non-specific immunotherapies and adjuvants, e.g., BCG, interleukin-2 (IL-2), and interferon alpha, immunomodulatory agents, e.g., thalidomide and lenalidomide (REVLIMID). It is within the ability of a skilled physician or oncologist to select an appropriate cancer therapy agent based on characteristics such as tumor type, cancer stage, previous exposure to cancer treatments or cancer therapeutic agents, and other characteristics of the cancer.

[0231] VI. Kit Embodiments Also disclosed herein are kits for carrying out the methods described above, including those relating to cancer classifiers. The kits may include one or more collection containers for collecting a specimen containing genetic material from an individual. The specimen may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. The kits may also include reagents for isolating nucleic acids from the specimen. The reagents may further include reagents for sequencing the nucleic acids, including buffers and detection agents. In one or more embodiments, the kits may include one or more sequencing panels containing probes for targeting specific genomic regions, specific mutations, specific gene variants, or some combination thereof. In other embodiments, the specimen collected by the kits is provided to a sequencing laboratory, which may sequence the nucleic acids in the specimen using the sequencing panel.

[0232] The kit may further include instructions for use of the reagents included in the kit. For example, the kit may include instructions for collecting a specimen and extracting nucleic acids from the test specimen. Exemplary instructions may include the order in which reagents should be added, the centrifugation speed to use for isolating nucleic acids from the test specimen, how to amplify the nucleic acids, how to sequence the nucleic acids, or any combination thereof. The instructions may also reveal how to operate a computing device as analytics system 200 to perform any steps of the described method or process.

[0233] In addition to the above components, the kit may include a computer-readable storage medium storing computer software for implementing the various methods described throughout this disclosure. One form in which these instructions may be present is as printed information on a suitable medium or substrate, such as on one or more sheets of paper on which the information is printed, in the kit's packaging, in a package insert, etc. Yet another means may be a computer-readable medium on which the instructions are stored in the form of computer code, such as a diskette, CD, hard drive, or network data storage. Yet another means may be a website address that can be used via the Internet to access the information at a removed site. [Example]

[0234] VII. Illustrative Results VII.A. Specimen Collection and Processing Study Design and Samples: CCGA (NCT02889978) is a prospective, multicenter, case-control, observational study with longitudinal follow-up. De-identified biospecimens were collected from approximately 15,000 participants at 342 centers. Samples were divided into a training set (1,785 cases) and a test set (1,015 cases); samples were selected to ensure a pre-specified distribution of cancer and non-cancer types across centers in each cohort, and cancer and non-cancer samples were age-frequency matched by sex.

[0235] Whole-genome bisulfite sequencing: cfDNA was isolated from plasma and analyzed using whole-genome bisulfite sequencing (WGBS; 30x depth). cfDNA was extracted from two plasma tubes per patient (maximum total volume of 10 ml) using a modified QIAamp Circulating Nucleic Acid Kit (Qiagen; Germantown, MD). Up to 75 ng of plasma cfDNA was subjected to bisulfite conversion using the EZ-96 DNA Methylation Kit (Zymo Research, D5003). Dual-indexed sequencing libraries were prepared using the converted cfDNA using the Accel-NGS Methyl-Seq DNA Library Preparation Kit (Swift BioSciences; Ann Arbor, MI), and the constructed libraries were quantified using the KAPA Library Quantification Kit for Illumina platforms (Kapa Biosystems; Wilmington, MA). The four libraries were pooled with a 10% PhiX v3 library (Illumina, FC-110-3001) and clustered on an Illumina NovaSeq 7000 S2 flow cell followed by 150bp paired-end sequencing (30x).

[0236] For each sample, the WGBS fragment set was reduced to a small subset of fragments with abnormal methylation patterns. Additionally, hypermethylated or hypomethylated cfDNA fragments were selected. cfDNA fragments were selected because they had abnormal methylation patterns and were both hypermethylated or hypermethylated, i.e., UFXM. Fragments frequently present in cancer-free individuals or fragments with unstable methylation are unlikely to yield highly discriminatory features for cancer status classification. Therefore, we used an independent reference set of 108 cancer-free, non-smoking participants from the CCGA study (age: 58±14 years, 79 [73%] women) (i.e., the reference genome) to create a statistical model and data structure of representative fragments. These samples were used to train a Markov chain model (order 3) that estimates the likelihood of a given sequence of CpG methylation status within a fragment. This model was demonstrated to calibrate within the range of normal fragments (p-value > 0.001) and was used to reject fragments with p-values ​​≥ 0.001 from the Markov model as not being sufficiently abnormal.

[0237] As described above, a further data reduction step selected only fragments that covered at least five CpGs and had a mean methylation of either >0.9 (hypermethylated) or <0.1 (hypomethylated). This procedure resulted in a median (range) of 2,800 UFXM fragments (1,500-12,000) for participants without cancer at training and a median (range) of 3,000 UFXM fragments (1,200-420,000) for participants with cancer at training. Because this data reduction procedure used only the reference set data, this stage only needed to be applied once for each sample.

[0238] VII.B. Mixed Model Results The following mixture model results were obtained from a mixture model trained and deployed by the methods described in Figures 4A and 4B and throughout this disclosure. The mixture model is trained using training samples containing methylated sequence reads. The analytics system may tune or fix the number of components in the mixture model as a hyperparameter. Once trained, the analytics system validates the accuracy of the mixture model using samples from a validation set.

[0239] 10 illustrates exemplary results of a two-component mixture model, according to an exemplary embodiment. The mixture model is trained to distinguish between two components. The analytics system utilizes a validation set of 40 samples with various known component ratios. The samples have an average sequencing depth of 10. The mixture model evaluates over 5,000 methylation variants in the genome. The analytics system applies the trained two-component mixture model to the methylation signatures of the validation set of 40 samples.

[0240] The top graph 1010 represents the accuracy of component proportion prediction by the mixture model. In the top graph 1010, the x-axis represents the true component proportions between two components for the 40 validation samples, while the y-axis represents the component proportions predicted by the mixture model. The diagonal line from (0.00,0.00) to (1.00,1.00) represents the accuracy prediction of the component proportions. Qualitatively, all 40 validation samples lie on or touch this diagonal line.

[0241] The bottom graph 1020 represents the ability of the mixture model to predict a methylation signature based on known component proportions. In graph 1020, the x-axis represents the true methylation variant allele ratio, while the y-axis represents the methylation variant allele ratio predicted by the mixture model. The diagonal line from (0.00,0.00) to (1.00,1.00) represents the precision prediction of the methylation variant allele proportion. For a given value on graph 1020, the precision of the prediction can be determined based on the deviation from the diagonal line. Qualitatively, a small number of predicted methylation variant allele ratios deviate significantly from the true methylation variant allele ratio.

[0242] FIG. 11 shows exemplary results of a five-component mixture model according to an exemplary embodiment. The mixture model is trained to distinguish component proportions among five components. The analytics system utilizes a validation set containing 16 samples from five classes. The samples from each class have the same known component proportions. Specifically, class 1 contains 60% of the first component (dark green) and 10% of each of the remaining components; class 2 contains 60% of the second component (orange) and 10% of each of the remaining components; class 3 contains 60% of the third component (blue) and 10% of each of the remaining components; class 4 contains 60% of the fourth component (pink) and 10% of each of the remaining components; and class 5 contains 60% of the fifth component (light green) and 10% of each of the remaining components. The mixture model evaluates over 5,000 methylation variants in the genome. The analytics system applies the trained five-component mixture model to the methylation signatures of the validation set of samples from the five classes. The top graph 1110 illustrates five classes of samples with their known breakdown of component proportions. Each class is clustered along the x-axis. The y-axis represents proportions or ratios. The bottom graph 1120 illustrates the mixture model's prediction of the five class component proportions. The mixture model accurately predicts that all class 1 samples have more than 50% of the first component (dark green), all class 2 samples have more than 50% of the second component (orange), all class 3 samples have more than 50% of the third component (blue), all class 4 samples have more than 50% of the fourth component (pink), and all class 5 samples have more than 50% of the fifth component (light green). While validation samples can be artificially generated to have exact component proportions, the mixture model was able to accurately identify the dominant component proportions in each of the 80 samples.

[0243] FIG. 12 shows exemplary results of a mixture model deconvolving breast tissue-positive hormone receptor status, breast tissue-negative hormone receptor status, prostate tissue, and uterine tissue, according to an exemplary embodiment. The mixture model was trained to allow the number of components to be self-defined. The mixture model was applied to validation samples known to have one of the cancer types corresponding to the components. Classes are clustered along the x-axis. The y-axis represents proportions or ratios. Red components represent non-cancerous or general signals common to many samples. For the prostate class (all samples diagnosed with prostate cancer) and the uterus class (all samples diagnosed with uterine cancer), the mixture model predicts that the samples will predominantly have a predominant proportion of orange components (for the prostate class) and purple components (for the uterus class) (apart from the red component). Although this mixture model is independent of the component labels, the mixture model successfully deconvolved and attributed similar methylation signatures to the same components. Although there were a few uterine specimens that could have been predicted to have less of the purple component than the other non-red components, the overwhelming majority of uterine specimens were predicted by the mixture model to have a significant contribution from the purple component. For breast HER2- and breast HER2+, the mixture model determined that there was enough difference between the two classes to justify having two separate components. This mixture model had some challenges; some specimens were predicted to have a significant contribution from the other component. This could be due to mislabeling by the healthcare provider who diagnosed the specimen, or it could also be due to the two tissues having similar methylation signatures.

[0244] FIG. 13A illustrates exemplary results of a mixture model for deconvolving non-cancer impurities from colorectal tissue in a colorectal specimen, according to an exemplary embodiment. When training the mixture model, the analytics system tunes the number of components to three for the colorectal specimen illustrated in graph 1310. These three components are represented by red, green, and blue colors. The red component appears to correspond to putative non-cancer impurities present in most specimens. Each column corresponds to a colorectal specimen. The gray column corresponds to a specimen enriched in disseminated tumor cells. Notably, the mixture model is capable of determining the purity of such specimens. Using the mixture model to determine tissue purity resulted in an improvement in purity determination from 50% to 87%. The two components, blue and green, likely indicate heterogeneity in the colorectal cancer tissue.

[0245] 13B illustrates exemplary results of a mixture model for deconvolving non-cancer impurities from bladder tissue in a bladder specimen, according to an exemplary embodiment. When training the mixture model, the analytics system tunes the number of components to two for the bladder specimen illustrated in graph 1320. These two components are represented by red and blue colors. The red component appears to correspond to putative non-cancer impurities present in most specimens. Each column corresponds to a bladder specimen. The gray columns correspond to specimens enriched in disseminated tumor cells. Notably, the mixture model is capable of determining the purity of such specimens. Using the mixture model to determine tissue purity resulted in an improvement in purity determination from 19% to 67%.

[0246] VIII. Additional Considerations The foregoing detailed description of the embodiments refers to the accompanying drawings, in which specific embodiments of the present disclosure are illustrated. Other embodiments having different structure and operation do not depart from the scope of the present disclosure. The term "the present invention" and the like are used with reference to certain specific examples of the many alternative aspects or embodiments of Applicants' invention set forth herein, and neither its use nor its absence is intended to limit the scope of Applicants' invention or the claims.

[0247] Embodiments of the present invention also relate to apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes and / or it may comprise a general-purpose computing device selectively enabled or reconfigured by a computer program stored in the computer. Such computer programs may be stored on a non-transitory, tangible, computer-readable storage medium or any type of medium suitable for storing electronic instructions, which medium may be coupled to a computer system bus. Furthermore, any computing system referred to herein may include a single processor or may be an architecture employing a multiple-processor design for increased computing power.

[0248] Any of the steps, operations, or processes described herein as being performed by the analytics system may be performed or implemented by one or more hardware or software modules of the device, alone or in combination with other computing devices. In one embodiment, the software modules are implemented using a computer program product that includes a computer-readable medium with computer program code executable by a computer processor to perform each and every step, operation, or process described.

Claims

1. 1. A method of training a machine-learned mixture model to identify tissue types, comprising: obtaining a set of training samples comprising at least 1000 methylated sequence reads obtained from sequencing deoxyribonucleic acid (DNA) fragments; modifying each training specimen to generate a corresponding specimen methylation signature by determining, for each genomic region among a plurality of genomic regions, a first set of methylated sequence reads that overlap the genomic region and a second set of methylated sequence reads that include methylation variations in the genomic region, wherein the specimen methylation signature is generated based, at least in part, on the first set of methylated sequence reads and the second set of methylated sequence reads; generating a training dataset comprising the sample methylation signatures; and training the machine-learned mixture model using the training dataset, wherein the machine-learned mixture model is configured to identify contributions of each of a plurality of tissue-of-origin types to DNA fragments in a specimen. A method comprising:

2. 2. The method of claim 1 , wherein at least one training sample is known to contain a first tissue type of the plurality of tissue types of origin, and wherein training the mixture model comprises training the mixture model to identify a contribution of the first tissue type of origin to DNA fragments in the one training sample.

3. 3. The method of claim 1, wherein at least one training sample is known to have a contribution of DNA fragments from each of the plurality of tissue-of-origin types, and training the mixture model comprises training the mixture model to identify a contribution of each of the plurality of tissue-of-origin types to DNA fragments in the one training sample.

4. The method of any one of claims 1 to 3, wherein at least one training sample is one of a liquid biopsy sample, a tissue biopsy sample, and a purified sample.

5. The method according to any one of claims 1 to 4, wherein at least one genomic region consists of one CpG site.

6. The method of any one of claims 1 to 5, wherein at least one genomic region comprises a plurality of CpG sites.

7. determining an average sequencing depth for each of an initial set of genomic regions based on the methylated sequence reads of the training sample; and selecting the plurality of genomic regions by filtering out genomic regions having an average sequencing depth below a threshold depth; The method of any one of claims 1 to 6, further comprising:

8. The method of any one of claims 1 to 7, wherein the methylation variation in a genomic region is one of two methylation patterns in the genomic region.

9. The method of claim 8, wherein the two methylation patterns in one genomic region include methylation and unmethylation.

10. The method of any one of claims 1 to 9, wherein the methylation variation in a genomic region is one of more than three methylation patterns in the genomic region.

11. modifying each training sample to generate the corresponding sample methylation signature; determining, for each genomic region among the plurality of genomic regions, a third set of methylated sequence reads having a reference state for the genomic region, wherein the reference state is any methylation pattern that does not belong to the methylation variant; and generating the specimen methylation signature further based on the third set of methylated sequence reads. The method of any one of claims 1 to 10, further comprising:

12. 12. The method of any one of claims 1 to 11, wherein the tissue types include a combination of non-cancerous impurities; squamous cell carcinoma tissue; skin cancer tissue; melanoma tissue; lung cancer tissue; lung adenocarcinoma tissue; lung squamous cell carcinoma tissue; peritoneal cancer tissue; gastrointestinal cancer tissue; pancreatic cancer tissue; cervical cancer tissue; ovarian cancer tissue; liver cancer tissue; hepatocellular carcinoma tissue; liver cancer tissue; bladder cancer tissue; testicular cancer tissue; breast cancer tissue; brain cancer tissue; colon cancer tissue; rectal cancer tissue; colorectal cancer tissue; endometrial or uterine cancer tissue; salivary gland cancer tissue; kidney or renal cancer tissue; prostate cancer tissue; vulvar cancer tissue; thyroid cancer tissue; anal cancer tissue; penile cancer tissue; head and neck cancer tissue; esophageal cancer tissue; and nasopharyngeal carcinoma (NPC) tissue.

13. 13. The method of claim 12, wherein the non-cancerous impurities include one or more of lymphocytes, macrophages, fibroblasts, vascular endothelial cells, or non-cancerous tissue.

14. 14. The method of claim 12 or 13, wherein the methylation signature of the non-cancer impurity is obtained from a reference database comprising a plurality of methylation signatures for non-cancer impurities.

15. The method of any one of claims 1 to 14, wherein training the machine-learned mixture model is by maximum likelihood estimation.

16. 16. The method of claim 1, wherein training the machine-learned mixture model comprises tuning a number of tissue types as a hyperparameter of the machine-learned mixture model.

17. Tuning the number of tissue types as one hyperparameter of the machine-learned mixture model; For each number of tissue of origin types within the range: training the machine-learned mixture model having the number as the hyperparameter; determining maximum likelihood by cross-validating the trained machine-learned mixture model with a holdout sample set; and implementing a penalty on the maximum likelihood based on the number; and selecting an optimal number from the range as the hyperparameter based on penalized maximum likelihood; 17. The method of claim 16, comprising:

18. 18. The method of claim 1, wherein training the machine-learned mixture model comprises training with one or more machine learning algorithms.

19. 19. The method of any one of claims 1 to 18, wherein the machine-learned mixture model comprises a first set of tissue type models, each tissue type modeling a methylation signature of DNA fragments of a tissue type of origin, and training the machine-learned mixture model comprises training the first set of tissue type models.

20. 20. The method of claim 19, wherein training the first sub-model comprises training each tissue type model according to a beta distribution.

21. 21. The method of claim 1, wherein the machine-learned mixture model comprises a deconvolution model for deconvolving the tissue-of-origin contributions for each training sample, and training the machine-learned mixture model comprises training the deconvolution model.

22. 22. The method of claim 21, wherein training the deconvolution model comprises training the deconvolution model according to a binomial distribution.

23. 23. The method of claim 1, wherein the machine-learned mixture model comprises a first hierarchical hierarchy of one or more sub-models for predicting the contribution of a macro tissue-of-origin type and a second hierarchical hierarchy of one or more sub-models for predicting the contribution of a tissue-of-origin type below the macro tissue type, wherein one first hierarchical sub-model predicts the contribution of one macro tissue type and a set of one or more second hierarchical sub-models predict the contribution of a set of tissue types below the one macro tissue type that is equal to the contribution of the one macro tissue type.

24. 1. A method for identifying the contribution of tissue type of origin to DNA fragments in a test specimen, comprising: obtaining a test specimen comprising at least 1000 methylation sequence reads for DNA fragments in the test specimen; For each of the plurality of genomic regions: determining a first set of methylated sequence reads that overlap the genomic region; and determining a second set of methylated sequence reads having alternative methylation signatures for the genomic region; generating a sample methylation signature for the training sample based on the first set of methylated sequence reads and the second set of methylated sequence reads across the plurality of genomic regions; and Identifying contributions of each tissue type of origin for DNA fragments in the test specimen by applying a machine learning mixture model to the specimen methylation signature, optionally wherein the machine learning mixture model is trained by a method according to any one of claims 1 to 23. A method comprising:

25. reporting the identified contribution of the tissue type of origin to a client device.

25. The method of claim 24, further comprising:

26. identifying the tissue type of origin of the highest contribution; and reporting treatment recommendations for the treatment of cancers originating primarily from said identified tissue type of origin; 26. The method of claim 24 or 25, further comprising:

27. comparing the tissue-of-origin contribution to the tissue-of-origin predicted prior to treatment; determining the efficacy of the treatment based on said comparison; and Reporting treatment recommendations based on said treatment effectiveness The method of any one of claims 24 to 26, further comprising:

28. determining the treatment as successful if the contribution of the first tissue type of origin is lower than a prior contribution of the first tissue type of origin predicted prior to the treatment; and reporting said success of said treatment. The method of any one of claims 24 to 27, further comprising:

29. determining the treatment as a failure if the contribution rate of the first tissue type of origin is higher than the prior contribution rate of the first tissue type of origin predicted before the treatment; and reporting said failure of said treatment; The method of any one of claims 24 to 28, further comprising:

30. reporting a treatment modification based on said failure of said treatment; 30. The method of claim 29, further comprising:

31. The method of any one of claims 24 to 30, wherein the test specimen is suspected to have cancer.

32. 32. The method of claim 31 , wherein the test specimen is predicted to have cancer by a cancer classifier.

33. 33. The method of claim 32, wherein the cancer classifier is trained on a cancer cohort of training samples and a non-cancer cohort of training samples, and each training sample from the cancer cohort and the non-cancer cohort comprises at least 1000 methylated sequence reads for DNA fragments in the training sample.

34. determining a subset of methylated sequence reads originating from non-cancer impurity types by applying the machine learning mixture model to the sample methylation signature; removing a subset of the methylated sequence reads originating from the non-cancer impurity type, thereby obtaining a feature set of methylated sequence reads; and predicting cancer in the test specimen by applying a cancer classifier to the feature set of the methylated sequence reads.

25. The method of claim 24, further comprising:

35. 1. A method for training a cancer classifier, comprising: obtaining a cancer cohort of training samples and a non-cancer cohort of training samples, wherein each training sample from the cancer cohort and the non-cancer cohort comprises at least 1000 methylated sequence reads for DNA fragments in the training sample; For each training sample: For each of the plurality of genomic regions: determining a first set of methylated sequence reads that overlap the genomic region; determining a second set of methylated sequence reads having alternative methylation signatures for the genomic region; generating a sample methylation by the specimen methylation signature is based, in part, on the first set of methylated sequence reads and the second set of methylated sequence reads; and applying a machine learning mixture model, optionally trained by the method of any one of claims 1 to 23, to the sample methylation signatures of each training sample in the cancer cohort to identify a subset of methylated sequence reads originating from non-cancer impurity types; for each training sample in the cancer cohort, filtering out a subset of the methylated sequence reads that originate from the non-cancer impurity type, thereby obtaining a resulting feature set of methylated sequence reads; generating a feature set of methylated sequence reads for each training sample in the non-cancer cohort; and training the cancer classifier with a feature set of the methylated sequence reads for the training samples from the cancer cohort and a feature set of methylated sequence reads for the training samples from the non-cancer cohort; A method comprising:

36. 36. The method of claim 35, wherein the cancer classifier is trained as a machine learning model.

37. 1. A method for training a cancer classifier, comprising: obtaining, for each training sample of a plurality of training samples, a set of methylation sequence reads obtained from sequencing DNA fragments of the training sample, wherein the plurality of training samples includes a training sample obtained from a first cohort of subjects diagnosed with cancer and a training sample obtained from a second cohort of subjects not diagnosed with cancer; For each training sample in the first cohort, and for each methylated sequence read of the training sample: applying the trained tissue type model to predict the likelihood that the methylated sequence reads are obtained from non-cancerous tissue; and filtering out the methylated sequence read if the predicted likelihood is above a threshold; generating, for each training sample in the first cohort, a methylation signature based on the non-excluded methylated sequence reads of the training sample; generating, for each training sample in the second cohort, a methylation signature based on the methylated sequence reads of the training sample; and and training a cancer classifier to detect the presence of a cancer signal having said methylation signature in said training samples. A method comprising:

38. 38. The method of claim 37, wherein the tissue-type model is trained with a beta distribution using methylation signatures of non-cancer training specimens generated from methylated sequence reads obtained by sequencing DNA fragments of the non-cancer training specimens.

39. 39. The method of claim 38, wherein the methylation signatures of the non-cancer training specimens are obtained from a reference database.

40. 40. The method of any one of claims 37 to 39, wherein the cancer classifier is trained as a machine learning model.

41. 41. The method of any one of claims 37 to 40, wherein the methylation signature is based on the identification of one or more methylation mutations in each genomic region of a plurality of genomic regions.

42. 1. A method for identifying a tissue type of origin for a DNA fragment, comprising: obtaining methylation sequence reads for the DNA fragments, the methylation sequence reads comprising a methylation signature across one or more genomic regions; predicting the likelihood that the methylation signature of the methylated sequence read is from each of a plurality of tissue types of origin by applying each of a plurality of tissue type models to the methylation signature, wherein the plurality of tissue type models are optionally part of a mixture model trained by the method of any one of claims 1 to 23; determining that the DNA fragment is most likely derived from the tissue type of origin; and Return a report of said decision A method comprising:

43. A non-transitory computer readable storage medium storing instructions which, when executed by a computer processor, cause the computer processor to perform the method of any one of claims 1 to 42.

44. 44. A system comprising: a computer processor; and the non-transitory computer-readable storage medium of claim 43.

45. 24. A computer program product comprising a non-transitory computer-readable storage medium storing a machine learning mixture model for predicting tissue type proportions from a test specimen, the product being made by the method of any one of claims 1 to 23.

46. 42. A computer program product comprising a non-transitory computer-readable storage medium storing a machine learning cancer classifier for predicting cancer in a test specimen, the product being made by the method of any one of claims 33-41.

47. A treatment kit comprising: a collection container for collecting a DNA sample from a subject; Optionally, one or more reagents for isolating DNA fragments in said DNA preparation; Optionally, one or more probes targeting one or more genomic loci determined to be indicative of a cancerous condition; and A non-transitory computer-readable storage medium according to claim 43 or a computer program product according to claim 45 or 46. A medical kit containing: