Model-Based Characterization and Classification Optimization

By training separate classifiers for different populations based on covariate characteristics and using tissue models, the method optimizes cancer classification from nucleic acid samples, addressing variations in covariate characteristics to enhance detection accuracy.

JP2025539741APending Publication Date: 2025-12-09GRAIL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025526856
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-16
Filing Date
2023-11-16
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing cancer classification methods struggle to accurately predict disease states from nucleic acid samples due to variations in covariate characteristics, such as age and biological sex, leading to suboptimal detection of cancer signals in different populations.

Method used

The method involves obtaining training samples with methylation-sequence reads from individuals, subdividing these samples based on covariate characteristics, and training separate classifiers for each population to determine cutoff thresholds and generate signal vectors, using tissue models and machine learning algorithms to optimize cancer classification.

Benefits of technology

This approach enhances the accuracy of cancer detection by accounting for individual differences, enabling early identification of cancer through improved classification of disease states based on DNA methylation patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025539741000001_ABST
    Figure 2025539741000001_ABST
Patent Text Reader

Abstract

One or more techniques are disclosed for optimizing cancer classification based on covariate characteristics. In a first approach, the analysis system can determine separate cutoff thresholds for positively detecting disease signals for various labels of the covariate characteristics. The system can subdivide the training samples based on the labels of the covariate characteristics to separately determine the cutoff thresholds. In another approach, the system can train different classifiers for each population. The system separates the training samples based on the labels of the covariate characteristics and trains the classifiers separately to generate signal vectors representing the amount of disease signal detected in the samples. The classifiers can be trained for different feature sets determined based on mutual information gain, genomic region coverage, and healthy activation fraction.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 425,889, filed November 16, 2022, which is incorporated by reference in its entirety.

[0002] 1. Field The present disclosure relates generally to model-based characterization and classifiers for predicting disease states from nucleic acid samples. [Background technology]

[0003] 2. Description of Related Technology DNA methylation plays a role in the regulation of gene expression. Abnormal DNA methylation is involved in many disease processes, including cancer. DNA methylation profiling using methylation sequencing (e.g., whole-genome bisulfite sequencing (WGBS)) is increasingly recognized as a valuable diagnostic tool for the detection, diagnosis, and / or monitoring of cancer. For example, specific patterns of differentially methylated regions can be useful as molecular markers for various disease states. Summary of the Invention [Problem to be solved by the invention]

[0004] Cancer classification generally involves applying one or more predictive models to features derived from gene sequencing data to predict an individual's disease state. In particular, cancer classification can utilize tissue models to characterize sequencing data by determining the tissue origin of sequence reads associated with nucleic acid fragments. The disease state can be a binary prediction that generally indicates the presence of a disease signal, or a multi-class prediction that indicates the presence of a specific disease signal.

[0005] One or more techniques can be implemented to optimize cancer classification based on covariate characteristics. In a first approach, the analysis system can determine separate cutoff thresholds for positively detecting disease signals for various labels of the covariate characteristics. The system can subdivide the training samples based on the labels of the covariate characteristics to separately determine the cutoff thresholds. In another approach, the system can train different classifiers for each population. The system separates the training samples based on the labels of the covariate characteristics and trains the classifiers separately to generate signal vectors representing the amount of disease signal detected in the samples. The classifiers can be trained for different feature sets determined based on mutual information gain, genomic region coverage, and healthy activation fraction. [Means for solving the problem]

[0006] Clause 1. Obtaining a plurality of training samples from individuals, each having one of a plurality of disease states and each associated with one of a plurality of labels for a covariate characteristic, wherein each training sample comprises methylation-sequence reads for at least 1,000 cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; generating a first set of training samples and a second set of training samples by subdividing the plurality of training samples; for each training sample, determining a feature vector based on the methylation-sequence reads of the training sample; for each methylation-sequence read of the training sample, applying each of a plurality of tissue models to the methylation-sequence reads to determine a likelihood that the methylation-sequence read is informative of the presence of one of the disease states associated with the tissue model; and determining a likelihood output by the tissue model that is most informative of the presence of one of the disease states. 1. A method comprising: generating a cancer signal vector for each training sample in the second set of training samples by assigning methylated sequence reads to one of a high disease state and determining a feature vector based on the methylated sequence reads assigned to each disease state; training a classifier with the feature vectors of a first set of training samples to generate a signal vector based on the input feature vectors, wherein the signal vector includes a value for each disease state; applying a cancer classifier to the feature vector of each training sample in the second set of training samples to generate a cancer signal vector for each training sample in the second set of training samples; and determining a cutoff threshold for each label of a plurality of labels for a covariate feature based on the signal vectors for training samples in the second set of training samples having the labels.

[0007] Clause 2. The method of clause 1, wherein the methylation sequence reads are obtained from a targeted methylation sequencing assay or a whole-genome bisulfite sequencing assay.

[0008] Clause 3. The method of clause 1 or 2, wherein each training sample comprises methylation sequence reads for at least 10,000 cfDNA fragments.

[0009] Clause 4. The method of any of clauses 1-3, wherein the first set of training samples and the second set of training samples comprise training samples with similar proportions of disease states.

[0010] Clause 5. The method of any of clauses 1-4, wherein the plurality of disease states comprises a non-cancerous state and one or more cancer states for one or more cancers of different origins.

[0011] Clause 6. The method of Clause 5, wherein the one or more cancers of different origins comprise breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis and ureter, non-urothelial renal cancer, prostate cancer, anorectal cancer, colorectal cancer, squamous cell carcinoma of the esophagus, non-squamous esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from hepatocytes, hepatobiliary carcinoma arising from cells other than hepatocytes, pancreatic cancer, human papillomavirus-associated head and neck cancer, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung carcinoma, squamous cell lung carcinoma, lung cancer other than adenocarcinoma or small cell lung carcinoma, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. In some embodiments, the type of cancer is additionally selected from the group comprising brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma, and gallbladder cancer.

[0012] Clause 7. The method of any of clauses 1-6, wherein the covariate characteristic is one of age, smoker status, and biological sex.

[0013] Clause 8. The method of any of clauses 1-7, further comprising the step of: for each methylated sequence read, determining whether the methylated sequence read has an informative methylation pattern by p-value filtering; and wherein a feature vector for each training sample is generated based on the methylated sequence reads having an informative methylation pattern.

[0014] Clause 9. The method of any of clauses 1-8, wherein each tissue model of the plurality of tissue models is trained by: obtaining a second plurality of training samples, each having one of the plurality of disease states, wherein each training sample comprises methylation sequence reads for cfDNA fragments obtained from an individual; generating, for each tissue model, a training dataset comprising at least 10,000 methylation sequence reads of training samples having the disease state associated with the tissue model; and training each tissue model with the associated training dataset such that the methylation sequence reads predict the likelihood of the presence of the associated disease state being informative.

[0015] Clause 10. The method of any of clauses 1 to 9, wherein each tissue model is one of a binomial model, an independent site model, a Markov model, or a mixed model.

[0016] Clause 11. The method of any of clauses 1 to 9, wherein the classifier is a machine learning model.

[0017] Clause 12. The method of any of clauses 1-11, further comprising: obtaining a test sample from a test individual having an unknown disease state and associated with a first label of a plurality of labels for the covariate feature, the test sample comprising methylated sequence reads for cfDNA fragments obtained from the test individual; generating a test feature vector based on the methylated sequence reads of the test sample by, for each methylated sequence read of the test sample, applying each of a plurality of tissue models to the methylated sequence read to determine a likelihood that the methylated sequence read is informative for the presence of a disease state associated with the tissue model, assigning the methylated sequence read to one of the disease states most likely output by the tissue model, and determining a test feature vector based on the methylated sequence reads assigned to each disease state; applying a classifier to the test feature vector of the test to generate a signal vector for the test sample; and detecting a positive disease signal for one or more of the disease states by applying a cutoff threshold associated with the first label for the covariate feature to the signal vector for the test sample.

[0018] Clause 13. The method of Clause 12, wherein detecting a positive disease signal for one or more of the disease states comprises, for each disease state, applying a cutoff threshold for the disease state to values ​​in the signal vector for the test sample corresponding to the disease state.

[0019] Clause 14. The method of any of clauses 1-13, wherein each individual is further associated with one of a second plurality of labels for a second covariate characteristic; determining a cutoff threshold comprises determining a cutoff threshold for each combination of one label from the plurality of labels for the covariate characteristic and one label from the second plurality of labels for the second covariate characteristic; the test sample is further associated with a second label for the second covariate characteristic; and detecting a positive disease signal for one or more of the disease states comprises applying a cutoff threshold associated with the combination of the first label for the covariate characteristic and the second label for the second covariate characteristic.

[0020] Clause 15. The method of any of clauses 12-14, further comprising the step of reporting a positive disease signal to a healthcare provider for further diagnostic workup steps.

[0021] Clause 16. Obtaining a plurality of training samples from individuals, each having one of a plurality of disease states and each associated with one of a plurality of labels for a covariate characteristic, wherein each training sample comprises methylation sequence reads for at least 1,000 cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; for each training sample, determining a feature vector based on the methylation sequence reads of the training sample; and for each methylation sequence read of the training sample, applying each of a plurality of tissue models to the methylation sequence read to determine a likelihood that the methylation sequence read is informative for the presence of one of the disease states associated with the tissue model. generating a signal vector based on the input feature vector, wherein the signal vector includes a value for each disease state, the signal vector including a value for each disease state, the value for each disease state, and a value for each disease state; generating a set of training samples for each label of a plurality of labels for the covariate characteristics, the set of training samples including a label for the covariate characteristics and a feature vector of the training samples associated with the label; and training a classifier for each label with the feature vector of the set of training samples corresponding to the label to generate a signal vector based on the input feature vector, the signal vector including a value for each disease state.

[0022] Clause 17. The method of clause 16, wherein the methylation sequence reads are obtained from a targeted methylation sequencing assay or a whole genome bisulfite sequencing assay.

[0023] Clause 18. The method of clause 16 or 17, wherein each training sample comprises methylation sequence reads for at least 10,000 cfDNA fragments.

[0024] Clause 19. The method of any of clauses 16-18, wherein the plurality of disease states comprises a non-cancerous state and one or more cancer states for one or more cancers of different origins.

[0025] Clause 20. The method of Clause 19, wherein the one or more cancers of different origins comprise breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis and ureter, non-urothelial renal cancer, prostate cancer, anorectal cancer, colorectal cancer, squamous cell carcinoma of the esophagus, non-squamous esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from hepatocytes, hepatobiliary carcinoma arising from cells other than hepatocytes, pancreatic cancer, human papillomavirus-associated head and neck cancer, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung carcinoma, squamous cell lung carcinoma, lung cancer other than adenocarcinoma or small cell lung carcinoma, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. In some embodiments, the type of cancer is additionally selected from the group comprising brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma, and gallbladder cancer.

[0026] Clause 21. The method of any of clauses 16 to 20, wherein the covariate characteristic is one of age, smoking status, and biological sex.

[0027] Clause 22. The method of any of clauses 16-21, further comprising the step of: for each methylated sequence read, determining whether the methylated sequence read has an informative methylation pattern by p-value filtering, wherein a feature vector for each training sample is generated based on the methylated sequence reads having an informative methylation pattern.

[0028] Clause 23. The method of any of clauses 16-22, wherein each tissue model of the plurality of tissue models is trained by: obtaining a second plurality of training samples, each having one of the plurality of disease states, wherein each training sample comprises methylation sequence reads for cfDNA fragments obtained from an individual; generating, for each tissue model, a training dataset comprising at least 10,000 methylation sequence reads of training samples having the disease state associated with the tissue model; and training each tissue model with the associated training dataset such that the methylation sequence reads predict the likelihood of the presence of the associated disease state being informative.

[0029] Clause 24. The method of any of clauses 16 to 23, wherein each tissue model is one of a binomial model, an independent site model, a Markov model, or a mixed model.

[0030] Clause 25. The method of any of clauses 16 to 24, wherein each classifier is a machine learning model.

[0031] Clause 26. The method of any of clauses 16-25, further comprising: obtaining a test sample from a test individual having an unknown disease state and associated with a first label of a plurality of labels for the covariate feature, the test sample comprising methylated sequence reads for cfDNA fragments obtained from the test individual; generating a test feature vector based on the methylated sequence reads of the test sample by, for each methylated sequence read of the test sample, applying each of a plurality of tissue models to the methylated sequence read to determine a likelihood that the methylated sequence read is informative for the presence of one of the disease states associated with the tissue model, assigning the methylated sequence read to one of the disease states most likely output by the tissue model, and determining a test feature vector based on the methylated sequence read assigned to each disease state; applying a classifier corresponding to the first label to the test test feature vector of the test sample; and detecting a positive disease signal for one or more of the disease states based on the signal vector for the test sample.

[0032] Clause 27. The method of Clause 26, wherein detecting a positive disease signal for one or more of the disease states comprises identifying the disease state with the greatest value in the signal vector as the disease state having a detected positive disease signal.

[0033] Clause 28. The method of any of clauses 16-27, wherein each individual is further associated with one of a second plurality of labels for the second covariate characteristic; generating the set of training samples includes generating a set of training samples for each combination of one label from the plurality of labels for the covariate characteristic and one label from the second plurality of labels for the second covariate characteristic; training the classifier includes training a classifier for each combination of a feature vector from the set of training samples corresponding to the combination; the test sample is further associated with a second label for the second covariate characteristic; and applying the classifier includes applying the classifier trained on the combination of the first label and the second label.

[0034] Clause 29. The method of any of clauses 12-14, further comprising the step of reporting a positive disease signal to a health care provider for further diagnostic workup steps.

[0035] Clause 30. Obtaining a plurality of training samples from individuals, each having one of a plurality of disease states, and each associated with one of a plurality of labels for a covariate characteristic, wherein each training sample comprises methylation sequence reads for at least 1,000 cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; generating, for each training sample, a feature vector based on the methylation sequence reads of the training sample by: determining, for each of a plurality of genomic regions, a methylation feature value based on one or more methylation sequence reads of the training sample that overlap with the genomic region; and determining a feature vector based on the methylation feature values ​​across the plurality of genomic regions; for the covariate characteristic. generating a set of training samples for each label of a plurality of labels for a covariate feature, the set including feature vectors of the training samples associated with the labels; determining, for each label, a mutual information score for each genomic region based on the feature vectors of the set of training samples corresponding to the labels; ranking the genomic regions based on the mutual information scores; selecting a set of features from the ranked genomic regions; modifying the feature vector to include methylation feature values ​​for the set of features; and training a classifier with the modified feature vectors to generate a signal vector based on the input feature vector, the signal vector including a value for each disease state.

[0036] Clause 31. The method of clause 30, wherein the methylation sequence reads are obtained from a targeted methylation sequencing assay or a whole genome bisulfite sequencing assay.

[0037] Clause 32. The method of clause 30 or 31, wherein each training sample comprises methylation sequence reads for at least 10,000 cfDNA fragments.

[0038] Clause 33. The method of any of clauses 30 to 32, wherein each genomic region comprises one or more CpG sites.

[0039] Clause 34. The method of any of clauses 30 to 33, wherein the methylation feature value is a methylation density in a genomic region.

[0040] Clause 35. The method of any of clauses 30-33, further comprising the step of: for each methylated sequence read, determining whether the methylated sequence read has an informative methylation pattern by p-value filtering, wherein a feature vector for each training sample is generated based on the methylated sequence reads having an informative methylation pattern.

[0041] Clause 36. The method of Clause 35, wherein the methylation feature value is a count of methylated sequence reads that overlap with the genomic region.

[0042] Clause 37. The method of any of clauses 30 to 36, wherein the plurality of disease states comprises a non-cancerous state and one or more cancer states for one or more cancers of different origins.

[0043] Clause 38. The method of Clause 37, wherein the one or more cancers of different origins comprise breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis and ureter, non-urothelial renal cancer, prostate cancer, anorectal cancer, colorectal cancer, squamous cell carcinoma of the esophagus, non-squamous esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from hepatocytes, hepatobiliary carcinoma arising from cells other than hepatocytes, pancreatic cancer, human papillomavirus-associated head and neck cancer, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung carcinoma, squamous cell lung carcinoma, lung cancer other than adenocarcinoma or small cell lung carcinoma, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. In some embodiments, the type of cancer is additionally selected from the group comprising brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma, and gallbladder cancer.

[0044] Clause 38. The method of any of clauses 30 to 37, wherein the covariate characteristic is one of age, smoker status, and biological sex.

[0045] Clause 39. The method of any of clauses 30-38, wherein determining a mutual information score for each genomic region for each label comprises: determining a pairwise information gain for each pairwise combination of disease states based on the discriminatory power of the genomic regions to classify the two disease states; and determining a mutual information score by combining the pairwise information gains across the pairwise combinations of disease states.

[0046] Clause 40. The method of any of clauses 30-39, further comprising: for each label, determining a coverage for each genomic region based on the sequencing depth of a set of training samples corresponding to the label; and excluding one or more genomic regions from selection of the set of features based on the coverage being below a threshold.

[0047] Clause 41. The method of any of clauses 30-40, further comprising: determining, for each label, for each genomic region, an activation score for the non-cancer disease state indicative of the number of training samples having the non-cancer disease state that exhibit activation of the genomic region; and excluding one or more genomic regions from selection of the set of features based on the activation score exceeding a threshold.

[0048] Clause 42. The method of any of clauses 30 to 41, wherein the set of features selected for each label is determined to optimize the accuracy of the classifier.

[0049] Clause 43. Any of the methods of clauses 30 to 42, in which a set of characteristics for one label is different from another set of characteristics for another label.

[0050] Clause 44. The method of clause 43, wherein one set of features includes one or more features that are different from another set of features.

[0051] Clause 45. The method of clause 43, wherein one set of features includes a different number of features than another set of features.

[0052] Clause 46. The method of any of clauses 30 to 45, wherein each classifier is a machine learning model.

[0053] Clause 47. The method of any of clauses 16-25, further comprising: obtaining a test sample from a test individual having an unknown disease state and associated with a first label of the plurality of labels for the covariate characteristic, the test sample comprising methylation sequence reads for cfDNA fragments obtained from the test individual; generating a test feature vector based on the methylation sequence reads of the test sample by, for each of the plurality of genomic regions, determining a methylation feature value based on one or more methylation sequence reads of the test sample that overlap with the genomic region, and determining the test feature vector based on the methylation feature values ​​across the plurality of genomic regions; applying a classifier corresponding to the first label to the test test feature vector to generate a signal vector for the test sample; and detecting a positive disease signal for one or more of the disease states based on the signal vector for the test sample.

[0054] Clause 48. The method of Clause 47, wherein detecting a positive disease signal for one or more of the disease states comprises identifying the disease state with the greatest value in the signal vector as the disease state having a detected positive disease signal.

[0055] Clause 49. The method of any of clauses 12-14, further comprising the step of reporting a positive disease signal to a health care provider for further diagnostic workup steps.

[0056] Clause 50. A method for optimizing feature selection, comprising: obtaining, via an analytical system, a plurality of sequence reads from a sample; generating, via the analytical system, a plurality of features based on the plurality of sequence reads; generating, via the analytical system, one or more of an activation fraction, a mutual information score, or a coverage based on the plurality of sequence reads; determining, via the analytical system, a rank of the plurality of features based on the one or more of the activation fraction, mutual information score, or coverage; and selecting, via the analytical system, a subset of features from the plurality of features based on the rank of the plurality of features.

[0057] Clause 51. The method of clause 50, wherein the sample comprises a cell-free nucleic acid sample.

[0058] Clause 52. The method of clause 50 or 51, wherein the plurality of features comprises a sequence read count of sequence reads that exceed a ratio threshold.

[0059] Clause 53. The method of any of clauses 50-52, wherein generating a plurality of features comprises determining methylation rates for a plurality of CpG sites within a plurality of sequence reads.

[0060] Clause 54. The method of any of clauses 50 to 53, wherein the activation fraction includes a non-cancer activation fraction.

[0061] Clause 55. The method of any of clauses 50-54, wherein the mutual information scores of the mutual information scores corresponding to features of the plurality of features are inversely proportional to the ranks of the features.

[0062] Clause 56. The method of any of clauses 50-55, wherein lower coverage corresponds to noisier features.

[0063] Clause 57. The method of any of clauses 50-56, further comprising training at least one classifier based on the subset of features, wherein the at least one classifier is trained to predict the presence or absence of a disease, a disease type, and / or a tissue of origin of the disease.

[0064] Clause 58. The method of clause 57, wherein at least one classifier is trained to determine the tissue of origin of the disease based on one or more characteristics.

[0065] Clause 59. The method of clause 58, wherein the one or more characteristics include at least one of age, smoking status, and sex.

[0066] Clause 60. The method of any of clauses 50-59, further comprising: determining a first activation fraction for a disease-positive prediction; and determining a second activation fraction for a disease-negative prediction; and wherein the subset of features is selected at least in part by minimizing the difference between the first activation fraction and the second activation fraction.

[0067] Clause 61. The method of any of clauses 50-60, further comprising: generating a set of training data corresponding to the subset of features, the training data being generated from sequence data labeled by disease type or tissue of origin, or labeled as healthy; and training a machine learning model using the generated set of training data, the machine learning model being configured to predict disease state and type based on the sequence data corresponding to the subset of features.

[0068] Clause 62. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the processor to perform the method of any of clauses 1 to 61.

[0069] Clause 63. A system comprising: a computer processor; and the non-transitory computer-readable storage medium of clause 62.

[0070] Clause 64. A treatment kit comprising: a collection container for storing a biological sample containing cfDNA fragments; optionally, one or more reagents for isolating cfDNA fragments from the biological sample; optionally, a sequencing panel comprising probes for targeting specific genomic regions; and the non-transitory computer-readable storage medium of Clause 62. [Brief explanation of the drawings]

[0071] [Figure 1A] 1 is a flowchart of a method for generating a classifier for predicting a disease state, according to various embodiments. [Figure 1B] 1 is a flowchart of a method for generating a classifier for predicting a disease state, according to various embodiments. [Figure 2A] 1 shows a flow chart of a device for sequencing a nucleic acid sample according to one embodiment. [Figure 2B] FIG. 1 is a block diagram of an analytical system for processing sequence reads, according to various embodiments. [Figure 3] 1 is a flowchart illustrating a process for sequencing a nucleic acid, according to various embodiments. [Figure 4A] 4 illustrates a portion of the process of FIG. 3 for sequencing nucleic acids to obtain methylation information and methylation state vectors, according to various embodiments. [Figure 4B] 10 illustrates the generation of a data structure for a control group, according to various embodiments. [Figure 4C] 1 shows a flowchart illustrating a process for determining informative fragments from a sample, according to various embodiments. [Figure 5] 1 illustrates a block of a reference genome, according to various embodiments. [Figure 6] 1 illustrates a process for determining features for training a classifier, according to various embodiments. [Figure 7A] 1 includes a confusion matrix showing the accuracy of a classifier, according to various embodiments. [Figure 7B]1 includes a confusion matrix showing the accuracy of a classifier, according to various embodiments. [Figure 7C] 1 includes a confusion matrix showing the accuracy of a classifier, according to various embodiments. [Figure 7D] 1 includes a confusion matrix showing the accuracy of a classifier, according to various embodiments. [Figure 7E] 1 includes a confusion matrix showing the accuracy of a classifier, according to various embodiments. [Figure 7F] 1 includes a confusion matrix showing the accuracy of a classifier, according to various embodiments. [Figure 8] 1 is a flowchart of a method for model-based characterization, according to various embodiments. [Figure 9A] 10 illustrates the sensitivity of a tissue-of-origin classifier, according to one embodiment. [Figure 9B] 10 illustrates the sensitivity of a tissue-of-origin classifier, according to one embodiment. [Figure 10A] 10 illustrates the sensitivity of the tissue-of-origin classifier at different cancer stages, according to one embodiment. [Figure 10B] 10 illustrates the sensitivity of the tissue-of-origin classifier at different cancer stages, according to one embodiment. [Figure 11] 1 shows a performance grid representing the accuracy of tissue-of-origin localization, according to one embodiment. [Figure 12] 1 shows the accuracy and sensitivity of a tissue-of-origin classifier at various cancer stages, according to one embodiment. [Figure 13A] 10 shows a ROC curve for a tissue-of-origin classifier according to one embodiment. [Figure 13B] 10 shows a ROC curve for a tissue-of-origin classifier according to one embodiment. [Figure 14] FIG. 1 is a data flow diagram for training models, according to various embodiments. [Figure 15] 1 shows precision-recall curves for uncertain call thresholds, according to various embodiments. [Figure 16]1 is a flowchart of a method for determining the probability that a sample has a disease state, according to various embodiments. [Figure 17] 1 illustrates the performance gain in sensitivity of a multi-layer perceptron model according to one embodiment. [Figure 18] 1 shows experimental results of a multi-layer perceptron model in determining tissue of origin, according to one embodiment. [Figure 19] 1 shows experimental results of a multi-layer perceptron model in determining tissue of origin by cancer stage, according to one embodiment. [Figure 20] 1 shows experimental results of a multi-layer perceptron model across cancer types, according to one embodiment. [Figure 21] 1 shows a graph of the likelihood of cancer type for non-cancer samples above 95% specificity. [Figure 22] 1 shows a graph of methylation sequencing data for non-cancer samples and hematological subtype cancer samples. [Figure 23A] 1 shows a flowchart illustrating a process for determining a binary threshold cutoff for binary cancer classification, according to one or more embodiments. [Figure 23B] 1 shows a flowchart illustrating the process of thresholding tissue-of-origin labels to determine a binary threshold cutoff for binary cancer classification, according to one or more embodiments. [Figure 24A] 1 shows a confusion matrix illustrating the performance of a cancer tissue-of-origin classifier trained with additional hematological cancer subtypes. [Figure 24B] 10 shows a confusion matrix illustrating the performance of a cancer tissue-of-origin classifier trained with additional hematological cancer subtypes. [Figure 25A] 1 shows a graph illustrating cancer prediction accuracy for cancer classifiers with and without adjusting threshold cutoffs for multiple cancer types across cancer stages. [Figure 25B] 1 shows a graph illustrating cancer prediction accuracy for cancer classifiers with and without adjusting threshold cutoffs for multiple cancer types across cancer stages. [Figure 26A] Receiver operating curves (ROC) showing the sensitivity and specificity of cancer detection using methylation data for the target genomic regions of assay panel A are shown. [Figure 26B] 1 is a confusion matrix showing the accuracy of cancer type classification for subjects determined to have cancer using methylation data for the target genomic regions of assay panel A. [Figure 27A] Receiver operating curves (ROC) showing the sensitivity and specificity of cancer detection using methylation data for the target genomic regions of assay panel B are shown. [Figure 27B] 10 is a confusion matrix showing the accuracy of cancer type classification for subjects determined to have cancer using methylation data for the target genomic regions of assay panel B. [Figure 28] 1 shows classifier performance for a proprietary cancer assay panel (Assay Panel C), according to one embodiment. [Figure 29A] 1 shows a tissue of origin (TOO) confusion matrix depicting the accuracy of cancer tissue of origin localization for assay panel C, according to one embodiment. [Figure 29B] 1 shows a tissue of origin (TOO) confusion matrix depicting the accuracy of cancer tissue of origin localization for assay panel C, according to one embodiment. [Figure 30] 10 shows classifier sensitivity performance in individual tumors by stage for assay panel C, according to one embodiment. [Figure 31] 10 illustrates tissue-of-origin accuracy for multiple iterations of a trained model according to various embodiments. [Figure 32] 1 illustrates a process for stratifying hematological signals into two tiers, according to various embodiments. [Figure 33] 1 shows predicted and actual labels of cancer signal origin for samples of male and female individuals, according to one embodiment. [Figure 34] 1 shows predicted and actual labels of cancer signal origin for a sample from a male individual, according to one embodiment. [Figure 35] 1 shows predicted and actual labels of cancer signal origin for a sample of a female individual, according to one embodiment. [Figure 36] 10 shows experimental results of fragment coverage versus feature rank, according to one embodiment. [Figure 37A] 1 shows experimental results of fragment coverage versus non-cancer activation fraction, according to one embodiment. [Figure 37B] 10 shows experimental results of non-cancer activation fraction versus feature rank, according to one embodiment. [Figure 38A] 10 shows experimental results of non-cancer activation fraction versus feature weight, according to one embodiment. [Figure 38B] 38B shows another view of the experimental results shown in FIG. 38A, according to one embodiment. [Figure 39A] 1 shows experimental results of activation fraction versus coverage, according to one embodiment. [Figure 39B] 39B shows another view of the experimental results shown in FIG. 39A, according to one embodiment. [Figure 40] 10 shows experimental results of various thresholding methods described herein, according to various embodiments. [Figure 41] 10 shows experimental results of various thresholding methods described herein, according to various embodiments. [Figure 42] 10 shows experimental results of various thresholding methods described herein, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0072] Early detection and classification of cancer is an important technology. Being able to detect cancer before it becomes symptomatic benefits all parties involved, including patients, physicians, and loved ones. For patients, early cancer detection may increase the likelihood of a beneficial outcome. For physicians, early cancer detection may increase the number of treatment pathways that can lead to beneficial outcomes. For loved ones, early cancer detection increases the likelihood that friends and family will not be lost to the disease.

[0073] In recent years, early cancer detection technology has advanced toward analyzing genetic fragments (e.g., DNA) in, for example, a person's blood to determine whether any of those genetic fragments originate from cancer cells. These new technologies allow physicians to identify the presence of cancer in patients that might otherwise go undetected, for example, by traditional screening processes. For example, consider the example of a person at high risk for breast cancer. Traditionally, this person would visit a doctor periodically for mammograms, which create images (e.g., X-rays) of breast tissue that physicians use to identify cancerous tissue. Unfortunately, even with the highest resolution mammograms, physicians can only identify tumors when they are approximately one millimeter in size. This means that the cancer has been present in the person for some time, remaining undiagnosed and untreated. Such visual determination is typical of most cancers, i.e., they are only identifiable after they have grown large enough to be detected by some imaging technology.

[0074] Cancer detection using analysis of patient gene fragments, such as blood, can alleviate this problem. For example, cancer cells will begin to shed DNA fragments into a person's bloodstream as soon as the DNA fragments are formed. This occurs when there are very few cancer cells and before they can be visualized by imaging technology. Therefore, with appropriate methods, a system that analyzes DNA fragments in the bloodstream can identify the presence of cancer in a person based on the shed cancer DNA fragments, and more importantly, it can do so before the cancer is identified using more traditional cancer detection techniques.

[0075] Cancer detection based on the analysis of DNA fragments is made possible by next-generation sequencing ("NGS") technology. NGS is broadly a group of technologies that enable high-throughput sequencing of genetic material. As discussed in more detail herein, NGS mainly consists of (1) sample preparation, (2) DNA sequencing, and (3) data analysis. Sample preparation is the essential laboratory method for preparing DNA fragments for sequencing, sequencing is the process of reading the ordered nucleotides in the sample, and data analysis is the processing and analysis of the genetic information in the sequencing data to identify the presence of cancer.

[0076] While these steps of NGS may help enable early cancer detection, they also introduce their own complex and detrimental problems to cancer detection, so any improvement in sample preparation, DNA sequencing, and / or data analysis (including pre-processing, algorithmic processing, and summarizing or presenting predictions or conclusions) will result in improvements in cancer detection technology, and more generally in early cancer detection.

[0077] To illustrate, (1) problems introduced in sample preparation include DNA sample quality, sample contamination, fragmentation bias, and accurate indexing. Resolving these problems will yield better genetic data for cancer detection.

[0078] Similarly, (2) problems introduced in sequencing include, for example, errors in the precise transcription of fragments (e.g., leading "A" instead of "C"), incorrect or difficult fragment assembly and overlap, different coverage uniformity, sequencing depth vs. cost vs. specificity, and insufficient sequencing length. Again, correcting any of these problems would result in improved genetic data for cancer detection.

[0079] (3) The challenges in data analysis are the most difficult and complex. The challenges posed stem from the sheer volume of data generated by NGS sequencing technology. Sequencing data for a single sample can be on the order of hundreds of thousands (up to millions) of sequence reads, amounting to terabytes of data. Multiply this by the thousands (up to tens of thousands) of samples collected for use in training analytical models. Effectively and efficiently analyzing that amount of data is procedurally and computationally demanding. For example, analyzing NGS sequencing involves several baseline processing steps, such as aligning reads to each other, aligning and mapping reads to a reference genome, de-duplication of duplicate reads, detecting sample contamination, identifying and calling variant genes, identifying and calling aberrantly methylated genes, and generating functional annotations. Performing any of these processes on terabytes of genetic data is computationally expensive even for the most powerful computer architectures and completely impossible for the average human mind. Additionally, genetic sequencing data derived from the error-prone processes of sample preparation and sequence reading may result in a large proportion of the resulting genetic data being of low quality or unusable for cancer identification. For example, large amounts of genetic data may contain contaminated samples, transcription errors, mismatched regions, over-represented regions, etc., and may not be suitable for highly accurate cancer detection. Furthermore, identifying and accounting for low-quality genetic data throughout the vast amount of genetic data obtained from NGS sequencing is procedurally and computationally demanding, and virtually impossible for the human mind to achieve. Overall, any process that results in more efficient processing of large-scale array sequencing data would improve cancer detection using NGS sequencing. Furthermore, such a process is a non-routine and non-traditional activity in a highly technical field, since it was devised as a solution to various challenges encountered in NGS sequencing.

[0080] In particular, (3) accurate identification of information-rich DNA from NGS data to identify the presence of cancer under data analysis is also a challenging issue at hand. To be effective, algorithms are required to compensate for errors generated by, for example, sample preparation and sequencing, and overcome the problems of large-scale data analysis associated with NGS technology. That is, designing machine learning models or other computational algorithms that enable early cancer detection based on next-generation sequencing technologies must be designed to take into account the challenges created by these technologies. Some of these technologies and models are discussed below, and specific improvements to the state-of-the-art technologies and models are further discussed. Furthermore, such technologies are non-routine and non-traditional activities in the technical field under study.

[0081] One particular challenge arises when predicting the presence of cancer signals in individuals with different backgrounds. For example, the likelihood of a biological female developing breast cancer is much higher than the likelihood of a biological male. In another example, the likelihood of a non-smoker developing lung cancer is much lower than the likelihood of a heavy smoker. Accounting for such statistical imbalances can improve predictions, thereby enabling more accurate diagnoses by healthcare providers. The present disclosure aims to address this challenge through one or more different approaches. In some embodiments, the system trains one cancer classifier that applies broadly to all samples, but adjusts the cutoff threshold for positively detecting the presence of one or more disease states for each subpopulation, as defined by one or more covariate characteristics. In other embodiments, the system trains a different classifier for each subpopulation. The training can vary according to the training samples used, the features evaluated, or some combination thereof.

[0082] Training of the machine learning models described herein (e.g., one or more cancer classifiers, tissue models, and other models referred to herein) involves, at least in part, the performance by a machine or computing system of one or more non-mathematical operations or the implementation of non-mathematical functions, including, but not limited to, data loading operations, data storage operations, data toggling or modification operations, non-transitory computer-readable storage medium modification operations, metadata removal or data cleansing operations, data compression operations, modification operations, image modification operations, noise application operations, and noise removal operations. Thus, training of the machine learning models described herein may be based on or include mathematical concepts, but is not simply limited to the performance of mathematical calculations, mathematical operations, or the act of calculating variables or numbers using mathematical methods.

[0083] Similarly, it should be noted that training the models described herein cannot in practice be performed by the human mind alone. Models are inherently complex, including vast amounts of weights and parameters related through one or more complex functions. For example, the number of weights in a model may be on the order of thousands, tens of thousands, hundreds of thousands, millions, or billions. Training may involve utilizing training samples on the order of thousands, tens of thousands, hundreds of thousands, millions, or billions. Training and / or developing such models involves a large number of operations that are infeasible for the human mind alone, even with the aid of pen and paper. In such embodiments, the operations may range from hundreds, thousands, tens of thousands, hundreds of thousands, millions, billions, or trillions. Thus, such models necessarily rely on computer technology for their implementation and use.

[0084] I. Definition Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. As used herein, the following terms have the meanings set forth below.

[0085] The term "individual" can refer to any living organism, such as a human or animal individual. The term "healthy individual" refers to an individual who is presumed to be cancer or disease free.

[0086] The term "subject" refers to an individual whose DNA is to be analyzed. A subject can be a test subject whose DNA is evaluated using whole genome sequencing or a targeted panel described herein to assess whether a person has a disease state (e.g., cancer, a type of cancer, or a tissue of origin of a cancer). A subject can also be part of a control group known to not have cancer or another disease. A subject can also be part of a cancer or other disease group known to have cancer or another disease. Control groups and cancer / disease groups can be used to aid in the design or validation of a targeted panel.

[0087] The term "reference sample" refers to a sample obtained from a subject with a known disease state.

[0088] The term "training sample" refers to samples obtained from known disease states that can be used to generate sequence reads. The training samples can be applied to a probabilistic model to generate features that can be utilized for disease state classification.

[0089] The term "test sample" refers to a sample that may have an unknown disease state.

[0090] The term "sequence read" refers to a nucleotide sequence read from a sample obtained from an individual. A sequence read can be generated from a nucleic acid fragment in a sample. A sequence read can be a collapsed sequence read generated from multiple sequence reads derived from multiple amplicons from a single original nucleic acid molecule. In some embodiments, a sequence read can be a de-duplicated sequence read. A sequence read can be obtained by various methods known in the art.

[0091] The term "disease state" refers to the presence or absence of a disease, a type of disease, and / or a tissue of origin of a disease. For example, in one embodiment, the present disclosure provides a method, system, and non-transitory computer-readable medium for detecting cancer (i.e., the presence or absence of cancer), a type of cancer, or a tissue of origin of a cancer.

[0092] The term "tissue of origin" or "TOO" or "cancer signal origin" refers to the organ, group of organs, body region, or cell type in which a disease state may occur or develop. For example, identifying the tissue of origin or cancer cell type typically allows for the identification of appropriate next steps for further diagnosis, staging, and treatment decisions.

[0093] As used herein, the term "methylation" refers to the chemical process by which a methyl group is added to a DNA molecule. Two of the four bases in DNA, cytosine ("C") and adenine ("A"), can be methylated. For example, a hydrogen atom on the pyrimidine ring of a cytosine base can be converted to a methyl group to form 5-methylcytosine. Methylation tends to occur at the dinucleotide of cytosine and guanine, referred to herein as a "CpG site." In other instances, methylation can occur at a cytosine that is not part of a CpG site or at another nucleotide that is not cytosine. However, these are rarer occurrences. In this disclosure, methylation is discussed with reference to CpG sites for clarity. However, the principles described herein are equally applicable to detecting methylation in non-CpG contexts, including non-cytosine methylation. For example, adenine methylation has been observed in bacterial, plant, and mammalian DNA, but has received little attention.

[0094] In such embodiments, the wet-lab assays used to detect methylation may differ from those described herein as being well known in the art. Furthermore, the methylation status vector may generally contain elements that are vectors of methylated or unmethylated sites (which may not be CpG sites in particular). With such substitutions, the remainder of the processes described herein remain the same, and consequently, the inventive concepts described herein are applicable to these other forms of methylation.

[0095] The term "CpG site" refers to a region of a DNA molecule in which a cytosine nucleotide is followed by a guanine nucleotide in the linear sequence of bases along its 5' to 3' direction. "CpG" is shorthand for 5'-C-phosphate-G-3', a cytosine and guanine separated by only one phosphate group; the phosphate links any two nucleotides together in DNA. The cytosine in a CpG dinucleotide can be methylated to form 5-methylcytosine.

[0096] The term "methylation site" refers to a single site on a DNA molecule where a methyl group can be added. While "CpG" sites are the most common methylation sites, methylation sites are not limited to CpG sites. For example, DNA methylation can occur at cytosine in CHG and CHH, where H is adenine, cytosine, or thymine. Cytosine methylation in the form of 5-hydroxymethylcytosine and its characteristics can also be assessed using the methods and procedures disclosed herein (see, e.g., WO 2010 / 037001 and WO 2011 / 127136, which are incorporated herein by reference). The terms "hypomethylated" or "hypermethylated" refer to the methylation state of a DNA molecule containing multiple CpG sites (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.), in which a high percentage (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50% to 100%) of the CpG sites are unmethylated or methylated, respectively.

[0097] The terms "cell-free deoxyribonucleic acid," "cell-free DNA," or "cfDNA" refer to deoxyribonucleic acid fragments that circulate in bodily fluids, such as blood, sweat, urine, or saliva, and that originate from one or more healthy cells and / or one or more cancer cells.

[0098] The term "circulating tumor DNA" or "ctDNA" refers to deoxyribonucleic acid fragments derived from tumor cells or other types of cancer cells, which may be released into an individual's bodily fluids, such as blood, sweat, urine, or saliva, as a result of biological processes such as apoptosis or necrosis of dying cells, or which may be actively released by viable tumor cells.

[0099] II. Method Overview FIG. 1A is an exemplary flowchart illustrating an overall workflow 100 for cancer classification of a sample, according to one or more embodiments. Workflow 100 may be performed by one or more entities, including, for example, a healthcare provider, a sequencing device, an analytical system, etc. The purpose of the workflow includes detecting and / or monitoring cancer in an individual. From a healthcare perspective, workflow 100 may serve to complement other existing cancer diagnostic tools. Workflow 100 may serve to enable early cancer detection and / or routine cancer surveillance to better inform treatment plans for individuals diagnosed with cancer. Overall workflow 100 may include additional / fewer steps than those shown in FIG. 1 .

[0100] A healthcare provider performs sample collection 110. An individual undergoing cancer classification visits a healthcare provider. The healthcare provider collects a sample for performing cancer classification. Examples of biological samples include, but are not limited to, a subject's tissue biopsy, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or ascites. The sample contains genetic material belonging to the individual, which can be extracted and sequenced for cancer classification. Once the sample is collected, the sample is provided to a sequencing device. Along with the sample, the healthcare provider can collect other information about the individual, such as biological sex, age, racial classification, smoking status, any previous diagnoses, etc.

[0101] The sequencing device performs sample sequencing 120. A laboratory clinician may perform one or more processing steps on the sample in preparation for sequencing. Once prepared, the clinician loads the sample into the sequencing device. An example of a device utilized for sequencing is further described in conjunction with FIGS. 2A and 2B. The sequencing device generally extracts and isolates fragments of nucleic acid to be sequenced and determines the sequence of nucleic acid bases corresponding to the fragments. Sequencing may also include amplification of the nucleic acid material. Various sequencing processes include Sanger sequencing, fragment analysis, and next-generation sequencing. Sequencing may be whole genome sequencing or targeted sequencing using a targeted panel. In the context of DNA methylation, bisulfite sequencing (e.g., as further described in FIGS. 3A and 3B) can determine the methylation status by bisulfite conversion of unmethylated cytosines at CpG sites. Sample sequencing 120 generates sequences for multiple nucleic acid fragments in the sample. In one or more embodiments, the sequence may include methylation state vectors, each vector of methylation states describing the methylation state for a CpG site on the fragment.

[0102] The analytical system performs pre-analytical processing 130. An exemplary analytical system is shown in Figure 2B. Pre-analytical processing 130 may include, but is not limited to, de-duplication of sequence reads, determining coverage metrics, determining whether a sample is contaminated, removing contaminating fragments, calling sequencing errors, etc.

[0103] The analysis system performs one or more analyses 140. An analysis is a statistical analysis or application of one or more trained models to predict at least the cancer status of the individual from whom the sample is derived. Different genetic features, such as methylation of CpG sites, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), and other types of genetic mutations, can be evaluated and considered. In the context of methylation, the analyses 140 can include a tissue purity assessment 142 (further described, for example, in Figures 4A, 4B, 5A, 5B, 6, and 7), feature extraction 144, and applying a cancer classifier 146 to determine a cancer prediction (further described, for example, in Figures 9A and 9B). The tissue purity assessment 142 involves applying a mixture model to deconvolute the proportions of tissue components that contribute DNA fragments to the sample. Generally, the tissue purity assessment 142 can be used to determine what proportion of methylated sequence reads for a sample are derived from cancerous tissue compared to the proportion of methylated sequence reads derived from non-cancerous cells. The tissue purity assessment 142 may be particularly useful for deconvolving heterogeneous tumors containing multiple clonal populations with potentially different genetic signatures. A mixture model may also be applied to quantify the cancer signal or aggressiveness of a cancer state based on the determined proportions. The cancer classifier 146 inputs the extracted features to determine a cancer prediction. The cancer prediction may be a label or a value. The label may indicate a specific cancer state; for example, a binary label may indicate the presence or absence of cancer, while a multi-class label may indicate one or more cancer types, cancer stages, etc., from multiple cancer types being screened for. The value may indicate the likelihood of a specific cancer state, for example, the likelihood of cancer and / or the likelihood of a specific cancer type.

[0104] The analysis system returns a prediction 150 to the healthcare provider. The prediction 150 may include a binary prediction of the presence or absence of cancer, a specific cancer type, cancer stage, tissue proportion, etc. The healthcare provider can establish or adjust a treatment plan based on the returned prediction 150. Optimization of processing is further described in Section VC Processing.

[0105] 1B is a flowchart of a method 160 for identifying a plurality of features for generating a classifier for predicting a disease state (e.g., the presence or absence of a disease, the type of disease, and / or the tissue of origin of the disease), according to various embodiments. In some embodiments, an analysis system 200 performs method 160 to process sequence reads of fragments from a nucleic acid sample. Method 160 includes, but is not limited to, the following steps: generating sequence reads 165; training a probability model 170 associated with each of a plurality of different disease states (e.g., different cancer types); applying the probability model to determine a value 175 based on the probability that the sequence read originated from a sample associated with each of the plurality of disease states associated with each probability model; identifying features 180 by determining a count of sequence reads having a value above a threshold; generating a classifier using the features 185; and, optionally, applying the classifier to predict a disease state and / or a tissue of origin associated with the disease state 190.

[0106] II.A. Overview of Methylation According to the present specification, cfDNA fragments from an individual are processed, for example, by converting unmethylated cytosines to uracils, and sequenced. The sequence reads are compared to a reference genome to identify the methylation status at specific CpG sites within the DNA fragments. Each CpG site may be methylated or unmethylated. Identifying fragments that are more informative than those from healthy individuals can provide insight into the subject's cancer state. In some embodiments, a sequence read or fragment is considered more informative (e.g., more informative about the presence of a disease state) if the sequence read or fragment is determined to be abnormally methylated. As is well known in the art, DNA methylation abnormalities (compared to healthy controls) can cause various effects that may contribute to cancer. Identifying cfDNA fragments that are more informative presents various challenges. First, determining a DNA fragment as more informative may require weighting compared to a group of control individuals, such that the determination becomes less reliable if the number of controls is small due to statistical variation within the smaller control group. In addition, methylation status may vary among groups of control individuals, which may be difficult to account for when determining whether a DNA fragment of interest is information-rich. Alternatively, methylation of cytosines at CpG sites may causally influence methylation at subsequent CpG sites. Encapsulating this dependency may be a challenge in itself.

[0107] Methylation typically occurs in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. In particular, methylation can occur at the dinucleotide of cytosine and guanine, referred to herein as a "CpG site." In other instances, methylation can occur at a cytosine that is not part of a CpG site, or at another nucleotide that is not cytosine. However, these are rarer occurrences. In this disclosure, methylation is discussed with reference to CpG sites for clarity. Informative DNA methylation can be identified as hypermethylated or hypomethylated, both of which can indicate a cancerous state. Throughout this disclosure, hypermethylation and hypomethylation can be characterized for a DNA fragment if the DNA fragment contains multiple thresholds of CpG sites where the threshold percentages are methylated or unmethylated.

[0108] The principles described herein may be equally applicable to detecting methylation in non-CpG contexts, including non-cytosine methylation. In such embodiments, the wet-lab assay used to detect methylation may differ from that described herein. Furthermore, the methylation state vectors discussed herein may generally contain elements that are methylated or unmethylated sites (specifically, may not be CpG sites). With such substitutions, the remainder of the processes described herein may remain the same, and consequently, the inventive concepts described herein may be applicable to these other forms of methylation.

[0109] II.B. Exemplary Sequencer and Analysis System 2A and 2B are flow charts of systems and devices for sequencing nucleic acid samples according to one embodiment. This exemplary flow chart includes devices such as a sequencer 270 and an analysis system 200. The sequencer 270 and the analysis system 200 can function in conjunction to perform one or more steps in the processes described herein.

[0110] In various embodiments, sequencer 270 receives enriched nucleic acid sample 260. As shown in FIG. 2A , sequencer 270 may include a graphical user interface 275 that allows user interaction with specific tasks (e.g., start sequencing or end sequencing), as well as another loading station 280 for loading a sequencing cartridge containing the enriched fragment sample and / or for loading buffers necessary to perform a sequencing assay. Thus, once a user of sequencer 270 provides the necessary reagents and sequencing cartridge to loading station 280 of sequencer 270, the user can initiate sequencing by interacting with graphical user interface 275 of sequencer 270. Once initiated, sequencer 270 performs sequencing and outputs sequence reads of the enriched fragments from nucleic acid sample 260.

[0111] In some embodiments, sequencer 270 is communicatively coupled to analysis system 200. Analysis system 200 includes a number of computing devices used to process sequence reads for various applications, such as assessment of methylation status at one or more CpG sites, variant calling, or quality control. Sequencer 270 can provide sequence reads in BAM file format to analysis system 200. Analysis system 200 can be communicatively coupled to sequencer 270 via wireless, wired, or a combination of wireless and wired communication technologies. Generally, analysis system 200 is comprised of a processor and a non-transitory computer-readable storage medium storing computer instructions that, when executed by the processor, cause the processor to process sequence reads or perform one or more steps of any of the methods or processes disclosed herein.

[0112] In some embodiments, sequence reads can be aligned to a reference genome using methods known in the art to determine alignment position information. Alignment positions generally represent the start and end positions of a region in the reference genome corresponding to the starting and ending nucleotide bases of a given sequence read. Corresponding to methylation sequencing, alignment position information can be generalized to indicate the first and last CpG sites contained within the sequence read according to the alignment to the reference genome. Alignment position information can further indicate the methylation status and location of all CpG sites within a given sequence read. Regions within the reference genome can be associated with genes or gene segments; therefore, analysis system 200 can label sequence reads with one or more genes that align to the sequence read. In one embodiment, fragment length (or size) is determined from the start and end positions.

[0113] In various embodiments, for example, when a paired-end sequencing process is used, the sequence reads are composed of a read pair denoted as R_1 and R_2. For example, the first read R_1 may be sequenced from a first end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 may be sequenced from a second end of the double-stranded DNA (dsDNA). Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 may be aligned consistently (e.g., in opposite directions) with the nucleotide bases of the reference genome. The alignment position information derived from the read pair R_1 and R_2 may include a start position (e.g., R_1) in the reference genome corresponding to the end of the first read and an end position (e.g., R_2) in the reference genome corresponding to the end of the second read. In other words, the start and end positions in the reference genome represent the approximate locations in the reference genome to which the nucleic acid fragments correspond. In one embodiment, the read pair R_1 and R_2 can be assembled into fragments, which are used for subsequent analysis and / or classification. An output file having a SAM (Sequence Alignment Map) format or a BAM (Binary) format can be generated and output for further analysis.

[0114] 2B, which is a block diagram of an analytical system 200 for processing DNA samples according to one embodiment. The analytical system implements one or more computing devices for use in analyzing DNA samples. The analytical system 200 comprises a sequence processor 210, a sequence database 215, a model database 225, one or more probabilistic models 230 and / or one or more classifiers 240, and a parameter database 235. In some embodiments, the analytical system 200 performs one or more steps in a method or process disclosed herein.

[0115] The sequence processor 210 generates methylation state vectors for fragments from the sample. At each CpG site on the fragment, the sequence processor 210 generates a methylation state vector for each fragment, via process 360 of FIG. 4B, that specifies the location of the fragment in the reference genome, the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment, whether methylated, unmethylated, or uncertain. The sequence processor 210 may store the methylation state vectors for the fragments in a sequence database 215. The data in the sequence database 215 may be organized so that the methylation state vectors from the sample are related to each other.

[0116] Additionally, multiple different models 230 can be stored in model database 225 or retrieved for use with a test sample. In one example, a model is a trained cancer classifier 240 that uses feature vectors derived from informative fragments to determine a cancer prediction for a test sample. The training and use of cancer classifiers is discussed elsewhere herein. Analysis system 200 can train one or more models 230 and / or one or more classifiers 240 and store the various trained parameters in parameter database 235. Analysis system 200 stores models 230 and / or classifiers along with their functions in model database 225.

[0117] During inference, the machine learning engine 220 uses one or more models 230 and / or classifiers 240 to return outputs. The machine learning engine accesses the models 230 and / or classifiers 240 in the model database 225 along with trained parameters from the parameter database 235. According to each model, the machine learning engine 220 receives inputs appropriate for the model and calculates outputs based on the received inputs, parameters, and each model's functions relating to the inputs and outputs. In some use cases, the machine learning engine 220 also calculates metrics that correlate to the confidence of the calculated outputs from the models. In other use cases, the machine learning engine 220 calculates other intermediate values ​​for use with the models.

[0118] II.C. Assay Protocol 3 is a flowchart illustrating a process 300 for sequencing nucleic acids, according to one embodiment. In some embodiments, process 300 is performed to generate sequence reads as part of step 110 of method 100 of FIG.

[0119] In step 310, a nucleic acid sample (e.g., DNA or RNA) is extracted from a subject. In this disclosure, DNA and RNA can be used interchangeably unless otherwise indicated. That is, the embodiments described herein may be applicable to both DNA and RNA types of nucleic acid sequences. However, the examples described herein may be clarified and explained with a focus on DNA. The sample may include nucleic acid molecules derived from any subset of the human genome, including the entire genome. The sample may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. In some embodiments, methods of collecting a blood sample (e.g., syringe or finger prick) may be less invasive than procedures for obtaining a tissue biopsy, which may require surgery. The extracted sample may include cfDNA and / or ctDNA. If the subject has a disease condition such as cancer, cell-free nucleic acids (e.g., cfDNA) in a sample extracted from the subject generally contain detectable levels of nucleic acids that can be used to assess the disease state.

[0120] In step 315, the extracted nucleic acids (e.g., including cfDNA fragments) are treated to convert unmethylated cytosines to uracil. In some embodiments, method 300 uses bisulfite treatment of the sample to convert unmethylated cytosines to uracil without converting methylated cytosines. For example, a commercially available kit such as EZ DNA Methylation™-Gold, EZ DNA Methylation™-Direct, or EZ DNA Methylation™-Lightning kit (available from Zymo Research Corp, Irvine, CA) is used for bisulfite conversion. In another embodiment, the conversion of unmethylated cytosines to uracil is achieved using an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for converting unmethylated cytosines to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).

[0121] In step 320, a sequencing library is prepared. In some embodiments, the preparation includes at least two steps. In the first step, an ssDNA ligation reaction is used to add ssDNA adapters to the 3'-OH ends of bisulfite-converted ssDNA molecules. In some embodiments, the ssDNA ligation reaction uses CircLigase II (Epicentre) to ligate ssDNA adapters to the 3'-OH ends of bisulfite-converted ssDNA molecules, where the 5' ends of the adapters are phosphorylated and the bisulfite-converted ssDNA is dephosphorylated (i.e., the 3' ends have a hydroxyl group). In another embodiment, the ssDNA ligation reaction uses Thermostable 5' AppDNA / RNA Ligase (available from New England BioLabs, Ipswich, MA) to ligate ssDNA adapters to the 3'-OH ends of bisulfite-converted ssDNA molecules. In this example, the first UMI adaptor is adenylated at the 5' end and blocked at the 3' end. In another embodiment, the ssDNA ligation reaction uses T4 RNA ligase (available from New England BioLabs) to ligate an ssDNA adaptor to the 3'-OH end of the bisulfite converted ssDNA molecule.

[0122] In the second step, a second strand DNA is synthesized in an extension reaction. For example, an extension primer that hybridizes to the primer sequence contained in the ssDNA adapter is used in a primer extension reaction to form a double-stranded bisulfite-converted DNA molecule. Optionally, in some embodiments, the extension reaction uses an enzyme that can lead through the uracil residue in the bisulfite-converted template strand.

[0123] Optionally, in the third step, dsDNA adapters are added to the double-stranded bisulfite-converted DNA molecules. The double-stranded bisulfite-converted DNA can then be amplified to add sequencing adapters. For example, PCR amplification using a forward primer containing a P5 sequence and a reverse primer containing a P7 sequence is used to add P5 and P7 sequences to the bisulfite-converted DNA. Optionally, during library preparation, unique molecular identifiers (UMIs) can be added to nucleic acid molecules (e.g., DNA molecules) by adapter ligation. UMIs are short nucleic acid sequences (e.g., 4-10 base pairs) that are added to the ends of DNA fragments during adapter ligation. In some embodiments, UMIs are degenerate base pairs that function as unique tags that can be used to identify sequence reads derived from specific DNA fragments. During PCR amplification after adapter ligation, the UMIs are replicated along with the attached DNA fragments, thereby providing a method for identifying sequence reads derived from the same original fragment in downstream analysis.

[0124] In optional step 325, nucleic acids (e.g., fragments) can be hybridized. Hybridization probes (also referred to herein as "probes") can be used to target and pull down nucleic acid fragments that are informative about a disease state. For a given workflow, probes can be designed to anneal (or hybridize) to a target (complementary) strand of DNA or RNA. The target strand can be the "plus" strand (e.g., the strand transcribed into mRNA and then translated into protein) or the complementary "minus" strand. Probes can span lengths of 10s, 100s, or 1000s of base pairs. Additionally, probes can cover overlapping portions of the target region.

[0125] In optional step 330, the hybridized nucleic acid fragments can be captured and enriched, e.g., amplified, using PCR. In some embodiments, target DNA sequences can be enriched from the library. This is used, for example, when a target panel assay is performed on the sample. For example, the target sequences can be enriched to obtain enriched sequences that can then be sequenced. Generally, any method known in the art can be used to isolate and enrich the target nucleic acid to which the probe is hybridized. For example, as is well known in the art, a biotin moiety can be added to the 5' end of the probe (i.e., biotinylation), and a streptavidin-coated surface (e.g., streptavidin-coated beads) can be used to facilitate isolation of the target nucleic acid hybridized to the probe.

[0126] In step 335, sequence reads, e.g., enriched sequences, are generated from the nucleic acid sample. Sequencing data can be obtained from the enriched DNA sequences by means known in the art. For example, the method can include next-generation sequencing (NGS) techniques, such as synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing by synthesis with reversible dye terminators.

[0127] In step 340, the sequence processor 210 can generate methylation information using the sequence reads. The methylation information determined from the sequence reads can then be used to generate a methylation state vector. Figure 4B is an illustration of a process 360, starting from process 300 of Figure 3, of sequencing a cfDNA molecule to obtain a methylation state vector 352, according to one embodiment. As an example, an analysis system receives a cfDNA molecule 312 (which contains three CpG sites in this example). As shown, the first and third CpG sites of the cfDNA molecule 312 are methylated 314. During processing step 315, the cfDNA molecule 312 is converted to generate a converted cfDNA molecule 322. During processing 315, the second CpG site, which was previously unmethylated, has its cytosine converted to uracil. However, the first and third CpG sites were not converted.

[0128] After conversion, the sequencing library 330 is prepared and sequenced to generate sequence reads 342. The analysis system aligns the sequence reads 342 to a reference genome 344 (not shown), which provides context for where in the human genome the fragment cfDNA originates. In this simplified example, the analysis system aligns the sequence reads 342 so that three CpG sites correlate to CpG sites 23, 24, and 25 (arbitrary reference identifiers used for convenience of illustration). Thus, the analysis system generates information about both the methylation status of all CpG sites on the cfDNA molecule 312 and the location in the human genome to which the CpG sites map. As shown, the methylated CpG sites on the sequence read 342 are read as cytosines. In this example, cytosines appear only in the first and third CpG sites in the sequence read 342, and it can be inferred that the first and third CpG sites in the original cfDNA molecule were methylated. On the other hand, since the second CpG site is read as a thymine (U is converted to T during the sequencing process), it can be inferred that the second CpG site was unmethylated in the original cfDNA molecule. With these two pieces of information, methylation state and location, the analysis system generates 200 a methylation state vector 352 for the fragment cfDNA 312. In this example, the resulting methylation state vector 352 is <M 23 , U 24 , M 25 >, where M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript numbers correspond to the position of each CpG site in the reference genome.

[0129] II.D. Identifying Information-Dense Fragments In some embodiments, the analysis system uses the methylation state vector of the sample to determine informative fragments for the sample. For example, for each nucleic acid molecule or fragment in the sample, the analysis system determines (through analysis of sequence reads derived therefrom) whether the nucleic acid molecule or fragment is an informative molecule or fragment relative to an expected methylation state vector from a healthy sample using the methylation state vector corresponding to the nucleic acid molecule. In one embodiment, the analysis system calculates a p-value score for each methylation state vector that describes the probability of observing that methylation state vector or another methylation state vector that is less likely in a healthy control group (e.g., as described in U.S. Patent Application Publication No. 2019 / 0287652, incorporated herein by reference). The process of calculating p-value scores is also described in Section II. BiP Value Filtering below. The analysis system can determine, and optionally exclude, sequence reads for nucleic acid molecules or fragments having methylation state vectors with p-value scores below a threshold as informative fragments. In another embodiment, the analysis system further labels fragments having at least some CpG sites above some threshold percentage of methylation or unmethylation as hypermethylated and hypomethylated fragments, respectively. Hypermethylated and hypomethylated fragments may also be referred to as unusual fragments with extreme methylation (UFXM). In other embodiments, the analysis system may implement various other probabilistic models for determining informative molecules or fragments. Examples of other probabilistic models include mixture models, deep probabilistic models, etc. In some embodiments, the analysis system may use any combination of the processes described below to identify informative fragments. The identified informative fragments may enable the analysis system to filter the set of methylation state vectors of samples for use in other processes, such as training and developing a cancer classifier.

[0130] II.DIP Value Filtering In one embodiment, the analysis system calculates a p-value score for each methylation state vector compared to methylation state vectors from fragments in a healthy control group. The p-value score represents the probability of observing a nucleic acid molecule whose methylation state matches the methylation state vector in the healthy control group. To determine informative DNA fragments, the analysis system uses a healthy control group in which the majority of fragments are normally methylated. When performing this probability analysis to determine informative fragments, the determination carries weight relative to the group of control subjects that make up the healthy control group. To ensure robustness in the healthy control group, the analysis system may select a threshold number of healthy individuals to provide samples containing DNA fragments. Figure 4B below illustrates how the analysis system generates a data structure for the healthy control group from which the p-value score can be calculated. Figure 4C illustrates how the p-value score is calculated using the generated data structure.

[0131] 4B is a flowchart illustrating a process 400 for generating a data structure for a healthy control group, according to one embodiment. To create the healthy control group data structure, the analysis system receives multiple DNA fragments (e.g., cfDNA) from multiple healthy individuals. A methylation state vector is determined for each fragment, e.g., via process 360.

[0132] Using the methylation state vector of each fragment, the analysis system subdivides the methylation state vector into strings of CpG sites 405. In one embodiment, the analysis system subdivides the methylation state vector 405 so that the resulting strings are all less than a given length. For example, a methylation state vector of length 11 may be subdivided into strings of length 3 or less, resulting in 9 length 3 strings, 10 length 2 strings, and 11 length 1 strings. In another example, a methylation state vector of length 7 may be subdivided into strings of length 4 or less, resulting in 4 length 4 strings, 5 length 3 strings, 6 length 2 strings, and 7 length 1 strings. If the methylation state vector is shorter than or equal to the specified string length, the methylation state vector may be converted into a single string containing all of the CpG sites in the vector.

[0133] The analysis system 200 tallies 410 strings by counting, for each possible CpG site and methylation state possibility in the vector, the number of strings present in the control set that have the specified CpG site as the first CpG site in the string and have that methylation state possibility. For example, at a given CpG site, considering a string length of 3, there are 2^3 or 8 possible string configurations. At that given CpG site, for each of the 8 possible string configurations, the analysis system tallies 410 the number of occurrences of each methylation state vector possibility in the control set. Continuing with this example, this may involve tallying the following quantities: for each starting CpG site x in the reference genome, <M x , M x+1 , M x+2 >, <M x , M x+1 , U x+2 >,... x , U x+1 , U x+2 The analysis system creates a data structure 415 that stores the counts tallied for each starting CpG site and string possibility.

[0134] ​Setting an upper limit on string length has several advantages. First, depending on the maximum string length, the size of the data structure created by the analysis system can increase dramatically. For example, a maximum string length of 4 means that all CpG sites must count at least 2^4 for strings of length 4. Increasing the maximum string length to 5 means that all CpG sites must count an additional 2^4 or 16, doubling the number of counts (and the computer memory required) compared to the previous string length. Reducing the string size helps keep the creation and performance of the data structure (e.g., for use in later access, as described below) reasonable from a computational and storage standpoint. Second, a statistical consideration for limiting the maximum string length is to avoid overfitting of downstream models that use string counts. If long strings of CpG sites do not biologically have a strong impact on outcomes (e.g., anomaly predictions that predict the presence of cancer), calculating probabilities based on large strings of CpG sites can be problematic because it requires a significant amount of data that may not be available, making the model too sparse for proper performance. For example, calculating the probability of an anomaly / cancer conditional on the previous 100 CpG sites requires counts of strings in a data structure of length 100, ideally some of which exactly match the methylation states of the previous 100. If only sparse counts of strings of length 100 are available, there is insufficient data to determine whether a given string of length 100 in a test sample is informative.

[0135] 4C is a flowchart illustrating a process 420 for identifying informative fragments from an individual, according to one embodiment. In process 420, the analysis system generates methylation state vectors 352 from the subject's cfDNA fragments. The analysis system treats each methylation state vector as follows:

[0136] For a given methylation state vector, the analysis system enumerates all possible methylation state vectors with the same starting CpG site and length (i.e., set of CpG sites) in the methylation state vector 430. Since each methylation state is generally either methylated or unmethylated, and each CpG site effectively has two possible states, a methylation state vector of length n is generated from the 2 nd of the methylation state vector. n As related to the likelihood of , the count of distinct possibilities for a methylation state vector depends on a power of 2. A methylation state vector that includes an uncertain state for one or more CpG sites allows the analysis system to enumerate the possibilities for the methylation state vector by considering only the CpG sites for which the state was observed 430.

[0137] Analysis system 200 calculates 440 the probability of observing each possible methylation state vector for the identified starting CpG site and methylation state vector length by accessing the healthy control data structure. In one embodiment, calculating the probability of observing a given probability uses Markov chain probability to model the joint probability calculation. In other embodiments, calculation methods other than Markov chain probability are used to determine the probability of observing each possible methylation state vector.

[0138] The analysis system 450 uses the calculated probabilities for each possibility to calculate a p-value score for the methylation state vector. In one embodiment, this involves identifying a calculated probability that corresponds to the likelihood of a match with the methylation state vector in question. Specifically, this is the likelihood of having the same set of CpG sites, or equivalently, the same starting CpG site and length as the methylation state vector. The analysis system sums the calculated probabilities of all possibilities whose probability is less than or equal to the identified probability to generate a p-value score.

[0139] This p-value represents the probability that the methylation state vector of the fragment or other methylation state vectors would be observed to be less likely in a healthy control group. Thus, a low p-value score generally corresponds to a methylation state vector that is rare in healthy individuals and labels the fragment as more informative compared to the healthy control group. A high p-value score generally relates to a methylation state vector that is expected to be present in a relative sense in healthy individuals. If the healthy control group is a non-cancerous group, for example, a low p-value indicates that the fragment is more informative compared to the non-cancerous group, and therefore likely indicates the presence of cancer in the subject.

[0140] As described above, the analysis system calculates a p-value score for each of a plurality of methylation state vectors, each representing a cfDNA fragment in the test sample. To identify which fragments are informative, the analysis system can filter the set of methylation state vectors based on their p-value scores 460. In one embodiment, filtering is performed by comparing the p-value scores to a threshold and keeping only fragments below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001, or 0.0001, etc.

[0141] According to exemplary results from this process, the analysis system yields a median (range) of 2,800 (1,500-12,000) fragments with informative methylation patterns for participants without cancer during training, and a median (range) of 3,000 (1,200-220,000) fragments with informative methylation patterns for participants with cancer during training. These filtered sets of fragments with informative methylation patterns can be used for downstream analysis, as described below.

[0142] In one embodiment, the analysis system uses a sliding window to determine the likelihood of the methylation state vector and calculate a p-value 455. Rather than enumerating the possibilities and calculating a p-value for the entire methylation state vector, the analysis system enumerates the possibilities and calculates a p-value only for a window of contiguous CpG sites, where the window is shorter in length (of CpG sites) than at least some of the fragment (otherwise the window would serve no purpose). The length of the window may be static, user-determined, dynamic, or otherwise selected.

[0143] When calculating p-values ​​for a methylation state vector larger than a window, the window identifies a contiguous set of CpG sites from the vector within the window, starting at the first CpG site in the vector. The analysis system calculates a p-value score for the window containing the first CpG site. The analysis system then "slides" the window to a second CpG site in the vector and calculates another p-value score for the second window. Thus, for a window size l and a methylation vector length m, each methylation state vector will generate m-l+1 p-value scores. After the p-value calculations for each portion of the vector are completed, the lowest p-value score from all sliding windows is considered the overall p-value score for the methylation state vector. In another embodiment, the analysis system aggregates the p-value scores for the methylation state vector to generate an overall p-value score.

[0144] Using a sliding window helps reduce the number of enumerated possibilities for methylation state vectors and their corresponding probability calculations that would otherwise need to be performed. To take a realistic example, consider a fragment with at least 54 CpG sites. Instead of calculating probabilities for 2^54 (approximately 1.8 x 10^16) possibilities to generate a single p-score, the analysis system can use a window of size 5 (for example), resulting in 50 p-value calculations for each of the 50 windows of methylation state vectors for the fragment. Each of the 50 calculations enumerates 2^5 (32) possibilities for the methylation state vector, totaling 50 x 2^5 (1.6 x 10^3) ​​probability calculations. This significantly reduces the number of calculations that would be performed without meaningful hits to accurately identify the informative fragment.

[0145] In embodiments where there is an uncertain state, the analysis system may calculate a p-value score that sums the CpG sites in the methylation state vector of the fragment with an uncertain state. The analysis system identifies all possibilities that have a consensus with all methylation states in the methylation state vector excluding the uncertain states. The analysis system may assign a probability to the methylation state vector as the sum of the probabilities of the identified possibilities. As an example, since the methylation states of CpG sites 1 and 3 are observed and are in consensus with the methylation state of the fragment at CpG sites 1 and 3, the analysis system may<M1、I2、U3> The probability of the methylation state vector of<M1、M2、U3> and<M1、U2、U3> The probability of a methylation state vector having one or more uncertain states is calculated as the sum of probabilities for the possible methylation state vectors of the CpG sites. This method of summing up CpG sites with uncertain states uses a calculation of probabilities for up to 2^i possibilities, where i represents the number of uncertain states in the methylation state vector. In additional embodiments, a dynamic programming algorithm can be implemented to calculate the probability of a methylation state vector having one or more uncertain states. Advantageously, the dynamic programming algorithm operates in linear computation time.

[0146] In one embodiment, the computational burden of calculating probabilities and / or p-value scores can be further reduced by caching at least some of the calculations. For example, the analysis system can cache in temporary or persistent memory calculations of probabilities for the likelihood of a methylation state vector (or a window thereof). If other fragments have the same CpG sites, caching the likelihood probabilities allows for efficient calculation of p-score values ​​without the need to recalculate the underlying likelihood probabilities. Equivalently, the analysis system can calculate a p-value score for each possibility of a methylation state vector associated with a set of CpG sites from the vector (or a window thereof). The analysis system can cache the p-value scores used to determine p-value scores for other fragments containing the same CpG sites. In general, the p-value scores of possibilities for methylation state vectors with the same CpG sites can be used to determine a different p-value score for a possibility from the same set of CpG sites.

[0147] II.D.II. Hypermethylated and Hypomethylated Fragments In some embodiments, the analysis system determines informative fragments as those having a number of CpG sites above a threshold and either a percentage of CpG sites that are methylated above a threshold or a percentage of CpG sites that are unmethylated above a threshold; the analysis system identifies such fragments as hypermethylated or hypomethylated fragments. Exemplary thresholds for fragment (or CpG site) length include values ​​above 3, 4, 5, 6, 7, 8, 9, 10, etc. Exemplary percentage thresholds for methylation or unmethylation include above 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%.

[0148] II.E. Reference Genome Blocks Figure 5 illustrates blocks of a reference genome according to one embodiment. The sequence processor 210 can partition the reference genome (or a subset of the reference genome) in one or more stages, e.g., for use cases involving targeted methylation assays. For example, the sequence processor 210 separates the reference genome into blocks of CpG sites. Each block is defined when there is a separation between two adjacent CpG sites that exceeds a threshold, e.g., greater than 200 base pairs (bp), 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1,000 bp, among other values. Thus, blocks can vary in base pair size. For each block, the sequence processor 210 can subdivide the block into windows of a particular length, e.g., 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1,100 bp, 1,200 bp, 1,300 bp, 1,400 bp, or 1,500 bp, among other values. In other embodiments, the windows can be 200 bp to 10 kilobase pairs (kbp), 500 bp to 2 kbp, or about 1 kbp long. (e.g., adjacent) windows can overlap by a percentage in number of base pairs or length, e.g., 10%, 20%, 30%, 40%, 50%, or 60%, among other values. The window can be separated between two adjacent CpG sites by more than a threshold value, for example, more than 200 base pairs (bp), 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1,000 bp, among other values.

[0149] The sequence processor 210 can analyze sequence reads derived from DNA fragments using a windowing process. Specifically, the sequence processor 210 scans blocks by window and reads fragments within each window. The fragments can be derived from tissues and / or high-signal cfDNA. High-signal cfDNA samples can be determined by a binary classification model, cancer stage, or another metric. By dividing the reference genome (e.g., using blocks and windows), the sequence processor 210 can facilitate computational parallelization. Furthermore, the sequence processor 210 can reduce computational resources for processing the reference genome by targeting sections of base pairs containing CpG sites while skipping other sections that do not contain CpG sites.

[0150] III. Model-Based Feature Engineering and Classification III.A. Model-Based Feature Engineering According to one embodiment, as shown in Figure 8, the present disclosure relates to model-based feature engineering to derive features useful for disease state classification. As described elsewhere herein, the disease state can be the presence or absence of a disease, the type of disease, and / or the tissue or origin of the disease. For example, as described herein, the disease state can be the presence or absence of cancer, the type of cancer, and / or the tissue of origin of the cancer. The type of cancer and / or tissue of origin of the cancer may be selected from the group comprising breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis, non-urothelial kidney cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from hepatocytes, hepatobiliary carcinoma arising from cells other than hepatocytes, pancreatic cancer, squamous cell carcinoma of the upper gastrointestinal tract, upper gastrointestinal carcinoma other than squamous cell carcinoma, head and neck cancer, lung cancer, e.g. lung adenocarcinoma, small cell lung carcinoma, squamous cell carcinoma and carcinoma other than adenocarcinoma or small cell lung carcinoma, neuroendocrine carcinoma, melanoma, thyroid carcinoma, sarcoma, multiple myeloma, lymphoma, and leukemia, among other cancer types.

[0151] In step 810, a first plurality of sequence reads is generated from a first reference sample having a first disease state, as described elsewhere herein, and a second plurality of sequence reads is generated from a second reference sample having a second disease state. The first plurality of sequence reads and / or the second plurality of sequence reads can be more than 10,000, more than 50,000, more than 100,000, more than 200,000, more than 500,000, more than 1,000,000, more than 2,000,000, more than 5,000,000, or more than 10,000,000 sequence reads. As used herein, a "reference sample" is a sample obtained from a subject with a known disease state.

[0152] In some embodiments, one or more reference samples with one or more known disease states can be used to train one or more probabilistic models, which in turn can be used to derive features for classifying the disease state of an unknown test sample.

[0153] The sample may be a genomic DNA (gDNA) sample or a cell-free DNA (cfDNA) sample. The reference sample may be a blood, plasma, serum, urine, feces, or saliva sample. Alternatively, the reference sample may be whole blood, a blood fraction, a tissue biopsy, a pleural effusion, a pericardial effusion, a cerebrospinal fluid, or an ascites fluid. In some embodiments, a first reference sample is obtained from a subject known to have cancer, and a second reference sample is obtained from a healthy or non-cancer subject. In some embodiments, a first reference sample is obtained from a subject known to have a first type of cancer (e.g., lung cancer), and a second reference sample is obtained from a subject known to have a second type of cancer (e.g., breast cancer). In yet other embodiments, a first reference sample is obtained from a subject known to have a tissue of origin of a first disease (e.g., lung disease), and a second reference sample is obtained from a tissue of origin of a second disease state (e.g., liver disease).

[0154] In step 815, the machine learning engine 220 trains a first probabilistic model 230 and a second probabilistic model 230 from the first and second plurality of sequence reads (generated in step 110), respectively, with each probabilistic model associated with a different one or more possible disease states. As described above, the disease state may be the presence or absence of cancer, the type of cancer, and / or the tissue of origin of the cancer. In various embodiments, the training data is divided into K subsets (convolutions) for K-fold cross-validation. The convolutions may be balanced for cancer / non-cancerous status, tissue of origin, cancer stage, age (e.g., grouped into 10-year buckets), gender, racial classification, and smoking status, among other factors. Data from K-1 of the convolutions may be used as training data for the probabilistic models, and held-out convolutions may be used as test data.

[0155] The machine learning engine 220 trains first and second probabilistic models 230 for first and second disease states, respectively, by fitting each of the probabilistic models 230 to a first plurality of sequence reads and a second plurality of sequence reads. For example, in one embodiment, the first probabilistic model is trained using a first plurality of sequence reads derived from one or more samples from subjects known to have cancer, and the second probabilistic model is trained using a second plurality of sequence reads derived from one or more samples from healthy or non-cancer subjects. In other embodiments, the first probabilistic model can be trained for a first type of cancer or a first tissue of origin, and the second probabilistic model can be trained for a second type of cancer or a second tissue of origin. As one skilled in the art will appreciate, any number of disease state probabilistic models can be trained using sequence reads derived from one or more samples taken from subjects with any one of several possible disease states. For example, in some embodiments, as described elsewhere herein, additional cancer-specific probability models (i.e., for additional cancer types and / or tissues of origin) can be trained for a third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, etc. (e.g., up to 20, 30, or more) specific cancer types and used to determine the probability that sequence reads from the training set or unknown cancer types are more likely to originate from one cancer type (or cancer tissue of origin) than another cancer type (or cancer tissue of origin).

[0156] As used herein, a "probability model" is any mathematical model that can assign a probability to a sequence read based on the methylation state at one or more sites on the read. During training, the machine learning engine 220 can be used to match sequence reads from one or more samples from subjects with known diseases and utilize methylation information or methylation state vectors (e.g., as described above with respect to Figures 3-4) to determine the sequence read probability indicative of the disease state. In particular, in one embodiment, the machine learning engine 220 determines the observed methylation rate for each CpG site within the sequence read. The methylation rate represents the fraction or percentage of base pairs that are methylated within the CpG site. The trained probability model 230 can be parameterized by a product of methylation rates. In general, any known probability model for assigning probabilities to sequence reads from a sample can be used. For example, the probability model can be a binomial model in which every site (e.g., CpG site) on a nucleic acid fragment is assigned a probability of methylation, or an independent site model in which methylation of each CpG is specified by a separate methylation probability, with methylation at one site assumed to be independent of methylation at one or more other sites on the nucleic acid fragment.

[0157] In some embodiments, the probabilistic model 230 is a Markov model, where the probability of methylation at each CpG site depends on the methylation state at several preceding CpG sites in the sequence read or in the nucleic acid molecule from which the sequence read is derived (see, e.g., U.S. Patent Application Serial No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification," filed March 13, 2019).

[0158] In some embodiments, the probabilistic model 230 is a "mixture model" fitted using a mixture of components from the underlying models. For example, in some embodiments, the mixture components can be determined using multiple independent site models, where methylation (e.g., methylation rate) at each CpG site is assumed to be independent of methylation at other CpG sites. Using the independent site model, the probability assigned to a sequence read, or the nucleic acid molecule from which it is derived, is the product of the methylation probability at each CpG site where the sequence read is methylated minus one for each CpG site where the sequence read is unmethylated. According to this embodiment, the machine learning engine 220 determines the methylation rate for each mixture component. The mixture model is parameterized by the sum of the product of the methylation rates and their associated mixture components. A probabilistic model Pr of n mixture components can be expressed as follows:

number

number

[0159] In some embodiments, the machine learning engine 220 fits the probabilistic model 230 using maximum likelihood estimation to find a set of parameters {β ki , f kThe maximum amount for the N total fragments can be expressed as:

number

[0160] As those skilled in the art will appreciate, other means can be used to fit a probability model or identify parameters that maximize the log-likelihood of all sequence reads derived from a reference sample. For example, in one embodiment, Bayesian fitting (e.g., using Markov chain Monte Carlo) is used, in which each parameter is not assigned a single value but is instead associated with a distribution. In another embodiment, gradient-based optimization is used, in which the gradient of the likelihood (or log-likelihood) versus parameter value is used to step through the parameter space toward an optimal value. In another embodiment, expectation maximization is used, in which a set of latent parameters (such as the identity of the mixture component from which each fragment originates) is set to its expected value under the previous model parameters, and then the model parameters are assigned to maximize the likelihood conditional on the assumed values ​​of the latent variables. The two-step process is then repeated until convergence is achieved.

[0161] In step 820, a plurality of training sequence reads is generated from a training sample. The plurality of training sequence reads may be greater than 10,000, greater than 50,000, greater than 100,000, greater than 200,000, greater than 500,000, greater than 1,000,000, greater than 2,000,000, greater than 5,000,000, or greater than 10,000,000 sequence reads. As used herein, a "training sample" is a sample obtained from a known disease state that can be used to generate sequence reads that are applied to the first and / or second probability models to generate features that can be used for disease state classification. In step 825, analysis system 200 applies first and second probability models 230 to determine first and second probability values ​​for each sequence read of the plurality of training sequence reads. The first and second probability values ​​are determined based on the probability that the sequence read is derived from a sample associated with the first and second disease states, respectively. Analysis system 200 can repeat step 130 for any additional probability models 230 (eg, trained from sequence reads from third, fourth, fifth, etc. reference samples) (not shown).

[0162] In step 830, one or more features are identified by comparing the first probability value and the second probability value for each of the plurality of training sequence reads. Generally, a variety of methods can be used to compare the first and second probability values ​​to identify features. For example, in one embodiment, the one or more features include a count of outlier sequence reads among the plurality of training sequence reads for which the first probability value is greater than the second probability value. The count can be a binary count, a total count of outlier sequence reads, or a total count of anonymously methylated sequence reads. In another embodiment, the one or more features include a count of sequence reads or fragments containing a particular methylation pattern. For example, the one or more features can be a count of sequence reads or fragments that are fully methylated at each CpG site, or a count of sequence reads or fragments that are partially methylated (e.g., at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% methylated). In another embodiment, the one or more features are identified using the output of a discriminative classifier trained within the single genomic region (e.g., the discriminative classifier can be a multilayer perceptron or a convolutional neural net model). In another embodiment, comparing the first probability value to the second probability value includes determining a ratio between the first probability value and the second probability value, and the one or more features include sequence read counts of sequence reads that exceed a ratio threshold.

[0163] In another embodiment, the first probability value or the second probability value is a log-likelihood value. For example, analysis system 200 can calculate a log-likelihood ratio R using fitted probability models associated with the first and second disease states, respectively. Specifically, the log-likelihood ratio can be calculated using the probability Pr of observing a methylation pattern on a fragment for a sample associated with the first disease state and the second disease state.

number

[0164] Analysis system 200 can identify features using multiple tiers of thresholds. For example, tiers include thresholds of 1, 2, 3, 4, 5, 6, 7, 8, and 9. In some embodiments, a smoothing function can be applied. For example, in response to determining that R is (e.g., significantly) less than a tier value, analysis system 200 assigns a feature value of approximately 0; in response to determining that R is equal to a tier value, analysis system 200 assigns a feature value of 0.5; and in response to determining that R is (e.g., significantly) greater than a tier value, analysis system 200 assigns a feature value of approximately 1. Each tier indicates a varying threshold above which a fragment (from which a sequence read was generated) is more likely to have originated from a sample associated with a disease state than from a healthy sample. Analysis system 200 can use the thresholds to determine a count of outlier fragments that can be used as features.

[0165] By filtering by a threshold, analysis system 200 can consider certain fragments to be outliers because they are unlikely to be present in a healthy sample. Therefore, outlier fragments can be considered more likely to be associated with (e.g., originate from) a disease state or cancer sample. The number of features can vary between different tiers; for example, one tier may have a different number of features than another tier based on the corresponding threshold. In other embodiments, analysis system 200 uses a different number of tiers or other thresholds. Other means for identifying features or ranking identified features based on a measure of their ability to distinguish between different disease states (e.g., using mutual information to determine a measure of the information content of a feature in distinguishing between two disease states) are described elsewhere herein.

[0166] In other embodiments, analysis system 200 can identify multiple features using different types of ratios or formulas. Machine learning engine 220 can determine a segment indicative of a disease state (e.g., cancer) based on whether at least one of the log-likelihood ratios considered for various disease states exceeds a threshold.

[0167] The plurality of features can then be used to train a disease state classifier, as described in more detail elsewhere herein, for example, in some embodiments, the plurality of features can be used to train a classifier for classification of the presence or absence of cancer, the type of cancer, and / or the tissue of origin of the cancer.

[0168] III.B. Classification of Tissue of Origin of Disease States According to another embodiment, the machine learning engine 220 trains probabilistic models 230 each associated with a different disease state of a set of disease states.

[0169] The machine learning engine 220 uses one or more sets of sequence reads generated from different disease states of the set of multiple disease states to train the probabilistic model 230. The disease states may include any number of cancer types or tissues of origin selected from the group including breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis, non-urothelial kidney cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from liver cells, hepatobiliary carcinoma arising from cells other than liver cells, pancreatic cancer, squamous cell carcinoma of the upper gastrointestinal tract, upper gastrointestinal carcinoma other than squamous cell carcinoma, head and neck cancer, lung cancer, e.g., lung adenocarcinoma, small cell lung cancer, squamous cell carcinoma and carcinoma other than adenocarcinoma or small cell lung cancer, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia, among other cancer types.

[0170] The machine learning engine 220 trains the probabilistic model 230 for each of a plurality of disease states by fitting the probabilistic model 230 to sequence reads derived from each sample corresponding to each disease state. For example, in some embodiments, the probabilistic model can be trained for a specific cancer type. According to this embodiment, a cancer-specific probabilistic model can be trained for a first, second, third, etc. specific cancer type and used to assess the cancer type (e.g., of an unknown test sample). For example, a set of sequence reads derived from one or more samples associated with lung cancer can be used to fit a lung cancer-specific probabilistic model. As another example, a set of sequence reads derived from one or more samples associated with breast cancer can be used to fit a breast cancer-specific probabilistic model. In some embodiments, a tissue-specific probabilistic model can be trained for a first, second, third, etc. tissue type and used to assess the tissue of origin of a disease state. For example, a set of sequence reads derived from a first tissue type (e.g., from a lung tissue sample such as a lung biopsy) can be used to train a probabilistic model for a first tissue of origin, and a set of sequence reads derived from a second tissue type (e.g., from a liver tissue sample such as a liver biopsy) can be used to train a probabilistic model for a second tissue of origin. Additionally, in some embodiments, a cancer probability model is trained using a set of sequence reads derived from one or more samples from subjects known to have cancer, and a non-cancer-specific probability model is trained using a set of sequence reads derived from one or more samples from healthy or non-cancer subjects. As one skilled in the art will appreciate, any number of disease state probability models can be trained utilizing sequence reads derived from one or more samples taken from subjects with any one of several possible disease states. For example, in some embodiments, multiple sequence reads may be generated from 3, 4, 5, 6, 7, 8, 9, 10, or more reference samples, each obtained from one or more subjects with different disease states (e.g., different types of cancer), and used to train 3, 4, 5, 6, 7, 8, 9, 10, or more probability models.

[0171] During training, the machine learning engine 220 can utilize methylation information or methylation state vectors (e.g., as described above with respect to Figures 3-4) to train on sequence reads indicative of a disease state. In particular, the machine learning engine 220 determines the observed methylation rate for each CpG site within the sequence read. The methylation rate represents the fraction or percentage of base pairs that are methylated within a CpG site. The trained probability model 230 can be parameterized by a product of methylation rates. As previously described, any known probability model for assigning probabilities to sequence reads from a sample can be used. For example, the probability model can be a binomial model in which all sites (e.g., CpG sites) on a nucleic acid fragment are assigned a methylation probability, or an independent site model in which methylation at each CpG is specified by a separate methylation probability, with methylation at one site assumed to be independent of methylation at one or more other sites on the nucleic acid fragment.

[0172] In some embodiments, in a Markov model, the probability of methylation at each CpG site depends on the methylation state at several preceding CpG sites in the sequence read or in the nucleic acid molecule from which the sequence read is derived (see, e.g., U.S. Patent Application Serial No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification," filed March 13, 2019).

[0173] In some embodiments, the probabilistic model 230 is a "mixture model" fitted using a mixture of components from the underlying models. For example, in some embodiments, the mixture components can be determined using multiple independent site models, where methylation (e.g., methylation rate) at each CpG site is assumed to be independent of methylation at other CpG sites. Using the independent site model, the probability assigned to a sequence read, or the nucleic acid molecule from which it is derived, is the product of the methylation probability at each CpG site where the sequence read is methylated minus one minus the methylation probability at each CpG site where the sequence read is unmethylated. According to this embodiment, the machine learning engine 220 determines the methylation rate for each mixture component. The mixture model is parameterized by the sum of the product of the methylation rates and their associated mixture components. A probabilistic model Pr of n mixture components can be expressed as follows:

number

number

[0174] In some embodiments, the machine learning engine 220 fits the probabilistic model 230 using maximum likelihood estimation to find a set of parameters {β ki , f kThe maximum amount for the N total fragments can be expressed as:

number

[0175] In step 130, analysis system 200 applies a probabilistic model 230 to calculate a value for each sequence read in a second set of sequence reads that is different from, for example, the first set of sequence reads generated in step 110. The value is calculated based on at least the probability that the sequence read (and corresponding fragment) originates from a sample associated with the disease state according to the probabilistic model 230. Analysis system 200 can repeat step 130 for each of the different probabilistic models 230. In some embodiments, analysis system 200 calculates the value using a log-likelihood ratio R according to the fitted probability model associated with a particular disease state. Specifically, the log-likelihood ratio can be calculated using the probability Pr of observing a methylation pattern on a fragment for a sample associated with the disease state and for a healthy sample.

number

[0176] III.C. Feature Selection FIG. 6 illustrates a process for determining features for training a classifier, according to one embodiment. As described above, machine learning engine 220 trains probabilistic models 230 associated with disease states. In the example shown in FIG. 6, probabilistic models 230 ("tissue models") are associated with non-cancer (healthy), breast cancer, and lung cancer. Analysis system 200 processes one or more cfDNA and / or tumor samples to obtain fragments and uses probabilistic models 230 to assign values ​​to the fragments associated with non-cancer (healthy), breast cancer, and lung cancer. Analysis system 200 can identify features for the classifier using information from sequence reads from the cfDNA and / or tumor samples. In some embodiments, analysis system 200 can obtain and assign fragments from each window of a divided reference genome, as shown in FIG. 5. Analysis system 200 aggregates the fragments from the windows into sequences to determine features for the classifier.

[0177] Analysis system 200 identifies features by determining a count of sequence reads having a value above a threshold. In embodiments where the value is based on a log-likelihood ratio R, the threshold is a threshold ratio. Analysis system 200 can identify features using multiple tiers of thresholds. For example, tiers include thresholds of 1, 2, 3, 4, 5, 6, 7, 8, and 9. Each tier represents a varying threshold above which the fragment (from which the sequence read was generated) is more likely to have originated from a sample associated with a disease state than from a healthy sample. Analysis system 200 can use the thresholds to determine a count of outlier fragments that can be used as features.

[0178] By filtering by a threshold, analysis system 200 can consider certain fragments to be outliers because they are less likely to be present in healthy samples. Therefore, outlier fragments can be considered more likely to be associated with (e.g., originate from) a disease state or cancer sample. The number of features can vary between various strata. In other embodiments, analysis system 200 uses a different number of strata or other thresholds. In other embodiments, analysis system 200 can filter fragments using other methods or scoring, such as p-values. In some embodiments, analysis system 200 calculates a p-value for a methylation state vector that describes the probability of observing a methylation state vector or other methylation state vector that is less likely in a healthy control group. To determine informative fragments, analysis system 200 uses a healthy control group that has a majority of fragments that are typically methylated (see, e.g., U.S. Patent Application No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification," filed March 13, 2019).

[0179] Analysis system 200 can repeat the process of identifying features for each probabilistic model. As a result, analysis system 200 can identify features for one or more disease states associated with the probabilistic models. In the example shown in Figure 6, analysis system 200 identifies one or more features for breast cancer and lung cancer.

[0180] In some embodiments, analysis system 200 ranks the identified features based on a measure of the feature's ability to distinguish between different disease states. For example, a feature is informative if it can distinguish a particular type of cancer from other types of cancer or from healthy samples. Analysis system 200 can use mutual information to determine a measure of the feature's information content in distinguishing between two disease states. For each pair of different disease states, analysis system 200 can designate one disease state, e.g., cancer type A, as a positive type and the other disease state, e.g., cancer type B, as a negative type.

[0181] Mutual information can be calculated using estimated fractions of positive and negative samples (e.g., cancer types A and B) for which a feature is expected to be non-zero in the resulting assay. For example, if a feature occurs frequently in healthy cfDNA, the analysis system 200 determines that the feature is unlikely to occur frequently in cfDNA associated with various types of cancer. As a result, the feature may be a weak measure in distinguishing disease states. In calculating mutual information I, variable X is a particular feature (e.g., binary), and variable Y represents a disease state, e.g., cancer type A or B.

number

[0182] In some embodiments, f A The value of is estimated by the fraction of cancer patients whose cfDNA is expected to contain a non-zero feature value. If the training data for cancer type A consists of cfDNA samples, this fraction can be estimated simply as the fraction of cfDNA samples in which the feature is observed. If the training data includes tumor samples, a correction can be applied to account for the lower fraction of tumor-derived fragments in cfDNA compared to the tumor. For N fragments in tumor samples whose values ​​are determined to be greater than a threshold (e.g., from step 140), analysis system 200 calculates the likelihood r of detecting each of those fragments in cfDNA from that patient as follows:

number

[0183] III.D.Classification Analysis system 200 uses the features to generate a classifier. The classifier is trained to predict the tissue of origin associated with a disease state for input sequences read from a test sample of a test subject. Analysis system 200 can select a predetermined number (e.g., 128, 256, 512, 1024) of top-ranking features for each disease state pair to train the classifier, for example, based on a mutual information calculation or another computational measure. The predetermined number can be treated as a hyperparameter selected based on performance in cross-validation. Analysis system 200 can also select features from regions of the reference genome determined to be more informative in distinguishing between disease state pairs. In various embodiments, analysis system 200 retains the best-performing tier for each region and for each cancer type pair (including non-cancer as the negative type).

[0184] In some embodiments, analysis system 200 trains a classifier by inputting a set of training samples having feature vectors into the classifier and adjusting classification parameters so that the classifier's function accurately associates the training feature vectors with their corresponding labels. Analysis system 200 can group the training samples into one or more sets of training samples for repeated batch training of the classifier. After inputting all sets of training samples containing its training feature vectors and adjusting the classification parameters, the classifier can be trained sufficiently to label test samples according to their feature vectors within some margin of error. Analysis system 200 can train the classifier according to any one of several methods, such as L1 regularized logistic regression or L2 regularized logistic regression (e.g., with a logarithmic loss function), generalized linear models (GLMs), random forests, multinomial logistic regression, multilayer perceptrons, support vector machines, neural nets, or any other suitable machine learning technique.

[0185] In various embodiments, analysis system 200 converts feature values ​​by binarization. In particular, feature values ​​greater than 0 are set to 1, and feature values ​​are either 0 or 1 (indicating the presence or absence of a disease state). In other embodiments, a smoothing function may be implemented instead of binarizing to 0 or 1 (e.g., to provide more granular values). As shown in FIG. 14, analysis system 200 can binarize features in cross-validation before training a classifier with the features.

[0186] In various embodiments, analysis system 200 trains a multinomial logistic regression classifier on the training data for the convolutions to generate predictions for the held-out data. For each of the K convolutions, analysis system 200 trains one logistic regression for each combination of hyperparameters. An example hyperparameter is the L2 penalty, a form of regularization applied to the logistic regression weights. Another example hyperparameter is topK, the number of top regions to retain for each tissue-type pair (including non-cancer). For example, if topK=16, analysis system 200 retains the top 16 regions for each tissue-type pair, as ranked by the mutual information procedure described herein. By following this procedure, analysis system 200 can generate predictions for each sample in the training set while ensuring that the classifier is not trained on the data for which the predictions are generated.

[0187] In various embodiments, for each set of hyperparameters, analysis system 200 evaluates the performance of the cross-validated predictions on the full performance set, and analysis system 200 selects the best-performing set of hyperparameters for retraining on the full training set. Performance can be determined based on a log-loss metric. Analysis system 200 can calculate log-loss by taking the negative logarithm of the prediction for the correct label for each sample and then summing across samples. For example, a perfect prediction of 1.0 for the correct label would result in a log-loss of 0 (lower is more accurate). To generate predictions for new samples, analysis system 200 can calculate feature values ​​using the methods described above, but limited to selected features (region / positive class combinations) under a selected topK value. Analysis system 200 can use the generated features to create predictions using a trained logistic regression model.

[0188] During development, analysis system 200 applies a classifier to predict the tissue of origin of a test sample, where the tissue of origin is associated with one of the disease states. In some embodiments, the classifier can return predictions or likelihoods for multiple disease states or tissues of origin. For example, the classifier can return a prediction that the test sample has a 65% likelihood of having breast cancer tissue of origin, a 25% likelihood of having lung cancer tissue of origin, and a 10% likelihood of having healthy tissue of origin. Analysis system 200 can further process the predictions to generate a single disease state determination.

[0189] III.E. Uncertain Localization In various embodiments, tumor fraction can be a covariate in predictions made across samples by a trained classifier or model. As tumor fraction decreases, score assignments (e.g., based on the log-likelihood ratio R described above) can become less definite until the limit of classification detection is reached (i.e., there is a 50% chance that cancer / cancer type will be detected). Samples with high cfDNA tumor fraction tend to be classified definitively, while samples with low cfDNA tumor fraction tend to be more ambiguous. With ambiguous signals, assignments are less reliable and may be correct or incorrect by chance. In a single localization use case, analysis system 200 can identify ambiguous signals and isolate their predictions to an "uncertain localization class."

[0190] For example, in some embodiments, analysis system 200 can determine a posteriori uncertain assignments from a set of tissue-of-origin localization vectors for individuals with cancer scores greater than a specificity target threshold. Analysis system 200 can determine uncertain assignments under cross-validation. For each sample, analysis system 200 can calculate a metric to capture the localization uncertainty for that sample. As one exemplary approach, analysis system 200 calculates the metric using the information entropy (bits) of the tissue-of-origin localizations, with a bit value of 0 occurring if one prediction is certain. In the most ambiguous case (equal probability for all n classes), analysis system 200 calculates (n) bit values. As another exemplary approach, analysis system 200 determines the metric using the difference (delta value) between the top score and the second-top score. If one prediction is certain, a delta value of 1 occurs. A delta value of 0 occurs in the most ambiguous case. By including indeterminate results, analysis system 200 can eliminate weak calls that are correct only by chance and improve the accuracy of definitive localization calls (e.g., the fraction correct for tissue-of-origin assignment).

[0191] As an alternative to post-hoc uncertain assignment, analysis system 200 can use expectation-maximization during training to determine assignment to uncertain classes. Analysis system 200 can also add a second layer to the classifier output to classify cases into uncertain classes.

[0192] Given the metrics and records of whether each sample was correctly localized, analysis system 200 can calculate a precision-recall curve for the uncertain call threshold, as shown in Figure 18. The cutoff point can be selected based on a target precision level, such as 90% in the example shown in Figure 18. Analysis system 200 can calculate cutoff points for localization labels individually (e.g., for a particular cancer type) or for all cancer types as a whole. The tradeoff is subject to optimization and may depend on the cost of an incorrect localization call versus the number of calls assigned an uncertain result (e.g., precision and recall).

[0193] III.F. Guarding Against Class Imbalance In various embodiments, individual samples s i The element score vector for contains the posterior probability of signal localization for each predicted class (e.g., disease state). Each element is scaled by a prior probability proportional to the proportion of training examples per class.

number

[0194] If the classes are imbalanced, samples with weak signals may be shifted to the inappropriate class. For example, a training set may contain 99% of samples with liver cancer detections but few detections of different cancer types. As a result, a classifier trained on this set may be skewed toward liver cancer predictions (or always guess that class). Furthermore, if the class proportions in the classifier training do not match the population frequencies to which the classifier is applied (e.g., the class proportions are more balanced), erroneous predictions may be generated.

[0195] To evaluate the ability of a classifier to localize cfDNA samples from methylation and / or genomic and / or clinical features, analysis system 200 can target inter-class rate equivalence. Analysis system 200 can calibrate scores to the incidence of disease states in a screening population, optionally considering disease detectability via tumor fraction. By modifying the prior applied to a classifier trained using a general training set, analysis system 200 can customize the classifier to improve predictions for specific populations associated with the prior (e.g., indicative of the distribution of disease states in that specific population). Different geographic regions or countries may have different priors based on the prevalence of specific disease states or cancer types in corresponding subpopulations of individuals.

[0196] As an example approach, analysis system 200 performs post-hoc recalibration of model scores. Specifically, analysis system 200 corrects the scores for a class by dividing the assigned probability by the frequency of examples in the training set for that class. The correction can optionally be stabilized by adding pseudocounts. Analysis system 200 then calculates the probability of each score vector s i can be normalized to sum to 1.

[0197] As another approach, analysis system 200 can resample the low-frequency training examples to a desired proportion. As yet another approach, analysis system 200 can reweight the loss function in classifier training.

[0198] III.G. Conditional Cancer Signal Origins In some embodiments, the analysis system 200 trains a classifier to predict cancer signal origin (tissue of origin) conditional on one or more characteristics (e.g., cofactors) of an individual. Exemplary characteristics include age, smoker status (e.g., whether the individual smokes regularly, occasionally, or never), biological sex (e.g., male or female), racial classification, etc. By taking one or more characteristics into account, the classifier's cancer signal origin predictions are tailored to a specific demographic, including the patient, rather than the general population of all individuals. As a result, the classifier is less likely to generate false positives and can reduce crosstalk between different types of cancer signal origins, as further described below with reference to Figures 33-35. Similarly, false positives and crosstalk between cancer signal origin predictions can be more effectively distributed among classifier results for patients. For example, because breast cancer in men is approximately 100 times less likely than breast cancer in women, it may be important to reduce male breast cancer CSO predictions and reallocate those predictions to other CSO predictions where errors do not obscure true cancer incidence. By reducing false positives assigned to low-incidence cancers and shifting classification to high-incidence cancers, the classifier can improve prediction accuracy, i.e., the percentage of samples labeled with a cancer signal origin that actually has its predicted cancer signal origin. A highly accurate classifier increases the likelihood that a patient will receive an intervention appropriate for the type of cancer they actually have and will not receive unnecessary intervention for a false positive. Even small improvements in accuracy can be important for the detection and treatment of low-incidence cancers (cancers that occur rarely).

[0199] In one embodiment, analysis system 200 takes one or more characteristics into account by training separate classifiers for different combinations of characteristics. For example, analysis system 200 trains four different classifiers based on gender (male or female) and smoker status (smoker or non-smoker). In this example, (1) a first classifier is trained using data from a sample of female, never-smoking individuals, (2) a second classifier is trained using data from a sample of female, never-smoking individuals, (3) a third classifier is trained using data from a sample of male, never-smoking individuals, and (4) a fourth classifier is trained using data from a sample of male, never-smoking individuals. Due to the different sets of training data used to train the four classifiers, each trained classifier has different potential weights personalized to specific characteristics. Furthermore, analysis system 200 can adjust the weights of existing classifiers using a particular set of training data associated with one or more characteristics.

[0200] In different embodiments, analysis system 200 takes one or more characteristics into account by applying one or more post-processing steps to the classifier's output. As one example, analysis system 200 multiplies the classifier's output probability by a factor based on the characteristics. For example, the factor may be based on the subject's age, since certain cancer signal origins are more likely to occur in older individuals. As another example, the factor may increase the probability of a lung cancer signal occurring if the subject is a smoker and decrease the output probability if the subject is a never-smoker.

[0201] Figure 33 shows predicted and actual labels of cancer signal origin for a set of male and female individuals, according to one embodiment. The y-axis shows the cancer signal origin predicted by a classifier that was not trained to generate predictions conditional on one or more characteristics. The classifier generated predictions for a sample of 797 individuals, including both male and female individuals. The x-axis shows the actual known cancer signal origin for the 797 samples, resulting in an overall accuracy of 88.5%. The classifier has a high level of accuracy for certain cancers, but can improve its accuracy level for other cancers.

[0202] For example, of 12 samples predicted to have anal cancer, only 5 actually had anal cancer, while 6 of the incorrect predictions actually had cervical cancer. Because male individuals do not have cervixes, if the classifier is personalized by gender, the accuracy of the predictions can be improved. Specifically, a classifier conditioned on males trained using data from only male samples would not produce a prediction of cervical cancer. Because the classifier does not produce an incorrect prediction label 6 of the cervical cancer samples as anal cancer, the accuracy of the classifier for anal cancer increases from 5 / 12 (41.7%) to 5 / 6 (83.3%).

[0203] As another example, Figure 33 shows that a classifier made 140 lung cancer predictions, of which 127 were correct and 13 were incorrect. The likelihood of lung cancer for non-smokers is roughly an order of magnitude smaller than the likelihood for smokers. Thus, a classifier trained using samples of smokers will produce predictions with higher accuracy than a classifier trained using samples of non-smokers. The correct prediction of lung cancer in non-smokers (approximately 127 / 10 = 12.7) can be difficult to distinguish from the crosstalk of 13 incorrect predictions of lung cancer that actually had different cancer signal origins.

[0204] Figure 34 shows predicted cancer signal origins and actual labels for a set of male individuals, according to one embodiment. The classifier correctly predicted one breast cancer label, but also generated false-positive breast cancer predictions that were actually head and neck cancer. Because breast cancer is rare in male individuals, a conditional classifier conditional on males would improve performance for breast cancer prediction. A conditional classifier trained using samples from only male individuals (no samples from female individuals) would result in a classifier with stricter thresholds or requirements for returning a breast cancer prediction.

[0205] Figure 35 shows predicted cancer signal origins and actual labels for samples from female individuals, according to one embodiment. The classifier predicted cervical cancer with 100% accuracy (2 / 2), but the 16.7% accuracy (2 / 12) is low because the 10 samples that actually had cervical cancer were labeled with a different cancer signal origin (head and neck cancer or anal cancer). Because cervical cancer does not occur in male individuals, a classifier conditioned on females would likely improve the accuracy of predicting cervical cancer.

[0206] III.H. Feature Optimization In some embodiments, analysis system 200 performs one or more processes to optimize feature selection when training a classifier. By optimizing the selection of more informative and less noisy features, analysis system 200 can improve the sensitivity of the trained classifier, for example, for one or more or all cancer types. An exemplary process for optimizing feature selection is to use a subset of the top-ranked features for the classifier. Feature rank is inversely proportional to mutual information. Thus, lower-ranked features have higher mutual information and are useful for classifying various cancer types. In some embodiments, analysis system 200 selects a specific number of top features. For example, analysis system 200 selects the lowest (best) ranked 256 features per cancer pair group by positive type, while the remaining features outside the top 256 per cancer pair group are not used by the classifier. The threshold of 256 features per cancer pair group is a hyperparameter for training the classifier. In other embodiments, the threshold is a different number of features (e.g., 128, 256, 512, 1024), commonly referred to as "TopK." The TopK threshold can be determined experimentally, using machine learning, or using another type of process.

[0207] To determine the TopK threshold for a conditional classifier trained for a particular demographic of individuals, the analysis system can train parallel classifiers based on characterization of the training dataset with different numbers of top informative features. The analysis system can then modify each training sample feature vector based on the number of top informative features to be evaluated. For example, the analysis system can train a first classifier that evaluates the top 128 features, a second classifier that evaluates the top 256 features, a third classifier that evaluates the top 512 features, and a fourth classifier that evaluates the top 1024 features. The analysis system can validate the accuracy of each classifier with a validation set of training samples. The classifier with the best accuracy is selected for the particular demographic of individuals.

[0208] Figure 36 shows experimental results of fragment coverage versus feature rank, according to one embodiment. Analysis system 200 determines "coverage_total" as the sum of the average coverage of positive-type fragments and the average coverage of negative-type fragments. As shown in Figure 36, the average coverage per sample varies based on cancer type (e.g., circulating lymphoid, Hodgkin's lymphoma, non-Hodgkin's lymphoma, and plasma cell). In some situations not shown in Figure 36, samples with lower coverage correspond to noisier features.

[0209] Figure 37A shows experimental results of fragment coverage versus non-cancer activation fraction, according to one embodiment. Figure 37B shows experimental results of non-cancer activation fraction versus feature rank, according to one embodiment. Analysis system 200 determines the non-cancer activation fraction as the fraction of non-cancer samples in which a particular feature is turned on (e.g., nc activation = sum(binarized(featureVal)fornon-cancersamples) / len(non-cancersamples)). Analysis system 200 uses the non-cancer activation fraction as a measure of noise because a larger non-cancer activation fraction value means the sample has more noise, and the classifier may be less useful in distinguishing between cancer types. Ideally, the non-cancer activation fraction value would be close to 0, resulting in a low noise level. In another embodiment, analysis system 200 uses the difference between the cancer activation fraction and the non-cancer activation fraction as a measure of noise. It is advantageous to select features that are not turned on in non-cancer samples (but are turned on in cancer samples) so that the trained classifier can make more accurate predictions (distinguishing cancer from non-cancer) based on the top features. As shown in Figure 37A, the data samples for circulating lymphatic lymphoma and non-Hodgkin's lymphoma have more noise at lower coverage. In comparison, the data samples for Hodgkin's lymphoma and plasma cell are less noisy in both Figures 37A and 37B.

[0210] Figure 38A shows experimental results of non-cancer activation fraction versus feature weight, according to one embodiment. Figure 38B shows another view of the experimental results shown in Figure 38A, according to one embodiment. The weight indicates the degree to which a particular feature contributes to the p-cancer score output by the classifier.

[0211] FIG. 39A shows experimental results of activation fraction versus coverage, according to one embodiment. FIG. 39B shows another view of the experimental results shown in FIG. 39A, according to one embodiment. FIGS. 39A-B show that stage 4 cancers generally have larger activation fractions than non-cancer and other cancer stages, enabling the trained classifier to distinguish stage 4 cancers. In contrast, distinguishing between activation fractions of stage 1 cancers and borderline false-positive detections is more difficult because their activation fractions are similar. Analysis system 200 determines borderline false-positive detections based on a p-cancer threshold, such as 0.1, although the cutoff threshold can vary.

[0212] As described above, analysis system 200 generates features by cancer pair, including positive types (those with a particular type of cancer) and negative types (those without a particular type of cancer). Analysis system 200 calculates mutual information scores for the cancer pairs and determines the threshold at which the mutual information score is highest. The analysis system then returns a set of features for each region of the fragment, the set including features for each cancer pair. For example, given 21 different types of cancer (plus non-cancer), a maximum of 21 × 21 features per region are returned (non-cancer is not used as a positive type). Analysis system 200 ranks the features by mutual information score within the cancer pair group. Analysis system 200 then filters the ranked features using a TopK threshold and applies duplicate elimination.

[0213] In one embodiment, analysis system 200 adjusts the TopK threshold based on the positive cancer type. In one implementation, analysis system 200 uses a lower TopK threshold for positive cancer types with higher levels of noise. For example, analysis system 200 applies a TopK threshold of 192 to circulating lymphocyte-positive cancer types and a TopK threshold of 256 to Hodgkin's lymphoma, non-Hodgkin's lymphoma, and plasma cell-positive cancer types. Additionally, analysis system 200 adjusts the TopK threshold based on the availability of positive cancer type samples. Figure 40 shows experimental results of this embodiment.

[0214] In one embodiment, analysis system 200 prioritizes features using statistics including one or more of coverage (either coverage_total or coverage_negative_type), activation fraction (non-cancer and / or cancer minus non-cancer), and mutual information score. As noted above, features with high mutual information are useful for classifying various cancer types. Features with low coverage are less desirable for training classifiers because they may have a high level of noise compared to features with higher coverage. Features with low non-cancer activation fractions are more desirable due to their lower noise levels. Similarly, features with a large difference between cancer and non-cancer activation fractions are more desirable due to their stronger signal quality. Analysis system 200 takes this into account by ranking features using mutual information score, coverage, and activation fraction (non-cancer and / or cancer minus non-cancer). In one example, for each cancer pair group, analysis system 200 applies thresholds for a minimum coverage level and a maximum non-cancer activation fraction level. The minimum coverage level and activation fraction threshold can vary depending on the TopK within each positive cancer type. As another example, analysis system 200 ranks features using the function: rank = α * mutual information score + β * coverage + γ * non-cancer activation fraction + delta * (cancer minus non-cancer activation fraction), where α, β, γ, and delta are different weights. In embodiments where features are initially ranked using mutual information score alone or another criterion, analysis system 200 can re-rank the features using additional statistics. Analysis system 200 applies a TopK threshold to filter features after ranking or re-ranking. Figure 41 shows experimental results for this embodiment. Figure 42 shows experimental results using the various thresholding methods described above, according to various embodiments.

[0215] In one embodiment, analysis system 200 generates a statistical model (e.g., Z-score) to identify outlier features based on mutual information. For example, for each cancer pair group, analysis system 200 estimates the distribution of mutual information scores and determines a cutoff threshold based on that distribution. The statistical model can also account for statistics of region fragment coverage or activation fraction.

[0216] In one embodiment, analysis system 200 uses the same "optimized" features described above for both the binary and TOO classification stages. In another embodiment, different sets of "optimized" features are used for the binary and TOO classification stages.

[0217] IV. Multilayer Perceptron Model In some embodiments, a multilayer perceptron model (“MLP”) can be used as an alternative to logistic regression for classification. Similar to a classifier based on logistic regression, an MLP classifier can be a single multi-class classifier for both detecting cancer and determining the tissue of origin (TOO) or type of cancer. For example, the multi-class classifier can be trained to distinguish between two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different types of cancer. In one embodiment, the multi-class cancer MLP model can also include a class label for non-cancer, and cancer detection can be determined (e.g., as 1—non-cancer). In another embodiment, the multilayer perceptron model can be a two-stage classifier, e.g., with one or more hidden layers, having a first stage for binary classification (e.g., cancer or non-cancer) and a second-stage multilayer perceptron model for multi-class classification (e.g., TOO).

[0218] In one embodiment, the multilayer perceptron includes a two-stage classifier: a first-stage multilayer perceptron (MLP) binary classifier with no hidden layer; and a second-stage multilayer perceptron (MLP) multi-class classifier with a single hidden layer. In one embodiment, samples determined to have cancer using the first-stage classifier are subsequently analyzed by the second-stage classifier.

[0219] In the first stage of training, a binary (two-class) multilayer perceptron model with no hidden layers is trained to detect the presence of cancer and to distinguish cancer samples from non-cancer samples (regardless of TOO). For each sample, the binary classifier outputs a prediction score indicating the likelihood of cancer presence or absence.

[0220] In a second stage of training, a parallel multi-class multilayer perceptron model can be trained to determine the type of cancer or the tissue of origin of the cancer. In one embodiment, only cancer samples that receive a score above a cutoff threshold (e.g., the 95th percentile of non-cancer samples in the first stage classifier) ​​can be included in the training of this multi-class MLP classifier. For each cancer sample used in training and testing, the multi-class MLP classifier outputs a prediction value for the cancer type to be classified, where each prediction value is the likelihood that a given sample has a particular cancer type. For example, the cancer classifier can return a cancer prediction for the test sample, including a prediction score for breast cancer, a prediction score for lung cancer, and / or a prediction score for no cancer.

[0221] 16 is a flowchart of a method 1600 for determining the probability that a sample has a disease state, according to various embodiments. In some embodiments, an analysis system 200 performs method 1600 to process sequence reads of fragments from a nucleic acid sample. Method 1600 includes, but is not limited to, the following steps, which are described with respect to components of analysis system 200:

[0222] In step 1610, analysis system 200 generates sequence reads from one or more biological samples. In some embodiments, analysis system 200 filters the sequence reads according to their p-value scores. The p-value scores of the sequence reads indicate the probability that methylation is observed in the nucleic acid fragments of the one or more biological samples corresponding to the sequence reads.

[0223] In step 1620, analysis system 200 uses the sequence reads to determine, for each location of the set of chromosomal locations, a count of nucleic acid fragments of one or more biological samples within that location that have at least a threshold similarity to fragments associated with a disease state, e.g., cancerous fragments. The disease state may be associated with at least one type of cancer, a stage of cancer, or another type of disease or condition.

[0224] Each position may represent the number of consecutive base pairs in a chromosome. The number of base pairs may vary between different positions. Analysis system 200 may generate sequence reads for multiple regions of a genome. There may be up to tens of thousands of regions or more. Each region may contain hundreds, thousands, or more base pairs. Method 1600 may be performed for whole genome bisulfite sequencing (WGBS) or targeted panel assays.

[0225] In step 1630, analysis system 200 uses the location counts as features to train a machine learning model. In some embodiments, analysis system 200 binarizes the features to indicate the presence or absence (e.g., a Boolean value) of one of the disease states at each location. A count of at least one nucleic acid fragment at a location indicates the presence of one of the disease states at that location. A count of zero nucleic acid fragments at a location indicates the absence of one of the disease states at that location. In some embodiments, the machine learning model can be a logistic regression model. In some embodiments, the machine learning model can be a multilayer perceptron model (neural network). Those skilled in the art will readily recognize that other machine learning models can be used, including, for example, a generalized linear model (GLM), a multilayer perceptron, a support vector machine, a random forest, or a neural network classifier.

[0226] In step 1640, the trained machine learning model determines the probability that the test sample has a disease state. The test sample may be obtained from a patient and may include blood and / or tissue. In optional step 1650, treatment is provided to the patient according to the probability. For example, in response to determining that the probability is greater than a threshold, treatment (e.g., medication or an interventional procedure) may be provided to the patient. In another embodiment, optional step 1650 may generate a test report including the probability that the test sample has the disease, providing the test results to the patient.

[0227] The experimental results shown in Figures 17-20 were obtained by training the model using samples from the CCGA study, which is further described below.

[0228] 17 shows the performance gain in sensitivity of a multi-layer perceptron model according to one embodiment. Compared to a logistic regression model, the multi-layer perceptron model (MLP) demonstrates a performance gain in sensitivity for disease detection across cancer stages I, II, III, and IV.

[0229] 18 shows experimental results of a multilayer perceptron model in determining tissue of origin, according to one embodiment. Compared to the logistic regression models (LR: 1803 and 1804), the multilayer perceptron models (MLP: 1801 and 1802) show improved accuracy in determining tissue of origin. The improved accuracy is realized when processing sequence reads associated with all cancer types in the training set, and when processing sequence reads from a training set that includes more than 10 example sequence reads for each cancer type in the training set.

[0230] 19 shows experimental results of a multilayer perceptron model in determining tissue of origin by cancer stage, according to one embodiment. Compared to a logistic regression (LR) model, the multilayer perceptron model (MLP) demonstrates performance gains in accuracy of tissue of origin (TOO) detection across cancer stages I, II, III, and IV. Among cancer stages, the performance gain for the MLP model is greatest for stage I.

[0231] 20 shows experimental results of a multi-layer perceptron model across cancer types, according to one embodiment. For most of the cancer types shown in FIG. 20, the multi-layer perceptron model (MLP) achieves higher accuracy of tissue of origin (TOO) detection compared to the logistic regression model.

[0232] In some embodiments, the analysis system uses a two-stage model to determine the tissue of origin (TOO) of a cancer or another type of disease state. The analysis system generates sequence reads from nucleic acid fragments of a biological sample. The analysis system determines a first set of training data by processing the sequence reads, for example, using any of the processes described in Section II.A. Assay Protocol. The analysis system can determine the first set of training data using methylation information. For example, the analysis system determines sequence reads that are hypomethylated by determining that a threshold or percentage of CpG sites corresponding to the sequence reads are unmethylated. Additionally, the analysis system determines sequence reads that are hypermethylated by determining that a threshold or percentage of CpG sites corresponding to the sequence reads are methylated. The analysis system can also determine that the sequence reads are informative. In some embodiments, the analysis system filters the sequence reads by removing sequence reads with p-values ​​below a threshold p-value.

[0233] The analysis system uses a first set of training data to train a binary classifier, where the binary classifier is trained to predict, for an input sequence read from a first test biological sample, a binary output, i.e., the presence or absence of at least one disease state in the first test biological sample.

[0234] Using the binary classifier's prediction, the analytical system can determine that a subset of the biological samples has the presence of one or more disease states. The binary classifier can be used to train a tissue-of-origin classifier. In particular, the analytical system determines a second set of training data using sequence reads corresponding to nucleic acid fragments of the subset of biological samples. The analytical system uses the second set of training data to train the tissue-of-origin classifier. The tissue-of-origin classifier is trained to predict, for input sequence reads from the second test biological sample, a tissue of origin associated with a disease state present in the second test biological sample. The first and second test biological samples can be the same sample or different samples.

[0235] In some embodiments, the analysis system uses a tissue-of-origin classifier to determine a score indicating the probability that a tissue-of-origin associated with a disease state is present in the second test biological sample. The analysis system can calibrate the score, for example, to adjust for overconfident model output. For example, the analysis system performs a k-nearest neighbor (KNN) operation using the feature space output by the tissue-of-origin classifier to associate the score. In one embodiment, the feature space includes the top two predicted labels from the tissue-of-origin classifier (e.g., lung cancer and prostate cancer) and an indication of whether the correct classification was for a disease state different from the top two predictions. The analysis system can also calibrate the score by normalizing the probabilities using the output of a binary classifier indicating different probabilities of the presence of at least one disease state present in the second test biological sample.

[0236] In some embodiments, the tissue of origin classifier is a multilayer perceptron that includes at least one hidden layer. The tissue of origin classifier can also include a 100-unit hidden layer or a 200-unit hidden layer, among other sizes of hidden layers. The multilayer perceptron can be fully connected and use a rectified linear unit activation function. In some embodiments, the binary classifier is a multilayer perceptron that does not include a hidden layer. In a different embodiment, the binary classifier is a multilayer perceptron that includes at least one hidden layer. In other embodiments, these classifiers can be logistic regression models, multinomial logistic regression models, or other types of machine learning models.

[0237] Furthermore, the analysis system can train the tissue-of-origin classifier and the binary classifier using one or more machine learning techniques known to those skilled in the art, including, for example, no early stopping (instead selecting a given number of training epochs), stochastic gradient descent, weight decay, dropout regularization, Adam optimization, He initialization, and learning rate scheduling, modified linear unit activation function, leaky modified linear unit activation function, sigmoid activation function, and boosting, among others. As shown in Figure 31, the tissue-of-origin accuracy of the tissue-of-origin classifier improves with training iterations. Each iteration can include a different combination of machine learning techniques. In addition, the increase in tissue-of-origin accuracy exists across different cancer stages: I, II, and III.

[0238] In some embodiments, the analysis system performs cross-validation on one or both of the tissue-of-origin classifier and the binary classifier. The analysis system can retrain the classifier using hyperparameters selected based on the output of the cross-validation. The analysis system can select hyperparameters by aggregating results from all convolutions in the cross-validation. In one embodiment, the analysis system selects hyperparameters to train the tissue-of-origin classifier by optimizing tissue-of-origin accuracy instead of log-likelihood, because the classifier is more reliable for samples with stronger signals.

[0239] In some embodiments, the analysis system determines the probability that a tissue of origin associated with a disease state is present in the second test biological sample using a tissue-of-origin classifier. In response to determining that the probability is greater than the tissue-of-origin threshold, the analysis system predicts that a tissue of origin associated with a disease state is present in the second test biological sample. The analysis system can determine different tissue-of-origin thresholds associated with different tissues of origin. Additionally, the analysis system can determine a tissue-of-origin threshold associated with a given disease state by iterating over a range of probabilities for candidate tissue-of-origin thresholds. For each iteration, the analysis system determines the sensitivity at a given specificity of the tissue-of-origin classifier. The analysis system can optimize the trade-off between the sensitivity and specificity of the tissue-of-origin classifier for a given disease state. The analysis system can determine the sensitivity using the score output by a binary classifier or a tissue-of-origin classifier. Additionally, the analysis system can stratify samples using the score from the tissue-of-origin classifier.

[0240] In some embodiments, the analysis system trains the binary classifier and the tissue-of-origin classifier using binarized features, each having a value of 0 or 1. In binarization, values ​​greater than 1 are replaced with 1.

[0241] V. Adjusting the Binary Classification Threshold The analysis system can adjust the trained cancer classifier to prune samples used to train the cancer classifier. In particular, the analysis system may attempt to remove non-cancerous samples with high tissue signals that weaken the sensitivity of the cancer classifier in cancer prediction. A high tissue signal refers to samples with a significant fraction of cfDNA from the tissue of origin, as determined by, for example, a tissue-of-origin (TOO) classifier, a multi-class cancer classifier, or other means, compared to a healthy distribution. Non-cancerous samples with high tissue signals are outliers in the non-cancerous distribution, which may represent pre-cancer, early-stage cancer, or undiagnosed cancer. The analysis system can identify non-cancerous samples with high tissue signals in at least one cancer type. In some embodiments, certain cancer types are further divided into cancer subtypes. For example, blood cancer types can be further divided into combinations of, for example, circulating lymphoid subtypes, non-Hodgkin's lymphoma (NHL) indolent subtypes, NHL aggressive subtypes, Hodgkin's lymphoma (HL) subtypes, myeloid subtypes, and plasma cell subtypes.

[0242] Referring to FIG. 21, FIG. 21 shows a graph of the likelihood of cancer type for non-cancerous samples exceeding 95% specificity. A cancer score was calculated for each non-cancerous sample from multiple non-cancerous samples, i.e., samples from healthy individuals not currently diagnosed with cancer. The cancer score can be determined by a binary classifier as the likelihood that the sample has cancer, taking into account the sample's methylation sequencing data. In other embodiments, the cancer score can be calculated according to other methods that input at least sequencing data (e.g., methylation, single nucleotide polymorphisms (SNPs), DNA, RNA, etc.) and output the likelihood of the sample having cancer based on the input sequencing data. An example of a classifier is a mixed model classifier. A distribution of non-cancerous samples can be generated according to the cancer scores of the non-cancerous samples. A binary threshold cutoff can be set to ensure a certain level of binary classification specificity, e.g., a true negative rate. Typically, a high specificity cutoff is used in classifying cancer, e.g., 90% to 99.9%, or 99.5% or higher. However, many non-cancer samples used to train the cancer classifier and that fall just below the specificity cutoff may have high tissue signal, thereby biasing the binary threshold cutoff positively.

[0243] To demonstrate, non-cancerous samples exceeding 95% specificity were selected and then input into a multi-class cancer classifier to determine the probability for each cancer type or tissue of origin (TOO). The cancer types or TOO labels used in this embodiment of the multi-class cancer classifier include: circulating lymphoid, myeloid, NHL low-grade, colorectal, NHL aggressive, lung, uterine, breast, prostate, pancreatic and gallbladder, upper gastrointestinal, bladder and urothelial, plasma cell, head and neck, kidney, ovarian, sarcoma, liver and bile duct, cervix, other tissue, HL, anorectal, melanoma, and thyroid. The graph in Figure 21 shows many non-cancerous samples with high tissue signal from at least one tissue type. Each dot in the tissue type column corresponds to the tissue-of-origin likelihood for the non-cancerous sample exceeding the 95% specificity threshold. Notably, many tissue types have multiple non-cancerous sample outliers with significant tissue contributions that are not typical for non-cancerous samples. This may occur when such non-cancer samples have cfDNA signals driven by cancer-like methylation, clonal fraction, and / or growth / turnover rates. It can be inferred that many of the non-cancer samples used to train cancer classifiers may represent pre-cancer, early-stage cancer, or undiagnosed cancer. Nevertheless, these non-cancer samples with significant tissue contributions reduce the sensitivity of cancer classification by shifting up the binary classification cutoff threshold, particularly for samples with significant tissue signals just below the previously set binary classification cutoff threshold. In fact, such signals (e.g., corresponding to circulating lymphoid, myeloid, and low-grade NHL) may be major attractors of false-positive results. Notably, circulating lymphoid, myeloid, low-grade NHL, colorectal, aggressive NHL, lung, uterine, breast, prostate, pancreas and gallbladder, upper gastrointestinal, plasma cell, head and neck, cervical, and HL all had at least one non-cancer sample with a tissue origin probability greater than 0.1. In particular, circulating lymphoid, myeloid, NHL low-grade, and NHL aggressive (all hematological subtypes) had two or more non-cancerous samples with a tissue origin probability greater than 0.5.

[0244] Referring to FIG. 22, FIG. 22 shows a graph of hematological subtypes separated according to methylation sequencing data. The graph in FIG. 22 demonstrates the ability to model hematological subtypes. This may prove useful for providing greater granularity to multi-class cancer classification (e.g., additional classification by hematological subtype label) or as a way to refine cancer classification by pruning non-cancer samples with high hematological subtype signals before training a cancer classifier. As described above, methylation signals can create a high-dimensional vector space by covering multiple CpG sites. The hematological subtype samples and non-cancer samples enable the analysis system to perform principal component analysis (PCA). PCA identifies orthogonal principal components (or embeddings) of the vector space in order of methylation signal variance across samples. The first principal component, denoted V1 on the horizontal axis of the graph, has the highest variance, and the second principal component, denoted V2 on the vertical axis of the graph, has the second highest variance. Annotated in graph 900 are clusters of samples by hematological subtype and non-cancer. The hematological subtypes shown include circulating lymphoid, solid lymphoid, plasma cell, and myeloid. The solid lymphoid subtypes can be further divided into HL, NHL low-grade, and NHL aggressive. The graph shows the potential for classifying according to hematological subtype for the addition of hematological subtypes in multiclass cancer classification or for modeling each of the hematological subtypes for tuning cancer classifiers.

[0245] Removal of VA high signal non-cancer samples 23A shows a flowchart illustrating a process 1000 for determining a binary threshold cutoff for binary cancer classification according to one or more embodiments. Binary classification for predicting cancer versus non-cancer involves evaluating a sample's cancer score against a determined binary threshold cutoff, where samples with a cancer score below the binary threshold cutoff are determined to be non-cancerous and samples with a cancer score equal to or greater than the binary threshold cutoff are determined to be cancerous. A trained multi-class cancer classifier evaluates the sample's methylation signals (and / or other sequencing data) to determine probabilities for several TOO labels classified by the multi-class cancer classifier. The TOO labels used in the multi-class cancer classifier can be cancer tissue types or cancer tissue subtypes (e.g., hematological subtypes, as described above). Process 1000 can be performed or accomplished by an analysis system.

[0246] The analysis system receives sequencing data for a plurality of biological samples containing cfDNA fragments, the biological samples including cancer samples and non-cancerous samples 1010. The sequencing data can be methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.

[0247] For each non-cancer sample, the analysis system classifies 1020 the non-cancer sample using a multi-class cancer classifier based on the sequencing-derived features, where the multi-class cancer classifier predicts a probability for each of a plurality of TOO labels. The analysis system can generate a feature vector for the non-cancer sample and assign an anomaly score for each CpG site under consideration based on at least one informative cfDNA fragment that overlaps with the CpG site.

[0248] For each non-cancer sample, the analysis system determines whether the predicted probability likelihood exceeds a TOO threshold for one or more TOO labels 1030. TOO threshold determination is further described below in Figure 23B.

[0249] The analysis system determines 1040 a binary threshold cutoff for predicting the presence of cancer, where the binary threshold cutoff is determined based on the distribution of non-cancer samples, excluding one or more non-cancer samples identified as having a probability likelihood exceeding at least one TOO threshold. Non-cancer samples for which at least one probability likelihood for a TOO label exceeds the TOO threshold corresponding to that TOO label are excluded. The analysis system then calculates a distribution of the non-cancer samples according to the cancer score for each non-cancer sample, and then determines from the distribution a binary threshold cutoff at a desired level of specificity (e.g., 99.4-99.9% specificity). Note that each cancer score can be determined according to sequencing data; for example, the cancer score can be output by a binary cancer classifier that predicts the likelihood of cancer based on methylation sequencing data, as described herein. In other embodiments, the cancer score can be calculated according to other methods that input at least sequencing data (e.g., methylation, single nucleotide polymorphisms (SNPs), DNA, RNA, etc.) and output the likelihood of a sample having cancer based on the input sequencing data.

[0250] FIG. 23B shows a flowchart illustrating process 1005 of thresholding TOO labels to determine a binary threshold cutoff for binary cancer classification, according to one or more embodiments. This process 1005 may be one embodiment of process 1000. Binary classification to predict cancer versus non-cancer evaluates the sample's cancer score against the determined binary threshold cutoff, where samples with a cancer score below the binary threshold cutoff are determined to be non-cancerous and samples with a cancer score equal to or greater than the binary threshold cutoff are determined to be cancerous. A trained multi-class cancer classifier evaluates the sample's methylation signals (and / or other sequencing data) to determine probabilities for several TOO labels classified by the multi-class cancer classifier. The TOO labels may be cancer tissue types or, more specifically, cancer tissue subtypes (e.g., hematological subtypes, as described above). Process 1005 may be performed or accomplished by an analysis system.

[0251] The analysis system obtains 1015 a training set including a plurality of samples labeled as cancer or non-cancer, and a holdout set including a plurality of samples each labeled as either cancer or non-cancer, i.e., cancer samples or non-cancer samples. Each sample in the training set includes methylation sequencing data, for example, generated according to process 300 of FIG. 3. In other embodiments, each training sample has other sequencing data used in parallel or used to replace the methylation sequencing data. Additionally, each sample from the training set and the holdout set has a cancer score. As described above, the cancer score can be determined by a binary classifier as the likelihood that a sample has cancer, given the sample's methylation sequencing data. In other embodiments, the cancer score is calculated according to other methods, such as those exemplified by the mixture models described herein, that input at least sequencing data (e.g., methylation, single nucleotide polymorphisms (SNPs), DNA, RNA, etc.) and output the likelihood of the sample having cancer according to the input sequencing data.

[0252] The analysis system determines a feature vector for each non-cancer training sample based on the methylation sequencing data 1025. The analysis system can determine the feature vector for each non-cancer training sample, for example, by determining an anomaly score for each CpG site in the set of CpG sites considered. In some embodiments, the analysis system defines an anomaly score for the feature vector with a binary score based on whether there is an informative fragment in the set of informative fragments that encompass the CpG site. Once all anomaly scores have been determined for the sample, the analysis system determines a feature vector as a vector of anomaly scores associated with each CpG site considered. The analysis system can additionally normalize the anomaly scores of the feature vector based on sample coverage.

[0253] The analysis system inputs the feature vector for each non-cancer training sample into a multi-class cancer classifier to generate a TOO prediction 1035. The multi-class cancer classifier is trained on multiple TOO labels, which may include cancer type, cancer subtype, non-cancer, or any combination thereof. The multi-class cancer classifier can be trained as described herein. The trained multi-class cancer classifier determines multiple probabilities for the TOO labels as cancer predictions, where the probabilities for the TOO labels indicate the likelihood of having the cancer corresponding to the TOO label.

[0254] In some examples, the analysis system sweeps 1045 or iterates across a range of probabilities for the TOO label as a candidate TOO threshold, calculating specificity and sensitivity across the range of probabilities for the TOO label. The analysis system may sweep incrementally across the range of probabilities, e.g., 0.01, 0.02, 0.03, 0.04, 0.05, etc. As the analysis system sweeps across the range of probabilities, it filters non-cancer training samples whose probability of the TOO label is equal to or greater than the candidate TOO threshold according to the output of the multi-class cancer classifier. As a numerical example, the analysis system considers a candidate TOO threshold of 0.35. Non-cancer training samples whose probability of the TOO label is equal to or greater than 0.35 are excluded from the training set. The analysis system determines an adjusted binary threshold cutoff based on the filtered training set. The analysis system calculates the specificity of the prediction with the adjusted binary threshold cutoff for the holdout set. Specificity refers to the accuracy of identifying non-cancer samples as non-cancer labels. The analysis system also calculates the sensitivity of the prediction with the adjusted binary threshold cutoff for the holdout set. Sensitivity refers to the accuracy of identifying a cancer sample as a cancer label. In practice, specificity and / or sensitivity can be defined according to a true positive rate, a false positive rate, a true negative rate, a false negative rate, another statistical calculation, etc.

[0255] The analysis system determines a TOO threshold for the TOO label 1055. The analysis system selects a TOO threshold from the candidate TOO thresholds by optimizing the calculated specificity and / or sensitivity across a range of candidate TOO thresholds. In some examples, a TOO threshold is determined or otherwise applied to a specific TOO tissue type class or subtype class, such as a hematology class. By way of example only, an algorithm for calculating and applying a TOO-specific probability threshold can be used to remove non-cancerous samples that exceed the signal of a hematological disorder. The algorithm may include first searching through a grid of probability values ​​for each pre-specified TOO label, and then evaluating the clinical specificity and clinical sensitivity of the holdout set using a calculated binary detection threshold for all values ​​after removing non-cancerous samples that meet or exceed the probability of the specified TOO label. By iterating across the probability grid, the algorithm will identify a combination of TOO thresholds for the pre-specified TOO labels that optimizes the trade-off between clinical specificity and clinical sensitivity for the holdout set. The final optimized TOO probability threshold will be used to remove non-cancerous samples whose TOO labels exceed any of the given values. The cleaned set of non-cancer samples is then used to calculate the cancer / non-cancer detection threshold. Furthermore, in some cases, TOO-specific thresholding can be manually set at any cutpoint, such as a desired specificity level (e.g., 99.4-99.9% specificity).

[0256] The analysis system refines the binary cancer classification by pruning non-cancer training samples that exceed the TOO thresholding before determining the binary threshold cutoff 1065. The analysis system eliminates non-cancer training samples from the training set according to the determined TOO threshold for the TOO label. The analysis system sets the binary threshold cutoff according to the filtered training set. For example, the analysis system determines a new binary threshold cutoff based on the filtered distribution of scores. In additional embodiments, the analysis system can determine a TOO threshold for every TOO label according to steps 1010, 1020, 1030, and 1040 to refine the binary cancer classification.

[0257] Stratification of sample distribution according to VBTOO signals In one or more embodiments, the analysis system tunes the cancer classifier by stratifying the sample distribution according to a TOO signal to determine a stratum-specific binary threshold cutoff. The analysis system can stratify the sample distribution according to a signal for one or more TOO labels determined according to the TOO predictions output by the multi-class cancer classifier.

[0258] As used herein, "high tissue signal" refers to a sample whose tissue signal (also referred to as a TOO label) for any type of tissue in general or for a specific cancer type exceeds a certain threshold. The tissue signal can be determined by a multi-class cancer classifier or other approach compared to a healthy distribution. Non-cancerous samples with high tissue signals are outliers in the non-cancerous distribution. Some of these non-cancerous samples may be pre-cancer, early-stage cancer, or undiagnosed cancer. The analysis system can identify non-cancerous samples with high tissue signals in at least one TOO label. In one approach to determining high tissue signals, the predicted value for the TOO label output by the multi-class cancer classifier is compared against a tissue signal threshold. Samples with predicted values ​​above the tissue signal threshold are considered to have high tissue signals for that TOO label; while samples with predicted values ​​below the tissue signal threshold are considered not to have high (or low) tissue signals for that TOO label. In another approach, one or more top predictions in the TOO prediction are considered. For example, a TOO prediction for a sample may have a first prediction of a colorectal TOO label, a second prediction of a breast TOO label, and a third prediction of a head / neck TOO label. If the top predictions are considered, the sample is considered to have a high tissue signal for the TOO label in the first prediction, which in this example is the colorectal TOO label. If the top two predictions are considered, there is a high tissue signal for both the colorectal TOO label and the breast TOO label. Other approaches to determining tissue signals may include other models trained to determine tissue signals for one or more TOO labels. Such models may include classifiers trained to determine tissue signals for a subset of TOO labels. For example, a hematology-specific classifier may be trained and used to determine tissue signals for one or more hematological subtypes. Other models include deconvolution models that can deconvolute tissue signals from methylation sequencing data (and / or other types of sequencing data).

[0259] Referring now to Figure 32, Figure 32 illustrates a process for stratifying hematological signals into two tiers, according to one or more embodiments. The following description describes stratification by hematological signals, but the principles can be easily applied to other TOO signals.

[0260] The analytical system stratifies 1300A a holdout set of cancer and non-cancerous samples according to their hematological signals into a low signal tier 1310 and a high signal tier 1320. Each sample in the holdout set has a cancer score determined by a binary cancer classifier and a TOO prediction determined by a multi-class cancer classifier. In one embodiment, the hematological signal of the sample is determined according to the TOO prediction output by the multi-class cancer classifier. In one embodiment, when considering one or more top predictions (e.g., top 1, top 2, etc.), a high hematological signal is determined if at least one of the top predictions being considered is one of the hematological subtypes (e.g., lymphoid neoplasia subtype and myeloid neoplasia subtype). Other hematological subtypes may be included. Thus, if a sample has a TOO prediction where at least one of the top predictions is considered to be a lymphoid neoplasia subtype or a myeloid neoplasia subtype, the sample is determined to have a high hematological signal. Otherwise, the sample is determined not to have a high hematological signal.

[0261] The analysis system determines a binary threshold cutoff for each tier for predicting the presence or absence of cancer in samples. Samples in the low signal tier 1310 are used by the analysis system to determine 1305 a binary threshold cutoff for predicting the presence or absence of cancer in samples in the low signal tier 1310. The binary threshold cutoff is determined 1305 according to a false positive budget set for the low signal tier 1310. Using the cancer scores for samples in the low signal tier 1310, the analysis system sweeps over a range of candidate binary threshold cutoffs, evaluating the true positive rate (also referred to as sensitivity) and false positive rate at each candidate binary threshold cutoff. The candidate binary threshold cutoff with the closest false positive rate within the false positive budget is determined to be the candidate binary threshold cutoff. The analysis system performs similar operations to determine 1315 a binary threshold cutoff for the high signal tier 1320. The false positive budget for the low signal tier 1310 and the false positive budget for the high signal tier 1320 may be set according to the ratio of the statistical true positive rates of the tiers. This ratio aims to reduce the false positive rate in the high signal tier 1320.

[0262] For a test sample, the analytical system places the test sample into either a low signal tier 1310 or a high signal tier 1320 according to the hematological signal. If the test sample is placed in the low signal tier 1310, the analytical system applies 1315 a binary threshold cutoff for the low signal tier 1310 to the cancer score of the test sample. If the cancer score is equal to or greater than the binary threshold cutoff for the low signal tier 1310, the analytical system returns a prediction of the presence of cancer in the test sample, otherwise it returns a prediction of no cancer. If the test sample is placed in the high signal tier 1320, the binary threshold cutoff for the low signal tier 1320 is applied 1325 to the cancer score of the test sample. If the cancer score is equal to or greater than the binary threshold cutoff for the high signal tier 1320, the analytical system returns a prediction of the presence of cancer in the test sample, otherwise it returns a prediction of no cancer.

[0263] VI. Circulating Cell-Free Genome Atlas Study In various embodiments, each predictive cancer model is trained using a set of training data derived from a training subset of patients from the Circulating Cell-Free Genome Atlas (CCGA) study (see Clinical Trial.gov Identifier: NCT02889978 (https: / / www.clinicaltrials.gov / ct2 / show / NCT02889978)), and then tested using a set of testing or validation data derived from a testing or validation subset of patients from the CCGA study.

[0264] The predictive cancer models described herein were trained using multiple known cancer types from the Circulating Cell-Free Genome Atlas (CCGA) study. The CCGA sample set included the following cancer types: breast, lung, prostate, colorectal, kidney, uterine, pancreatic, esophageal, lymphoma, head and neck, ovarian, hepatobiliary, melanoma, cervical, multiple myeloma, leukemia, thyroid, bladder, stomach, and anorectal. Thus, the model can be a multi-cancer model (or multi-cancer classifier) ​​for detecting one or more, two or more, three or more, four or more, five or more, ten or more, or twenty or more different types of cancer.

[0265] The predictive cancer model can be trained using a refined set of training data derived from a first subset of patients from the CCGA study and then tested using a refined set of test data derived from a second subset of patients from the CCGA study.

[0266] VII. Cancer Assay Panels In various embodiments, the predictive cancer models described herein use samples enriched using a cancer assay panel containing multiple probes or multiple probe pairs. Several targeted cancer assay panels are known in the art, for example, as described in International Publication No. WO 2019 / 195268, filed April 2, 2019, PCT / US2019 / 053509, filed September 27, 2019, and PCT / US2020 / 015082, filed January 24, 2020 (which are incorporated herein by reference). For example, in some embodiments, a cancer assay panel can be designed to include multiple probes (or probe pairs) capable of capturing fragments that together can provide information relevant to the diagnosis of cancer. In some embodiments, a panel comprises at least 50, 100, 500, 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, 25,000, or 50,000 pairs of probes, while in other embodiments, a panel comprises at least 500, 1,000, 2,000, 5,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, or 100,000 probes. The plurality of probes together can comprise at least 100,000, 200,000, 400,000, 600,000, 800,000, 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million, or 10 million nucleotides. A probe (or probe pair) is specifically designed to target one or more genomic regions that are differentially methylated in cancer samples and non-cancer samples. Target genomic regions can be selected according to a size budget (determined by sequencing budget and desired sequencing depth) to maximize classification accuracy.

[0267] Samples enriched using a cancer assay panel can be subjected to targeted sequencing. Samples enriched using a cancer assay panel can be used to generally detect the presence or absence of cancer and / or provide a cancer classification, such as cancer type, cancer stage, such as I, II, III, or IV, or provide the tissue of origin from which the cancer is believed to originate. Depending on the purpose, the panel can include probes (or probe pairs) targeting genomic regions that are differentially methylated between general cancer (pan-cancer) samples and non-cancerous samples, or only in cancerous samples with a specific cancer type (e.g., lung cancer-specific target). Specifically, cancer assay panels are designed based on bisulfite sequencing data generated from cell-free DNA (cfDNA) or genomic DNA (gDNA) from cancer and / or non-cancerous individuals.

[0268] In some embodiments, a cancer assay panel designed by the methods provided herein comprises at least 1,000 pairs of probes, each pair comprising two probes configured to overlap each other by an overlapping sequence comprising a 30-nucleotide fragment. The 30-nucleotide fragment comprises at least five CpG sites, and at least 80% of the at least five CpG sites are either CpG or UpG. The 30-nucleotide fragment is configured to bind to one or more genomic regions in a cancerous sample, and the one or more genomic regions have at least five methylation sites with an aberrant methylation pattern. Another cancer assay panel comprises at least 2,000 probes, each of which is designed as a hybridization probe complementary to one or more genomic regions. Each genomic region is selected based on the criteria that it (i) comprises at least 30 nucleotides and (ii) at least five methylation sites, and the at least five methylation sites have an aberrant methylation pattern and are either hypomethylated or hypermethylated.

[0269] Each probe (or probe pair) is designed to target one or more target genomic regions. The target genomic regions are selected based on several criteria designed to increase the selective enrichment of relevant cfDNA fragments while reducing noise and non-specific binding. For example, a panel can include probes that can selectively bind to and enrich differentially methylated cfDNA fragments in cancerous samples. In this case, sequencing of the enriched fragments can provide information relevant to the diagnosis of cancer. Furthermore, probes can be designed to target genomic regions determined to have abnormal methylation patterns and / or hypermethylated or hypomethylated patterns to provide additional selectivity and specificity of detection. For example, genomic regions can be selected if they have a methylation pattern with a low p-value according to a Markov model trained on a set of non-cancerous samples that additionally covers at least five CpGs, 90% of which are either methylated or unmethylated. In other embodiments, genomic regions can be selected using a mixture model, as described herein.

[0270] Each probe (or probe pair) can target a genomic region containing at least 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 60 bp, 70 bp, 80 bp, or 90 bp. Genomic regions can be selected by containing less than 20, 15, 10, 8, or 6 methylation sites. Genomic regions can be selected if at least 80, 85, 90, 92, 95, or 98% of at least five methylation (e.g., CpG) sites are either methylated or unmethylated in non-cancerous or cancerous samples.

[0271] The genomic regions can be further filtered to select only those that are likely to be informative based on their methylation patterns, for example, CpG sites that are differentially methylated between cancerous and non-cancerous samples (e.g., aberrantly methylated or unmethylated in cancer versus non-cancer). For selection, a calculation can be performed for each CpG site. In some embodiments, a first count is determined, which is the number of cancer-containing samples containing fragments overlapping with the CpG (cancer_count), and a second count is determined, which is the number of total samples containing fragments overlapping with the CpG (total count). Genomic regions can be selected based on a criterion that is positively correlated with the number of cancer-containing samples containing fragments overlapping with the CpG (cancer_count) and inversely correlated with the number of total samples containing fragments overlapping with the CpG (total count).

[0272] In one embodiment, the number of non-cancerous samples having fragments overlapping the CpG site (n non-cancer ) and the number of cancerous samples (n cancer ) is counted. The probability that the sample is cancer is then calculated, for example, as (n cancer +1) / (n cancer +n non-cancer +2). CpG sites by this metric are ranked and greedily added to the panel until the panel size budget is exhausted.

[0273] Which samples are used for cancer counting can vary depending on whether the assay is intended to be a pan-cancer or single-cancer assay, or what kind of flexibility is desired in choosing which CpG sites contribute to the panel. A panel for diagnosing a specific cancer type (e.g., TOO) can be designed using a similar process. In this embodiment, for each cancer type and for each CpG site, information gain is calculated to determine whether to include a probe targeting that CpG site. Information gain is calculated for samples with a given cancer type compared to all other samples. For example, there are two random variables, "AF" and "CT." "AF" is a binary variable indicating whether there is an abnormal fragment overlapping a specific CpG site in a particular sample (yes or no). "CT" is a binary random variable indicating whether the cancer is of a particular type (e.g., lung cancer or a cancer other than lung cancer). Mutual information can be calculated for "CT" given "AF." That is, knowing whether there are any informative fragments that overlap with a particular CpG site provides how many bits of information about the cancer type (lung vs. non-lung in the example). This can be used to rank CpGs based on how specific they are to a particular cancer type (e.g., TOO). This procedure is repeated for multiple cancer types. For example, if a particular region is generally differentially methylated only in lung cancer (but not other cancer types or non-cancer), CpGs in that region will tend to have high information gain for lung cancer. For each cancer type, CpG sites were ranked by this information gain metric and then greedily added to the panel until the size budget for that cancer type was exhausted.

[0274] Further filtration can be performed to select target genomic regions with less than a threshold number of off-target genomic regions. For example, a genomic region is selected only if there are less than 15, 10, or 8 off-target genomic regions. In other cases, filtration is performed to remove genomic regions when the sequence of the target genomic region appears more than 5, 10, 15, 20, 25, or 30 times in the genome. Further filtration can be performed to select target genomic regions when sequences with 90%, 95%, 98%, or 99% homology to the target genomic region appear less than 15, 10, or 8 times in the genome, or to remove target genomic regions when sequences with 90%, 95%, 98%, or 99% homology to the target genomic region appear more than 5, 10, 15, 20, 25, or 30 times in the genome. This is to eliminate repetitive probes that may pull down undesired off-target fragments, which may affect assay efficiency.

[0275] In some embodiments, it has been demonstrated that a fragment-probe overlap of at least 45 bp is necessary to achieve a significant amount of pull-down (although this number may vary depending on the assay details). Furthermore, it has been suggested that a mismatch rate of more than 10% between the probe and fragment sequences in the overlap region is sufficient to significantly disrupt binding and therefore pull-down efficiency. Therefore, sequences that can align to the probe along at least 45 bp with at least a 90% match rate are candidates for off-target pull-down. Thus, in one embodiment, the number of such regions is scored. The best probes have a score of 1, meaning that they match in only one location (the intended target region). Probes with low scores (e.g., less than 5 or 10) are acceptable, but any probes with scores above this are discarded. Other cutoff values ​​can be used for specific samples.

[0276] In various embodiments, the selected target genomic region can be located within a variety of locations in the genome, including, but not limited to, exons, introns, intergenic regions, and other portions. In some embodiments, probes can be added that target non-human genomic regions, such as those that target viral genomic regions.

[0277] VIII. Kit Implementation Also disclosed herein are kits for carrying out the above-described methods, including those related to cancer classifiers. The kits may include one or more collection containers for collecting a sample containing genetic material from an individual. The sample may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. Such kits may include reagents for isolating nucleic acids from the sample. The reagents may further include reagents for sequencing the nucleic acid, including buffers and detection agents. In one or more embodiments, the kit may include one or more sequencing panels containing probes for targeting specific genomic regions, specific mutations, specific genetic variants, or any combination thereof. In one or more embodiments, the kit includes at least one panel containing contamination-targeting probes. In other embodiments, samples collected via the kit are provided to a sequencing laboratory, which can sequence the nucleic acids in the sample using the sequencing panels.

[0278] The kit may further include instructions for using the reagents included in the kit. For example, the kit may include instructions for collecting a sample and extracting nucleic acid from a test sample. Exemplary instructions may include the order in which to add reagents, the centrifugation speed used to isolate nucleic acid from a test sample, a method for amplifying nucleic acid, a method for sequencing nucleic acid, or any combination thereof. The instructions may also describe how to operate a computing device as analysis system 200 to perform any of the steps of the described methods.

[0279] In addition to the above components, the kit may include a computer-readable storage medium storing computer software for carrying out the various methods described throughout this disclosure. One form in which these instructions may reside is as printed information on a suitable medium or substrate, such as one or more sheets of paper printed with the kit packaging or printed information within the packaging. Yet another alternative may be a computer-readable medium, such as a diskette, CD, hard drive, or network data storage device, on which the instructions are stored in the form of computer code. Yet another alternative may be a website address that can be used via the Internet to access information about the site to be removed.

[0280] IX. Cancer Uses In some embodiments, the methods, analysis systems, and / or classifiers of the present disclosure can be used to detect the presence (or absence) of cancer, monitor the progression or recurrence of cancer, monitor treatment response or effectiveness, determine the presence or presence of minimal residual disease (MRD), or monitor MRD, or any combination thereof. In some embodiments, the analysis systems and / or classifiers can be used to identify the tissue or origin of the cancer. For example, the systems and / or classifiers can be used to identify the cancer as any of the following cancer types: head and neck cancer, liver / bile duct cancer, upper GI cancer, pancreatic / gallbladder cancer, colorectal cancer, ovarian cancer, lung cancer, multiple myeloma, lymphoid neoplasm, melanoma, sarcoma, breast cancer, and uterine cancer. For example, as described herein, a classifier can be used to generate a likelihood or probability score (e.g., 0 to 100) that a sample feature vector is from a subject with cancer. In some embodiments, the probability score is compared to a threshold probability to determine whether the subject has cancer. In other embodiments, the likelihood or probability score can be evaluated at different time points (e.g., before or after treatment) to monitor disease progression or to monitor treatment effectiveness (e.g., therapeutic efficacy). In still other embodiments, the likelihood or probability score can be used to make or influence clinical decisions (e.g., diagnosing cancer, selecting a treatment, assessing treatment effectiveness, etc.). For example, in one embodiment, if the likelihood or probability score exceeds a threshold, a physician can prescribe an appropriate treatment. In some embodiments, a test report can be generated to provide a patient with test results including, for example, a probability score that the patient has a disease state (e.g., cancer), a type of disease (e.g., type of cancer), and / or a tissue of origin of the disease (e.g., tissue of origin of the cancer).

[0281] IX.A. Early Cancer Detection In some embodiments, the methods and / or classifiers of the present disclosure are used to detect the presence or absence of cancer in a subject suspected of having cancer. For example, a classifier (described herein) can be used to determine the likelihood or probability score that a sample feature vector is from a subject with cancer.

[0282] In one embodiment, a probability score of 60 or greater may indicate that the subject has cancer. In yet other embodiments, a probability score of 65 or greater, 70 or greater, 75 or greater, 80 or greater, 85 or greater, 90 or greater, or 95 or greater may indicate that the subject has cancer. In other embodiments, the probability score may indicate the severity of the disease. For example, a probability score of 80 may indicate a more severe form or later stage of cancer compared to a score below 80 (e.g., a score of 70). Similarly, an increase in the probability score over time (e.g., at a second later time point) may indicate disease progression, or a decrease in the probability score over time (e.g., at a second later time point) may indicate successful treatment.

[0283] In another embodiment, a cancer log odds ratio can be calculated for a test subject by taking the log of the ratio of the probability of being cancerous to the probability of being non-cancerous (i.e., 1 minus the probability of being cancerous), as described herein. According to this embodiment, a cancer log odds ratio greater than 1 may indicate that the subject has cancer. In yet other embodiments, a cancer log odds ratio greater than 1.2, greater than 1.3, greater than 1.4, greater than 1.5, greater than 1.7, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4 indicates that the subject has cancer. In other embodiments, the cancer log odds ratio may indicate the severity of the disease. For example, a cancer log odds ratio greater than 2 may indicate a more severe form or later stage of cancer compared to a score less than 2 (e.g., a score of 1). Similarly, an increase in the cancer log odds ratio over time (e.g., at a second later time point) can indicate disease progression, or a decrease in the cancer log odds ratio over time (e.g., at a second later time point) can indicate successful treatment.

[0284] According to aspects of the present disclosure, the methods and systems of the present disclosure can be trained to detect or classify multiple cancer indications. For example, the methods, systems, and classifiers of the present disclosure can be used to detect the presence of one or more, two or more, three or more, five or more, or ten or more different types of cancer.

[0285] In some embodiments, the cancer is one or more of head and neck cancer, liver / bile duct cancer, upper GI cancer, pancreatic / gallbladder cancer, colorectal cancer, ovarian cancer, lung cancer, multiple myeloma, lymphoid neoplasms, melanoma, sarcoma, breast cancer, and uterine cancer.

[0286] IX.B. Cancer and Treatment Monitoring In certain embodiments, the first time point is before cancer treatment (e.g., before resection surgery or therapeutic intervention) and the second time point is after cancer treatment (e.g., after resection surgery or therapeutic intervention), and the method is utilized to monitor the effectiveness of the treatment. For example, if the second likelihood or probability score is decreased compared to the first likelihood or probability score, the treatment is deemed successful. However, if the second likelihood or probability score is increased compared to the first likelihood or probability score, the treatment is deemed unsuccessful. In other embodiments, both the first and second time points are before cancer treatment (e.g., before resection surgery or therapeutic intervention). In yet other embodiments, both the first and second time points are after cancer treatment (e.g., before resection surgery or therapeutic intervention), and the method is utilized to monitor the effectiveness of the treatment or loss of effectiveness of the treatment. In yet other embodiments, cfDNA samples can be obtained from a cancer patient at first and second time points and analyzed, for example, to monitor the progression of the cancer, to determine whether the cancer is in remission (e.g., after treatment), to monitor or detect residual disease or disease recurrence, or to monitor the effectiveness of a treatment (e.g., a therapeutic drug).

[0287] Those skilled in the art will readily appreciate that test samples can be obtained from a cancer patient over any desired set of time points and analyzed according to the methods of the present disclosure to monitor the cancer status in the patient. In some embodiments, the first and second time points are from about 15 minutes up to about 30 years, e.g., about 30 minutes, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or about 24 hours, e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, or about 30 days, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months, or e.g., about 1, 1.5, 2, 2.5, 3, 3.5, 20, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5, or about 30 years. In other embodiments, test samples may be obtained from the patient at least once every three months, at least once every six months, at least once every year, at least once every two years, at least once every three years, at least once every four years, or at least once every five years.

[0288] IX.C. Processing In yet another embodiment, information (e.g., likelihood or probability score) obtained from any of the methods described herein can be used to make or influence clinical decisions (e.g., diagnosing cancer, selecting a treatment, assessing the effectiveness of a treatment, etc.). For example, in one embodiment, if the likelihood or probability score exceeds a threshold, a physician can prescribe an appropriate treatment (e.g., resective surgery, radiation therapy, chemotherapy, and / or immunotherapy). In some embodiments, information such as the likelihood or probability score can be provided as a readout to a physician or subject.

[0289] A classifier (described herein) can be used to determine a likelihood or probability score that a sample feature vector is from a subject with cancer. In one embodiment, if the likelihood or probability exceeds a threshold, an appropriate treatment (e.g., resective surgery or therapy) is prescribed. For example, in one embodiment, if the likelihood or probability score is 60 or greater, one or more appropriate treatments are prescribed. In another embodiment, if the likelihood or probability score is 65 or greater, 70 or greater, 75 or greater, 80 or greater, 85 or greater, 90 or greater, or 95 or greater, one or more appropriate treatments are prescribed. In other embodiments, a cancer log-odds ratio can indicate the effectiveness of a cancer treatment. For example, an increase in the log-odds ratio of cancer over time (e.g., at a second time point after treatment) can indicate that the treatment was ineffective. Similarly, a decrease in the log-odds ratio of cancer over time (e.g., at a second time point after treatment) can indicate successful treatment. In another embodiment, if the cancer log odds ratio is greater than 1, greater than 1.5, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4, one or more appropriate treatments are prescribed.

[0290] In some embodiments, the treatment is one or more cancer therapeutic agents selected from the group including chemotherapeutic agents, targeted cancer therapeutic agents, differentiation therapeutic agents, hormonal therapeutic agents, and immunotherapeutic agents. For example, the treatment can be one or more chemotherapeutic agents selected from the group including alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeleton-disrupting agents (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapeutic agents selected from the group including signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapeutic agents, including retinoids, such as tretinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormonal therapy agents selected from the group including antiestrogens, aromatase inhibitors, progestins, estrogens, antiandrogens, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapeutic agents selected from the group including monoclonal antibody therapy, e.g., rituximab (RITUXAN) and alemtuzumab (CAMPATH), non-specific immunotherapy and adjuvants, e.g., BCG, interleukin-2 (IL-2), and interferon-alpha, immunomodulatory agents, e.g., thalidomide and lenalidomide (REVLIMID). It is within the ability of a skilled artisan or oncologist to select an appropriate cancer therapy agent based on characteristics such as tumor type, stage of cancer, previous exposure to cancer treatments or therapeutic agents, and other characteristics of the cancer. [Example]

[0291] X. Working Example XA Example 1 - Whole Genome Bisulfite Sequencing (WBGS) First CCGA Substudy: The data shown in Figures 7A–7F were obtained from the first CCGA substudy. Here, for plasma cfDNA extraction, training data blood samples (N = 1,785) were collected from untreated individuals diagnosed with cancer (including 20 tumor types and all stages of cancer) and healthy individuals without a cancer diagnosis (controls). Another set of blood samples (N = 1,010) was collected and used for validation. Unless otherwise indicated, cell-free DNA (cfDNA) and genomic DNA (gDNA) extracted from the first CCGA substudy samples were subjected to whole-genome bisulfite sequencing.

[0292] In the classification process, analysis system 200 treats fragment methylation states as drawn from a mixture of potential methylation patterns. Analysis system 200 assigns to observed fragments a relative probability of originating from a particular cancer tissue of origin.

[0293] More specifically, as described herein, a probabilistic model was fitted to sequence reads from multiple regions (or windows) from each cancer type (and for non-cancer or healthy samples). In this case, a mixture model was used, where each mixture component was an independent site model (methylation at each CpG is independent of methylation at other CpGs). The model was fitted using maximum likelihood estimation to identify a set of parameters that maximizes the total log-likelihood of all fragments from one cancer type (or non-cancer).

[0294] For each region and each cancer type pair (including non-cancer as the negative type), a multinomial logistic regression classifier was trained using the best-performing hierarchy. For each sample (regardless of label), in each region, for each cancer type, for each fragment, the log-likelihood ratio was calculated as described above, and for each set of "hierarchy" values, R cancer type The number of fragments per stratum was quantified. The quantified reads per stratum were binarized and used as features to train a classifier.

[0295] Finally, where indicated, to generate predictions for unknown samples, feature values ​​were determined (as described above) and the generated features were used to generate predictions of cancer and / or tissue of origin utilizing a trained multinomial logistic regression classifier.

[0296] 7A, 7B, and 7C include confusion matrices showing the accuracy of a classifier, according to various embodiments. In some embodiments, analysis system 200 uses a confusion matrix to determine the accuracy of a classifier. The confusion matrix contains information describing the success rate for a classifier in identifying each disease state.

[0297] As shown in FIG. 7A, matrix 710 includes example performance of a multinomial model-based classifier trained using a set of cfDNA samples (no tissue samples). Matrix 720 includes example performance of a mixture model-based classifier trained by analysis system 200 using the same set of cfDNA samples. Scores along the diagonal of the matrix indicate correct predictions, i.e., where the predicted tissue of origin for a fragment matches the true tissue of origin. Compared to the baseline multinomial model-based classifier, the mixture model-based classifier has a higher overall accuracy in predicting the presence of the cancer types shown in the matrix.

[0298] The samples in the training set can be filtered based on one or more criteria (e.g., a particular specificity level). For example, the training set includes samples that were determined to have cancer based on 98% specificity according to the m-score. The remaining (e.g., 2%) non-cancer samples that were (incorrectly) identified as having cancer were excluded from display in the confusion matrix for clarity.

[0299] As shown in Figure 7B, matrix 730 contains example performance of a classifier based on a mixture model trained using a cross-validation training set of cfDNA samples (no tissue samples), and matrix 740 contains example performance of a classifier based on a mixture model trained using a cross-validation training set of cfDNA and tissue samples.

[0300] As shown in FIG. 7C, matrix 750 includes example performance of a mixture model-based classifier trained using a set of cfDNA samples (no tissue samples) from a clinical study titled Circulating Cell-free Genome Atlas Study ("CCGA"). Matrix 740 includes example performance of a mixture model-based classifier trained using a set of cfDNA and tissue samples from CCGA. The CCGA study was described by Clinical Trial.gov Identifier: NCT02889978 (https: / / www.clinicaltrials.gov / ct2 / show / NCT02889978).

[0301] XB Example 2 - Classification of cancers using targeted bisulfite sequencing from early breakouts of the second CCGA substudy Second CCGA Substudy: The data shown in Figures 9A-9B, 10A-10B, 11, and 12 were obtained from an early breakout from the second CCGA substudy. Here, training data blood samples (N=3,132) were collected from untreated individuals diagnosed with cancer (including 20 tumor types and all stages of cancer) and healthy individuals without a cancer diagnosis (controls) for plasma cfDNA extraction. A separate set of blood samples (N=1,354) was collected and used for validation. In some embodiments, where indicated, the training set also included training data from tissue samples (i.e., gDNA). To determine the analysis population, the training data blood samples were filtered based on several factors. For example, 105 samples were excluded as clinically unlocked; 11 samples were excluded based on eligibility criteria; 58 samples were excluded for undetermined cancer or treatment status (unevaluable); 4 untreated samples and 72 unevaluable assays were excluded (unanalyzable); 581 samples were saved for future analysis. The resulting analytical population of 2,301 samples included 1,422 cancer samples and 879 non-cancer samples.

[0302] Participant demographics for individuals in the sub-studies are shown in Table 1 below.

[0303] [Table 1]

[0304] To identify cancer- and tissue-defining methylation signals, extracted cfDNA was subjected to bisulfite sequencing assays targeting the most informative regions of the methylome, as identified from the GRAIL-proprietary whole-genome bisulfite sequencing assay and methylation database.

[0305] We used a methylation database that examines genome-wide fragment-level methylation patterns across 811 cancer cell methylomes representing 21 tumor types (97% of SEER cancer incidence). To generate a methylation database of cancer-defining methylation signals, genomic DNA from formalin-fixed, paraffin-embedded (FFPE) tumor tissues and isolated cells from tumors was subjected to whole-genome bisulfite sequencing. The methylation database was used for panel design and training to optimize classifier performance, as described herein. A large cancer and non-cancer methylation sequence database was generated to classify multiple cancers with high specificity and enable target selection for a single test capable of identifying the tissue of origin.

[0306] Target selection and panel design: Target genomic regions were selected using the methylation sequence database from the CCGA study, as described herein. Specifically, the cfDNA sequences in the database were filtered based on p-value using the non-cancer distribution, and only fragments with p<0.001 were retained. The selected cfDNA was further filtered to retain only those that were at least 90% methylated or 90% unmethylated. Next, for each CpG site in the selected fragment, the number of cancer or non-cancer samples containing fragments overlapping with the CpG site was counted. Specifically, P(cancer|overlapping fragment) was calculated for each CpG, and genomic sites with high P-values ​​were selected as general cancer targets. By design, the selected fragments had very low noise (i.e., there were few overlapping non-cancer fragments).

[0307] To find cancer type-specific targets, a similar selection process was performed: CpG sites were ranked based on their information gain, and one cancer type was compared to all other samples (i.e., non-cancer + other cancer types).

[0308] A cancer assay panel comprising probes targeting selected genomic regions was generated as described herein. Specifically, the panel was designed to detect the presence of cancer generally (i.e., versus non-cancer) or a specific type of cancer (e.g., TOO). The panel includes a probe set targeting each of the selected genomic regions.

[0309] Probes were designed to overlap any of the CpG sites contained within the start / stop range of any of the target regions (eg, informative fragments).

[0310] Classification: In the classification process, analysis system 200 treats fragment methylation states as drawn from a mixture of potential methylation patterns. Analysis system 200 assigns observed fragments a relative probability of originating from cancer. For tissue-of-origin classification, analysis system 200 assigns observed fragments a relative probability of originating from a particular tissue. Analysis system 200 combines fragments characteristic of cancer and tissue-of-origin across target regions to classify cancer vs. non-cancerous and / or identify tissue-of-origin. For binary cancer classification, analysis system 200 estimates sensitivity at 99% specificity.

[0311] More specifically, as described in Example VI.a, a probabilistic model was fitted to sequence reads derived from multiple regions (or windows) from each cancer type (and for non-cancerous or healthy samples), the identified features, and a trained multinomial logistic regression classifier. To generate predictions for unknown samples, feature values ​​were determined (as described above), and the generated features were used to generate predictions for cancer and / or tissue of origin utilizing a trained multinomial logistic regression classifier.

[0312] Figures 9A and 9B show the sensitivity of the tissue-of-origin classifier generated by the methods described in this disclosure. Sensitivity is reported at 99% specificity, with 95% confidence intervals shown. Figure 9A shows model predictions for a pre-specified list of cancers. Figure 9B shows model predictions for other cancers included in the CCGA study. Demographic information alone (baseline modeling) correctly classified <5% of participants. For the pre-specified list of cancers (anorectal, breast [HR-negative], colorectal, esophageal, gastric, head and neck, hepatobiliary, lung, lymphoid neoplasms [chronic lymphocytic leukemia, lymphoma], multiple myeloma, ovarian, and pancreatic), the overall sensitivity was 76.1% (95% CI: 73.1–78.9%). Sensitivity for early stage (I–III) cancers in this cohort was 68.8% (95% CI: 64.8–72.6%). The overall sensitivity was 55.1% (95% CI: 52.5-57.7%) across all cancer types and stages. For early stage (I-III) cancers, the sensitivity was 43.8% (95% CI: 40.7-46.8%).

[0313] Figures 10A and 10B show the sensitivity of the tissue-of-origin classifier at different cancer stages. As indicated in the legend, sensitivity by individual stage for all pre-specified cancers of interest is reported at 99% specificity. Numbers in boxes represent the total number of samples included at each stage. 95% confidence intervals are shown. "Lymphoid neoplasms" include lymphomas (stages I-IV) and chronic lymphocytic leukemia (unstaged, included as "NI").

[0314] Figure 11 shows a performance grid depicting the accuracy of tissue-of-origin localization. Using the tissue-of-origin classifier with the methylation database for stage I-IV samples, there is agreement between the true (x-axis) and predicted (y-axis) tissue-of-origin per sample. The slope of the graph corresponds to the percentage of predicted tissue-of-origin (y-axis) that were correct (x-axis). Analysis showed that the accuracy of tissue-of-origin localization (fraction of all TOO predictions that were correct) was higher with the methylation database (p = 0.0066). This was consistent in predicting stages I-III: 89.9% (384 / 427), as further demonstrated in Table 2.

[0315] [Table 2]

[0316] An effective multi-cancer test should ideally simultaneously detect clinically significant cancers across stages with very high specificity (and thus have a single, fixed, low false positive rate) and accurately determine tissue of origin. To demonstrate the potential of this approach, simultaneous detection (reported sensitivity at 99% specificity) and tissue of origin determination for a pre-specified list of cancer types at individual stages are summarized in Figure 12. Figure 12 thus shows the accuracy and sensitivity of the tissue of origin classifier at various cancer stages.

[0317] Figures 13A and 13B show receiver operating characteristic (ROC) curves for the tissue-of-origin classifier, which demonstrate classifier performance with a specificity of 99%, a sensitivity of 55% for all cancers, and a sensitivity of 76% for multiple cancers.

[0318] These data demonstrate that a classification method using targeted methylation features can simultaneously detect multiple cancer types with a specificity (99%) appropriate for population screening. Detection of multiple cancers was achieved with a single, fixed, low false-positive rate. This approach also precisely localizes the tissue of origin, simplifying downstream diagnostic workup. Additionally, incorporating data from a large methylation database improved classifier performance.

[0319] Together, this supports the potential clinical applicability of the method described in this disclosure as an early multi-cancer detection test for a number of clinically important cancer types.

[0320] XC Example 3 - Cancer Classification Using Targeted Bisulfite Sequencing from the Complete Second CCGA Substudy Generation of Mixture Model Classifier: To maximize performance, the predictive cancer model described in this example was trained using sequence data from multiple samples from known cancer types and non-cancer samples from both CCGA substudies (CCGA1 and CCGA2), multiple tissue samples for known cancers from CCGA1, and multiple non-cancer samples from the STRIVE study (see Clinical Trail.gov Identifier: NCT03085888 (http: / / clinicaltrials.gov / ct2 / show / NCT03085888)). The STRIVE study is a prospective, multicenter, observational cohort study to validate assays for the early detection of breast cancer and other invasive cancers that provided additional non-cancer training samples to train the classifier described herein. Known cancer types included from the CCGA sample set included: breast, lung, prostate, colorectal, kidney, uterine, pancreatic, esophageal, lymphoma, head and neck, ovarian, hepatobiliary, melanoma, cervical, multiple myeloma, leukemia, thyroid, bladder, stomach, and anorectal. Thus, a model can be a multi-cancer model (or multi-cancer classifier) ​​for detecting one or more, two or more, three or more, four or more, five or more, ten or more, or twenty or more different types of cancer. 4,841 participants (2,836 cancer; 2,005 non-cancer) from the CCGA study and 2,202 non-cancer participants from the STRIVE study were included in this pre-specified analysis. Of these, 3,133 samples from CCGA were assigned to training (1,742 cancer; 1,391 non-cancer) and 1,354 to validation (740 cancer, 614 non-cancer). 1,587 samples from STRIVE were assigned to training and 615 to validation. Participant allocation is shown. Overall, 3,052 samples in training (1,531 cancer; 1,521 non-cancer) and 1,264 samples in validation (654 cancer; 610 non-cancer) were analyzable in the pre-specified primary analysis population.Additional details of the CCGA2 substudy and the analyses detailed in this example were described in an article in the Annals of Oncology, titled "Sensitive and specific multi-cancer detection and localization using methylation signatures in cell-free DNA," published online on March 30, 2020 (https: / / www.annalsofoncology.org / article / S0923-7534(20)36058-0 / fulltext).

[0321] The following classifier performance data was reported for the lock classifier trained on cancer and non-cancer samples from CCGA2, the CCGA substudy, and non-cancer samples from STRIVE. Individuals in the CCGA2 substudy were different from those in the CCGA1 substudy, which used cfDNA to select target genomes (described in International Publication No. WO 2019 / 195268, filed April 2, 2019; PCT / US2019 / 053509, filed September 27, 2019; and PCT / US2020 / 015082, filed January 24, 2020, which are incorporated herein by reference). From the CCGA2 study, blood samples were collected from individuals diagnosed with untreated cancer (including 20 tumor types and all stages of cancer) and healthy individuals diagnosed without cancer (controls). For STRIVE, blood samples were collected from women within 28 days of their screening mammogram. Cell-free DNA (cfDNA) was extracted from each sample and treated with bisulfite to convert unmethylated cytosine to uracil. The bisulfite-treated cfDNA was enriched for information-rich cfDNA molecules in three cancer assay panels: (1) the pan-cancer assay panel #4 described and disclosed in WO2019 / 195268 (labeled herein as assay panel A); (2) the pan-cancer assay panel #5 described and disclosed in WO2019 / 195268 (labeled herein as assay panel B); and (3) a large-scale proprietary pan-cancer assay panel (assay panel C, described below) using hybridization probes designed to enrich for bisulfite-converted nucleic acids from each of multiple target genomic regions. The enriched bisulfite-converted nucleic acid molecules were sequenced using paired-end sequencing on an Illumina platform (San Diego, CA) to obtain one set of sequence reads per training sample, and the resulting read pairs were aligned to the reference genome and assembled into fragments to identify methylated and unmethylated CpG sites.

[0322] Mixture model-based featurization For each cancer type (including non-cancer), a probabilistic mixture model was trained and used to assign a probability to each fragment from each cancer and non-cancer sample based on how likely the fragment was to be observed in a given sample type.

[0323] Fragment-level analysis Briefly, for each sample type (cancer and non-cancer samples), a probabilistic model was fitted to fragments from training samples for each cancer and non-cancer type for each region (each region was used as is if it was less than 1 kb; otherwise, it was subdivided into 1 kb-long regions (e.g., overlapping 500 base pairs) with 50% overlap between adjacent regions). The probabilistic model trained for each sample type was a mixture model, where each of the three mixture components was an independent site model, in which methylation at each CpG was assumed to be independent of methylation at other CpGs. Fragments were excluded from the model if their p-value (from the non-cancer Markov model) was greater than 0.01; if they were marked as overlapping fragments; if their bag size (for target-methylated samples only) was greater than 1; if they did not cover at least one CpG site; or if they were greater than 1000 bases in length. Retained training fragments were assigned to a region if they overlapped with at least one CpG from that region. If a fragment overlapped CpGs in multiple regions, it was assigned to all of the multiple regions.

[0324] Local Source Model Each probability model was fitted using maximum likelihood estimation to identify a set of parameters that maximized the log-likelihood of all fragments derived from each sample type, subject to a regularization penalty. Specifically, in each classification domain, a set of probability models was trained, one for each training label (i.e., one for each cancer type and one for non-cancer). Each model took the form of a Bernoulli mixture model with three components. Mathematically,

number

number

[0325] Characterization Once the probabilistic model was trained, a set of numerical features was calculated for each sample. Specifically, for each fragment from each training sample in each region, features were extracted for each cancer type and non-cancer sample. The extracted features were a count of outlier fragments (i.e., highly informative fragments), defined as those whose log-likelihood under a first cancer model exceeded the log-likelihood under a second cancer or non-cancer model by at least a threshold stratum value. The outlier fragments were separately counted by genomic region, sample model (i.e., cancer type), and stratum (for stratums 1, 2, 3, 4, 5, 6, 7, 8, and 9), resulting in nine features per region per sample type. Thus, each feature was defined by three characteristics: genomic region; a "positive" cancer type label (excluding non-cancer); and a stratum value selected from the set {1, 2, 3, 4, 5, 6, 7, 8, 9}. The numerical value of each feature was calculated as follows:

number

[0326] Feature Ranking For each set of pairwise features, we ranked the features using mutual information based on their ability to distinguish a first cancer type (which defined the log-likelihood model from which the features were derived) from a second cancer type or from non-cancer. Specifically, two ranked lists of features were compiled for each unique pair of class labels: one in which the first label was assigned as "positive" and the second label as "negative," and the other in which the positive / negative assignments were swapped (except for the "non-cancer" label, which was only allowed as a negative label). For each of these ranked lists, only features whose positive cancer type label (as in Equation (3)) matched the positive label under consideration were included in the ranking. For each such feature, the fraction of training samples with non-zero feature values ​​was calculated separately for the positive and negative labels. Features for which this fraction was greater in the positive label were ranked by mutual information for the class-label pair.

[0327] The top 256 features from each pairwise comparison were identified and added to the final feature set for each cancer type and non-cancer. To avoid redundancies, if multiple features were selected from the same positive type and genomic region (i.e., for multiple negative types), only the one assigned the lowest (most informative) rank for that cancer type pair was kept, and the relationship was broken by selecting a higher hierarchical value. Features in the final feature set for each sample (cancer type and non-cancer) were binarized (any feature value greater than 0 was set to 1, so all features were either 0 or 1).

[0328] Classifier Training The training samples were then split into separate 5-fold cross-validation training sets, and in each case a two-stage classifier was trained for each convolution, training on 4 / 5 of the training samples and using the remaining 1 / 5 for validation.

[0329] In the first stage of training, we trained a binary (two-class) logistic regression model to detect the presence of cancer and distinguish cancer samples from non-cancer samples (regardless of TOO). When training this binary classifier, we assigned a sample weight to male non-cancer samples to counteract gender imbalance in the training set. For each sample, the binary classifier outputs a predicted score indicating the likelihood of cancer presence or absence.

[0330] In the second stage of training, a parallel multi-class logistic regression model for determining cancer tissue of origin was trained with TOO as the target label. Only cancer samples that received a score above the 95th percentile of non-cancer samples in the first stage classifier were included in training this multi-class classifier. For each cancer sample used in training and the multi-class classifier, the multi-class classifier outputs a prediction for the cancer type to be classified, where each prediction is the likelihood that a given sample has a particular cancer type. For example, the cancer classifier can return a cancer prediction for a test sample, including a prediction score for breast cancer, a prediction score for lung cancer, and / or a prediction score for no cancer.

[0331] Both binary and multiclass classifiers were trained by mini-batch stochastic gradient descent, and in each case, training was stopped early when the performance of the validation convolution (assessed by cross-entropy loss) began to decline. To predict samples outside the training set, at each stage, the scores assigned by the five cross-validated classifiers were averaged. Scores assigned to sex-inappropriate cancer types were set to 0, and the remaining values ​​were renormalized to sum to 1.

[0332] The scores assigned to the validation convolutions in the training set are retained for use in assigning cutoff values ​​(thresholds) that target specific performance metrics.In particular, the probability scores assigned to the non-cancer samples in the training set are used to define the threshold corresponding to a specific specificity level.For example, for a desired specificity target of 99.4%, the threshold is set at the 99.4th percentile of the cross-validated cancer detection probability scores assigned to the non-cancer samples in the training set.The training samples with probability scores exceeding the threshold are called cancer positive.

[0333] Each training sample that tested positive for cancer was then assigned a TOO or cancer type assessment from the multi-class classifier. First, a multi-class logistic regression classifier assigned a set of probability scores (one for each predicted cancer type) to each sample. The confidence of these scores was then assessed as the difference between the highest and second-highest scores assigned by the multi-class classifier for each sample. The cross-validated training set scores were then used to identify a minimum threshold such that 90% of the cancer samples in the training set whose top two scores differed by more than the threshold were assigned the correct TOO label as their highest score. Thus, the scores assigned to the validation convolution during training were further used to determine a second threshold for distinguishing between confident and uncertain TOO calls.

[0334] During prediction, samples that received a score from the binary (first-stage) classifier below a predefined specificity threshold were assigned a "non-cancer" label. For the remaining samples, those whose top-two TOO score difference from the second-stage classifier was below a second predefined threshold were assigned an "indeterminate cancer" label. The remaining samples were assigned the cancer label with the highest score assigned by the TOO classifier.

[0335] Classifier performance of targeted genomic region panels The discriminatory value of the target genomic regions in assay panels A-C was evaluated by testing the ability of the cancer classifier to detect cancer and any of 20 different cancer types according to the methylation status of these target genomic regions. For assay panels A-B, performance was evaluated across the entire training set of 1,531 cancer samples and 1,521 non-cancer samples used to train the classifier, as shown in Table 1. For assay panel C, performance was evaluated using 1,264 samples (654 cancer; 610 non-cancer) in validation against a classifier trained using the same set of 3,052 samples (1,531 cancer; 1,521 non-cancer) used to train assay panels A-B. For each sample, differentially methylated cfDNA was enriched using a bait set containing all of the target genomic regions included in assay panels A-C. The classifier was then constrained to provide a cancer diagnosis based solely on the methylation status of the target genomic regions in the list being evaluated. As previously described in this example, a two-stage classifier embodiment was trained with TOO as the target label, including a binary (two-class) logistic regression classifier model to detect the presence of cancer, trained to distinguish cancer samples from non-cancerous samples (regardless of TOO), and a second stage in which a multi-class logistic regression classifier model was trained to determine the tissue of origin of the cancer. Also, as previously described, both classifier models were trained and validated using model-based featurization.

[0336] [Table 3]

[0337] Assay Panels A and B: The results of classifier performance analysis for assay panels A and B are shown in Figures 26A and 27A. In each figure, part A is the receiver operating curve (ROC) showing true and false positive results for cancer or non-cancer determination. The asymmetric shape of these ROC curves indicates that the classifiers were designed to minimize false positive results. The area under the curve for assay panels A and B was 0.83 for both assay panels.

[0338] A determination of cancer type (i.e., TOO) was made using the classifier for all samples that tested positive for cancer. Figures 26B and 27B contain confusion matrices showing the accuracy of TOO accuracy for assay panels A and B, respectively. The confusion matrices contain information describing the success rate for the classifiers in identifying each of the cancer types and in ruling out indeterminate cancer calls.

[0339] As shown in Figures 26B and 27B, the TOO confusion matrices demonstrate the performance for the multiclass logistic regression classifier, as described above. The agreement between the actual tissue of origin (x-axis) and the predicted tissue of origin (y-axis) per sample using the targeted methylation classifier is shown. Scores along the diagonal of the matrix indicate correct predictions, i.e., where the predicted tissue of origin for a fragment matches the true tissue of origin. As shown in Figure 26B, cancer assay panel A had a TOO accuracy of approximately 90.8% (711 / 783) when excluding indeterminate cancer calls. Figure 27B also shows that assay panel B had a TOO accuracy of approximately 90.3% (705 / 781) when excluding indeterminate cancer calls.

[0340] The results of these classifiers are further summarized in Tables 2-3, which show the accuracy of cancer detection and cancer type determination at a specificity of 0.990, indicating a false positive rate of 1%. The results are plotted by cancer stage. The results show improved cancer detection and cancer type determination for samples from individuals with late-stage cancer (e.g., stage II) compared to samples from individuals with early-stage cancer (e.g., stage III). For all cancer stages (without separation by stage), cancer type determination was approximately 89% accurate for both assay panels A and B (including indeterminate cancer calls).

[0341] Table 2. Classification accuracy using genomic regions of assay panel A. Data for cancer presence and cancer type at a specificity of 0.990 show the accuracy percentage, 95% confidence interval in parentheses, and number correctly assigned relative to the total in parentheses.

[0342] [Table 4]

[0343] Table 3. Classification accuracy using genomic regions of assay panel B. Data for the presence of cancer and cancer type at a specificity of 0.990 show the accuracy percentage, 95% confidence interval in parentheses, and number correctly assigned relative to the total in parentheses.

[0344] [Table 5]

[0345] Assay Panel C: As described above, a third large, proprietary pan-cancer assay panel was also tested. Assay Panel C was designed from WGBS data from the first CCGA substudy, CCGA1, using the feature selection methods disclosed in PCT / US2019 / 053509, filed September 27, 2019, and PCT / US2020 / 015082, filed January 24, 2020, which are incorporated herein by reference. The large, proprietary targeted methylation panel covered 103,456 distinct regions (17.2 Mb) and 1,116,720 CpGs. Assay panel C contained 363,033 CpGs in 68,059 regions (7.5 Mb) covered by probes targeting hypomethylated fragments, 585,181 CpGs in 28,521 regions (7.4 Mb) covered by probes targeting hypermethylated fragments, and 218,506 CpGs in 6,876 regions (2.3 Mb) targeting both types of fragments. Each aberrant target region contained 1 to 590 CpGs, with a median CpG count of 3 for hypomethylated target regions and 6 for hypermethylated target regions. CpGs were present in the following genomic regions: 193,818 (17%) within the region 1-5 kbp upstream of the transcription start site (TSS); 278,872 (24%) within promoters (<1 kbp upstream of the TSS); 500,996 (43%) within introns; 292,789 (25%) within exons; 247,752 (21%) within intron-exon boundaries; 134,144 (11%) within 5' untranslated regions; 182,174 (16%) between genes; and the remaining 1,817 (<1%) were not annotated. Percentages are based on the total number of CpGs and do not add up to 100% because each CpG may receive multiple annotations due to overlapping genes and / or transcripts.

[0346] For this evaluation, the sample was divided into a training set (n = 4,720) and an independent validation set (n = 1,969). A total of 4,316 participants (training: 3,052 [1,531 cancer: Stage I: 28%; Stage II: 25%; Stage III: 20%; Stage IV: 24%; Missing / Unexpected: 3%; 1,521 non-cancer]; validation: 1,264 [654 cancer: Stage I: 28%; Stage II: 25%; Stage III: 21%; Stage IV: 23%; Missing / Unexpected: 3%; 610 non-cancer]) were analyzable and included in the primary analysis population.

[0347] The results of the classifier performance analysis for the training and validation sets are shown in Figures 28-30. Panel A of Figure 28 shows the specificity results for both the training and validation sets, and Panel B shows the sensitivity for pre-specified cancers (a subset of 12 high-signal cancers based on the results of the first substudy and mortality data: anal, bladder, colon / rectum, esophagus, head and neck, liver / bile duct, lung, lymphoma, ovarian, pancreas, plasma cell neoplasm, and stomach) and all cancer types (>20) from stages I to IV. Panel C of Figure 28 shows the tissue of origin (TOO) accuracy results for both the training and validation sets, and Panel B shows the sensitivity for pre-specified cancers and all cancer types from stages I to IV. Figure 29 shows the TOO confusion matrix for both the training and validation sets, and Figure 30 shows the sensitivity results for pre-specified cancer types from both the training and validation sets.

[0348] In Figure 28, sensitivity (y-axis) is reported by clinical stage (x-axis) for pre-specified cancer types (left panel) for training (orange) and validation (greenish-blue), and for all cancer types (right panel). Accuracy for tissue of origin (y-axis) is reported by clinical stage (x-axis) for pre-specified cancer types (left panel) for training (orange) and validation (greenish-blue), and for all cancer types (right panel). Numbers indicate samples in the training|validation set.

[0349] As shown in Figure 28, the classifier achieved consistently high specificity between the cross-validated training set and the independent validation set (99.8% [95% CI: 99.4-99.9%] vs. 99.3% [98.3-99.8%], respectively; P = 0.095), reflecting a single, consistent false positive rate (FPR) of less than 1% across all 20 cancer types. Specificity in the validation set was similar for CCGA and STRIVE non-cancer samples (99.3% [97.4-99.9%] vs. 99.4% [97.9-99.9%], respectively), confirming that performance was not biased by site or sample selection. Sensitivity was consistent in the training and validation sets. For all cancers, the sensitivity for stages I–III was 44.2% (95% CI: 41.3–47.2%) vs. 43.9% (39.4–48.5%), respectively (P = 1.000). For a prespecified set of 12 high-signal cancers, the sensitivity for stages I–III was 69.8% (65.6–73.7%) vs. 67.3% (60.7–73.3%), respectively (P = 0.988). Similarly, the sensitivity for stages I–IV across all cancer types was 55.2% (52.7–57.7%) vs. 54.9% (51.0–58.8%), respectively (P = 0.897), and for prespecified cancers, it was 77.9% (75.0–80.7%) vs. 76.4% (71.6–80.7%), respectively (P = 0.573).

[0350] Additionally, as shown in Figure 28, sensitivity increased with increasing disease stage. In validation, sensitivity across prespecified cancer types was 39% (27-52%) for stage I (n=62), 69% (56-80%) for stage II (n=62), 83% (75-90%) for stage III (n=102), and 92% (86-96%) for stage IV (n=130). Among all cancer types, sensitivity was 18% (13-25%) for stage I (n=185), 43% (35-51%) for stage II (n=166), 81% (73-87%) for stage III (n=134), and 93% (87-96%) for stage IV (n=148).

[0351] Performance for individual tumor types is shown in Figure 30. Sensitivity of 99.8% specificity (training, orange) or 99.3% specificity (validation, teal) with 95% confidence intervals is reported for each cancer type with at least 50 samples. Clinical stage is shown below the plot, as are the number of samples in training and validation.

[0352] As shown in Figure 28, a pre-specified analysis of TOO accuracy (fraction of all correct TOO predictions) found that cancer-like signals in the validation set predicted TOO in 96% (344 / 359) of samples; of these, the accuracy was 93% (321 / 344). Accuracy was consistent between the training and validation sets and across disease durations. The classifier distinguished between >20 cancer types included in the study, with consistent performance within each cancer type.

[0353] Figure 29 shows confusion matrices representing the accuracy of tissue-of-origin localization in (A) the training set and (B) the validation set. The agreement between the actual tissue-of-origin (x-axis) and the predicted tissue-of-origin (y-axis) per sample using the targeted methylation classifier is shown. Color corresponds to the percentage of predicted tissue-of-origin calls. Included participants (training: n = 844, validation: n = 359) were participants with cancer who were predicted to have cancer with 99.8% specificity (training) or 99.3% specificity (validation). Tissue-of-origin calls were assigned to 95% (806 / 844) of cases during training and 96% (344 / 359) of cases during validation; calls were correct in 92% (744 / 806) of cases during training and 93% (321 / 344) of cases during validation.

[0354] XD Example 4 - Adjusting the Binary Classification Threshold According to a generalized embodiment of binary cancer classification, an analysis system determines a cancer score for the test sample based on the test sample's sequencing data (e.g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.). The analysis system compares the cancer score for the test sample to a binary threshold cutoff for predicting whether the test sample is likely to have cancer. The binary threshold cutoff can be adjusted using TOO thresholding based on one or more TOO subtype classes. The analysis system can further generate a feature vector for the test sample that is used in a multi-class cancer classifier to determine a cancer prediction indicative of one or more likely cancer types.

[0355] Figure 24A shows a confusion matrix illustrating the performance of a trained cancer classifier according to an exemplary implementation. The cancer classifier was trained according to the principles described above. TOO labels include lymphoid neoplasm, lung, kidney, non-cancer, head and neck, prostate, breast, upper gastrointestinal tract, liver and bile duct, colorectal, cervix, pancreas and gallbladder, uterus, sarcoma, bladder and urothelium, ovary, anorectal, unknown, melanoma, multiple myeloma, myeloid neoplasm, and thyroid. Notably, the classification accuracy is 89.1% for the 1,151 samples considered in this holdout set.

[0356] Figure 24B shows a confusion matrix illustrating the performance of a cancer classifier trained with additional hematological cancer subtypes. The cancer classifier was trained according to the principles described above. In contrast to Figure 24A, the TOO labels for hematological subtypes have been adjusted. In Figure 24A, the hematological subtypes include lymphoid neoplasm, multiple myeloma, and myeloid neoplasm. In Figure 24B, the hematological subtypes include Hodgkin's lymphoma (HL), NHL aggressive, NHL low-grade, myeloid, circulating lymphoma (or lymphoid), and plasma cell. Notably, the classification accuracy is 87.5% for 1,076 cases.

[0357] 25A and 25B show graphs illustrating cancer prediction accuracy for multiple cancer types across cancer stages. In this example, a cancer classifier is trained after pruning non-cancerous samples according to process 1000 described above. The analysis system determined multiple TOO thresholds for hematological subtypes. The analysis system excluded non-cancerous samples with at least one TOO probability equal to or greater than the corresponding TOO threshold for the hematological subtype. The graphs shown show classification sensitivity across various stages for the following cancer types: rectal, bladder and urothelial, breast, cervix, colorectal, head and neck, liver and bile duct, lung, melanoma, ovarian, pancreas and gallbladder, prostate, kidney, sarcoma, thyroid, upper gastrointestinal tract, and uterus. The graphs for each cancer type show the predictive sensitivity for each stage of cancer type by the first cancer classifier without TOO thresholding, labeled "locked_v1_orgi," and the second cancer classifier with TOO thresholding, labeled "v2_custom." Notably, for many cancer types, the second cancer classifier has higher predictive accuracy while maintaining narrow confidence intervals, given the larger number of samples available for validation. Of particular note is the higher predictive accuracy for many cancer types at the stage I and II levels, indicating the improved predictive potential of TOO thresholding in early-stage cancers.

[0358] XI. Additional Considerations The foregoing description of embodiments of the present disclosure has been presented for purposes of illustration; it is not intended to be exhaustive or to limit the invention to the precise form disclosed. Those skilled in the art will recognize that many modifications and variations are possible in light of the above disclosure.

[0359] Some portions of this specification describe embodiments of the present disclosure in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. While such operations are described in functional, computational, or logical terms, they will be understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Further, without loss of generality, it has proven convenient at times to refer to arrangements of these operations as modules. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.

[0360] Any of the steps, operations, or processes described herein can be performed or implemented in one or more hardware or software modules, alone or in combination with other devices. In some embodiments, the software modules are implemented in a computer program product, including a computer-readable non-transitory medium containing computer program code, that can be executed by a computer processor to perform any or all of the steps, operations, or processes described.

[0361] Embodiments may also relate to products produced by the computational processes described herein. Such products may include information resulting from the computational processes, which information may be stored on a non-transitory, tangible, computer-readable storage medium and may include any embodiment of a computer program product or other data combination described herein.

[0362] Finally, the language used herein has been selected primarily for ease of reading and descriptive purposes, and not to delineate or limit the subject matter of the present invention. Accordingly, it is intended that the scope of the invention be limited not by this detailed description, but by any claims that issue on this application based thereon. Accordingly, the disclosure of embodiments herein is intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.

Claims

1. obtaining a plurality of training samples from individuals each having one of a plurality of disease states and each associated with one of a plurality of labels for a covariate characteristic, each training sample comprising methylation sequence reads for at least 1,000 cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; generating a first set of training samples and a second set of training samples by subdividing the plurality of training samples; For each training sample, a feature vector based on the methylated sequence reads of the training sample is For each methylated sequence read of the training sample, applying each of a plurality of tissue models to the methylation sequence reads to determine a likelihood that the methylation sequence read is informative for the presence of a disease state associated with the tissue model; assigning the methylated sequence read to one of the disease states most likely output by the tissue model; determining the feature vector based on the methylated sequence reads assigned to each disease state; and training a classifier with the feature vectors of the first set of training samples to generate a signal vector based on the input feature vectors, the signal vector including a value for each disease state; applying the cancer classifier to the feature vector of each training sample in the second set of training samples to generate a cancer signal vector for each training sample in the second set of training samples; determining, for each label of the plurality of labels for the covariate characteristic, a cutoff threshold for each disease state based on the signal vector for the training sample in the second set of training samples having the label; A method comprising:

2. 10. The method of claim 1, wherein the methylated sequence reads are obtained from a targeted methylation sequencing assay or a whole-genome bisulfite sequencing assay.

3. 3. The method of claim 1 or 2, wherein each training sample comprises methylated sequence reads for at least 10,000 cfDNA fragments.

4. The method of any one of claims 1 to 3, wherein the first set of training samples and the second set of training samples comprise similar proportions of training samples for the disease states.

5. The method of any one of claims 1 to 4, wherein the plurality of disease states comprises a non-cancerous state and one or more cancer states for one or more cancers of different origins.

6. 6. The method of claim 5, wherein the one or more cancers of different origins comprise breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis and ureter, non-urothelial renal cancer, prostate cancer, anorectal cancer, colorectal cancer, squamous cell carcinoma of the esophagus, non-squamous esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from hepatocytes, hepatobiliary carcinoma arising from cells other than hepatocytes, pancreatic cancer, human papillomavirus-associated head and neck cancer, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung carcinoma, squamous cell lung carcinoma, lung cancer other than adenocarcinoma or small cell lung carcinoma, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. In some embodiments, the type of cancer is additionally selected from the group including brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma, and gallbladder cancer.

7. The method of any one of claims 1 to 6, wherein the covariate characteristic is one of age, smoker status, and biological sex.

8. determining, for each methylated sequence read, whether the methylated sequence read has an informative methylation pattern by p-value filtering, wherein the feature vector for each training sample is generated based on the methylated sequence reads with informative methylation patterns. The method of any one of claims 1 to 7, further comprising:

9. Each of the plurality of tissue models comprises: obtaining a second plurality of training samples, each having one of a plurality of disease states, each training sample comprising methylated sequence reads for cfDNA fragments obtained from the individual; generating, for each tissue model, a training dataset comprising at least 10,000 methylated sequence reads of training samples having the disease state associated with the tissue model; training each tissue model with the associated training dataset so that methylation sequence reads predict the likelihood of being informative for the presence of the associated disease state; The method according to any one of claims 1 to 8, wherein the training is carried out by

10. The method of any one of claims 1 to 9, wherein each tissue model is one of a binomial model, an independent site model, a Markov model, or a mixed model.

11. The method of any one of claims 1 to 9, wherein the classifier is a machine learning model.

12. obtaining a test sample from a test individual having an unknown disease state and associated with a first label of the plurality of labels for the covariate characteristic, the test sample comprising methylation sequence reads for cfDNA fragments obtained from the test individual; a test feature vector based on the methylated sequence reads of the test sample; For each methylated sequence read of the test sample, applying each of the plurality of tissue models to the methylation sequence reads to determine a likelihood that the methylation sequence read is informative for the presence of a disease state associated with the tissue model; assigning the methylated sequence read to one of the disease states most likely output by the tissue model; determining the test feature vector based on the methylated sequence reads assigned to each disease state; and applying the classifier to the test feature vector of the test to generate a signal vector for the test sample; detecting a positive disease signal for one or more of the disease states by applying the cutoff threshold associated with the first label for the covariate characteristic to the signal vector for the test sample; The method of any one of claims 1 to 11, further comprising:

13. detecting said positive disease signal for one or more of said disease states, for each disease state, applying the cutoff threshold for that disease state to the value in the signal vector for the test sample corresponding to that disease state; 13. The method of claim 12, comprising:

14. each individual is further associated with one of a second plurality of labels for a second covariate characteristic; determining the cutoff threshold for each combination of one label from the plurality of labels for the covariate characteristic and one label from the second plurality of labels for the second covariate characteristic; the test sample is further associated with a second label for the second covariate characteristic; detecting the positive disease signal for one or more of the disease conditions comprises applying the cutoff threshold associated with the combination of the first label for the covariate characteristic and the second label for the second covariate characteristic. The method according to any one of claims 1 to 13.

15. reporting said positive disease signal to a health care provider for further diagnostic workup steps. The method of any one of claims 12 to 14, further comprising:

16. obtaining a plurality of training samples from individuals each having one of a plurality of disease states and each associated with one of a plurality of labels for a covariate characteristic, each training sample comprising methylation sequence reads for at least 1,000 cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; For each training sample, a feature vector is generated based on the methylated sequence reads of the training sample. For each methylated sequence read of the training sample, applying each of a plurality of tissue models to the methylation sequence reads to determine a likelihood that the methylation sequence read is informative for the presence of a disease state associated with the tissue model; assigning the methylated sequence read to one of the disease states most likely output by the tissue model; determining the feature vector based on the methylated sequence reads assigned to each disease state; and generating a set of training samples for each label of the plurality of labels for the covariate characteristic including the feature vector of the training sample associated with the label for the covariate characteristic; For each label, training a classifier with the feature vectors of the set of training samples corresponding to the label to generate a signal vector based on the input feature vector, the signal vector including a value for each disease state; A method comprising:

17. 17. The method of claim 16, wherein the methylated sequence reads are obtained from a targeted methylation sequencing assay or a whole genome bisulfite sequencing assay.

18. 18. The method of claim 16 or 17, wherein each training sample comprises methylated sequence reads for at least 10,000 cfDNA fragments.

19. The method of any one of claims 16 to 18, wherein the plurality of disease states comprises a non-cancerous state and one or more cancer states for one or more cancers of different origins.

20. 20. The method of claim 19, wherein the one or more cancers of different origins comprise breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis and ureter, non-urothelial renal cancer, prostate cancer, anorectal cancer, colorectal cancer, squamous cell carcinoma of the esophagus, non-squamous esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from hepatocytes, hepatobiliary carcinoma arising from cells other than hepatocytes, pancreatic cancer, human papillomavirus-associated head and neck cancer, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung carcinoma, squamous cell lung carcinoma, lung cancer other than adenocarcinoma or small cell lung carcinoma, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. In some embodiments, the type of cancer is additionally selected from the group including brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma, and gallbladder cancer.

21. 21. The method of any one of claims 16 to 20, wherein the covariate characteristic is one of age, smoker status, and biological sex.

22. determining, for each methylated sequence read, whether the methylated sequence read has an informative methylation pattern by p-value filtering, wherein the feature vector for each training sample is generated based on the methylated sequence reads with informative methylation patterns. The method of any one of claims 16 to 21, further comprising:

23. Each of the plurality of tissue models comprises: obtaining a second plurality of training samples, each having one of a plurality of disease states, each training sample comprising methylated sequence reads for cfDNA fragments obtained from the individual; generating, for each tissue model, a training dataset comprising at least 10,000 methylated sequence reads of training samples having the disease state associated with the tissue model; training each tissue model with the associated training dataset so that methylation sequence reads predict the likelihood of being informative for the presence of the associated disease state; The method according to any one of claims 16 to 22, wherein the method is trained by

24. The method of any one of claims 16 to 23, wherein each tissue model is one of a binomial model, an independent site model, a Markov model, or a mixed model.

25. The method of any one of claims 16 to 24, wherein each classifier is a machine learning model.

26. obtaining a test sample from a test individual having an unknown disease state and associated with a first label of the plurality of labels for the covariate characteristic, the test sample comprising methylation sequence reads for cfDNA fragments obtained from the test individual; a test feature vector based on the methylated sequence reads of the test sample; For each methylated sequence read of the test sample, applying each of the plurality of tissue models to the methylation sequence reads to determine a likelihood that the methylation sequence read is informative for the presence of a disease state associated with the tissue model; assigning the methylated sequence read to one of the disease states most likely output by the tissue model; determining the test feature vector based on the methylated sequence reads assigned to each disease state; and applying the classifier corresponding to the first label to the test feature vector of the test to generate a signal vector for the test sample; detecting a positive disease signal for one or more of said disease states based on said signal vector for said test sample; The method of any one of claims 16 to 25, further comprising:

27. detecting said positive disease signal for one or more of said disease states, identifying the disease state with the highest value in the signal vector as the disease state having the positive disease signal detected.

27. The method of claim 26, comprising:

28. each individual is further associated with one of a second plurality of labels for a second covariate characteristic; generating the set of training samples includes generating a set of training samples for each combination of one label from the plurality of labels for the covariate characteristic and one label from the second plurality of labels for the second covariate characteristic; training the classifier includes training a classifier for each combination of the feature vectors from the set of training samples corresponding to the combination; the test sample is further associated with a second label for the second covariate characteristic; applying the classifier includes applying the classifier trained on the combination of the first label and the second label. The method according to any one of claims 16 to 27.

29. reporting said positive disease signal to a health care provider for further diagnostic workup steps. The method of any one of claims 12 to 14, further comprising:

30. obtaining a plurality of training samples from individuals each having one of a plurality of disease states and each associated with one of a plurality of labels for a covariate characteristic, each training sample comprising methylation sequence reads for at least 1,000 cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; For each training sample, a feature vector is generated based on the methylated sequence reads of the training sample. determining, for each of a plurality of genomic regions, a methylation feature value based on one or more methylated sequence reads of the training sample that overlap with the genomic region; determining the feature vector based on the methylation feature values ​​across the plurality of genomic regions; and generating a set of training samples for each label of the plurality of labels for the covariate characteristic including the feature vector of the training sample associated with the label for the covariate characteristic; For each label, determining a mutual information score for each genomic region based on the feature vectors of the set of training samples corresponding to the labels; ranking the genomic regions based on the mutual information scores; selecting a set of features from the ranked genomic regions; modifying the feature vector to include methylation feature values ​​for the set of features; training a classifier with the modified feature vector to generate a signal vector based on the input feature vector, the signal vector including a value for each disease state; A method comprising:

31. 31. The method of claim 30, wherein the methylated sequence reads are obtained from a targeted methylation sequencing assay or a whole genome bisulfite sequencing assay.

32. 32. The method of claim 30 or 31, wherein each training sample comprises methylated sequence reads for at least 10,000 cfDNA fragments.

33. The method of any one of claims 30 to 32, wherein each genomic region comprises one or more CpG sites.

34. The method according to any one of claims 30 to 33, wherein the methylation feature value is a methylation density in the genomic region.

35. determining, for each methylated sequence read, whether the methylated sequence read has an informative methylation pattern by p-value filtering, wherein the feature vector for each training sample is generated based on the methylated sequence reads with informative methylation patterns. The method of any one of claims 30 to 33, further comprising:

36. 36. The method of claim 35, wherein the methylation feature value is a count of methylated sequence reads that overlap with the genomic region.

37. 37. The method of any one of claims 30 to 36, wherein the plurality of disease states comprises a non-cancerous state and one or more cancer states for one or more cancers of different origins.

38. 38. The method of claim 37, wherein the one or more cancers of different origins comprise breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis and ureter, non-urothelial renal cancer, prostate cancer, anorectal cancer, colorectal cancer, squamous cell carcinoma of the esophagus, non-squamous esophageal cancer, gastric cancer, hepatobiliary carcinoma arising from hepatocytes, hepatobiliary carcinoma arising from cells other than hepatocytes, pancreatic cancer, human papillomavirus-associated head and neck cancer, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung carcinoma, squamous cell lung carcinoma, lung cancer other than adenocarcinoma or small cell lung carcinoma, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. In some embodiments, the type of cancer is additionally selected from the group including brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma, and gallbladder cancer.

39. 38. The method of any one of claims 30 to 37, wherein the covariate characteristic is one of age, smoker status, and biological sex.

40. determining the mutual information score for each genomic region for each label, determining a pairwise information gain for each pairwise combination of disease states based on the discriminatory power of the genomic regions to classify the two disease states; determining the mutual information score by combining the pairwise information gains across the pairwise combinations of disease states; The method of any one of claims 30 to 38, comprising:

41. For each label, determining coverage for each genomic region based on the sequencing depth of the set of training samples corresponding to the labels; excluding one or more genomic regions from the selection of the set of features based on the coverage being below a threshold; 40. The method of any one of claims 30 to 39, further comprising:

42. For each label, determining, for each genomic region, an activation score for the non-cancer disease state indicative of the number of training samples having the non-cancer disease state that exhibit activation of the genomic region; excluding one or more genomic regions from selection of said set of features based on an activation score exceeding a threshold; The method of any one of claims 30 to 40, further comprising:

43. The method of any one of claims 30 to 41, wherein the set of selected features for each label is determined to optimize the accuracy of the classifier.

44. A method according to any one of claims 30 to 42, wherein one set of features for one label is different from another set of features for another label.

45. 44. The method of claim 43, wherein the one set of features includes one or more features that are different from the other set of features.

46. 44. The method of claim 43, wherein the one set of features includes a different number of features than the other set of features.

47. The method of any one of claims 30 to 45, wherein each classifier is a machine learning model.

48. obtaining a test sample from a test individual having an unknown disease state and associated with a first label of the plurality of labels for the covariate characteristic, the test sample comprising methylation sequence reads for cfDNA fragments obtained from the test individual; a test feature vector based on the methylation sequence reads of the test sample; determining, for each of the plurality of genomic regions, a methylation feature value based on one or more methylated sequence reads of the test sample that overlap with the genomic region; determining the test feature vector based on the methylation feature values ​​across the plurality of genomic regions; and applying the classifier corresponding to the first label to the test feature vector of the test to generate a signal vector for the test sample; detecting a positive disease signal for one or more of said disease states based on said signal vector for said test sample; The method of any one of claims 16 to 25, further comprising:

49. detecting said positive disease signal for one or more of said disease states, identifying the disease state with the highest value in the signal vector as the disease state having the positive disease signal detected.

48. The method of claim 47, comprising:

50. reporting said positive disease signal to a health care provider for further diagnostic workup steps. The method of any one of claims 12 to 14, further comprising:

51. 1. A method for optimizing feature selection, comprising: obtaining a plurality of sequence reads from the sample via an analysis system; generating, via the analysis system, a plurality of features based on the plurality of sequence reads; generating, via the analysis system, one or more of an activation fraction, a mutual information score, or a coverage based on the plurality of sequence reads; determining, via the analysis system, a rank for the plurality of features based on one or more of the activation fraction, mutual information score, or coverage; selecting, via the analysis system, a subset of features from the plurality of features based on the rankings of the plurality of features; A method comprising:

52. 51. The method of claim 50, wherein the sample comprises a cell-free nucleic acid sample.

53. 52. The method of claim 50 or 51, wherein the plurality of features comprises a sequence read count of sequence reads that exceed a ratio threshold.

54. 53. The method of any one of claims 50-52, wherein generating the plurality of features comprises determining methylation rates for a plurality of CpG sites in the plurality of sequence reads.

55. 54. The method of any one of claims 50 to 53, wherein the activation fraction comprises a non-cancer activation fraction.

56. 55. The method of any one of claims 50 to 54, wherein the mutual information scores of the mutual information scores corresponding to a feature of the plurality of features are inversely proportional to the rank of the feature.

57. The method of any one of claims 50 to 55, wherein lower coverage corresponds to noisier features.

58. 57. The method of any one of claims 50 to 56, further comprising training at least one classifier based on said subset of features, wherein said at least one classifier is trained to predict the presence or absence of said disease, disease type, and / or disease tissue of origin.

59. 58. The method of claim 57, wherein the at least one classifier is trained to determine the tissue of origin of the disease based on one or more characteristics.

60. 59. The method of claim 58, wherein the one or more characteristics include at least one of age, smoker status, and gender.

61. determining a first activation fraction for a positive disease prediction; determining a second activation fraction for a negative disease prediction; Further comprising: the subset of features is selected at least in part by minimizing a difference between the first activation fraction and the second activation fraction.

60. The method of any one of claims 50 to 59.

62. generating a set of training data corresponding to the subset of features, wherein the training data is generated from sequence data labeled by disease type or tissue of origin, or labeled as healthy; training a machine learning model using the set of generated training data, the machine learning model configured to predict disease state and type based on sequence data corresponding to the subset of features; 61. The method of any one of claims 50 to 60, further comprising:

63. A non-transitory computer readable storage medium storing instructions that, when executed by a computer processor, cause the processor to perform the method of any one of claims 1 to 61.

64. the computer processor; 63. A non-transitory computer-readable storage medium according to claim 62; A system comprising:

65. a collection container for storing the biological sample containing the cfDNA fragments; Optionally, one or more reagents for isolating the cfDNA fragments from the biological sample; Optionally, a sequencing panel comprising probes for targeting specific genomic regions; 63. A non-transitory computer-readable storage medium according to claim 62; A treatment kit comprising: