Optimizing model-based feature extraction and classification
By analyzing the methylation sequence of cfDNA fragments in the patient's blood, and optimizing cancer classification using covariate characteristics and machine learning models, the shortcomings of traditional cancer detection technology in the early stages are solved and an earlier accurate diagnosis is achieved.
Patent Information
- Application Number
- CN202380089852.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-16
- Filing Date
- 2023-11-16
- Publication Date
- 2025-08-08
AI Technical Summary
Existing cancer detection technologies are difficult to accurately diagnose in the early stages of cancer. Traditional methods can often only be detected when the tumor reaches a certain size, and cannot be effectively identified before cancer cells fall out into the bloodstream.
By analyzing the methylation sequence of free DNA (cfDNA) fragments in the patient's blood, optimizing cancer classification using covariate characteristics, using machine learning models and feature extraction technology, and training classifiers based on methylation patterns and genomic region coverage to determine the possibility of disease signals.
Accurate prediction and classification of disease status in the early stages of cancer is achieved, the sensitivity and specificity of cancer detection is improved, and the existence of disease can be identified before traditional methods cannot detect cancer.
Smart Images

Figure CN120457490A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 425,889, filed on November 16, 2022, the entire contents of which are incorporated by reference. Technical Field
[0002] The present disclosure generally relates to model-based feature extraction and classifiers for predicting disease states from nucleic acid samples. Background Art
[0003] DNA methylation plays a role in regulating gene expression. Abnormal DNA methylation has been associated with many disease processes (including cancer). DNA methylation profiling using methylation sequencing (e.g., whole genome bisulfite sequencing (WGBS)) is increasingly being considered an important diagnostic tool for detecting, diagnosing, and / or monitoring cancer. For example, specific patterns of differentially methylated regions can be used as molecular markers for various disease states. Summary of the Invention
[0004] Cancer classification typically involves applying one or more predictive models to features derived from genetic sequencing data to predict an individual's disease status. Specifically, cancer classification can be performed by extracting features from sequencing data using tissue models to determine the tissue of origin of sequence reads associated with nucleic acid fragments. Disease status can be a binary prediction generally indicating the presence of a disease signature, or a multi-class prediction indicating the presence of a specific disease signature.
[0005] One or more techniques can be implemented to optimize cancer classification based on covariate characteristics. In a first approach, the analysis system can determine separate cutoff thresholds for detecting a positive disease signal for different labels of the covariate characteristics. The system can segment the training samples based on the labels of their covariate characteristics to determine the cutoff thresholds separately. In other approaches, the system can train a different classifier for each population. The system separates the training samples based on the labels of their covariate characteristics and trains the classifiers separately to generate a signal vector representing the amount of disease signal detected in the sample. The classifiers can be trained on different feature sets, such as those determined based on mutual information gain, genomic region coverage, and healthy activation scores.
[0006] Clause 1. A method comprising: obtaining a plurality of training samples from an individual, each training sample having one of a plurality of disease states and each training sample being associated with one of a plurality of labels of a covariate characteristic, wherein each training sample comprises methylation sequence reads of at least 1,000 cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; generating a first training sample set and a second training sample set by subdividing the plurality of training samples; for each training sample, generating a feature vector based on the methylation sequence reads of the training sample by: for each methylation sequence read of the training sample: applying each of a plurality of tissue models to the methylation sequence reads to determine whether the methylation sequence read indicates one of the disease states associated with the tissue model The invention also provides a method for determining a cancer signal vector for each training sample in the second training sample set, wherein the method further comprises: determining a probability of the presence of a disease state, assigning the methylation sequence read to one of the disease states having the highest probability output by the tissue models, and determining the feature vector based on the methylation sequence reads assigned to each disease state; training a classifier using the feature vector of the first training sample set to generate a signal vector based on the input feature vector, wherein the signal vector includes a value for each disease state; applying a cancer classifier to the feature vector of each training sample in the second training sample set to generate a cancer signal vector for each training sample in the second training sample set; and determining a cutoff threshold for each disease state for each label in the multiple labels of the covariate characteristic based on the signal vector of the training samples with the label in the second training sample set.
[0007] Clause 2. The method of clause 1, wherein the methylation sequence reads are obtained by a targeted methylation sequencing assay or a whole-genome bisulfite sequencing assay.
[0008] Clause 3. The method of any one of clauses 1 to 2, wherein each training sample comprises at least 10,000 methylation sequence reads of cfDNA fragments.
[0009] Clause 4. The method according to any one of clauses 1 to 3, wherein the first training sample set and the second training sample set comprise similar proportions of training samples across the disease states.
[0010] Clause 5. The method according to any one of Clauses 1 to 4, wherein the plurality of disease states comprises a non-cancerous state and one or more cancerous states of one or more cancers of different origin.
[0011] Clause 6. The method of clause 5, wherein the one or more cancers of different origin include: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial cancer of the renal pelvis and ureter, renal cancer other than urothelial cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal squamous cell carcinoma, esophageal cancer other than squamous, gastric cancer, hepatobiliary cancer arising from hepatocellular cells, hepatobiliary cancer arising from cells other than hepatocellular cells, pancreatic cancer, head and neck cancer associated with human papillomavirus, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung cancer, lung squamous cell carcinoma, and lung cancer other than adenocarcinoma or small cell lung cancer, neuroendocrine cancer, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. In some embodiments, the cancer type is further selected from the group consisting of: brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma, and gallbladder cancer.
[0012] Clause 7. The method of any one of clauses 1 to 6, wherein the covariate characteristic is one of: age, smoker status, and biological sex.
[0013] Clause 8. The method according to any one of clauses 1 to 7 further comprises: for each methylation sequence read, using p-value screening to determine whether the methylation sequence read has an informative methylation pattern, wherein the feature vector of each training sample is generated based on the methylation sequence read with an informative methylation pattern.
[0014] Clause 9. A method according to any one of clauses 1 to 8, wherein each of the multiple tissue models is trained by: obtaining a second plurality of training samples, each training sample having one of a plurality of disease states, wherein each training sample comprises methylation sequence reads of cfDNA fragments obtained from the individual; generating a training data set for each tissue model, the training data set comprising at least 10,000 methylation sequence reads of training samples having a disease state associated with the tissue model; and training each tissue model using the associated training data set to predict the likelihood that the methylation sequence reads indicate the presence of the associated disease state.
[0015] Clause 10. The method of any one of clauses 1 to 9, wherein each tissue model is one of: a binomial model, an independent site model, a Markov model, or a mixture model.
[0016] Clause 11. A method according to any one of clauses 1 to 9, wherein the classifier is a machine learning model.
[0017] Clause 12. The method according to any one of clauses 1 to 11 further comprises: obtaining a test sample having an unknown disease state from a test individual and associated with a first label among the multiple labels of the covariate characteristic, wherein the test sample comprises methylation sequence reads of cfDNA fragments obtained from the test individual; generating a test feature vector based on the methylation sequence reads of the test sample in the following manner: for each methylation sequence read of the test sample: applying each of the multiple tissue models to the methylation sequence reads to determine the likelihood that the methylation sequence read indicates the presence of a disease state associated with the tissue model, assigning the methylation sequence read to one of the disease states having the highest likelihood output by the tissue models, and determining the test feature vector based on the methylation sequence reads assigned to each disease state; applying the classifier to the test feature vector of the test to generate a signal vector for the test sample; and detecting a positive disease signal for one or more of the disease states by applying a cutoff threshold associated with the first label of the covariate characteristic to the signal vector of the test sample.
[0018] Clause 13. The method of clause 12, wherein detecting a positive disease signal for one or more of the disease states comprises, for each disease state, applying a cutoff threshold for the disease state to the values in the signal vector of the test sample corresponding to the disease state.
[0019] Clause 14. A method according to any one of clauses 1 to 13, wherein each individual is further associated with one of a second plurality of labels of a second covariate characteristic; wherein determining the cutoff thresholds comprises determining for each combination of one of the plurality of labels from the covariate characteristic and one of the second plurality of labels from the second covariate characteristic; wherein the test sample is further associated with a second label of the second covariate characteristic; and wherein detecting a positive disease signal for one or more of the disease states comprises applying a cutoff threshold associated with a combination of the first label of the covariate characteristic and the second label of the second covariate characteristic.
[0020] Clause 15. The method of any one of Clauses 12 to 14, further comprising reporting the positive disease signal to a healthcare provider for additional diagnostic steps.
[0021] Clause 16. A method comprising: obtaining a plurality of training samples from an individual, each training sample having one of a plurality of disease states and each training sample being associated with one of a plurality of labels of a covariate characteristic, wherein each training sample comprises at least 1,000 methylated sequence reads of cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; generating, for each training sample, a feature vector based on the methylated sequence reads of the training sample by: for each methylated sequence read of the training sample: applying each of a plurality of tissue models to the methylated sequence reads to determine a feature vector indicative of the methylated sequence read The method comprises the steps of: generating a classifier comprising: determining a probability of the existence of a disease state associated with the tissue model, assigning the methylation sequence read to one of the disease states having the highest probability output by the tissue models, and determining the feature vector based on the methylation sequence reads assigned to each disease state; generating a training sample set for each of the multiple labels of the covariate characteristic, including a feature vector of the training sample associated with the label of the covariate characteristic; and training a classifier, for each label, using the feature vector of the training sample set corresponding to the label to generate a signal vector based on the input feature vector, wherein the signal vector includes a value for each disease state.
[0022] Clause 17. The method of clause 16, wherein the methylation sequence reads are obtained by a targeted methylation sequencing assay or a whole-genome bisulfite sequencing assay.
[0023] Clause 18. The method of any one of clauses 16 to 17, wherein each training sample comprises at least 10,000 methylation sequence reads of cfDNA fragments.
[0024] Clause 19. The method according to any one of Clauses 16 to 18, wherein the plurality of disease states comprises a non-cancerous state and one or more cancerous states of one or more cancers of different origin.
[0025] Clause 20. The method of Clause 19, wherein the one or more cancers of different origin include: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial cancer of the renal pelvis and ureter, renal cancer other than urothelial cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal squamous cell carcinoma, esophageal cancer other than squamous, gastric cancer, hepatobiliary cancer arising from hepatocellular cells, hepatobiliary cancer arising from cells other than hepatocellular cells, pancreatic cancer, head and neck cancer associated with human papillomavirus, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung cancer, lung squamous cell carcinoma, and lung cancer other than adenocarcinoma or small cell lung cancer, neuroendocrine cancer, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. In some embodiments, the cancer type is further selected from the group consisting of: brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma, and gallbladder cancer.
[0026] Clause 21. The method of any one of clauses 16 to 20, wherein the covariate characteristic is one of: age, smoker status, and biological sex.
[0027] Clause 22. The method according to any one of clauses 16 to 21 further comprises: for each methylation sequence read, using p-value screening to determine whether the methylation sequence read has an informative methylation pattern, wherein the feature vector of each training sample is generated based on the methylation sequence read with an informative methylation pattern.
[0028] Clause 23. A method according to any one of clauses 16 to 22, wherein each of the multiple tissue models is trained by: obtaining a second plurality of training samples, each training sample having one of a plurality of disease states, wherein each training sample comprises methylation sequence reads of cfDNA fragments obtained from the individual; generating, for each tissue model, a training dataset comprising at least 10,000 methylation sequence reads of training samples having a disease state associated with the tissue model; and training each tissue model using the associated training dataset to predict the likelihood that the methylation sequence reads indicate the presence of the associated disease state.
[0029] Clause 24. The method of any one of clauses 16 to 23, wherein each tissue model is one of: a binomial model, an independent site model, a Markov model, or a hybrid model.
[0030] Clause 25. A method according to any one of clauses 16 to 24, wherein each classifier is a machine learning model.
[0031] Clause 26. The method according to any one of clauses 16 to 25 further includes: obtaining a test sample having an unknown disease state and associated with a first label among the multiple labels of the covariate characteristic from a test individual, wherein the test sample includes methylation sequence reads of cfDNA fragments obtained from the test individual; generating a test feature vector based on the methylation sequence reads of the test sample in the following manner: for each methylation sequence read of the test sample: applying each of the multiple tissue models to these methylation sequence reads to determine the likelihood that the methylation sequence read indicates the presence of a disease state associated with the tissue model, assigning the methylation sequence read to one of the disease states having the highest likelihood output by the tissue models, and determining the test feature vector based on the methylation sequence reads assigned to each disease state; applying a classifier corresponding to the first label to the test feature vector of the test to generate a signal vector for the test sample; and detecting positive disease signals for one or more of the disease states based on the signal vector of the test sample.
[0032] Clause 27. The method of Clause 26, wherein detecting a positive disease signal for one or more of the disease states comprises identifying the disease state having a maximum value in the signal vector as the disease state having the detected positive disease signal.
[0033] Clause 28. A method according to any one of clauses 16 to 27, wherein each individual is further associated with one label from a second plurality of labels of a second covariate characteristic; wherein generating multiple training sample sets includes: generating a training sample set for each combination of one label from the plurality of labels of the covariate characteristic and one label from the second plurality of labels of the second covariate characteristic; wherein training the classifiers includes training a classifier for each combination using a feature vector from the training sample set corresponding to the combination; wherein the test sample is further associated with a second label of the second covariate characteristic; and wherein applying the classifier includes applying a classifier trained for the combination of the first label and the second label.
[0034] Clause 29. The method of any one of Clauses 12 to 14, further comprising reporting the positive disease signal to a healthcare provider for additional diagnostic steps.
[0035] Clause 30. A method comprising: obtaining a plurality of training samples from an individual, each training sample having one of a plurality of disease states and each training sample being associated with one of a plurality of labels of a covariate characteristic, wherein each training sample comprises at least 1,000 methylated sequence reads of cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; generating, for each training sample, a feature vector based on the methylated sequence reads of the training sample by: determining, for each of a plurality of genomic regions, a methylation feature value based on one or more methylated sequence reads of the training sample that overlap with the genomic region, and generating a feature vector based on the plurality of genomic regions. The method comprises the steps of: determining a feature vector based on the methylation feature value of the domain; generating a training sample set for each of the multiple labels of the covariate characteristic, including a feature vector of the training sample associated with the label of the covariate characteristic; for each label: determining a mutual information score for each genomic region based on the feature vector of the training sample set corresponding to the label; ranking the genomic regions based on the mutual information score; selecting a feature set from the ranked genomic regions; modifying the feature vectors to include the methylation feature value of the feature set; and training a classifier using the modified feature vector to generate a signal vector based on the input feature vector, wherein the signal vector includes a value for each disease state.
[0036] Clause 31. The method of clause 30, wherein the methylation sequence reads are obtained by a targeted methylation sequencing assay or a whole-genome bisulfite sequencing assay.
[0037] Clause 32. The method of any one of clauses 30 to 31, wherein each training sample comprises at least 10,000 methylation sequence reads of cfDNA fragments.
[0038] Clause 33. The method of any one of clauses 30 to 32, wherein each genomic region comprises one or more CpG sites.
[0039] Clause 34. The method according to any one of clauses 30 to 33, wherein the methylation signature value is the methylation density of the genomic region.
[0040] Clause 35. The method according to any one of clauses 30 to 33 further comprises: for each methylation sequence read, using p-value screening to determine whether the methylation sequence read has an informative methylation pattern, wherein the feature vector of each training sample is generated based on the methylation sequence read with an informative methylation pattern.
[0041] Clause 36. The method of clause 35, wherein the methylation signature value is a count of methylated sequence reads overlapping the genomic region.
[0042] Clause 37. The method according to any one of Clauses 30 to 36, wherein the plurality of disease states comprises a non-cancerous state and one or more cancerous states of one or more cancers of different origin.
[0043] Clause 38. The method of clause 37, wherein the one or more cancers of different origin include: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial cancer of the renal pelvis and ureter, renal cancer other than urothelial cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal squamous cell carcinoma, esophageal cancer other than squamous, gastric cancer, hepatobiliary cancer arising from hepatocellular cells, hepatobiliary cancer arising from cells other than hepatocellular cells, pancreatic cancer, head and neck cancer associated with human papillomavirus, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung cancer, lung squamous cell carcinoma, and lung cancer other than adenocarcinoma or small cell lung cancer, neuroendocrine cancer, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. In some embodiments, the cancer type is further selected from the group consisting of: brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma, and gallbladder cancer.
[0044] Clause 38. The method of any one of clauses 30 to 37, wherein the covariate characteristic is one of: age, smoker status, and biological sex.
[0045] Clause 39. A method according to any one of clauses 30 to 38, wherein determining the mutual information score for each genomic region for each label includes: determining the pairwise information gain of each pairwise disease state combination based on the discriminatory ability of the genomic region to classify two disease states; and determining the mutual information score by combining the pairwise information gains of these pairwise disease state combinations.
[0046] Clause 40. The method according to any one of clauses 30 to 39, further comprising: for each label: determining the coverage of each genomic region based on the sequencing depth of the training sample set corresponding to the label; and excluding one or more genomic regions from selection for the feature set based on the coverage being lower than a threshold.
[0047] Clause 41. The method of any one of clauses 30 to 40, further comprising: for each label: for each genomic region, determining an activation score for the non-cancer disease state, the activation score indicating the number of training samples with the non-cancer disease state that exhibit activation of the genomic region; and excluding one or more genomic regions from selection for the feature set based on the activation score being above a threshold.
[0048] Clause 42. The method of any one of clauses 30 to 41, wherein the set of features selected for each label is determined to optimize the precision of the classifier.
[0049] Clause 43. The method of any one of clauses 30 to 42, wherein a feature set of one tag is different from another feature set of another tag.
[0050] Clause 44. The method of clause 43, wherein the one feature set comprises one or more features that are different from the other feature set.
[0051] Clause 45. The method of clause 43, wherein the one feature set comprises a different number of features than the other feature set.
[0052] Clause 46. A method according to any one of clauses 30 to 45, wherein each classifier is a machine learning model.
[0053] Clause 47. The method according to any one of clauses 16 to 25 further comprises: obtaining a test sample having an unknown disease state and associated with a first label among the multiple labels of the covariate characteristic from a test individual, wherein the test sample comprises methylation sequence reads of cfDNA fragments obtained from the test individual; generating a test feature vector based on the methylation sequence reads of the test sample by: determining, for each of the multiple genomic regions, a methylation feature value based on one or more methylation sequence reads of the test sample overlapping with the genomic region, and determining the test feature vector based on the methylation feature values of the multiple genomic regions; applying a classifier corresponding to the first label to the test feature vector of the test to generate a signal vector for the test sample; and detecting a positive disease signal for one or more of the disease states based on the signal vector of the test sample.
[0054] Clause 48. The method of clause 47, wherein detecting a positive disease signal for one or more of the disease states comprises identifying the disease state having a maximum value in the signal vector as the disease state having the detected positive disease signal.
[0055] Clause 49. The method of any one of Clauses 12 to 14, further comprising reporting the positive disease signal to a healthcare provider for additional diagnostic steps.
[0056] Item 50. A method for optimizing feature selection, the method comprising: obtaining a plurality of sequence reads from a sample via an analysis system; generating a plurality of features based on the plurality of sequence reads via the analysis system; generating one or more of activation scores, mutual information scores, or coverages based on the plurality of sequence reads via the analysis system; determining a ranking of the plurality of features based on the one or more activation scores, mutual information scores, or coverages via the analysis system; and selecting a feature subset from the plurality of features based on the ranking of the plurality of features via the analysis system.
[0057] Clause 51. The method of clause 50, wherein the sample comprises a free nucleic acid sample.
[0058] Clause 52. The method of any one of clauses 50 to 51, wherein the plurality of features comprises sequence read counts of sequence reads exceeding a ratio threshold.
[0059] Clause 53. The method of any one of clauses 50 to 52, wherein generating the plurality of features comprises determining methylation rates of a plurality of CpG sites in the plurality of sequence reads.
[0060] Clause 54. The method of any one of clauses 50 to 53, wherein the activation score comprises a non-cancer activation score.
[0061] Clause 55. The method of any one of clauses 50 to 54, wherein a mutual information score of the mutual information scores corresponding to a feature of the plurality of features is inversely proportional to the rank of the feature.
[0062] Clause 56. The method of any one of clauses 50 to 55, wherein lower coverage corresponds to noisier features.
[0063] Clause 57. The method of any one of clauses 50 to 56, further comprising training at least one classifier based on the feature subset, the at least one classifier being trained to predict the presence or absence of a disease, the type of disease, and / or the tissue of origin of the disease.
[0064] Clause 58. The method of clause 57, wherein the at least one classifier is trained to determine the tissue of origin of the disease based on one or more characteristics.
[0065] Clause 59. The method of clause 58, wherein the one or more characteristics include at least one of age, smoker status, and gender.
[0066] Clause 60. The method of any one of clauses 50 to 59, further comprising: determining a first activation score for positive prediction of disease; and determining a second activation score for negative prediction of disease; wherein the feature subset is selected at least in part by minimizing a difference between the first activation score and the second activation score.
[0067] Clause 61. The method according to any one of clauses 50 to 60 further comprises: generating a training data set corresponding to the feature subset, the training data being generated from sequence data labeled with disease type or tissue of origin or marked as healthy; and training a machine learning model using the generated training data set, the machine learning model being configured to predict disease status and type based on sequence data corresponding to the feature subset.
[0068] Clause 62. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the processor to perform the method of any one of clauses 1 to 61.
[0069] Clause 63. A system comprising: a computer processor; and the non-transitory computer-readable storage medium of clause 62.
[0070] Item 64. A therapeutic kit comprising: a collection container for storing a biological sample comprising cfDNA fragments; optionally, one or more reagents for isolating cfDNA fragments from the biological sample; optionally, a sequencing set comprising probes for targeting specific genomic regions; and a non-transitory computer-readable storage medium as described in Item 62. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1A is a flow chart of a method for generating a classifier to predict a disease state, according to various embodiments.
[0072] Figure 1B is a flow chart of a method for generating a classifier to predict a disease state, according to various embodiments.
[0073] Figure 2A According to one embodiment, a flow chart of a device for sequencing a nucleic acid sample is shown.
[0074] Figure 2B is a block diagram of an analysis system for processing sequence reads according to various embodiments.
[0075] Figure 3 is a flow chart describing a process for sequencing nucleic acids, according to various embodiments.
[0076] Figure 4AAccording to various embodiments Figure 3 Illustration of part of the process of sequencing nucleic acids to obtain methylation information and methylation state vectors.
[0077] Figure 4B Generation of a data structure for a control group is presented according to various embodiments.
[0078] Figure 4C A flow chart describing a process for determining informative fragments from a sample is presented according to various embodiments.
[0079] Figure 5 is an illustration of blocks of a reference genome according to various embodiments.
[0080] Figure 6 is an illustration of a process for determining features for training a classifier, according to various embodiments.
[0081] Figure 7A 、 Figure 7B 、 Figure 7C 、 Figure 7D 、 Figure 7E and Figure 7F Included is a confusion matrix indicating the accuracy of a classifier according to various embodiments.
[0082] Figure 8 is a flow chart of a model-based feature extraction method according to various embodiments.
[0083] Figure 9A and Figure 9B The sensitivity of the tissue of origin classifier is demonstrated according to the embodiments.
[0084] Figure 10A and Figure 10B According to the embodiment, the sensitivity of the tissue of origin classifier at different cancer stages is demonstrated.
[0085] Figure 11 A performance grid representing the accuracy of tissue of origin localization is shown according to an embodiment.
[0086] Figure 12 According to the embodiment, the accuracy and sensitivity of the tissue of origin classifier in different cancer stages are demonstrated.
[0087] Figure 13A and Figure 13B The ROC curve of the tissue of origin classifier is shown according to an embodiment.
[0088] Figure 14 A data flow diagram for training a model is depicted according to various embodiments.
[0089] Figure 15Precision-recall curves for uncertain base call thresholds are shown according to various embodiments.
[0090] Figure 16 is a flow chart of a method for determining a probability that a sample has a disease state, according to various embodiments.
[0091] Figure 17 Performance gains in sensitivity of a multilayer perceptron model are demonstrated according to an embodiment.
[0092] Figure 18 According to the embodiment, experimental results of the multilayer perceptron model in determining the tissue of origin are presented.
[0093] Figure 19 According to the embodiment, experimental results of the multilayer perceptron model in determining the tissue of origin by cancer stage are presented.
[0094] Figure 20 According to the embodiment, experimental results of the multilayer perceptron model on different types of cancer are presented.
[0095] Figure 21 A graph showing the likelihood of cancer type for non-cancer samples at a specificity greater than 95%.
[0096] Figure 22 A graph showing methylation sequencing data for non-cancer samples and hematological subtype cancer samples.
[0097] Figure 23A A flow chart describing a process for determining a binary threshold cutoff value for binary cancer classification is presented in accordance with one or more embodiments.
[0098] Figure 23B A flow chart describing a process for thresholding tissue-of-origin labels to determine a binary threshold cutoff value for binary cancer classification is presented in accordance with one or more embodiments.
[0099] Figure 24A and Figure 24B Confusion matrices showing the performance of trained cancer tissue-of-origin classifiers for additional hematological cancer subtypes are presented.
[0100] Figure 25A and Figure 25B A graph showing the cancer prediction accuracy of the cancer classifier at different cancer stages for multiple cancer types with and without adjusting the threshold cutoff.
[0101] Figure 26A Receiver operating characteristic curves (ROC) showing the sensitivity and specificity of cancer detection using methylation data of target genomic regions of detection panel A are depicted.
[0102] Figure 26B is a confusion matrix that depicts the accuracy of cancer type classification using the methylation data of the target genomic regions of test set A for subjects with confirmed cancer.
[0103] Figure 27A Receiver operating characteristic curves (ROC) showing the sensitivity and specificity of cancer detection using methylation data of target genomic regions of detection panel B are depicted.
[0104] Figure 27B is a confusion matrix that depicts the accuracy of cancer type classification using the methylation data of the target genomic regions of test set B for subjects with confirmed cancer.
[0105] Figure 28 The classifier performance of a proprietary cancer detection panel (Detection Panel C) is shown according to the examples.
[0106] Figure 29A and Figure 29B According to an embodiment, a tissue of origin (TOO) confusion matrix representing the cancer tissue of origin localization accuracy of detection group C is shown.
[0107] Figure 30 The classifier sensitivity performance for detecting individual tumors of group C by stage is shown according to an embodiment.
[0108] Figure 31 Tissue of origin accuracy for multiple iterations of a trained model is shown according to various embodiments.
[0109] Figure 32 A process of stratifying hematology signals into two layers is presented according to various embodiments.
[0110] Figure 33 According to one embodiment, predicted and actual signatures of cancer signal origins for samples from male and female individuals are presented.
[0111] Figure 34 According to one embodiment, predicted and actual signatures of the origin of cancer signals in samples from male individuals are presented.
[0112] Figure 35 According to one embodiment, predicted and actual signatures of cancer signal origins for samples from female individuals are presented.
[0113] Figure 36 According to one embodiment, experimental results of segment coverage and feature ranking are presented.
[0114] Figure 37A According to one embodiment, experimental results of fragment coverage and non-cancer activation scores are presented.
[0115] Figure 37B Experimental results of non-cancer activation scores and feature rankings are presented according to one embodiment.
[0116] Figure 38A Experimental results of non-cancer activation scores and feature weights are presented according to one embodiment.
[0117] Figure 38B According to one embodiment, Figure 38A Another view of the experimental results shown in .
[0118] Figure 39A Experimental results of activation score and coverage are presented according to one embodiment.
[0119] Figure 39B According to one embodiment, Figure 39A Another view of the experimental results shown in .
[0120] Figures 40 to 42 Experimental results for the various thresholding methods described herein are presented according to various embodiments. DETAILED DESCRIPTION
[0121] Early detection and classification of cancer is an important technology. Being able to detect cancer before it causes symptoms benefits all parties involved, including patients, doctors, and loved ones. For patients, early cancer detection allows for a greater chance of a beneficial outcome; for doctors, early cancer detection allows for more treatment pathways that may lead to a beneficial outcome; and for loved ones, early cancer detection increases the likelihood that friends and family members will not be lost to the disease.
[0122] Recently, early cancer detection technology has been moving toward analyzing genetic fragments (e.g., DNA) in (e.g., human blood) to determine whether any of those fragments originate from cancer cells. These new technologies allow doctors to identify cancers in patients that might not otherwise be detected (e.g., using conventional screening methods). For example, consider the case of a person at high risk for breast cancer. Traditionally, this person would visit their doctor regularly for mammograms, which produce images (e.g., x-rays) of their breast tissue, allowing doctors to identify cancerous tissue. Unfortunately, even with the highest-resolution mammograms, doctors can only identify tumors when they are approximately one millimeter in size. This means that the cancer may have been present in the person for some time, undiagnosed and untreated. For most cancers, visual identification like this is common—that is, the tumor can only be identified once it has grown large enough to be detected using some imaging technology.
[0123] Cancer detection by analyzing genetic fragments from a patient (e.g., blood) alleviates this problem. For illustration, assume that once cancer cells form, they shed DNA fragments into a person's bloodstream. This occurs when the number of cancer cells is very small and before they can be observed using imaging techniques. Therefore, with the appropriate method, a system that analyzes DNA fragments in the bloodstream can identify the presence of cancer in a person based on the shed cancer DNA fragments. More importantly, the system can identify the cancer before more traditional cancer detection technologies can detect it.
[0124] Cancer detection based on DNA fragment analysis can be achieved through next-generation sequencing ("NGS") technology. Broadly speaking, NGS is a group of technologies that can perform high-throughput sequencing of genetic material. As discussed in more detail herein, NGS primarily consists of (1) sample preparation, (2) DNA sequencing, and (3) data analysis. Sample preparation is the laboratory method required to prepare DNA fragments for sequencing, sequencing is the process of reading the ordered nucleotides in a sample, and data analysis is the processing and analysis of the genetic information in the sequencing data to identify the presence of cancer.
[0125] While these steps of NGS may help achieve early cancer detection, they themselves introduce complex and potentially harmful issues to cancer detection. Therefore, any improvements to sample preparation, DNA sequencing, and / or data analysis (including preprocessing, algorithmic processing, and summarization or presentation of predictions or conclusions) will improve cancer detection technology and early cancer detection more broadly.
[0126] For example, (1) sample preparation issues include DNA sample quality, sample contamination, fragmentation bias, and accurate indexing. Solving these issues will provide better genetic data for cancer detection.
[0127] Similarly, (2) problems associated with sequencing include, for example, errors in accurate transcription of fragments (e.g., reading "C" as "A", etc.), incorrect or difficult fragment assembly and overlap, inconsistent coverage uniformity, the trade-off between sequencing depth and cost and specificity, and insufficient sequencing length. Similarly, resolving any of these issues will provide improved genetic data for cancer detection.
[0128] (3) The problems in data analysis are the most arduous and complex. The large amount of data generated by NGS sequencing technology brings many challenges. The sequencing data of a single sample can be on the order of hundreds of thousands (up to millions) of sequence reads, totaling terabytes of data. Multiply this by the thousands (up to tens of thousands) of samples collected for training analytical models. Effectively and efficiently analyzing such a large amount of data has high requirements in terms of both procedures and computing. For example, analyzing NGS sequencing involves several baseline processing steps, such as aligning reads with each other, aligning and mapping reads to the reference genome, deduplicating duplicate reads, detecting sample contamination, identifying and base-calling variant genes, identifying and base-calling abnormally methylated genes, generating functional annotations, etc. Performing any of these processing steps on terabytes of genetic data is computationally expensive even for the most powerful computer architecture, and is completely impossible for a normal human brain. In addition, because the sample preparation and sequence reading processes are prone to errors, the genetic sequencing data derived from these processes, that is, a large portion of the resulting genetic data, may be of low quality or unusable for cancer identification. For example, large amounts of genetic data may include contaminated samples, transcription errors, mismatched regions, overrepresented regions, and so on, making them unsuitable for highly accurate cancer detection. Identifying and processing low-quality genetic data within the vast amounts of genetic data obtained from NGS sequencing is also extremely difficult procedurally and computationally, and impractical for the human brain to perform. Overall, any process that facilitates more efficient processing of large-scale sequencing data will improve cancer detection using NGS sequencing. Furthermore, such processes are designed to address various obstacles encountered with NGS sequencing, making them unconventional and innovative activities in the field of technological exploration.
[0129] In particular, under (3) data analysis, accurately identifying informative DNA from NGS data to identify the presence of cancer is also a difficult task. In order to achieve effective detection, algorithms are being sought to compensate for errors such as those caused by sample preparation and sequencing, and to overcome the large-scale data analysis problems brought about by NGS technology. That is, when designing one or more machine learning models or other computational processing algorithms to achieve early cancer detection based on next-generation sequencing technology, they must be configured to take into account the problems brought about by those technologies. Some of those technologies and models are discussed below, and specific improvements to the most advanced technologies and models are further discussed. In addition, such technologies are unconventional and innovative activities in the field of technology exploration.
[0130] A particular challenge arises when predicting the presence or absence of a cancer signature in individuals of different backgrounds. For example, a biological female is much more likely to develop breast cancer than a biological male. Or, in another example, a non-smoker is much less likely to develop lung cancer than a heavy smoker. Taking these differences into account statistically can improve predictions, allowing healthcare providers to make more accurate diagnoses. The present disclosure aims to address this challenge through one or more different approaches. In some embodiments, the system trains a cancer classifier that is broadly applied to all samples, but customizes the cutoff threshold for positively detecting the presence of one or more disease states for each subpopulation as defined by one or more covariate characteristics. In other embodiments, the system trains a different classifier for each subpopulation. The training may vary depending on the training samples used, the features evaluated, or some combination thereof.
[0131] Training of the machine learning models described herein (such as one or more cancer classifiers, tissue models, and other models mentioned herein) includes, at least in part, performing one or more non-mathematical operations or implementing non-mathematical functions by a machine or computing system, examples of which include, but are not limited to, data loading operations, data storage operations, data switching or modification operations, non-transitory computer-readable storage medium modification operations, metadata deletion or data cleaning operations, data compression operations, modification operations, image modification operations, noise application operations, noise removal operations, etc. Thus, training of the machine learning models described herein can be based on or can involve mathematical concepts, but is not limited to the act of performing mathematical calculations, mathematical operations, or calculating variables or numbers using mathematical methods.
[0132] Similarly, it should be noted that the training of the models described herein cannot actually be accomplished by the human brain alone. These models are complex in nature, comprising a large number of weights and parameters associated through one or more complex functions. For example, the number of weights in the model may be on the order of thousands, tens of thousands, hundreds of thousands, millions, or billions. Training may require the use of training samples on the order of thousands, tens of thousands, hundreds of thousands, millions, or billions. The training and / or deployment of such models involves a large number of operations that cannot be accomplished by the human brain alone or with the aid of pen and paper. In such embodiments, the number of operations may be hundreds, thousands, tens of thousands, hundreds of thousands, millions, billions, or trillions. Therefore, the implementation and use of such models are necessarily rooted in computer technology. I. Definition
[0133] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by those skilled in the art to which this specification belongs. As used herein, the following terms have the meanings ascribed to them below.
[0134] The term "subject" may refer to any living organism, such as a human subject or an animal subject. The term "healthy subject" refers to an individual who is assumed not to have cancer or disease.
[0135] The term "subject" refers to an individual whose DNA is being analyzed. A subject can be a test subject whose DNA will be evaluated using whole genome sequencing or a targeted panel as described herein to assess whether the person has a disease state (e.g., cancer, cancer type, or cancer tissue of origin). A subject can also be part of a control group that is known not to have cancer or other diseases. A subject can also be part of a cancer or other disease group that is known to have cancer or other diseases. Control groups and cancer / disease panels can be used to assist in designing or validating targeted panels.
[0136] The term "reference sample" refers to a sample obtained from a subject with a known disease state.
[0137] The term "training sample" refers to a sample obtained from a known disease state that can be used to generate sequence reads. The training sample can be applied to a probabilistic model to generate features that can be used to classify the disease state.
[0138] The term "test sample" refers to a sample that may have an unknown disease state.
[0139] The term "sequence read" refers to a nucleotide sequence read from a sample obtained from an individual. A sequence read can be generated from nucleic acid fragments in a sample. A sequence read can be a collapsed sequence read generated from multiple sequence reads obtained from multiple amplicons of a single original nucleic acid molecule. In some embodiments, a sequence read can be a deduplicated sequence read. Sequence reads can be obtained by various methods known in the art.
[0140] The term "disease state" refers to the presence or absence of a disease, the type of disease, and / or the tissue of origin of a disease. For example, in one embodiment, the present disclosure provides methods, systems, and non-transitory computer-readable media for detecting cancer (i.e., the presence or absence of cancer), the type of cancer, or the tissue of origin of a cancer.
[0141] The term "tissue of origin" or "Tissue of Origin" or "Cancer Signal of Origin" refers to the organ, group of organs, body part, or cell type from which a disease state may arise or originate. For example, identifying the tissue of origin or cancer cell type often allows for determination of appropriate next steps for further diagnosis, staging, and treatment decisions.
[0142] As used herein, the term "methylation" refers to the chemical process of adding a methyl group to a DNA molecule. Two of the four bases of DNA, cytosine ("C") and adenine ("A"), can be methylated. For example, the hydrogen atoms on the pyrimidine ring of the cytosine base can be converted into methyl groups, thereby forming 5-methylcytosine. Methylation tends to occur at dinucleotides consisting of cytosine and guanine, which are referred to herein as "CpG sites." In other cases, methylation may also occur on cytosine that is not part of a CpG site or on another nucleotide that is not cytosine; however, these occurrences are relatively rare. In this disclosure, for the sake of clarity, methylation is discussed with reference to CpG sites. However, the principles described herein are equally applicable to methylation detection in non-CpG contexts, including methylation of non-cytosines. For example, adenine methylation has been observed in bacterial, plant, and mammalian DNA, although it has received much less attention.
[0143] In such embodiments, the laboratory wet assays used to detect methylation may differ from those described herein and known in the art. Further, the methylation state vector may contain some elements, which are generally vectors of sites that have or have not undergone methylation (even if those sites are not specifically CpG sites). With this substitution, the rest of the process described herein remains unchanged, and therefore, the inventive concepts described herein are also applicable to these other forms of methylation.
[0144] The term "CpG site" refers to a region of a DNA molecule where a cytosine nucleotide is followed by a guanine nucleotide in its linear base sequence along the 5' to 3' direction. "CpG" is shorthand for 5'-C-phosphate-G-3', which represents cytosine and guanine separated by only a phosphate group; the phosphate links any two nucleotides in DNA. The cytosine in a CpG dinucleotide can be methylated to form 5-methylcytosine.
[0145] The term "methylation site" refers to a single site of a DNA molecule to which a methyl group can be added. A "CpG" site is the most common methylation site, but methylation sites are not limited to CpG sites. For example, DNA methylation can occur on the cytosine in CHG and CHH, where H is adenine, cytosine, or thymine. The cytosine methylation of 5-hydroxymethylcytosine can also be assessed using the methods and procedures disclosed herein (see, for example, WO 2010 / 037001 and WO 2011 / 127136, which are incorporated herein by reference) and features thereof. The term "hypomethylated" or "hypermethylated" refers to the methylation state of a DNA molecule containing multiple CpG sites (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.), wherein a high proportion (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage within the range of 50%-100%) of the CpG sites are unmethylated or methylated, respectively.
[0146] The terms "cell-free deoxyribonucleic acid," "cell-free DNA," or "cfDNA" refer to fragments of deoxyribonucleic acid that circulate in body fluids (such as blood, sweat, urine, or saliva) and originate from one or more healthy cells and / or from one or more cancer cells.
[0147] The term "circulating tumor DNA" or "ctDNA" refers to fragments of deoxyribonucleic acid originating from tumor cells or other types of cancer cells that are released into an individual's bodily fluids (e.g., blood, sweat, urine, or saliva) as a result of biological processes such as apoptosis or necrosis of dying cells, or by active release from viable tumor cells. II. Method Overview
[0148] Figure 1A 1 is an exemplary flow chart describing an overall workflow 100 for classifying cancer in a sample according to one or more embodiments. The workflow 100 is performed by one or more entities (e.g., including medical personnel, sequencing equipment, analytical systems, etc.). The purpose of the workflow includes detecting and / or monitoring cancer in an individual. From a healthcare perspective, the workflow 100 can be used to supplement other existing cancer diagnostic tools. The workflow 100 can be used to provide early cancer detection and / or conventional cancer monitoring to better provide information for a treatment plan for an individual diagnosed with cancer. The overall workflow 100 can include more / fewer steps than those shown in FIG1 .
[0149] A healthcare provider performs sample collection 110. An individual requiring cancer classification visits their healthcare provider. The healthcare provider collects a sample for cancer classification. Examples of biological samples include, but are not limited to, a tissue biopsy, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial effusion, or peritoneal fluid from the subject. The sample includes genetic material belonging to the individual, which can be extracted and sequenced for cancer classification. Once the sample is collected, the sample is provided to a sequencing device. In addition to the sample, the healthcare provider may also collect other information related to the individual, such as biological sex, age, race, smoking status, any previous diagnosis, and the like.
[0150] The sequencing device performs sample sequencing 120. The laboratory clinician may perform one or more processing steps on the sample to prepare for sequencing. Once prepared, the clinician loads the sample into the sequencing device. Figure 2A and Figure 2B , further describes an example of an apparatus for sequencing. The sequencing apparatus typically extracts and separates nucleic acid fragments and sequences the fragments to determine the nucleobase sequences corresponding to the fragments. Sequencing can also include amplification of the nucleic acid material. Different sequencing processes include Sanger sequencing, fragment analysis, and next generation sequencing. Sequencing can be whole genome sequencing or targeted sequencing using a target panel. In the context of DNA methylation, bisulfite sequencing (e.g., in Figure 3 A and Figure 3 B) can determine the methylation status by bisulfite conversion of unmethylated cytosine at CpG sites. Sample sequencing 120 generates sequences of multiple nucleic acid fragments in the sample. In one or more embodiments, these sequences may include methylation state vectors, where each methylation state vector describes the methylation state of a CpG site on the fragment.
[0151] The analysis system performs pre-analysis processing 130. Figure 2B An example analysis system is described in . Pre-analysis processing 130 may include, but is not limited to, deduplicating sequence reads, determining coverage-related metrics, determining whether a sample is contaminated, removing contaminated fragments, base calling sequencing errors, and the like.
[0152] The analysis system performs one or more analyses 140. These analyses are statistical analyses or the application of one or more trained models to at least predict the cancer status of the individual from whom the sample was derived. Different genetic features can be evaluated and considered, such as methylation of CpG sites, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), other types of genetic mutations, etc. In the context of methylation, the analysis 140 can include a tissue purity assessment 142 (e.g., Figure 4A 、 Figure 4B 、 Figure 5 A. Figure 5 B. Figure 6 7 ), feature extraction 144 and applying a cancer classifier 146 to determine a cancer prediction (e.g., Figure 9A and Figure 9B ). Tissue purity assessment 142 involves applying a mixture model to deconvolute the proportion of tissue components that contribute DNA fragments to the sample. In general, tissue purity assessment 142 can be used to determine the proportion of methylated sequence reads shed from cancerous tissue in a sample compared to the proportion of methylated sequence reads shed from non-cancerous cells. Tissue purity assessment 142 may be particularly useful in deconvoluting heterogeneous tumors that contain multiple clonal populations with potentially different genetic characteristics. Mixture models can also be used to quantify cancer signals or classify cancer states based on determined proportions. Cancer classifier 146 inputs the extracted features to determine a cancer prediction. The cancer prediction can be a label or a value. A label can indicate a specific cancer state, for example, a binary label can indicate the presence or absence of cancer, a multi-class label can indicate one or more cancer types, cancer stages, etc. among the multiple cancer types screened. A value can indicate the likelihood of a specific cancer state, for example, the likelihood of having cancer and / or the likelihood of a specific cancer type.
[0153] The analysis system returns a prediction 150 to the healthcare provider. Prediction 150 can include a binary prediction of the presence or absence of cancer, a specific cancer type, cancer stage, tissue proportion, etc. The healthcare provider can develop or adjust the treatment plan based on the returned prediction 150. Treatment optimization is further described in Section VC, "Treatment."
[0154] Figure 1B is a flow chart of a method 160 for identifying multiple features for use in generating a classifier to predict a disease state (e.g., the presence or absence of a disease, the type of disease, and / or the tissue of origin of a disease) according to various embodiments. In some embodiments, the analysis system 200 performs the method 160 to process sequence reads of fragments from a nucleic acid sample. The method 160 includes, but is not limited to, the following steps: generating 165 sequence reads; training 170 probability models associated with each of a plurality of different disease states (e.g., different cancer types); applying 175 the probability models to determine a value based on the probability that the sequence read is derived from a sample associated with each of the plurality of disease states associated with each probability model; identifying 180 features by determining a count of sequence reads having a value exceeding a threshold; generating 185 a classifier using the features, and optionally applying 190 the classifier to predict a disease state and / or a tissue of origin associated with the disease state. II.A. Overview of Methylation
[0155] According to this specification, cfDNA fragments from individuals are processed, for example, unmethylated cytosine is converted into uracil, and then sequenced, and the sequence reads are compared with the reference genome to identify the methylation status of specific CpG sites in these DNA fragments. Each CpG site may be methylated or unmethylated. Compared with healthy individuals, the identification of informative fragments can provide insights into the cancer status of the subject. In some embodiments, when a sequence read or fragment is determined to be abnormally methylated, the sequence read or fragment is informative (for example, indicating the presence of a disease state). As is well known in the art, abnormal DNA methylation (compared to healthy controls) can have different effects, which may cause cancer. In the identification of informative cfDNA fragments, various challenges have arisen. First, the need to determine that a DNA fragment is informative is convincing only if it is compared with a group of control individuals. Therefore, if the number of individuals in the control group is small, the statistical variability caused by the small size of the control group will cause the determination result to lose credibility. Furthermore, methylation status may vary among a group of control individuals, making it difficult to fully account for these differences when determining whether a subject's DNA fragment is informative. On the other hand, methylation of cytosine at one CpG site can have a causal effect on methylation at subsequent CpG sites. Accounting for this dependency presents its own challenges.
[0156] Methylation typically occurs in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted into a methyl group, forming 5-methylcytosine. In particular, methylation can occur at a dinucleotide consisting of cytosine and guanine, which is referred to herein as a "CpG site." In other cases, methylation may also occur on a cytosine that is not part of a CpG site or on another nucleotide that is not cytosine; however, these occurrences are relatively rare. In this disclosure, for the sake of clarity, methylation is discussed with reference to CpG sites. Informative DNA methylation can be identified as hypermethylation or hypomethylation, both of which can indicate a cancerous state. Throughout this disclosure, if a DNA fragment contains more than a threshold number of CpG sites, and more than a threshold percentage of those CpG sites are methylated or unmethylated, then the hypermethylation and hypomethylation of the DNA fragment can be characterized.
[0157] The principles described herein are equally applicable to methylation detection in non-CpG backgrounds, including methylation of non-cytosine. In such embodiments, the laboratory wet assays for detecting methylation may be different from those described herein. Further, the methylation state vectors discussed herein may contain elements that are typically sites where methylation has or has not occurred (even if those sites are not specifically CpG sites). After this adjustment, the remaining processes described herein can remain unchanged, and therefore, the inventive concepts described herein are also applicable to those other forms of methylation. II.B. Exemplary Sequencers and Analysis Systems
[0158] Figure 2A and Figure 2B 2 is a flow chart illustrating a system and apparatus for sequencing a nucleic acid sample according to one embodiment. The illustrative flow chart includes an apparatus such as a sequencer 270 and an analysis system 200. The sequencer 270 and the analysis system 200 can work in tandem to perform one or more steps in the process described herein.
[0159] In various embodiments, the sequencer 270 receives the enriched nucleic acid sample 260. Figure 2A As shown, the sequencer 270 can include a graphical user interface 275 that enables a user to interact with specific tasks (e.g., start sequencing or terminate sequencing) and one or more loading stations 280 for loading sequencing cartridges containing the enriched fragment samples and / or for loading buffers necessary to perform sequencing assays. Thus, once the user of the sequencer 270 has provided the necessary reagents and sequencing cartridges to the loading stations 280 of the sequencer 270, the user can initiate sequencing by interacting with the graphical user interface 275 of the sequencer 270. Once initiated, the sequencer 270 performs sequencing and outputs sequence reads of the enriched fragments in the nucleic acid sample 260.
[0160] In some embodiments, the sequencer 270 is communicatively coupled to the analysis system 200. The analysis system 200 includes a certain number of computing devices for processing sequence reads to meet various application requirements, such as evaluating the methylation status of one or more CpG sites, variant base identification, or quality control. The sequencer 270 can provide the sequence reads to the analysis system 200 in a BAM file format. The analysis system 200 can be communicatively coupled to the sequencer 270 via wireless, wired, or a combination of wireless and wired communication technologies. Typically, the analysis system 200 is configured with a processor and a non-transitory computer-readable storage medium storing computer instructions, which, when the processor executes the computer instructions, causes the processor to process the sequence reads or perform one or more steps of any method or process disclosed herein.
[0161] In some embodiments, the sequence reads can be aligned with the reference genome using methods known in the art to determine alignment position information. The alignment position can generally describe the starting position and ending position of a region in the reference genome, which corresponds to the starting nucleotide base and the ending nucleotide base of a given sequence read. For methylation sequencing, the concept of alignment position information can be generalized to indicate the first CpG site and the last CpG site included in the sequence read based on the alignment with the reference genome. The alignment position information can further indicate the methylation status and position of all CpG sites in a given sequence read. The region in the reference genome may be associated with a gene or a fragment of a gene; therefore, the analysis system 200 can mark the sequence read with one or more genes aligned with the sequence read. In one embodiment, the fragment length (or size) is determined by the starting position and the ending position.
[0162] In various embodiments, for example, when a paired-end sequencing process is used, the sequence read consists of a pair of reads, represented by R_1 and R_2. For example, the first read R_1 can be sequenced from one end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 can be sequenced from the other end of the double-stranded DNA (dsDNA). Therefore, the nucleotide base pairs of the first read R_1 and the second read R_2 can be aligned in a consistent manner (e.g., in the opposite direction) with the nucleotide bases of the reference genome. The alignment position information derived from the read pairs R_1 and R_2 may include a starting position in the reference genome corresponding to one end of the first read (e.g., R_1), and an end position in the reference genome corresponding to one end of the second read (e.g., R_2). In other words, the starting position and end position in the reference genome represent positions that the nucleic acid fragment may correspond to in the reference genome. In one embodiment, the read pairs R_1 and R_2 can be assembled into fragments, and the fragments are used for subsequent analysis and / or classification. Output files in SAM (sequence alignment map) format or BAM (binary) format can be generated and exported for further analysis.
[0163] Now refer to Figure 2B , Figure 2B is a block diagram of an analysis system 200 for processing a DNA sample, according to one embodiment. The analysis system implements one or more computing devices to analyze the DNA sample. The analysis system 200 includes a sequence processor 210, a sequence database 215, a model database 225, one or more probability models 230 and / or one or more classifiers 240, and a parameter database 235. In some embodiments, the analysis system 200 performs one or more steps of the methods or processes disclosed herein.
[0164] The sequence processor 210 generates a methylation state vector for each fragment from the sample. For each CpG site on the fragment, the sequence processor 210 generates a methylation state vector for each fragment, which indicates the position of the fragment in the reference genome, the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment (methylated, unmethylated, or uncertain). This process is done by Figure 4B The sequence processor 210 may store the methylation state vector of each fragment in the sequence database 215. The data in the sequence database 215 may be organized such that the methylation state vectors from the same sample are associated.
[0165] Furthermore, a plurality of different models 230 can be stored in the model database 225 or retrieved for use with a test sample. In one example, the model is a trained cancer classifier 240 that uses feature vectors derived from informative segments to determine a cancer prediction for a test sample. The training and use of cancer classifiers are discussed elsewhere herein. The analysis system 200 can train one or more models 230 and / or one or more classifiers 240 and store various trained parameters in the parameter database 235. The analysis system 200 stores the models 230 and / or classifiers along with associated functions in the model database 225.
[0166] During inference, the machine learning engine 220 uses one or more models 230 and / or classifiers 240 to return outputs. The machine learning engine accesses the models 230 and / or classifiers 240 in the model database 225 and the trained parameters from the parameter database 235. For each model, the machine learning engine 220 receives inputs appropriate for that model and computes outputs based on the received inputs, parameters, and functions in each model that relate the inputs and outputs. In some use cases, the machine learning engine 220 further computes metrics related to the confidence level of the outputs computed by the model. In other use cases, the machine learning engine 220 computes other intermediate values used in the model. II.C. Detection Protocol
[0167] Figure 3 is a flow chart describing a process 300 for sequencing a nucleic acid, according to an embodiment. In some embodiments, process 300 is performed to generate sequence reads as part of step 110 of method 100 of FIG.
[0168] In step 310, a nucleic acid sample (e.g., DNA or RNA) is extracted from the subject. In this disclosure, unless otherwise indicated, DNA and RNA can be used interchangeably. That is, the embodiments described herein can be applicable to both DNA and RNA types of nucleic acid sequences. However, for the purpose of clarity and explanation, the examples described herein can focus on DNA. The sample can include nucleic acid molecules derived from any subset of the human genome (including the entire genome). The sample can include blood, plasma, serum, urine, feces, saliva, other types of body fluids, or any combination thereof. In some embodiments, the method for extracting a blood sample (e.g., syringe or finger prick) can be less invasive than the procedure for obtaining a tissue biopsy (which may require surgery). The extracted sample can include cfDNA and / or ctDNA. If the subject has a disease state, such as cancer, the free nucleic acids (e.g., cfDNA) in the sample extracted from the subject typically include detectable levels of nucleic acids that can be used to assess the disease state.
[0169] In step 315, the extracted nucleic acid (e.g., including cfDNA fragments) is treated to convert unmethylated cytosine to uracil. In some embodiments, method 300 uses bisulfite treatment of the sample, which converts unmethylated cytosine to uracil without converting methylated cytosine. For example, commercially available kits can be used to perform bisulfite conversion, such as EZ DNA Methylation TM -Gold, EZ DNA Methylation TM -Direct or EZ DNA Methylation TM -Lightning kit (from Zymo Research, Inc. (Irvine, CA)). In another embodiment, the conversion of unmethylated cytosine to uracil is accomplished by an enzymatic reaction. For example, the conversion can be accomplished using a commercially available kit, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).
[0170] In step 320, a sequencing library is prepared. In some embodiments, this preparation includes at least two steps. In the first step, an ssDNA adapter is added to the 3′-OH end of the bisulfite-converted ssDNA molecule using an ssDNA ligation reaction. In some embodiments, the ssDNA ligation reaction uses Circular Ligase II (Epicentre) to ligate the ssDNA adapter to the 3′-OH end of the bisulfite-converted ssDNA molecule, wherein the 5′ end of the adapter is phosphorylated and the bisulfite-converted ssDNA has been dephosphorylated (i.e., the 3′ end has a hydroxyl group). In another embodiment, the ssDNA ligation reaction uses thermostable 5′ App DNA / RNA ligase (available from New England BioLabs, Ipswich, Massachusetts, USA) to ligate the ssDNA adapter to the 3′-OH end of the bisulfite-converted ssDNA molecule. In this example, the 5′ end of the first UMI adapter is adenylated and the 3′ end is blocked. In another embodiment, the ssDNA ligation reaction uses T4 RNA ligase (available from New England Biolabs) to ligate ssDNA linkers to the 3'-OH ends of bisulfite-converted ssDNA molecules.
[0171] In the second step, the second strand DNA is synthesized in an extension reaction. For example, an extension primer that hybridizes to a primer sequence included in the ssDNA adapter is used in a primer extension reaction to form a double-stranded bisulfite-converted DNA molecule. Optionally, in some embodiments, the extension reaction uses an enzyme that can read uracil residues in the bisulfite-converted template strand.
[0172] Optionally, in a third step, dsDNA adapters are added to the double-stranded bisulfite-converted DNA molecules. The double-stranded bisulfite-converted DNA can then be amplified to add sequencing adapters. For example, PCR amplification using a forward primer comprising a P5 sequence and a reverse primer comprising a P7 sequence is used to add P5 and P7 sequences to the bisulfite-converted DNA. Optionally, during library preparation, a unique molecular identifier (UMI) can be added to a nucleic acid molecule (e.g., a DNA molecule) by adapter ligation. A UMI is a short nucleic acid sequence (e.g., 4-10 base pairs) that is added to the end of a DNA fragment during adapter ligation. In some embodiments, a UMI is a degenerate base pair that serves as a unique tag for identifying sequence reads derived from a specific DNA fragment. During PCR amplification after adapter ligation, the UMI is replicated along with the attached DNA fragment, which provides a method for identifying sequence reads from the same initial fragment in downstream analysis.
[0173] In optional step 325, nucleic acids (e.g., fragments) can be hybridized. Hybridization probes (also referred to herein as "probes") can be used to target and obtain nucleic acid fragments that provide information about a disease state. For a given workflow, these probes can be designed to anneal (or hybridize) to a target (complementary) strand of DNA or RNA. The target strand can be a "plus" strand (e.g., a strand that is transcribed into mRNA and subsequently translated into protein) or a complementary "minus" strand. The probes can range in length from tens, hundreds, or thousands of base pairs. In addition, the probes can cover overlapping portions of the target region.
[0174] In optional step 330, the hybridized nucleic acid fragments are captured and can be enriched, for example, using PCR amplification. In certain embodiments, the target DNA sequence can be enriched from the library. For example, this method is used when a targeted panel assay is performed on a sample. For example, the target sequence can be enriched to obtain a subsequent enrichment sequence that can be sequenced. Generally, any method known in the art can be used to separate and enrich the target nucleic acid of the probe hybridization. For example, as is well known in the art, a biotin moiety can be added to the 5' end of the probe (i.e., biotinylation) to facilitate the separation of the target nucleic acid hybridized with the probe using a streptavidin-coated surface (e.g., streptavidin-coated beads).
[0175] In step 335, sequence reads are produced from nucleic acid samples (for example, the sequence of enrichment).Sequencing data can be obtained from the dna sequence of enrichment by means known in the art.For example, the method can include second generation sequencing (NGS) technology, including synthesis technology (Illumina (Illumina)), pyrophosphate sequencing (454 Life Sciences (454 Life Sciences)), ion semiconductor technology (Ion Torrent sequencing), single molecule real-time sequencing (Pacific Biosciences (Pacific Biosciences)), connection sequencing (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies (Oxford Nanopore Technologies)) or double-end sequencing.In certain embodiments, sequencing while synthesizing with reversible dye terminator is used to carry out large-scale parallel sequencing.
[0176] In step 340, the sequence processor 210 may generate methylation information using the sequence reads. The methylation information determined from the sequence reads may then be used to generate a methylation state vector. Figure 4B According to the embodiment Figure 3300 for sequencing a cfDNA molecule. This is a diagrammatic representation of process 360 for obtaining a methylation state vector 352, beginning with process 300 for sequencing a cfDNA molecule. For example, the analysis system receives a cfDNA molecule 312, which, in this example, contains three CpG sites. As shown, the first and third CpG sites of cfDNA molecule 312 are methylated 314. During processing step 315, cfDNA molecule 312 is converted to generate a converted cfDNA molecule 322. During processing 315, the unmethylated second CpG site has its cytosine converted to uracil. However, the first and third CpG sites are not converted.
[0177] After conversion, a sequencing library 330 is prepared and sequenced to generate sequence reads 342. An analysis system (not shown) compares the sequence reads 342 to a reference genome 344. The reference genome 344 provides context that explains where in the human genome the cfDNA fragment originates. In this simplified example, the analysis system compares the sequence reads 342 so that the three CpG sites correspond to CpG sites 23, 24, and 25 (these are arbitrary reference identifiers used for ease of description). Thus, the analysis system can generate information about the methylation status of all CpG sites on the cfDNA molecule 312 and the locations where these CpG sites are mapped in the human genome. As shown, the methylated CpG sites on the sequence reads 342 are read as cytosine. In this example, cytosine only appears at the first and third CpG sites of the sequence reads 342, from which it can be inferred that the first and third CpG sites in the original cfDNA molecule are methylated. The second CpG site is read as thymine (U is converted to T during the sequencing process), so it can be inferred that the second CpG site in the original cfDNA molecule is unmethylated. With the two pieces of information, methylation status and position, the analysis system generates 200 methylation state vectors 352 for the cfDNA fragment 312. In this example, the resulting methylation state vector 352 is <M 23 , U 24 , M 25 >, where M corresponds to methylated CpG sites, U corresponds to unmethylated CpG sites, and the subscript numbers correspond to the position of each CpG site in the reference genome. II.D. Identification of informative fragments
[0178] In some embodiments, the analysis system uses the methylation state vector of the sample to determine the informational fragments of the sample. For example, for each nucleic acid molecule or fragment in the sample, the analysis system uses the methylation state vector corresponding to the nucleic acid molecule relative to the expected methylation state vector of the healthy sample to determine whether the nucleic acid molecule or fragment is an informational molecule or fragment (via analysis of sequence reads derived therefrom). In one embodiment, the analysis system calculates a p-value score for each methylation state vector, which describes the probability of observing the methylation state vector or other less likely methylation state vectors in the healthy control group (as described in, for example, U.S. Patent Application Publication No. 2019 / 0287652, which is incorporated herein by reference). The process of calculating the p-value score is further discussed in the II.BiP value screening section below. The analysis system can determine the sequence reads of the nucleic acid molecules or fragments with a methylation state vector below a threshold p-value score as informational fragments, and optionally screen out these sequence reads. In another embodiment, the analysis system further marks fragments having at least a certain number of CpG sites and methylation or unmethylation percentages exceeding a certain threshold as hypermethylated fragments and hypomethylated fragments, respectively. Hypermethylated fragments or hypomethylated fragments can also be referred to as extreme methylation abnormal fragments (UFXM). In other embodiments, the analysis system can use a variety of other probability models to determine informative molecules or fragments. Examples of other probability models include hybrid models, deep probability models, etc. In some embodiments, the analysis system can use any combination of multiple processes as described below to identify informative fragments. Through the identified informative fragments, the analysis system can screen the methylation state vector set of the sample for other processes, such as for training and deploying cancer classifiers. II.DIP value screening
[0179] In one embodiment, the analysis system calculates a p-value score for each methylation state vector compared to the methylation state vector of the fragment in the healthy control group. The p-value score describes the probability of observing a nucleic acid molecule with a methylation state that matches the methylation state vector in the healthy control group. In order to determine that a DNA fragment is informative, the analysis system uses a healthy control group in which the majority of the fragments are normally methylated. When performing this probabilistic analysis for determining informative fragments, the determination needs to be compared to a group of control subjects that constitute the healthy control group to be convincing. In order to ensure the robustness of the healthy control group, the analysis system can select a certain threshold number of healthy individuals to obtain a sample including DNA fragments. The following Figure 4B A method is described for generating a data structure for a healthy control group that an analysis system can use to calculate a p-value score. Figure 4C Methods for computing p-value scores using the generated data structure are described.
[0180] Figure 4B 4 is a flow chart describing a process 400 for generating a data structure for a healthy control group, according to an embodiment. To create the data structure for the healthy control group, the analysis system receives a plurality of DNA fragments (e.g., cfDNA) from a plurality of healthy individuals. For example, a methylation state vector for each fragment is identified via process 360.
[0181] For each fragment's methylation state vector, the analysis system subdivides 405 the methylation state vector into strings of CpG sites. In one embodiment, the analysis system subdivides 405 the methylation state vector so that the resulting strings are all less than a given length. For example, a methylation state vector of length 11 may be subdivided into strings of length less than or equal to 3, resulting in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, a methylation state vector of length 7 may be subdivided into strings of length less than or equal to 4, resulting in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If a methylation state vector is shorter than or the same length as a specified string, then the methylation state vector may be converted into a single string containing all of the CpG sites in the vector.
[0182] The analysis system 200 counts 410 these strings by, for each possible CpG site and methylation state probability in the vector, calculating the number of strings in the control group that have the specified CpG site as the first CpG site in the string and have that methylation state probability. For example, at a given CpG site and considering a string length of 3, there are 2^3 or 8 possible string configurations. At that given CpG site, for each of the 8 possible string configurations, the analysis system counts 410 the number of times each methylation state vector probability occurs in the control group. Continuing with this example, this may involve counting the following quantities: for each starting CpG site x in the reference genome, <M x , M x+1 , M x+2 >、 <M x , M x+1 , U x+2 >、...、 x , U x+1 , U x+2 >. The analysis system creates 415 a data structure that stores statistical counts of each starting CpG site and string likelihood.
[0183] There are several benefits to setting an upper limit on string length. First, the size of the data structure created by the analysis system increases dramatically depending on the maximum string length. For example, a maximum string length of 4 means that there are at least 2^4 numbers to count for each CpG site for a string of length 4. Increasing the maximum string length to 5 means that there are an additional 2^4, or 16, numbers to count for each CpG site, doubling the number of numbers to count (and the computer memory required) compared to the previous string length. Reducing the string size helps keep data structure creation and performance (e.g., for subsequent access, as described below) reasonable in terms of computation and storage. Second, one statistical consideration for limiting the maximum string length is to avoid overfitting downstream models that use string counts. If long strings at CpG sites do not biologically strongly influence outcomes (e.g., prediction of abnormalities that are predictive of the presence of cancer), then calculating probabilities based on large strings at CpG sites can be problematic because it requires a large amount of data that may not be available and therefore may be too sparse for the model to perform properly. For example, calculating the probability of abnormality / cancer conditional on the first 100 CpG sites would require counts of strings of length 100 in the data structure, ideally some of which exactly match the methylation states of the first 100. If only sparse counts of strings of length 100 are available, the data will be insufficient to determine whether a given string of length 100 in a test sample is informative.
[0184] Figure 4C is a flow chart describing a process 420 for identifying informative fragments from an individual, according to an embodiment. In process 420, the analysis system generates methylation state vectors 352 from the subject's cfDNA fragments. The analysis system processes each methylation state vector as follows.
[0185] For a given methylation state vector, the analysis system enumerates 430 all possibilities of methylation state vectors having the same starting CpG site and the same length (i.e., the set of CpG sites) as the given methylation state vector. Since each methylation state is typically either methylated or unmethylated, there are actually two possible states at each CpG site, and therefore the number of different possibilities for the methylation state vector depends on a power of 2, such that a methylation state vector of length n will have associated 2 n For a methylation state vector containing an uncertain state for one or more CpG sites, the analysis system may enumerate 430 the likelihood of the methylation state vector by considering only those CpG sites with observed states.
[0186] The analysis system 200 accesses the data structure of the healthy control group and calculates 440 the probability of observing each methylation state vector likelihood for the identified starting CpG site and the methylation state vector length. In one embodiment, when calculating the probability of observing a given likelihood, Markov chain probability is used to model the joint probability calculation. In other embodiments, a calculation method other than Markov chain probability is used to determine the probability of observing each methylation state vector likelihood.
[0187] The analysis system uses the calculated probabilities for each possibility to calculate a p-value score 450 for the methylation state vector. In one embodiment, this includes identifying a calculated probability corresponding to a probability that matches the methylation state vector in question. Specifically, this is a probability that has the same set of CpG sites as the methylation state vector, or similarly, has the same starting CpG site and length as the methylation state vector. The analysis system sums the calculated probabilities for any possibilities with a probability less than or equal to the identified probability to generate a p-value score.
[0188] This p-value represents the probability of observing the methylation state vector of the fragment or other less likely methylation state vectors in the healthy control group. Therefore, a low p-value score generally corresponds to a methylation state vector that is rare in healthy individuals and causes the fragment to be marked as informative relative to the healthy control group. A high p-value score is generally associated with a methylation state vector that is expected to be present in relatively healthy individuals. For example, if the healthy control group is a non-cancerous group, a low p-value indicates that the fragment is informative relative to the non-cancerous group and therefore may indicate the presence of cancer in the test subject.
[0189] As described above, the analysis system calculates a p-value score for each of a plurality of methylation state vectors, each of which represents a cfDNA fragment in the test sample. To identify which fragments are informative, the analysis system may filter 460 the set of methylation state vectors based on their p-value scores. In one embodiment, filtering is performed by comparing the p-value score to a threshold and retaining only those fragments that fall below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001, 0.0001, or a similar order of magnitude.
[0190] Based on the results of the example process, the analysis system generated a median (range) of 2,800 (1,500-12,000) fragments with informative methylation patterns in participants without cancer during training, and a median (range) of 3,000 (1,200-220,000) fragments with informative methylation patterns in participants with cancer during training. These filtered fragment sets with informative methylation patterns can be used for downstream analysis as described below.
[0191] In one embodiment, the analysis system uses a 455 sliding window to determine the likelihood of a methylation state vector and calculate a p-value. The analysis system only needs to enumerate the likelihood and calculate the p-value for a window containing consecutive CpG sites, where the length of the window (CpG sites) is shorter than at least some fragments (otherwise, the window would be meaningless), rather than enumerating the likelihood and calculating the p-value for the entire methylation state vector. The window length can be static, user-defined, dynamic, or selected in some other way.
[0192] When calculating the p-value for a methylation state vector that is larger than a window, the window starts at the first CpG site in the vector and identifies a set of contiguous CpG sites in the vector that are within the window. The analysis system calculates a p-value score for the window that includes the first CpG site. The analysis system then "slides" the window to the second CpG site in the vector and calculates another p-value score for the second window. Therefore, for a window size of l and a methylation vector length of m, each methylation state vector will generate m–l+1 p-value scores. After the p-value calculation is completed for each part of the vector, the lowest p-value score in all sliding windows is used as the overall p-value score for the methylation state vector. In another embodiment, the analysis system aggregates the p-value scores of the methylation state vectors to generate an overall p-value score.
[0193] The use of sliding windows helps reduce the number of methylation state vector possibilities that need to be enumerated and the corresponding probability calculations that would otherwise have to be performed. As a practical example, a fragment may have up to 54 CpG sites. Instead of calculating the probabilities of 2^54 (approximately 1.8×10^16) possibilities to generate a single p-value, the analysis system can use (for example) a window size of 5 to perform 50 p-value calculations for each of the 50 windows of the methylation state vector for that fragment. Each of these 50 calculations enumerates 2^5 (32) methylation state vector possibilities, for a total of 50×2^5 (1.6×10^3) probability calculations. This significantly reduces the number of calculations that need to be performed without significantly affecting the accurate identification of informative fragments.
[0194] In the embodiment of an uncertain state, the analysis system can calculate a p-value score that sums the CpG sites with an uncertain state in the methylation state vector of the fragment. The analysis system identifies all possibilities that are consistent with all methylation states in the methylation state vector except the uncertain state. The analysis system can assign a probability to the methylation state vector that is the sum of the probabilities of these identified possibilities. For example, the analysis system calculates the methylation state vector<M1,I2,U3> The probability of methylation as the state vector<M1,M2,U3> and<M1,U2,U3> =The sum of the probabilities of the possibilities because the methylation states of CpG sites 1 and 3 have been observed and are consistent with the methylation states of the fragment at CpG sites 1 and 3. This method of summing the probabilities of CpG sites with uncertain states uses a probability calculation of up to 2^i possibilities, where i represents the number of uncertain states in the methylation state vector. In another embodiment, a dynamic programming algorithm can be used to calculate the probabilities of methylation state vectors with one or more uncertain states. Advantageously, the dynamic programming algorithm runs in linear computation time.
[0195] In one embodiment, by caching at least some calculations, the computational load of calculating probabilities and / or p-value scores can be further reduced. For example, the analysis system can cache the probability calculations of the methylation state vector (or its window) possibilities in a transient or persistent memory. If other fragments have the same CpG site, caching the probabilities of these possibilities allows the p-value scores to be calculated efficiently without recalculating the probabilities of the potential possibilities. Equivalently, the analysis system can calculate the p-value scores for each possibility of the methylation state vector associated with the CpG site set in the vector (or its window). The analysis system can cache these p-value scores for determining the p-value scores of other fragments including the same CpG site. In general, the p-value scores of the possibilities of the methylation state vector with the same CpG site can be used to determine the p-value scores of different possibilities under the same CpG site set. II.D.II. Hypermethylated and Hypomethylated Fragments
[0196] In some embodiments, the analysis system determines an informative fragment as a fragment having more than a threshold number of CpG sites and either having more than a threshold percentage of methylated CpG sites or having more than a threshold percentage of unmethylated CpG sites; the analysis system identifies such fragments as hypermethylated fragments or hypomethylated fragments. Example thresholds for fragment (or CpG site) length include more than 3, 4, 5, 6, 7, 8, 9, 10, etc. Example thresholds for methylation or unmethylation percentage include more than 80%, 85%, 90% or 95%, or any other percentage in the range of 50%-100%. II.E. Blocks of the Reference Genome
[0197] Figure 5 It is a diagram of a block of reference genome according to an embodiment. Sequence processor 210 can partition reference genome (or a subset of reference genome) in one or more stages, for example, for use cases related to targeted methylation assays. For example, sequence processor 210 is divided into CpG site blocks with reference genome. When the interval between two adjacent CpG sites exceeds a certain threshold value (for example, greater than 200 base pairs (bp), 300bp, 400bp, 500bp, 600bp, 700bp, 800bp, 900bp or 1,000bp equivalent values), a block is defined. Therefore, the base pair size in the block can be different. For each block, sequence processor 210 can be subdivided into a window of certain length, for example 500bp, 600bp, 700bp, 800bp, 900bp, 1,000bp, 1,100bp, 1,200bp, 1,300bp, 1,400bp or 1,500bp equivalent values. In other embodiments, the window length can be from 200bp to 10 kilobase pairs (kbp), from 500bp to 2kbp, or about 1kbp. Windows (e.g., adjacent windows) can overlap a certain number of base pairs or a certain percentage of length, such as 10%, 20%, 30%, 40%, 50% or 60% values. A window can be separated between two adjacent CpG sites that exceed a threshold value (e.g., greater than 200 base pairs (bp), 300bp, 400bp, 500bp, 600bp, 700bp, 800bp, 900bp or 1,000bp values).
[0198] The sequence processor 210 can use a windowing process to analyze sequence reads derived from DNA fragments. Specifically, the sequence processor 210 scans blocks window by window and reads the fragments within each window. These fragments can be derived from tissue and / or high signal cfDNA. High signal cfDNA samples can be determined by binary classification models, cancer staging or other indicators. By partitioning the reference genome (e.g., using blocks and windows), the sequence processor 210 can promote computational parallelization. In addition, the sequence processor 210 can reduce the computational resources used to process the reference genome by targeting the base pair portion that includes the CpG site while skipping other portions that do not include the CpG site. III. Model-based Feature Engineering and Classification III.A. Model-based Feature Engineering
[0199] According to one embodiment, Figure 8As shown, the present disclosure relates to feature engineering based on the model for deriving features that can be used to classify disease states. As described elsewhere herein, disease states can be the presence or absence of a disease, the type of disease, and / or the tissue of origin of a disease. For example, as described herein, disease states can be the presence or absence of cancer, the type of cancer, and / or the tissue of origin of a cancer. Cancer type and / or cancer tissue of origin can be selected from the group consisting of: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis, renal cancer other than urothelial cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal cancer, gastric cancer, hepatobiliary cancer caused by hepatocellular carcinoma, hepatobiliary cancer caused by cells other than hepatocellular cancer, pancreatic cancer, squamous cell carcinoma of the upper digestive tract, upper digestive tract cancer other than squamous, head and neck cancer, lung cancer (such as lung adenocarcinoma, small cell lung cancer, squamous cell lung cancer, and cancer other than adenocarcinoma or small cell lung cancer), neuroendocrine cancer, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia, as well as other cancer types.
[0200] In step 810, as described elsewhere herein, a first plurality of sequence reads is generated from a first reference sample having a first disease state, and a second plurality of sequence reads is generated from a second reference sample having a second disease state. The first plurality of sequence reads and / or the second plurality of sequence reads can be greater than 10,000, greater than 50,000, greater than 100,000, greater than 200,000, greater than 500,000, greater than 1,000,000, greater than 2,000,000, greater than 5,000,000, or greater than 10,000,000 sequence reads. As used herein, a "reference sample" is a sample obtained from a subject having a known disease state.
[0201] In some embodiments, one or more reference samples with one or more known disease states can be used to train one or more probabilistic models, which in turn can be used to derive features for classifying the disease state of an unknown test sample.
[0202] The sample can be a genomic DNA (gDNA) sample or a cell-free DNA (cfDNA) sample. The reference sample can be a blood, plasma, serum, urine, feces, and saliva sample. Alternatively, the reference sample can be whole blood, blood fraction, tissue biopsy, pleural fluid, pericardial effusion, cerebrospinal fluid, and peritoneal fluid. In some embodiments, the first reference sample is obtained from a subject known to have cancer, and the second reference sample is obtained from a healthy subject or a non-cancer subject. In some embodiments, the first reference sample is obtained from a subject known to have a first type of cancer (e.g., lung cancer), and the second reference sample is obtained from a subject known to have a second type of cancer (e.g., breast cancer). In yet other embodiments, the first reference sample is obtained from a subject known to have a first disease origin tissue (e.g., lung disease), and the second reference sample is obtained from a second disease state origin tissue (e.g., liver disease).
[0203] In step 815, the machine learning engine 220 trains a first probability model 230 and a second probability model 230 from the first plurality of sequence reads and the second plurality of sequence reads (generated in step 110), respectively, each probability model being associated with a different disease state in one or more possible disease states. As previously described, the disease state can be the presence or absence of cancer, the type of cancer, and / or the tissue of origin of the cancer. In various embodiments, the training data is divided into K subsets (folds) for K-fold cross-validation. The folds can be balanced for factors such as cancer / non-cancer status, tissue of origin, cancer stage, age (e.g., grouped into age groups every 10 years), gender, race, and smoking status. Data from the K-1 folds can be used as training data for the probability model, and the remaining folds can be used as test data.
[0204] The machine learning engine 220 trains a first probability model and a second probability model 230 for the first disease state and the second disease state, respectively, by fitting each probability model 230 to the first plurality of sequence reads and the second plurality of sequence reads, respectively. For example, in one embodiment, the first probability model is fit using a first plurality of sequence reads derived from one or more samples of subjects known to have cancer, and the second probability model is fit using a second plurality of sequence reads derived from one or more samples of healthy subjects or non-cancer subjects. In other embodiments, the first probability model can be trained for a first type of cancer or a first tissue of origin, and the second probability model can be trained for a second type of cancer or a second tissue of origin. It will be understood by those skilled in the art that any number of disease state probability models can be trained using sequence reads derived from one or more samples obtained from subjects having any of a variety of possible disease states. For example, in some embodiments, additional cancer-specific probability models (i.e., models for additional types of cancer and / or tissues of origin) can be trained for a third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, and so on (e.g., up to twenty, thirty, or more) specific types of cancer, and these probability models are used to determine the probability that a sequence read derived from a training set or an unknown cancer type is more likely to be derived from one cancer type (or cancer tissue of origin) rather than another cancer type (or cancer tissue of origin), as described elsewhere herein.
[0205] As used herein, a "probabilistic model" is any mathematical model that can assign a probability to a sequence read based on the methylation state of one or more sites on the read. During training, the machine learning engine 220 is fitted to sequence reads from one or more samples of subjects with a known disease and can be used to utilize methylation information or methylation state vectors (e.g., previously described Figure 3 4 ) determines the probability of a sequence read indicating a disease state. Specifically, in one embodiment, the machine learning engine 220 determines the observed methylation rate of each CpG site in the sequence read. The methylation rate represents the fraction or percentage of base pairs methylated within the CpG site. The trained probability model 230 can be parameterized by the product of the methylation rates. In general, any known probability model for assigning probabilities to sequence reads from a sample can be used. For example, the probability model can be a binomial model, in which each site (e.g., a CpG site) on a nucleic acid fragment is assigned a methylation probability; or it can be an independent site model, in which the methylation of each CpG is specified by a different methylation probability, in which the methylation of one site is considered to be independent of the methylation of one or more other sites on the nucleic acid fragment.
[0206] In some embodiments, the probability model 230 is a Markov model in which the probability of methylation of each CpG site depends on the methylation status of a certain number of previous CpG sites in the sequence read or in the nucleic acid molecule from which the sequence read is derived. See, for example, U.S. patent application Ser. No. 16 / 352,602, filed on March 13, 2019, entitled “Anomalous Fragment Detection and Classification.”
[0207] In some embodiments, the probability model 230 is a "hybrid model" that is fitted using a mixture of components from the underlying model. For example, in some embodiments, multiple independent site models can be used to determine a hybrid component, where the methylation (e.g., methylation rate) of each CpG site is considered to be independent of the methylation of other CpG sites. Using the independent site model, the probability assigned to a sequence read or the nucleic acid molecule from which it is derived is the product of the methylation probability of each CpG site that the sequence read is methylated and the methylation probability of each CpG site that the sequence read is not methylated. According to this embodiment, the machine learning engine 220 determines the methylation rate of each of the hybrid components. The hybrid model is parameterized by the sum of the hybrid components, each of which is associated with the product of the methylation rates. The probability model Pr of n hybrid components can be expressed as: For the input fragment, m i ∈{0,1} represents the observed methylation status of the fragment at position i in the reference genome, where 0 indicates unmethylated and 1 indicates methylated. The score assigned to each hybrid component k is f k , where f k ≥0 and The probability of methylation at position i in the CpG site of hybrid component k is β ki Therefore, the probability of being unmethylated is 1-β ki The number n of mixing components can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.
[0208] In some embodiments, the machine learning engine 220 uses maximum likelihood estimation to fit the probability model 230 to identify the parameter set {β ki ,f k}, this parameter set maximizes the log-likelihood of all segments derived from the disease state under the condition of applying a regularization penalty of strength r to each methylation probability. For a total of N segments, the quantity to be maximized can be expressed as:
[0209] It will be appreciated by those skilled in the art that other means can be used to fit a probability model or identify parameters that maximize the log-likelihood of all sequence reads derived from a reference sample. For example, in one embodiment, Bayesian fitting (e.g., using Markov chain Monte Carlo) is used, wherein each parameter is not assigned a single value, but is associated with a distribution. In other embodiments, gradient-based optimization is used, wherein the gradient of likelihood (or log-likelihood) is used to progressively reach an optimal value in parameter space relative to the parameter value. In other embodiments, expectation maximization is performed, wherein a set of potential parameters (such as, the identity of the hybrid component from which each fragment is derived) is set to its expected value under the previous model parameters, and then the parameters of the assigned model are maximized under the assumed value conditions of these potential variables. These two steps are then repeated until convergence.
[0210] In step 820, a plurality of training sequence reads are generated from a training sample. The plurality of training sequence reads may be more than 10,000, more than 50,000, more than 100,000, more than 200,000, more than 500,000, more than 1,000,000, more than 2,000,000, more than 5,000,000, or more than 10,000,000 sequence reads. As used herein, a "training sample" is a sample obtained from a known disease state that can be used to generate sequence reads, which are then applied to a first and / or second probability model to generate features that can be used for disease state classification. In step 825, the analysis system 200 applies the first probability model and the second probability model 230 to determine a first probability value and a second probability value for each sequence read in the plurality of training sequence reads. The first probability value and the second probability value are determined based on the probability that the sequence read is derived from a sample associated with the first disease state and the second disease state, respectively. Analysis system 200 can repeat step 130 for any additional probability models 230 (eg, probability models trained from sequence reads of a third, fourth, fifth, etc. reference sample) (not shown).
[0211] In step 830, one or more features are identified by comparing the first probability value and the second probability value of each of the multiple training sequence reads. In general, a variety of methods can be used to compare the first probability value and the second probability value and identify features. For example, in one embodiment, one or more features include a count of abnormal sequence reads in the multiple training sequence reads whose first probability value is greater than the second probability value. The count can be a binary count, a total count of abnormal sequence reads, or a total count of anonymous methylated sequence reads. In another embodiment, one or more features include a count of sequence reads or fragments containing a specific methylation pattern. For example, one or more features can be a count of sequence reads or fragments that are fully methylated at each CpG site, a count of sequence reads or fragments that are partially methylated (e.g., at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95% methylated). In another embodiment, one or more features are identified using the output of a discriminant classifier trained within a single genomic region (e.g., the discriminant classifier can be a multilayer perceptron or a convolutional neural network model). In another embodiment, comparing the first probability value and the second probability value includes determining a ratio of the first probability value and the second probability value, and the one or more features include sequence read counts of sequence reads exceeding a ratio threshold.
[0212] In another embodiment, the first probability value or the second probability value is a log-likelihood value. For example, the analysis system 200 can calculate the log-likelihood ratio R using the fitted probability models associated with the first disease state and the second disease state, respectively. Specifically, the log-likelihood ratio can be calculated using the probability Pr of observing the methylation pattern on the fragments of the samples associated with the first disease state and the second disease state:
[0213] The analysis system 200 can use multiple levels of thresholds to identify features. For example, the levels include thresholds 1, 2, 3, 4, 5, 6, 7, 8, and 9. In some embodiments, a smoothing function can be applied. For example, in response to determining that R is (e.g., significantly) less than the level value, the analysis system 200 assigns a feature value of ˜0; in response to determining that R is equal to the level value, the analysis system 200 assigns a feature value of 0.5; in response to determining that R is (e.g., significantly) greater than the level value, the analysis system 200 assigns a feature value of ˜1. Each level represents a different threshold at which the fragment (from which the sequence reads were generated) is more likely to originate from a sample associated with a disease state rather than a healthy sample. The analysis system 200 can use thresholds to determine the count of abnormal fragments that can be used as features.
[0214] By screening using thresholds, the analysis system 200 can treat certain fragments as outliers because these fragments are less likely to be present in healthy samples. Therefore, abnormal fragments can be considered more likely to be associated with a disease state or cancer sample (e.g., originate from a disease state or cancer sample). The number of features between different levels may vary, for example, based on corresponding thresholds, one level may have a different number of features than another level. In other embodiments, the analysis system 200 uses a different number of levels or other thresholds. Other means for identifying features, or other means for ranking identified features based on a measure of the features in distinguishing different disease states (e.g., using mutual information to determine a measure of the information content of a feature in distinguishing two disease states) are described elsewhere herein.
[0215] In other embodiments, the analysis system 200 can use different types of ratios or equations to identify multiple features. The machine learning engine 220 can determine that a segment indicates a disease state (e.g., cancer) based on whether at least one of the log-likelihood ratios considered for various disease states is above a threshold.
[0216] Subsequently, as described in further detail elsewhere herein, the plurality of features can be used to train a disease state classifier. For example, in some embodiments, the plurality of features can be used to train a classifier to classify the presence or absence of cancer, the type of cancer, and / or the tissue of origin of the cancer. III.B. Classification of Tissue of Origin of Disease States
[0217] According to another embodiment, the machine learning engine 220 trains probabilistic models 230, each model being associated with a different disease state from a set of multiple disease states.
[0218] The machine learning engine 220 trains the probability model 230 using one or more sets of sequence reads, wherein each set of the one or more sets of sequence reads is generated from a different disease state in the set of multiple disease states. The disease states can include any number of cancer types or tissues of origin selected from the group consisting of breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial carcinoma of the renal pelvis, renal cancer other than urothelial cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal cancer, gastric cancer, hepatobiliary cancer arising from hepatocellular cells, hepatobiliary cancer arising from cells other than hepatocellular cells, pancreatic cancer, squamous cell carcinoma of the upper gastrointestinal tract, upper gastrointestinal cancer other than squamous, head and neck cancer, lung cancer (such as lung adenocarcinoma, small cell lung cancer, squamous cell lung cancer, and cancers other than adenocarcinoma or small cell lung cancer), neuroendocrine cancer, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma and leukemia, as well as other cancer types.
[0219] The machine learning engine 220 trains the probability model 230 for each of the multiple disease states by fitting the probability model 230 to the sequence reads derived from each sample corresponding to each disease state. For example, in some embodiments, a probability model can be trained for a specific type of cancer. According to this embodiment, cancer-specific probability models can be trained for first, second, third, and other specific types of cancer, and these probability models are used to assess the cancer type (e.g., the cancer type of an unknown test sample). For example, a set of sequence reads derived from one or more samples associated with lung cancer is used to fit a lung cancer-specific probability model. As another example, a set of sequence reads derived from one or more samples associated with breast cancer is used to fit a breast cancer-specific probability model. In some embodiments, tissue-specific probability models can be trained for first, second, and third tissue types, and these probability models are used to assess the disease state tissue of origin. For example, a set of sequence reads derived from a first tissue type (e.g., derived from a lung tissue sample, such as a lung biopsy) can be used to fit a first tissue of origin probability model, and a set of sequence reads derived from a second tissue type (e.g., derived from a liver tissue sample, such as a liver biopsy) can be used to fit a second tissue of origin probability model. Alternatively, in some embodiments, a set of sequence reads from one or more samples of subjects known to have cancer is used to fit a cancer probability model, and a set of sequence reads from one or more samples of healthy subjects or non-cancer subjects is used to fit a non-cancer specific probability model. It will be understood by those skilled in the art that sequence reads derived from one or more samples obtained from subjects with any of a variety of possible disease states can be used to train any number of disease state probability models. For example, in some embodiments, multiple sequence reads can be generated from 3, 4, 5, 6, 7, 8, 9, 10 or more reference samples, each reference sample being obtained from one or more subjects with a different disease state (e.g., a different type of cancer), and these sequence reads are used to train 3, 4, 5, 6, 7, 8, 9, 10 or more probability models.
[0220] During training, methylation information or methylation state vectors (e.g., previous Figure 34 ) is trained on sequence reads indicating a disease state. Specifically, the machine learning engine 220 determines the observed methylation rate of each CpG site in the sequence reads. The methylation rate represents the fraction or percentage of base pairs methylated within the CpG site. The trained probability model 230 can be parameterized by the product of the methylation rates. As previously described, any known probability model for assigning probabilities to sequence reads from a sample can be used. For example, the probability model can be a binomial model, in which each site (e.g., a CpG site) on a nucleic acid fragment is assigned a methylation probability; or it can be an independent site model, in which the methylation of each CpG is specified by a different methylation probability, in which the methylation of one site is considered to be independent of the methylation of one or more other sites on the nucleic acid fragment.
[0221] In some embodiments, a Markov model is provided in which the methylation probability of each CpG site depends on the methylation status of a number of previous CpG sites in the sequence read or in the nucleic acid molecule from which the sequence read is derived. See, for example, U.S. patent application Ser. No. 16 / 352,602, filed on March 13, 2019, entitled “Anomalous Fragment Detection and Classification.”
[0222] In some embodiments, the probability model 230 is a "hybrid model" that is fitted using a mixture of components from the underlying model. For example, in some embodiments, multiple independent site models can be used to determine a hybrid component, where the methylation (e.g., methylation rate) of each CpG site is considered to be independent of the methylation of other CpG sites. Using the independent site model, the probability assigned to a sequence read or the nucleic acid molecule from which it is derived is the product of the methylation probability of each CpG site that the sequence read is methylated and the methylation probability of each CpG site that the sequence read is not methylated. According to this embodiment, the machine learning engine 220 determines the methylation rate of each of the hybrid components. The hybrid model is parameterized by the sum of the hybrid components, each of which is associated with the product of the methylation rates. The probability model Pr of n hybrid components can be expressed as: For the input fragment, m i ∈{0,1} represents the observed methylation status of the fragment at position i in the reference genome, where 0 indicates unmethylated and 1 indicates methylated. The score assigned to each hybrid component k is f k , where f k ≥0 and The probability of methylation at position i in the CpG site of hybrid component k is β kiTherefore, the probability of being unmethylated is 1-β ki The number n of mixing components can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.
[0223] In some embodiments, the machine learning engine 220 uses maximum likelihood estimation to fit the probability model 230 to identify the parameter set {β ki ,f k}, this parameter set maximizes the log-likelihood of all segments derived from the disease state under the condition of applying a regularization penalty of strength r to each methylation probability. For a total of N segments, the quantity to be maximized can be expressed as:
[0224] In step 130, the analysis system 200 applies the probability model 230 to calculate a value for each sequence read in the second set of sequence reads (e.g., different from the first set of sequence reads generated in step 110). These values are calculated based on at least the probability that the sequence read (and corresponding fragment) originates from a sample associated with the disease state of the probability model 230. The analysis system 200 can repeat step 130 for each different probability model 230. In some embodiments, the analysis system 200 calculates the value using the log-likelihood ratio R of the fitted probability model associated with certain disease states. Specifically, the log-likelihood ratio can be calculated using the probability Pr of observing a methylation pattern on fragments of samples associated with the disease state and healthy samples: In other embodiments, analysis system 200 may use different types of ratios or equations to calculate the value. Machine learning engine 220 may determine that a segment indicates a disease state (e.g., cancer) based on whether at least one of the log-likelihood ratios considered for various disease states is above a threshold. III.C. Feature Selection
[0225] Figure 6 is a diagram of a process for determining features for training a classifier according to an embodiment. As previously described, the machine learning engine 220 trains a probability model 230 associated with a disease state. Figure 6 In the example shown, the probability model 230 ("tissue model") is associated with non-cancer (healthy), breast cancer, and lung cancer. The analysis system 200 processes one or more cfDNA and / or tumor samples to obtain fragments and uses the probability model 230 to assign values to the fragments associated with non-cancer (healthy), breast cancer, and lung cancer. The analysis system 200 can use information from sequence reads of cfDNA and / or tumor samples to identify features for classifiers. In some embodiments, the analysis system 200 can obtain and assign fragments to each window from a partitioned reference genome, such as Figure 5 The analysis system 200 aggregates the segments from the windows into sequences to determine features for the classifier.
[0226] The analysis system 200 identifies features by determining the count of sequence reads having a value exceeding a threshold. In an embodiment where the value is based on a log-likelihood ratio R, the threshold is a threshold ratio. The analysis system 200 can use multiple levels of thresholds to identify features. For example, the levels include thresholds 1, 2, 3, 4, 5, 6, 7, 8, and 9. Each level represents a different threshold at which the fragment (from which the sequence read was generated) is more likely to originate from a sample associated with a disease state rather than a healthy sample. The analysis system 200 can use thresholds to determine the count of abnormal fragments that can be used as features.
[0227] By using thresholds for screening, the analysis system 200 can deem certain fragments as outliers because these fragments are unlikely to be present in healthy samples. Thus, abnormal fragments can be considered more likely to be associated with a disease state or cancer sample (e.g., to be derived from a disease state or cancer sample). The number of features between different levels may vary. In other embodiments, the analysis system 200 uses a different number of levels or other thresholds. In other embodiments, the analysis system 200 may use other methods or scores (such as p-values) to screen fragments. In some embodiments, the analysis system 200 calculates a p-value for a methylation state vector, which describes the probability of observing the methylation state vector or other less likely methylation state vectors in a healthy control group. To determine a fragment as informative, the analysis system 200 uses a healthy control group in which the majority of the fragments are normally methylated (see, for example, U.S. patent application Ser. No. 16 / 352,602, entitled “Anomalous Fragment Detection and Classification,” filed on March 13, 2019).
[0228] The analysis system 200 may repeat the steps of identifying features for each probability model. Thus, the analysis system 200 may identify features of one or more disease states associated with the probability model. Figure 6 In the example shown, analysis system 200 identifies one or more characteristics of breast cancer and lung cancer.
[0229] In some embodiments, the analysis system 200 ranks the identified features based on their ability to distinguish between different disease states. For example, a feature is informative if it can distinguish a certain type of cancer from other types of cancer or healthy samples. The analysis system 200 can use mutual information to determine a measure of the information content of a feature in distinguishing between two disease states. For each pair of different disease states, the analysis system 200 can designate one disease state (e.g., cancer type A) as a positive class and the other disease state (e.g., cancer type B) as a negative class.
[0230] Mutual information can be calculated using the estimated proportion of samples of positive and negative types (e.g., cancer types A and B) whose features are expected to be non-zero in the outcome assay. For example, if a feature frequently appears in healthy cfDNA, the analysis system 200 determines that the feature is unlikely to frequently appear in cfDNA associated with various types of cancer. Therefore, the feature may be a weaker metric in distinguishing disease states. When calculating mutual information I, variable X is a feature (e.g., binary), and variable Y represents the disease state, such as type A or type B cancer: p(1|A)=f A +f H -f H f A The joint probability mass function of X and Y is p(x, y), and the marginal probability mass functions are p(x) and p(y). The analysis system 200 can assume that feature missingness is uninformative and that the prior probabilities of the two disease states are equal, e.g., p(Y=A)=p(Y=B)=0.5. The probability of observing a given binary feature of type A cancer (e.g., in cfDNA) is denoted as p(1|A), where f A is the probability of observing the signature in a tumor ctDNA sample (or high signal cfDNA sample) associated with type A cancer, and f H is the probability of observing the feature in a healthy or non-cancer cfDNA sample.
[0231] In some embodiments, f is estimated by the proportion of cancer patients whose cfDNA would be expected to include non-zero feature values. AWhen the training data for type A cancer consists of cfDNA samples, the proportion can be simply estimated as the proportion of cfDNA samples in which the feature is observed. When the training data includes tumor samples, a correction can be applied to account for the lower proportion of tumor-derived fragments in cfDNA compared to tumor. For the N fragments in the tumor sample that are determined to have a value greater than the threshold (e.g., from step 140), the analysis system 200 calculates the probability r of detecting each of these fragments in cfDNA from the patient as: The probability of observing at least one fragment in the patient's cfDNA can then be calculated as p(N cfDNA >0)=1-(1-r) N In order to estimate f A , p(N cfDNA >0) can be averaged across all type A cancer training samples, where the probability is assigned to 1 for cfDNA samples with the feature, 0 for cfDNA samples lacking the feature, and 1-(1-r) for tumor samples. N In some embodiments, these estimates are based on a predetermined assumed value for the tumor fraction in cfDNA of patients with early-stage cancer (e.g., 0.1%), the cfDNA sequencing depth (e.g., 1000x), and the tumor sequencing depth (e.g., 25x) applied to the patient's final assay. H , the analysis system 200 uses a portion of the positive samples to determine how many additional samples would result in a positive detection classification at a greater sequencing depth. III.D. Classification
[0232] The analysis system 200 uses these features to generate a classifier. The classifier is trained to predict the tissue of origin associated with the disease state for the input sequence reads of the test sample from the test subject. The analysis system 200 can select a predetermined number (e.g., 128, 256, 512, 1024) of the top-ranked features for each pair of disease states to train the classifier, for example, based on a mutual information calculation or another calculated metric. The predetermined number can be regarded as a hyperparameter selected based on cross-validated performance. The analysis system 200 can also select features from the region of the reference genome that are determined to provide more information when distinguishing a pair of disease states. In various embodiments, the analysis system 200 retains the best performing hierarchy for each region and each cancer type pair (including non-cancer as a negative type).
[0233] In some embodiments, the analysis system 200 trains the classifier by inputting a training sample set with its feature vector into the classifier and adjusting the classification parameters so that the classifier function can accurately associate the training feature vector with its corresponding label. The analysis system 200 can group the training samples into one or more sets of training samples for iterative batch training of the classifier. After inputting all training sample sets including their training feature vectors and adjusting the classification parameters, the classifier can be fully trained to label the test sample according to its feature vector within a certain error limit. The analysis system 200 can train the classifier according to any of a variety of methods, such as L1 regularized logistic regression or L2 regularized logistic regression (for example, using a logarithmic loss function), a generalized linear model (GLM), a random forest, a multinomial logistic regression, a multilayer perceptron, a support vector machine, a neural network, or any other suitable machine learning technique.
[0234] In various embodiments, the analysis system 200 transforms the feature values by binarization. Specifically, feature values greater than 0 are set to 1, so that the feature value is either 0 or 1 (indicating the presence or absence of a disease state). In other embodiments, a smoothing function can be implemented (e.g., to provide a more granular value) rather than binarization to 0 or 1. Figure 14 As shown, the analysis system 200 can binarize the features in cross-validation and then use these features to train the classifier.
[0235] In various embodiments, the analysis system 200 trains a multinomial logistic regression classifier on the training data in one fold and generates predictions for the held-out data. For each of the K folds, the analysis system 200 trains a logistic regression for each combination of hyperparameters. An example hyperparameter is the L2 penalty, a form of regularization applied to the weights of the logistic regression. Another example hyperparameter is topK, the number of high-ranking regions to retain for each tissue type pair (including non-cancer). For example, when topK=16, the analysis system 200 retains the top 16 regions for each tissue type pair, ranked according to the mutual information process described herein. By following this process, the analysis system 200 can generate predictions for each sample in the training set while ensuring that the classifier is not trained on the data for which it generates predictions.
[0236] In various embodiments, for each set of hyperparameters, the analysis system 200 performs a performance evaluation on the cross-validated predictions of the complete training set, and the analysis system 200 selects the hyperparameter set with the best performance to retrain on the complete training set. Performance can be determined based on a logarithmic loss metric. The analysis system 200 can calculate the logarithmic loss by taking the negative logarithm of the prediction of the correct label for each sample and then summing the samples. For example, a perfect prediction of 1.0 for the correct label will result in a logarithmic loss of 0 (the lower the more accurate). To generate predictions for new samples, the analysis system 200 can calculate feature values using the above method, but only for the features (region / positive class combination) selected under the selected topK value. The analysis system 200 can use the generated features to create predictions using a trained logistic regression model.
[0237] During deployment, analysis system 200 applies a classifier to predict the tissue of origin of a test sample, where the tissue of origin is associated with one of the disease states. In some embodiments, the classifier can return predictions or probabilities for more than one disease state or tissue of origin. For example, a classifier can return a prediction that the test sample has a 65% probability of having breast cancer tissue of origin, a 25% probability of having lung cancer tissue of origin, and a 10% probability of having healthy tissue of origin. Analysis system 200 can further process the predictions to generate a single disease state determination. III.E. Uncertain Positioning
[0238] In various embodiments, the tumor score can be a covariate in the predictions made by a trained classifier or model for a sample. As the tumor score decreases, the score assignment (e.g., based on the previously described log-likelihood ratio R) may become less clear until the limit of classification detection is reached (i.e., a 50% probability of detecting a cancer / cancer type). Samples with high cfDNA tumor scores tend to be clearly classified, while samples with low cfDNA tumor scores tend to be more ambiguous. In cases where the signal is ambiguous, the assignment becomes less reliable and may be correct or incorrect by chance. In the use case of a single localization, the analysis system 200 can identify ambiguous signals and separate these predictions into an "uncertain localization class."
[0239] For example, in some embodiments, the analysis system 200 can determine a post hoc uncertain assignment from a set of tissue of origin location vectors for individuals whose cancer score is greater than a specific target threshold. The analysis system 200 can determine the uncertain assignment based on cross-validation. For each sample, the analysis system 200 can calculate an indicator to capture the location uncertainty of the sample. As an example method, the analysis system 200 uses the information entropy (bits) of the tissue of origin location to calculate the indicator, where the bit value is zero when a prediction is certain. In the most ambiguous case (the probability of all n categories is equal), the analysis system 200 calculates the bit value (n). As another example method, the analysis system 200 uses the difference (δ value) between the highest ranking score and the second highest ranking score to determine the indicator. When a prediction is certain, the δ value is 1. In the most ambiguous case, the δ value is 0. By including uncertain results, the analysis system 200 can filter out weak base calls that are only correct by chance and improve the accuracy of clear location base calls (e.g., the proportion of correct tissue of origin assignments).
[0240] As an alternative to post hoc uncertain assignments, the analysis system 200 can use expectation maximization during training to determine assignments to uncertain classes.The analysis system 200 can also add a second layer to the classifier output to classify cases into uncertain classes.
[0241] Given the metrics and a record of whether each sample was correctly located, the analysis system 200 can calculate a precision-recall curve for an uncertain base calling threshold, such as Figure 18 For example, the target accuracy level (such as Figure 18 The cutoff is selected by averaging 90% of the base calls in the example shown. The analysis system 200 can calculate the cutoff for the localization signature individually (e.g., for a particular cancer type) or globally for all cancer types. The tradeoff needs to be optimized and may depend on the cost of mislocalizing a base call versus the number of base calls assigned an uncertain result (e.g., precision and recall). III.F. Preventing Class Imbalance
[0242] In various embodiments, a single sample s i The elements of the score vector include the posterior probability of signal localization for each predicted class (e.g., disease state). Each element is scaled by the prior probability proportional to the proportion of training examples for each class:
[0243] If the classes are unbalanced, samples with weak signals may be shifted to an inappropriate class. For example, a training set may include 99% of liver cancer detection samples, but very few detection samples of other cancer types. As a result, a classifier trained on this set may be biased towards liver cancer predictions (or always guess that class). Furthermore, if the class ratios used in classifier training are incompatible with the population frequencies to which the classifier is applied (e.g., a more balanced class ratio), incorrect predictions may be produced.
[0244] To evaluate the ability of the classifier to localize cfDNA samples from methylation and / or genomic and / or clinical features, the analysis system 200 can set a goal of equal proportions between different classes. The analysis system 200 can calibrate the score based on the incidence of the disease state in the screening population, optionally taking into account the detectability of the disease by the tumor score. By modifying the priors of the classifier trained using a common training set, the analysis system 200 can customize the classifier to improve predictions for a specific population associated with a prior (e.g., indicating the distribution of disease states in that specific population). Different geographic regions or countries may have different priors based on the prevalence of a particular disease state or cancer type in the corresponding individual subpopulation.
[0245] As an example method, the analysis system 200 performs a post-hoc recalibration of the model scores. Specifically, the analysis system 200 corrects the scores for a class by dividing the assigned probability by the frequency of training set examples for that class. Optionally, the correction can be stabilized by adding pseudo counts. The analysis system 200 can then perform a post-hoc recalibration of the model scores for each score vector s. i Normalize so that the sum is one.
[0246] As another approach, the analysis system 200 may resample the low-frequency training examples to a desired ratio.As yet another approach, the analysis system 200 may reweight the loss function in classifier training. III.G. Origins of Conditional Cancer Signaling
[0247] In some embodiments, the analysis system 200 trains a classifier to predict the origin of a cancer signal (tissue of origin) conditioned on one or more characteristics of an individual (e.g., cofactors). Example characteristics include: age, smoker status (e.g., whether the individual is a regular smoker, a sometimes smoker, or a never smoker), biological sex (e.g., male or female), race, etc. By considering one or more characteristics, the classifier's prediction of the origin of a cancer signal can be customized for a specific population, including the patient, rather than for a general population of all individuals. As a result, the classifier may be less likely to generate false positives and may reduce crosstalk between different types of cancer signal origins, which will be referenced below. Figures 33 to 35Further described. Similarly, false positives and crosstalk between cancer signal origin predictions can be more effectively distributed across the classifier's results for patients. For example, the incidence of male breast cancer is approximately 100 times lower than that of female breast cancer, so it may be important to reduce predictions of male breast CSOs and reallocate these predictions to other CSO predictions so that the errors do not mask the incidence of true cancers. By reducing false positives assigned to low-incidence cancers and shifting classification to high-incidence cancers, the classifier can improve the precision of its predictions, that is, the percentage of samples labeled as cancer signal origins that actually have that predicted cancer signal origin. A high-precision classifier increases the likelihood that a patient will receive appropriate intervention for the type of cancer they actually have, rather than receiving unnecessary intervention due to false positives. Even a small improvement in precision or accuracy can be significant for the detection and treatment of low-frequency cancers (rare cancers).
[0248] In an embodiment, the analysis system 200 takes one or more characteristics into account by training separate classifiers for different combinations of characteristics. For example, the analysis system 200 trains four different classifiers based on gender (male or female) and smoker status (smoker or never smoker). In this example, (1) the first classifier is trained using data from individual samples of females and never smokers, (2) the second classifier is trained using data from individual samples of females and smokers, (3) the third classifier is trained using data from individual samples of males and never smokers, and (4) the fourth classifier is trained using data from individual samples of males and smokers. Because the training data sets used to train these four classifiers are different, each trained classifier has different potential weights personalized for a specific characteristic. In addition, the analysis system 200 can use specific training data sets associated with one or more characteristics to adjust the weights of existing classifiers.
[0249] In various embodiments, analysis system 200 considers one or more characteristics by applying one or more post-processing steps to the classifier's output. As one example, analysis system 200 multiplies the classifier's output probability by a characteristic-based factor. For example, the factor might be based on the subject's age, as a certain cancer signal origin is more likely in older individuals. As another example, if the subject is a smoker, the factor increases the probability of the signal originating from lung cancer, while if the subject is a non-smoker, the factor decreases the output probability.
[0250] Figure 33According to one embodiment, the predicted and actual labels for the origin of cancer signals for a collection of male and female individuals are shown. The y-axis represents the predicted origin of the cancer signal by a classifier that has not yet been trained to generate predictions based on one or more characteristics. The classifier generated predictions for samples from 797 individuals (including male and female individuals). The x-axis represents the actual known origin of the cancer signal for the 797 samples, with an overall accuracy of 88.5%. Although the classifier has a high accuracy for some cancers, its accuracy for other cancers needs to be improved.
[0251] For example, of 12 samples predicted to have anal cancer, only 5 actually had anal cancer, while 6 of the incorrect predictions were actually cervical cancer. Since male individuals do not have a cervix, the classifier can improve its prediction accuracy if it is personalized based on gender. Specifically, a classifier conditioned on male individuals trained only on data from male individuals will not generate a prediction of cervical cancer. Because the classifier will not incorrectly predict that the 6 cervical cancer samples were anal cancer, the classifier's accuracy for anal cancer will improve from 5 / 12 (41.7%) to 5 / 6 (83.3%).
[0252] As another example, Figure 33 The classifier is shown to have made 140 predictions for lung cancer, of which 127 were accurate and 13 were inaccurate. Non-smokers are about an order of magnitude less likely to develop lung cancer than smokers. Therefore, a classifier trained using samples from smokers will produce more accurate predictions than a classifier trained using samples from non-smokers. The accurate predictions for lung cancer in non-smokers (about 127 / 10 = 12.7) may be difficult to distinguish from the crosstalk of the 13 incorrect lung cancer predictions (which actually have different cancer signal origins).
[0253] Figure 34 According to one embodiment, the predicted labels and actual labels of the cancer signal origin of the collection of male individuals are shown. The classifier correctly predicted the breast cancer label, but also generated a false positive breast cancer prediction, which was actually head and neck cancer. Since the probability of male individuals suffering from breast cancer is very small, a conditional classifier conditioned on males will improve the performance of breast cancer prediction. A conditional classifier trained only with samples from male individuals (without samples from female individuals) will produce a classifier with stricter thresholds or requirements when returning breast cancer predictions.
[0254] Figure 35According to one embodiment, the predicted and actual labels for cancer signal origins for samples from female individuals are presented. The classifier's prediction accuracy for cervical cancer was 100% (2 out of 2), but its accuracy was lower at 16.7% (2 out of 12) because 10 samples that actually had cervical cancer were labeled with different cancer signal origins (head and neck cancer or anal cancer). Since cervical cancer does not occur in male individuals, a classifier conditioned on female individuals can improve the accuracy of cervical cancer prediction. III.H. Optimization Features
[0255] In some embodiments, the analysis system 200 performs one or more processes to optimize feature selection when training a classifier. By optimizing the selection of more informative and less noisy features, the analysis system 200 can improve the sensitivity of the trained classifier, for example, for one or more or all cancer types. An example process for optimizing feature selection is to use a subset of the top-ranked features for the classifier. The rank of a feature is inversely proportional to the mutual information. Therefore, low-ranked features have higher mutual information, making them useful for classifying different types of cancer. In some embodiments, the analysis system 200 selects a certain number of top-ranked features. For example, the analysis system 200 selects the 256 lowest-ranked (best) features for each cancer pair by positive type, and the remaining features outside the top 256 features for each cancer pair are not used in the classifier. The threshold of 256 features per cancer pair is a hyperparameter used to train the classifier. In other embodiments, the threshold is a different number of features (e.g., 128, 256, 512, 1024), commonly referred to as "TopK." The TopK threshold can be determined experimentally, using machine learning, or using other types of processes.
[0256] In order to determine the TopK threshold of a conditional classifier trained for a specific population of individuals, the analysis system can train parallel classifiers based on feature extraction of training data sets with different numbers of top-ranked informative features. The analysis system can modify the feature vector of each training sample based on the number of top-ranked informative features being evaluated. For example, the analysis system trains a first classifier that evaluates the top 128 features, a second classifier that evaluates the top 256 features, a third classifier that evaluates the top 512 features, and a fourth classifier that evaluates the top 1024 features. The analysis system can use a validation set of training samples to verify the accuracy of each classifier. The classifier with the best accuracy is selected for the specific population of individuals.
[0257] Figure 36 According to one embodiment, the experimental results of segment coverage and feature ranking are shown. The analysis system 200 determines "coverage_total" as the sum of the average coverage of positive type segments and the average coverage of negative type segments. Figure 36As shown, the average coverage per sample varies based on cancer type (e.g., circulating lymphoid neoplasms, Hodgkin lymphoma, non-Hodgkin lymphoma, and plasma cell neoplasms). Figure 36 In some cases not shown in , the samples with lower coverage correspond to noisier features.
[0258] Figure 37A According to one embodiment, experimental results of fragment coverage and non-cancer activation scores are presented. Figure 37B Experimental results of non-cancer activation scores and feature rankings are presented according to one embodiment. The analysis system 200 determines the non-cancer activation score as the portion of non-cancer samples where a feature appears (e.g., non-cancer activation = sum(binarized non-cancer samples (featureVals) / len(non-cancer samples))). The analysis system 200 uses the non-cancer activation score as a measure of noise because a larger non-cancer activation score value may mean that the sample has more noise, which is less useful for the classifier to distinguish between cancer types. Ideally, the non-cancer activation score value should be close to zero to have a lower noise level. In another embodiment, the analysis system 200 uses the difference between the cancer activation score and the non-cancer activation score as a measure of noise. It is advantageous to select features that do not appear in non-cancer samples (but appear in cancer samples) so that the trained classifier can make more accurate predictions (distinguishing between cancer and non-cancer) based on the top-ranked features. As Figure 37A As shown in Figure 2, the data samples for circulating lymphoid neoplasms and non-Hodgkin's lymphoma are noisier at lower coverage. In contrast, Figure 37A and Figure 37B The data samples for Hodgkin lymphoma and plasma cell neoplasms in the WT group have low noise.
[0259] Figure 38A Experimental results of non-cancer activation scores and feature weights are presented according to one embodiment. Figure 38B According to one embodiment, Figure 38A Another view of the experimental results shown in . The weight indicates how much a feature contributes to the p-cancer score output by the classifier.
[0260] Figure 39A Experimental results of activation score and coverage are presented according to one embodiment. Figure 39B According to one embodiment, Figure 39A Another view of the experimental results shown in . Figures 39A to 39BAs shown, stage 4 cancer generally has a higher activation score than non-cancer and other cancer stages, which enables the trained classifier to distinguish stage 4 cancer. In contrast, the activation scores between stage 1 cancer and borderline false positive detections are more difficult to distinguish because their activation scores are similar. The analysis system 200 determines borderline false positive detections based on a p-cancer threshold (such as 0.1), but the cutoff threshold can be different.
[0261] As previously described, the analysis system 200 generates features for cancer pairs, including positive types (having a certain type of cancer) and negative types (not having a certain type of cancer). The analysis system 200 calculates the mutual information score for the cancer pairs and determines the threshold with the highest mutual information score. The analysis system then returns a feature set for each fragment region, where the feature set includes one feature for each pair of cancers. For example, given 21 different types of cancer types (plus non-cancer), a maximum of 21×21 features are returned for each region (non-cancer is not considered a positive type). The analysis system 200 ranks the features based on the mutual information score within the cancer pair group. Next, the analysis system 200 uses the TopK threshold to filter the ranked features and apply deduplication.
[0262] In an embodiment, the analysis system 200 adjusts the TopK threshold based on the positive cancer type. As one embodiment, the analysis system 200 uses a lower TopK threshold for positive cancer types with higher noise levels. For example, the analysis system 200 applies a TopK threshold of 192 to the circulating lymphoid neoplasm-positive cancer type and a TopK threshold of 256 to the Hodgkin's lymphoma, non-Hodgkin's lymphoma, and plasma cell neoplasm-positive cancer types. Alternatively, the analysis system 200 adjusts the TopK threshold based on the availability of positive cancer type samples. Figure 40 The experimental results of this embodiment are shown.
[0263] In an embodiment, the analysis system 200 prioritizes features using statistics including one or more of coverage (coverage_total or coverage_negative_type), activation scores (non-cancer and / or cancer minus non-cancer), and mutual information scores. As described above, features with high mutual information help classify different types of cancer. Features with lower coverage may have higher noise levels than features with higher coverage and are therefore less suitable for training a classifier. Features with low non-cancer activation scores are more desirable due to their low noise levels. Similarly, features with a larger difference between cancer activation scores and non-cancer activation scores are more desirable due to their stronger signal quality. The analysis system 200 takes this into account by ranking features using mutual information scores, coverage, and activation scores (non-cancer and / or cancer minus non-cancer). In one example, for each cancer pair group, the analysis system 200 applies thresholds for the minimum coverage level and the maximum non-cancer activation score level. Within each positive cancer type, the minimum coverage level and activation score thresholds may vary from TopK to TopK. As another example, the analysis system 200 ranks features using the following function: rank = α * mutual information score + β * coverage + γ * non-cancer activation score + δ * (cancer minus non-cancer activation score), where α, β, γ, and δ are different weights. In embodiments where features are initially ranked using only mutual information scores or another criterion, the analysis system 200 can re-rank features using additional statistics. The analysis system 200 applies a TopK threshold to filter features after ranking or re-ranking. Figure 41 The experimental results of this embodiment are shown. Figure 42 Experimental results using the different thresholding methods described above are presented according to various embodiments.
[0264] In an embodiment, the analysis system 200 generates a statistical model (e.g., a Z-score) to identify abnormal features based on mutual information. For example, for each cancer pair, the analysis system 200 estimates the distribution of mutual information scores and determines a cutoff threshold based on the distribution. The statistical model can also take into account region segment coverage or activation score statistics.
[0265] In an embodiment, the analysis system 200 uses the same "optimized" features described above for both the binary classification stage and the TOO classification stage. In another embodiment, the binary classification stage and the TOO classification stage use different sets of "optimized" features. IV. Multilayer Perceptron Model
[0266] In some embodiments, a multilayer perceptron model ("MLP") can be used for classification instead of logistic regression. As with logistic regression-based classifiers, the MLP classifier can be a single multiclass classifier for both detecting cancer and determining the tissue of origin (TOO) or cancer type. For example, a multiclass classifier can be trained to distinguish between two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different types of cancer. In one embodiment, the multiclass cancer MLP model can also include a class label for non-cancer and can determine cancer detection (e.g., as 1-non-cancer). In another embodiment, the multilayer perceptron model can be a two-stage classifier having a first stage for binary classification (e.g., cancer or non-cancer) and a second stage multilayer perceptron model for multiclass classification (e.g., TOO), e.g., with one or more hidden layers.
[0267] In one embodiment, the multilayer perceptron includes two classifiers: a first-stage multilayer perceptron (MLP) binary classifier with no hidden layer; and a second-stage multilayer perceptron (MLP) multiclass classifier with a single hidden layer. In one embodiment, a sample determined to have cancer using the first-stage classifier is subsequently analyzed by the second-stage classifier.
[0268] In the first level of training, a binary (two-class) multilayer perceptron model without hidden layers for detecting the presence of cancer can be trained to distinguish between cancer samples (regardless of TOO) and non-cancer samples. For each sample, the binary classifier outputs a prediction score indicating the likelihood of the presence or absence of cancer.
[0269] In the second level of training, a parallel multi-class multilayer perceptron model can be trained to determine the cancer type or tissue of origin of the cancer. In one embodiment, only cancer samples with scores above a cutoff threshold (e.g., the 95th percentile of non-cancer samples in the first level classifier) can be included in the training of the multi-class MLP classifier. For each cancer sample used in training and testing, the multi-class MLP classifier outputs a predicted value for the classified cancer type, where each predicted value is the likelihood that a given sample has a certain cancer type. For example, the cancer classifier can return a cancer prediction for the test sample, including a prediction score for breast cancer, a prediction score for lung cancer, and / or a prediction score for no cancer.
[0270] Figure 16 16 is a flow chart of a method 1600 for determining a probability of a sample having a disease state, according to various embodiments. In some embodiments, analysis system 200 performs method 1600 to process sequence reads of fragments from a nucleic acid sample. Method 1600 includes, but is not limited to, the following steps, which are described with respect to components of analysis system 200.
[0271] In step 1610, the analysis system 200 generates sequence reads from one or more biological samples. In some embodiments, the analysis system 200 selects the sequence reads based on their p-value scores. The p-value scores of the sequence reads indicate the probability of observing methylation in the nucleic acid fragments of the one or more biological samples corresponding to the sequence reads.
[0272] In step 1620, analysis system 200 uses the sequence reads to determine, for each position in a set of positions on a chromosome, a count of nucleic acid fragments of one or more biological samples that are within the position and that have at least a threshold similarity to fragments associated with a disease state (e.g., cancer-like fragments). The disease state can be associated with at least one type of cancer, a stage of cancer, or another type of disease or condition.
[0273] Each position can represent a certain number of consecutive base pairs of a chromosome. The number of base pairs between different positions may vary. Analysis system 200 can generate sequence reads for multiple regions of the genome. There can be up to tens of thousands or more regions. Each region can include hundreds, thousands or more base pairs. Method 1600 can be performed for whole genome bisulfite sequencing (WGBS) or targeted group detection.
[0274] In step 1630, the analysis system 200 uses the position count as a feature to train the machine learning model. In some embodiments, the analysis system 200 binarizes the features to indicate the presence or absence of a certain disease state in each position (e.g., a Boolean value). The count of at least one nucleic acid fragment at a certain position indicates the presence of a certain disease state at that position. The count of zero nucleic acid fragments at a certain position indicates the absence of a certain disease state at that position. In some embodiments, the machine learning model can be a logistic regression model. In some embodiments, the machine learning model can be a multilayer perceptron model (neural network). Those skilled in the art will readily appreciate that other machine learning models can be used, including, for example, generalized linear models (GLMs), multilayer perceptrons, support vector machines, random forests, or neural network classifiers.
[0275] In step 1640, the trained machine learning model determines the probability that the test sample has a disease state. The test sample can be obtained from a patient and can include blood and / or tissue. In optional step 1650, treatment is provided to the patient based on the probability. For example, in response to determining that the probability is greater than a threshold, treatment (e.g., medication or interventional surgery) can be provided to the patient. In another embodiment, in optional step 1650, a test report can be generated to provide the patient with their test results, including the probability that the test sample has the disease.
[0276] Figures 17 to 20The experimental results shown were obtained by training the model using samples from the CCGA study, which is further described below.
[0277] Figure 17 The performance gain of the sensitivity of the multilayer perceptron model was demonstrated according to the embodiment. Compared with the logistic regression model, the multilayer perceptron model (MLP) showed performance gain in disease detection sensitivity for cancer stages I, II, III, and IV.
[0278] Figure 18 Experimental results using a multilayer perceptron model for tissue of origin determination are presented according to the embodiments. Compared to logistic regression models (LR: 1803 and 1804), the multilayer perceptron models (MLP: 1801 and 1802) showed improved accuracy in tissue of origin determination. This improvement in accuracy was achieved when processing sequence reads associated with all cancer types in the training set, as well as when processing sequence reads from a training set that included more than 10 example sequence reads for each cancer type in the training set.
[0279] Figure 19 The present invention demonstrates experimental results using a multilayer perceptron model for tissue of origin (TOO) determination by cancer stage. Compared to a logistic regression (LR) model, the multilayer perceptron (MLP) model demonstrated improved performance in tissue of origin (TOO) detection accuracy for cancer stages I, II, III, and IV. The MLP model achieved the greatest performance gain in stage I cancer.
[0280] Figure 20 The experimental results of the multilayer perceptron model on different types of cancer are shown according to the embodiment. Figure 20 For most cancer types shown, the multilayer perceptron model (MLP) achieved higher accuracy in tissue of origin (TOO) detection compared to the logistic regression model.
[0281] In some embodiments, the analysis system uses a two-level model to determine the tissue of origin (TOO) of a cancer or other type of disease state. The analysis system generates sequence reads from nucleic acid fragments of a biological sample. The analysis system determines a first training data set by processing the sequence reads, for example, using any of the processes described in Part II. A. Detection scheme. The analysis system can use methylation information to determine the first training data set. For example, the analysis system determines hypomethylated sequence reads by determining a threshold number or percentage of unmethylated CpG sites corresponding to the sequence reads. In addition, the analysis system determines hypermethylated sequence reads by determining a threshold number or percentage of methylated CpG sites corresponding to the sequence reads. The analysis system can also determine whether the sequence reads are informative. In some embodiments, the analysis system screens sequence reads by removing sequence reads with a p-value less than a threshold p-value.
[0282] The analysis system trains a binary classifier using a first training dataset.The binary classifier is trained to predict a binary output, ie, the presence or absence of at least one disease state in the first test biological sample, for input sequence reads from the first test biological sample.
[0283] Using the predictions of the binary classifier, the analysis system can determine that a subset of the biological samples has one or more disease states. The binary classifier can be used to train a tissue of origin classifier. Specifically, the analysis system uses sequence reads corresponding to nucleic acid fragments of the subset of the biological sample to determine a second training data set. The analysis system uses the second training data set to train the tissue of origin classifier. The tissue of origin classifier is trained to predict the tissue of origin associated with the disease state present in the second test biological sample based on the input sequence reads from the second test biological sample. The first test biological sample and the second test biological sample can be the same sample or different samples.
[0284] In some embodiments, the analysis system uses a tissue of origin classifier to determine a score indicating the probability of the presence of a tissue of origin associated with a disease state in the second test biological sample. The analysis system can calibrate the score, for example to adjust the output of an overconfident model. For example, the analysis system uses the feature space output by the tissue of origin classifier to perform a k-nearest neighbor (KNN) operation in combination with the score. In an embodiment, the feature space includes the top two predicted labels (e.g., lung cancer and prostate cancer) from the tissue of origin classifier, and an indication of whether the correct classification is a disease state different from the top two predictions. The analysis system can also calibrate the score by normalizing the probability using the output of a binary classifier (the output indicating the different probabilities of the presence of at least one disease state in the second test biological sample).
[0285] In some embodiments, the tissue of origin classifier is a multilayer perceptron that includes at least one hidden layer. The tissue of origin classifier can also include a hidden layer of 100 units or a hidden layer of 200 units, as well as other hidden layer sizes. The multilayer perceptron can be fully connected and use a rectified linear unit activation function. In some embodiments, the binary classifier is a multilayer perceptron that does not include a hidden layer. In various embodiments, the binary classifier is a multilayer perceptron that includes at least one hidden layer. In other embodiments, these classifiers can be logistic regression models, multinomial logistic regression models, or other types of machine learning models.
[0286] In addition, the analysis system can use one or more machine learning techniques known to those skilled in the art to train the tissue of origin classifier and the binary classifier, including, for example, not stopping early (but choosing a given number of training epochs), stochastic gradient descent, weight decay, dropout regularization, Adam optimization, He initialization and learning rate scheduling, rectified linear unit activation function, leaky rectified linear unit activation function, sigmoid activation function, and bootstrapping, etc. Figure 31 As shown, the tissue-of-origin classifier's accuracy improves with each training iteration. Each iteration can include a different combination of machine learning techniques. Furthermore, the accuracy improves for each cancer stage (I, II, and III).
[0287] In some embodiments, the analysis system performs cross-validation on one or both of the tissue-of-origin classifier and the binary classifier. The analysis system can retrain the classifier using hyperparameters selected based on the output of the cross-validation. The analysis system can select hyperparameters by aggregating the results from all folds in the cross-validation. In some embodiments, the analysis system selects hyperparameters to train the tissue-of-origin classifier by optimizing tissue-of-origin accuracy rather than log-likelihood, because the classifier can be more confident in samples with stronger signals.
[0288] In some embodiments, the analysis system determines the probability of the presence of a tissue of origin associated with a disease state in the second test biological sample using a tissue of origin classifier. The analysis system predicts the presence of a tissue of origin associated with a disease state in the second test biological sample in response to determining that the probability is greater than a tissue of origin threshold. The analysis system can determine different tissue of origin thresholds associated with different tissues of origin. In addition, the analysis system can determine the tissue of origin threshold associated with a given disease state by iterating a series of different probabilities of candidate tissue of origin thresholds. For each iteration, the analysis system determines the sensitivity rate of the tissue of origin classifier at a given specificity rate. The analysis system can optimize the trade-off between the sensitivity rate and the specificity rate of the tissue of origin classifier for a given disease state. The analysis system can determine the sensitivity rate using a score output by a binary classifier or a tissue of origin classifier. In addition, the analysis system can stratify samples using the score from the tissue of origin classifier.
[0289] In some embodiments, the analysis system trains the binary classifier and the tissue of origin classifier using binarized features, each having a value of 0 or 1. In the binarization, values greater than 1 will be replaced by 1. V. Adjusting the binary classification threshold
[0290] The analysis system can adjust the trained cancer classifier to prune the samples used to train the cancer classifier. Specifically, the analysis system can attempt to remove non-cancer samples with high tissue signals, which reduce the sensitivity of the cancer classifier in cancer prediction. High tissue signal means that the proportion of cfDNA from the tissue of origin (TOO) in the sample is relatively large compared to the healthy distribution, for example, as determined by a tissue of origin classifier, a multi-class cancer classifier, or other means. Non-cancer samples with high tissue signals are outliers in the non-cancer distribution, and they may be precancerous, early-stage cancers, or undiagnosed cancers. The analysis system can identify non-cancer samples with high tissue signals in at least one cancer type. In some embodiments, certain cancer types are further divided into cancer subtypes. For example, hematological cancer types can be further divided into combinations of, for example, circulating lymphoid tumor subtypes, indolent non-Hodgkin's lymphoma (NHL) subtypes, aggressive NHL subtypes, Hodgkin's lymphoma (HL) subtypes, myeloid tumor subtypes, and plasma cell tumor subtypes.
[0291] refer to Figure 21 , Figure 21 A graph showing the likelihood of cancer types for non-cancer samples at a specificity higher than 95% is presented. A cancer score is calculated for each non-cancer sample from a plurality of non-cancer samples (i.e., samples from healthy individuals who have not currently been diagnosed with cancer). The cancer score can be determined by a binary classifier as the likelihood that a sample has cancer given the methylation sequencing data of a given sample. In other embodiments, the cancer score can be calculated according to other methods that at least input sequencing data (e.g., methylation, single nucleotide polymorphism (SNP), DNA, RNA, etc.) and output the likelihood that a sample has cancer based on the input sequencing data. An example of a classifier is a hybrid model classifier. A distribution of non-cancer samples can be generated based on the cancer score of the non-cancer sample. A binary threshold cutoff value can be set to ensure a certain level of binary classification specificity, such as a true negative rate. Typically, a high specificity cutoff value is used when classifying cancer, e.g., a specificity between 90% and 99.9%, or 99.5% or higher. However, many non-cancer samples used to train a cancer classifier that are just below the specificity cutoff value may have high tissue signals, thereby producing a positive bias for the binary threshold cutoff value.
[0292] To demonstrate this, non-cancer samples with a specificity greater than 95% were selected and then input into a multi-class cancer classifier to determine the probability of each cancer type or tissue of origin (TOO). The cancer types or TOO labels used in this multi-class cancer classifier embodiment include circulating lymphoid neoplasms, myeloid neoplasms, indolent NHL, colorectal cancer, invasive NHL, lung cancer, uterine cancer, breast cancer, prostate cancer, pancreatic and gallbladder cancer, upper gastrointestinal cancer, bladder and urothelial cancer, plasma cell neoplasms, head and neck cancer, kidney cancer, ovarian cancer, sarcoma, hepatobiliary cancer, cervical cancer, other tissue tumors, HL, anorectal cancer, melanoma, and thyroid cancer. Figure 21 The graph in Figure 2 shows many non-cancer samples with high tissue signal for at least one tissue type. Each point in the row of tissue types corresponds to a tissue probability of origin for a non-cancer sample above the 95% specificity threshold. Notably, many tissue types have multiple non-cancer sample outliers with significant tissue contributions, which is unusual for non-cancer samples. This can occur when such non-cancer samples have cfDNA signals driven by cancer-like methylation, clonal fraction, and / or growth / turnover rates. It can be inferred that a large number of non-cancer samples used to train the cancer classifier may represent precancerous conditions, early-stage cancers, or undiagnosed cancers. Nevertheless, these non-cancer samples with significant tissue contributions can shift the binary classification cutoff upward, thereby reducing the sensitivity of cancer classification, especially for samples with significant tissue signals just below the previously set binary classification cutoff. In practice, such signals (e.g., corresponding to circulating lymphoid neoplasms, myeloid neoplasms, and indolent NHL) can be a major contributor to false-positive determinations. Of note, circulating lymphoid neoplasms, myeloid neoplasms, indolent NHL, colorectal cancer, aggressive NHL, lung cancer, uterine cancer, breast cancer, prostate cancer, pancreatic and gallbladder cancer, upper gastrointestinal cancer, plasma cell neoplasms, head and neck cancer, cervical cancer, and HL had at least one non-cancer sample with a tissue origin probability greater than 0.1. In particular, circulating lymphoid neoplasms, myeloid neoplasms, indolent NHL, and aggressive NHL (all hematological subtypes) had two or more non-cancer samples with a tissue origin probability greater than 0.5.
[0293] refer to Figure 22 , Figure 22 A diagram showing hematologic subtypes based on methylation sequencing data. Figure 22The graph demonstrates the ability to model hematological subtypes. It turns out that this can be beneficial for providing finer granularity for multi-class cancer classification (for example, using hematological subtype labels for additional classification), or as a way to adjust cancer classification by pruning non-cancer samples with high hematological subtype signals before training the cancer classifier. As described above, methylation signals can cover multiple CpG sites, thereby creating a high-dimensional vector space. Using hematological subtype samples and non-cancer samples, the analysis system can perform principal component analysis. Principal component analysis identifies orthogonal principal components (or embeddings) of the vector space in the order of the variance of the methylation signals in the sample. The first principal component is shown as V1 on the horizontal axis of the graph and has the highest variance, while the second principal component is shown as V2 on the vertical axis of the graph and has the second highest variance. Clusters of each hematological subtype and non-cancer sample are marked on the graph 900. The hematological subtypes shown include circulating lymphoid tumors, solid lymphoid tumors, plasma cell tumors, and myeloid tumors. Solid lymphoid tumor subtypes can be further divided into HL, indolent NHL, and aggressive NHL. This diagram illustrates the potential of classification based on hematological subtypes—either by adding hematological subtypes to a multi-class cancer classification or by modeling each hematological subtype to tune a cancer classifier. VA removes high signal non-cancer samples
[0294] Figure 23A A flowchart describing a process 1000 for determining a binary threshold cutoff for binary cancer classification is shown according to one or more embodiments. The binary classification for predicting between cancer and non-cancer evaluates the cancer score of the sample based on the determined binary threshold cutoff, wherein samples with a cancer score below the binary threshold cutoff are determined to be non-cancer, and samples with a cancer score equal to or above the binary threshold cutoff are determined to be cancer. A trained multi-class cancer classifier evaluates the methylation signals (and / or other sequencing data) of the sample to determine the probabilities of multiple TOO labels classified by the multi-class cancer classifier. The TOO labels used in the multi-class cancer classifier can be cancer tissue types or cancer tissue subtypes (e.g., the hematological subtypes described above). Process 1000 can be performed or completed by an analysis system.
[0295] The analysis system receives 1010 sequencing data of a plurality of biological samples containing cfDNA fragments, the biological samples including cancer samples and non-cancer samples. The sequencing data can be methylation sequencing data, SNP sequencing data, another type of DNA sequencing data, RNA sequencing data, etc.
[0296] For each non-cancer sample, the analysis system classifies the non-cancer sample based on the features derived from the sequencing using a multi-class cancer classifier, wherein the multi-class cancer classifier predicts the probability of each of the plurality of TOO signatures 1020. The analysis system can generate a feature vector for the non-cancer sample, assigning an abnormality score to each CpG site under consideration based on at least one informative cfDNA fragment overlapping with the CpG site.
[0297] For each non-cancer sample, the analysis system determines 1030 whether the predicted probability likelihood exceeds a TOO threshold for one or more TOO signatures. The TOO threshold determination will be described below. Figure 23B Further described in .
[0298] The analysis system determines 1040 a binary threshold cutoff value for predicting the presence of cancer, the binary threshold cutoff value being determined based on the distribution of non-cancer samples (excluding one or more non-cancer samples identified as having a probability likelihood exceeding at least one TOO threshold value). Non-cancer samples for which at least one probability likelihood of a TOO tag exceeds the TOO threshold value corresponding to the TOO tag will be excluded. The analysis system then calculates a distribution of non-cancer samples based on the cancer score of each non-cancer sample, and then determines a binary threshold cutoff value for a desired specificity level (e.g., 99.4-99.9% specificity) based on the distribution. It is noteworthy that each cancer score can be determined based on sequencing data, for example, a cancer score can be output by a binary cancer classifier that predicts the likelihood of cancer based on methylation sequencing data, as described herein. In other embodiments, the cancer score can be calculated according to other methods that at least input sequencing data (e.g., methylation, single nucleotide polymorphism (SNP), DNA, RNA, etc.) and output the likelihood of the sample having cancer based on the input sequencing data.
[0299] Figure 23B A flowchart describing a process 1005 for thresholding a TOO tag to determine a binary threshold cutoff for binary cancer classification is shown according to one or more embodiments. The process 1005 can be an embodiment of the process 1000. The binary classification for predicting between cancer and non-cancer evaluates the cancer score of the sample based on the determined binary threshold cutoff, wherein samples with a cancer score below the binary threshold cutoff are determined to be non-cancer, and samples with a cancer score equal to or above the binary threshold cutoff are determined to be cancer. A trained multi-class cancer classifier evaluates the methylation signals (and / or other sequencing data) of the sample to determine the probabilities of multiple TOO tags classified by the multi-class cancer classifier. The TOO tag can be a cancer tissue type, or more specifically a cancer tissue subtype (e.g., the hematological subtype described above). The process 1005 can be performed or completed by an analysis system.
[0300] The analysis system obtains 1015 a training set comprising a plurality of samples having a cancer or non-cancer label, and a holdout set comprising a plurality of samples having a cancer or non-cancer label, i.e., cancer samples or non-cancer samples, respectively. Each sample in the training set comprises, for example, Figure 3 The methylation sequencing data generated by process 300 is used. In other embodiments, each training sample has other sequencing data used in series or substituted for the methylation sequencing data. In addition, each sample in the training set and the holdout set has a cancer score. As described above, the cancer score can be determined by a binary classifier as the probability that the sample has cancer given the methylation sequencing data of the given sample. In other embodiments, the cancer score is calculated according to other methods, which at least input sequencing data (e.g., methylation, single nucleotide polymorphism (SNP), DNA, RNA, etc.) and output the probability that the sample has cancer based on the input sequencing data, taking the hybrid model described herein as an example.
[0301] For each non-cancer training sample, the analysis system determines 1025 a feature vector based on the methylation sequencing data. The analysis system can determine the feature vector for each non-cancer training sample, for example, by determining an abnormality score for each CpG site in the set of CpG sites under consideration. In some embodiments, the analysis system uses a binary score to define an abnormality score for the feature vector based on whether an informative segment is present in the set of informative segments covering the CpG site. Once all abnormality scores for the sample are determined, the analysis system determines a feature vector as a vector of abnormality scores associated with each CpG site under consideration. The analysis system can further normalize the abnormality scores of the feature vector based on the coverage of the sample.
[0302] The analysis system inputs 1035 the feature vector of each non-cancer training sample into a multi-class cancer classifier to generate a TOO prediction. The multi-class cancer classifier is trained on a plurality of TOO labels including cancer type, cancer subtype, non-cancer, or any combination thereof. The multi-class cancer classifier can be trained as described herein. The trained multi-class cancer classifier determines a plurality of probabilities for the TOO labels as cancer predictions, wherein the probabilities of the TOO labels indicate the likelihood of having the cancer corresponding to the TOO label.
[0303] In some examples, the analysis system scans 1045 or iterates through a range of TOO label probabilities as a candidate TOO threshold, thereby calculating specificity and sensitivity within the range of TOO label probabilities. The analysis system may scan the probability range incrementally (e.g., by 0.01, 0.02, 0.03, 0.04, 0.05, etc.). As the analysis system scans the probability range, the analysis system filters out non-cancer training samples with a TOO label probability equal to or greater than the candidate TOO threshold based on the output of the multi-class cancer classifier. As a numerical example, the analysis system considers a candidate TOO threshold of 0.35. Non-cancer training samples with a TOO label probability equal to or greater than 0.35 are filtered out from the training set. The analysis system determines an adjusted binary threshold cutoff value based on the filtered training set. The analysis system uses the adjusted binary threshold cutoff value to calculate a predicted specificity rate for the holdout set. Specificity refers to the accuracy with which non-cancer samples are identified as non-cancer labels. The analysis system also uses the adjusted binary threshold cutoff value to calculate a predicted sensitivity rate for the holdout set. Sensitivity refers to the accuracy with which cancer samples are identified as cancer labels. In practice, specificity and / or sensitivity rates can be defined in terms of true positive rate, false positive rate, true negative rate, false negative rate, other statistical calculations, etc.
[0304] The analysis system determines 1055 a TOO threshold for the TOO signature. The analysis system selects a TOO threshold from the candidate TOO thresholds by optimizing the specificity rate and / or sensitivity rate calculated within the candidate TOO threshold range. In some examples, a TOO threshold is determined or otherwise applied for certain TOO tissue type categories or subtype categories (such as hematology categories). By way of example only, an algorithm for calculating and applying a TOO-specific probability threshold can be used to remove non-cancer samples with excessively high blood disease signals. The algorithm can include, for each pre-specified TOO signature, first searching a grid of probability values, and for each value, evaluating the clinical specificity and clinical sensitivity of the holdout set using a binary detection threshold calculated after removing non-cancer samples with an equal or greater probability of the specified TOO signature. By iterating over the probability grid, the algorithm will identify a combination of TOO thresholds for the pre-specified TOO signature that optimizes the trade-off between clinical specificity and clinical sensitivity of the holdout set. The final optimized TOO probability threshold will be used to filter out non-cancer samples that exceed any value given by the TOO signature. The cleaned non-cancer sample set will be used to calculate a cancer-non-cancer detection threshold. However, in some examples, TOO-specific thresholding can be manually set to any cutoff point, such as a desired specificity level (eg, 99.4-99.9% specificity).
[0305] The analysis system adjusts 1065 the binary cancer classification by pruning non-cancer training samples that exceed the TOO threshold before determining the binary threshold cutoff. The analysis system filters out non-cancer training samples from the training set based on the TOO threshold determined for the TOO label. The analysis system sets the binary threshold cutoff based on the filtered training set. For example, the analysis system determines the new binary threshold cutoff based on the filtered score distribution. In other embodiments, the analysis system can determine the TOO threshold for any TOO label according to steps 1010, 1020, 1030, and 1040 to adjust the binary cancer classification. VB stratifies sample distribution based on TOO signal
[0306] In one or more embodiments, the analysis system adjusts the cancer classifier by stratifying the sample distribution according to TOO signal to determine a binary threshold cutoff value for each stratum. The analysis system can stratify the sample distribution according to the signal of one or more TOO signatures determined by the TOO predictions output by the multi-class cancer classifier.
[0307] As used herein, "high tissue signal" refers to a sample with a tissue signal exceeding a certain threshold, for example, generally for any type of tissue or for a specific cancer type (also known as a TOO signature). Tissue signals can be determined by a multi-class cancer classifier or other methods compared to a healthy distribution. Non-cancerous samples with high tissue signals are outliers in the non-cancerous distribution. Some of these non-cancerous samples may be precancerous, early-stage cancers, or undiagnosed cancers. The analysis system can identify non-cancerous samples with high tissue signals in at least one TOO signature. In one method of determining a high tissue signal, the predicted value of the TOO signature output by the multi-class cancer classifier is compared with a tissue signal threshold. Samples with predicted values higher than the tissue signal threshold are considered to have a high tissue signal for the TOO signature; samples with predicted values lower than the tissue signal threshold are considered not to have a high tissue signal for the TOO signature (or to have a low tissue signal). In another method, one or more top-ranked predictions in the TOO predictions are considered. For example, the TOO prediction for a sample has a first prediction for a colorectal TOO signature, a second prediction for a breast TOO signature, and a third prediction for a head / neck TOO signature. If the top-ranked predictions are considered, the sample is considered to have a high tissue signal for the TOO tag in the first prediction, i.e., the colorectal TOO tag in the example. If the first two predictions are considered, there are high tissue signals in both the colorectal TOO tag and the breast TOO tag. Other methods for determining tissue signals may include other models that are trained to determine the tissue signals of one or more TOO tags. Such models may include classifiers that are trained to determine the tissue signals of a subset of TOO tags. For example, a hematology-specific classifier can be trained and used to determine the tissue signals of one or more hematological subtypes. Other models include deconvolution models that can deconvolve tissue signals in methylation sequencing data (and / or other types of sequencing data).
[0308] Now refer to Figure 32 , Figure 32 A process of stratifying hematology signals into two layers is presented according to one or more embodiments. Although the following description describes stratification of hematology signals, the principles can be readily applied to other TOO signals.
[0309] The analysis system 1300A stratifies a holdout set of cancer samples and non-cancer samples into a low signal layer 1310 and a high signal layer 1320 based on the hematological signal. Each sample in the holdout set has a cancer score determined by a binary cancer classifier and a TOO prediction determined by a multi-class cancer classifier. In one embodiment, the hematological signal of the sample is determined based on the TOO prediction output by the multi-class cancer classifier. In one embodiment, when considering one or more top-ranked predictions (e.g., top one, top two, etc.), a high hematological signal is determined if at least one of the top-ranked predictions considered is one of the hematological subtypes (e.g., lymphoid tumor subtype and myeloid tumor subtype). Other hematological subtypes may be included. Therefore, if at least one of the top-ranked predictions in the TOO prediction for the sample is considered to be a lymphoid tumor subtype or a myeloid tumor subtype, the sample is determined to have a high hematological signal. Otherwise, the sample is determined not to have a high hematological signal.
[0310] The analysis system determines a binary threshold cutoff value for each stratum to predict the presence of cancer in a sample. The analysis system uses samples from the low-signal stratum 1310 to determine 1305 a binary threshold cutoff value for predicting the presence of cancer in samples from the low-signal stratum 1310. The binary threshold cutoff value 1305 is determined based on a false positive budget set for the low-signal stratum 1310. Using the cancer scores of the samples from the low-signal stratum 1310, the analysis system scans a series of candidate binary threshold cutoff values, thereby evaluating the true positive rate (also known as sensitivity) and false positive rate of each candidate binary threshold cutoff value. The candidate binary threshold cutoff value with the false positive rate closest to the false positive budget is determined as the candidate binary threshold cutoff value. The analysis system performs a similar operation to determine 1315 a binary threshold cutoff value for the high-signal stratum 1320. The false positive budget for the low-signal stratum 1310 and the false positive budget for the high-signal stratum 1320 can be set based on the ratio of the statistical true positive rates of each stratum. This ratio is intended to suppress the false positive rate in the high-signal stratum 1320.
[0311] For the test sample, the analysis system places the test sample into a low signal tier 1310 or a high signal tier 1320 based on the hematology signal. If the test sample is placed into the low signal tier 1310, the analysis system applies 1315 the binary threshold cutoff value of the low signal tier 1310 to the cancer score of the test sample. If the cancer score is greater than or equal to the binary threshold cutoff value of the low signal tier 1310, the analysis system returns a prediction of the presence of cancer in the test sample; otherwise, it returns a prediction of the absence of cancer. If the test sample is placed into the high signal tier 1320, the analysis system applies 1325 the binary threshold cutoff value of the low signal tier 1320 to the cancer score of the test sample. If the cancer score is greater than or equal to the binary threshold cutoff value of the high signal tier 1320, the analysis system returns a prediction of the presence of cancer in the test sample; otherwise, it returns a prediction of the absence of cancer. VI. Study of the free genome map of circulating cells
[0312] In various embodiments, the use of data from the Circulating Cell-Free Genome Atlas (CCGA) study (see ClinicalTrial.gov identifier: NCT02889978 ( https: / / www.clinicaltrials.gov / ct2 / show / NCT 02889978) ) were used to train each predictive cancer model and then subsequently tested using a set of testing or validation data derived from a testing or validation subset of patients from the CCGA study.
[0313] The prediction cancer model described herein is trained using a variety of known cancer types from the circulating cell free genome atlas (CCGA) study. The CCGA sample set includes the following cancer types: breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, and anorectal cancer. Therefore, the model can be a multi-cancer model (or multi-cancer classifier) for detecting one or more, two or more, three or more, four or more, five or more, ten or more, or 20 or more different types of cancer. The predictive cancer model can be trained using a refined training dataset of a first subset of patients from the CCGA study, and then tested using a refined testing dataset of a second subset of patients from the CCGA study. VII. Cancer Detection Panel
[0314] In various embodiments, the predictive cancer models described herein use samples enriched by a cancer detection panel comprising a plurality of probes or a plurality of probe pairs. A variety of targeted cancer detection panels are known in the art, for example, as described in WO 2019 / 195268, filed April 2, 2019, PCT / US2019 / 053509, filed September 27, 2019, and PCT / US2020 / 015082, filed January 24, 2020 (these patents are incorporated herein by reference). For example, in some embodiments, a cancer detection panel can be designed to include a plurality of probes (or probe pairs) that can capture fragments that can collectively provide information relevant to cancer diagnosis. In some embodiments, a detection panel comprises at least 50, 100, 500, 1,000, 2,000, 2,500, 5,000, 6,000, 7,500, 10,000, 15,000, 20,000, 25,000, or 50,000 pairs of probes. In other embodiments, a detection panel comprises at least 500, 1,000, 2,000, 5,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000, or 100,000 probes. The multiple probes may comprise at least 100,000, 200,000, 400,000, 600,000, 800,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, or 10,000,000 nucleotides in total. Probes (or probe pairs) are specifically designed to target one or more genomic regions of differential methylation in cancer samples and non-cancer samples. Target genomic regions can be selected to maximize classification accuracy under the constraints of a size budget (determined by sequencing budget and desired sequencing depth).
[0315] The sample enriched using the cancer detection group can be subjected to targeted sequencing. The sample enriched using the cancer detection group can be used to detect the presence or absence of cancer in general, and / or provide cancer classification, such as cancer type, cancer stage (such as I, II, III or IV), or provide the tissue of origin from which the cancer is believed to originate. Depending on the purpose, a detection group can include probes (or probe pairs) that target the genomic regions that are differentially methylated between general cancerous (pan-cancer) samples and non-cancerous samples, or only target the cancerous samples of specific cancer types (e.g., lung cancer-specific targets). Specifically, the cancer detection group is designed based on bisulfite sequencing data generated by free DNA (cfDNA) or genomic DNA (gDNA) of cancer and / or non-cancer individuals.
[0316] In some embodiments, the cancer detection group designed by the method provided herein includes at least 1,000 pairs of probes, wherein each pair includes two probes, and the two probes are configured to overlap each other through overlapping sequences including 30 nucleotide fragments. These 30 nucleotide fragments include at least five CpG sites, wherein at least 80% of the at least five CpG sites are CpG or UpG. These 30 nucleotide fragments are configured to bind to one or more genomic regions in a cancerous sample, wherein the one or more genomic regions have at least five methylation sites with abnormal methylation patterns. Another cancer detection group includes at least 2,000 probes, wherein each probe is designed as a hybridization probe complementary to one or more genomic regions. Each genomic region is selected based on the following criteria: it includes (i) at least 30 nucleotides and (ii) at least five methylation sites, wherein the at least five methylation sites have abnormal methylation patterns and are either hypomethylated or hypermethylated.
[0317] Each probe (or probe pair) is designed to target one or more target genomic regions. The selection of target genomic regions is based on several criteria that are intended to increase the selective enrichment of relevant cfDNA fragments while reducing noise and non-specific binding. For example, a detection group can include probes that can selectively bind to and enrich differentially methylated cfDNA fragments in cancerous samples. In this case, sequencing the enriched fragments can provide information relevant to cancer diagnosis. In addition, probes can be designed to target genomic regions that are determined to have abnormal methylation patterns and / or hypermethylated or hypomethylated patterns to provide additional detection selectivity and specificity. For example, when genomic regions have methylation patterns with low p-values according to a Markov model trained on a set of non-cancerous samples, and these genomic regions additionally cover at least 5 CpGs, 90% of which are methylated or unmethylated, these genomic regions can be selected. In other embodiments, a hybrid model can be used to select genomic regions, as described herein.
[0318] Each probe (or probe pair) can target a genomic region comprising at least 25bp, 30bp, 35bp, 40bp, 45bp, 50bp, 60bp, 70bp, 80bp, or 90bp. Genomic regions can be selected by comprising less than 20, 15, 10, 8, or 6 methylation sites. Genomic regions can be selected when at least 80%, 85%, 90%, 92%, 95%, or 98% of at least five methylation (e.g., CpG) sites in a non-cancerous or cancerous sample are methylated or unmethylated.
[0319] Genomic regions can be further screened to select only those regions that may be informative based on their methylation patterns, for example, CpG sites that are differentially methylated between cancerous and non-cancerous samples (e.g., aberrantly methylated or unmethylated in cancer and non-cancer). To make the selection, a calculation can be performed for each CpG site. In some embodiments, a first count is determined, i.e., the number of cancer-containing samples that include fragments overlapping with the CpG (cancer_count), and a second count is determined, i.e., the number of total samples that include fragments overlapping with the CpG (total). Genomic regions can be selected based on criteria that are positively correlated with the number of cancer samples that include fragments overlapping with the CpG (cancer_count) and negatively correlated with the number of total samples that include fragments overlapping with the CpG (total).
[0320] In one embodiment, the number of non-cancerous samples (n) with fragments overlapping with CpG sites is calculated. 非癌症 ) and the number of cancerous samples (n 癌症 ) are counted. Then the probability that the sample is cancer is estimated, for example (n 癌症 +1) / (n 癌症 +n 非癌症 +2). CpG sites are ranked according to this metric and greedily added to the test set until the test set size budget is exhausted.
[0321] Depending on whether the detection purpose is pan-cancer detection or single cancer detection, or depending on what kind of flexibility is expected when selecting which CpG sites contribute to the detection group, the samples used for cancer counting can be different. A similar process can be used to design a detection group for diagnosing a specific cancer type (e.g., TOO). In this embodiment, for each cancer type and each CpG site, information gain is calculated to determine whether a probe targeting the CpG site is included. Information gain is calculated for a comparison of samples with a given cancer type with all other samples. For example, two random variables "AF" and "CT". "AF" is a binary variable (yes or no) indicating whether there are abnormal fragments overlapping with a specific CpG site in a specific sample. "CT" is a binary random variable indicating whether the cancer belongs to a specific type (e.g., lung cancer or cancer other than lung cancer). The mutual information of "CT" when given "AF" can be calculated. That is, if you know whether there are informative fragments overlapping with a specific CpG site, how many information bits about the cancer type (lung cancer vs. non-lung cancer in this example) can be obtained. This can be used to rank the specificity of CpG for a specific cancer type (e.g., TOO). This process is repeated for multiple cancer types. For example, if a particular region is typically differentially methylated only in lung cancer (and not in other cancer types or non-cancer), the information gain of CpGs in that region tends to be higher for lung cancer. For each cancer type, CpG sites are ranked by this information gain metric and then greedily added to the test set until the size budget for that cancer type is exhausted.
[0322] Further screening can be performed to select a target genomic region with a non-target genomic region that is less than a threshold value. For example, a genomic region is selected only when the non-target genomic region is less than 15, 10, or 8. In other cases, when the sequence of the target genomic region occurs more than 5, 10, 15, 20, 25, or 30 times in the genome, screening is performed to remove the genomic region. When a sequence that is 90%, 95%, 98%, or 99% homologous to the target genomic region occurs less than 15, 10, or 8 times in the genome, further screening can be performed to select the target genomic region, or when a sequence that is 90%, 95%, 98%, or 99% homologous to the target genomic region occurs more than 5, 10, 15, 20, 25, or 30 times in the genome, further screening is performed to remove the target genomic region. This is to exclude duplicate probes that may pull down non-target fragments, which are unnecessary and can affect detection efficiency.
[0323] In some embodiments, it has been shown that a fragment-probe overlap of at least 45 bp is required to achieve a non-negligible amount of pull-down (although this number may vary depending on the specifics of the assay). In addition, studies have shown that a mismatch rate of more than 10% between the probe and the fragment sequence in the overlapping region is sufficient to severely disrupt binding, thereby affecting the pull-down efficiency. Therefore, sequences that can be aligned with the probe along at least 45 bp with a match rate of at least 90% are candidates for non-target pull-down. Therefore, in one embodiment, the number of such regions is scored. The best probes are scored as 1, which means that they match only in one place (the expected target area). Probes with lower scores (e.g., below 5 or 10) are acceptable, but any probes above this score will be discarded. Other cutoff values can be used for specific samples.
[0324] In various embodiments, the selected target genomic region can be located at various positions in the genome, including but not limited to exons, introns, intergenic regions, and other parts. In some embodiments, probes targeting non-human genomic regions can be added, such as probes targeting viral genomic regions. VIII. Kit Implementation
[0325] Also disclosed herein is a test kit for performing the above-mentioned method (including the method relevant to the cancer classifier). The test kit may include one or more collection containers for collecting samples comprising genetic material from an individual. The sample may include blood, plasma, serum, urine, feces, saliva, other types of body fluids, or any combination thereof. Such a test kit may include a reagent for isolating nucleic acid from the sample. The reagent may further include a reagent for sequencing the nucleic acid, including a buffer and a detection agent. In one or more embodiments, the test kit may include one or more sequencing groups, which include probes for targeting a specific genomic region, a specific mutation, a specific genetic variation, or some combination thereof. In one or more embodiments, the test kit includes at least one detection group comprising a contamination targeting probe. In other embodiments, the sample collected via the test kit is provided to a sequencing laboratory, which can use the sequencing group to sequence the nucleic acid in the sample.
[0326] The kit may further include instructions for use of the reagents contained in the kit. For example, the kit may include instructions for collecting a sample and extracting nucleic acid from a test sample. Exemplary instructions may include the order in which reagents are added, the centrifugation speed used to isolate nucleic acid from a test sample, how to amplify nucleic acid, how to sequence nucleic acid, or any combination thereof. The instructions may further explain how to operate a computing device serving as analysis system 200 to perform the steps of any of the described methods.
[0327] In addition to the above components, the test kit can also include a computer-readable storage medium, which is stored with computer software for executing the various methods described in this disclosure. A form in which these instructions can exist is as a suitable medium or substrate (for example, one or more sheets of paper, information is printed on these papers), in the packaging of the test kit, in the printed information in the packaging insert. Another way can be a computer-readable medium, such as a floppy disk, a CD, a hard disk, a network data storage, wherein the instructions are stored in the form of computer code. Another means that can exist is a website, which can be used to access the information of a remote site via the Internet. IX. Cancer Applications
[0328] In some embodiments, the method, analysis system and / or classifier of the present disclosure can be used to detect the presence (or absence) of cancer, monitor cancer progression or recurrence, monitor treatment response or effectiveness, determine the presence or monitor minimal residual disease (MRD) or any combination thereof. In some embodiments, the analysis system and / or classifier can be used to identify the tissue of origin of cancer. For example, these systems and / or classifiers can be used to identify cancer as any of the following cancer types: head and neck cancer, liver / bile duct cancer, upper digestive tract cancer, pancreatic / gallbladder cancer; colorectal cancer, ovarian cancer, lung cancer, multiple myeloma, lymphoid tumors, melanoma, sarcoma, breast cancer and uterine cancer. For example, as described herein, a classifier can be used to generate a sample feature vector from the possibility or probability score (e.g., from 0 to 100) of a subject suffering from cancer. In some embodiments, the probability score is compared with a threshold probability to determine whether the subject suffers from cancer. In other embodiments, the possibility or probability score can be assessed at different time points (e.g., before or after treatment) to monitor disease progression or to monitor treatment effectiveness (e.g., therapeutic effect). In yet other embodiments, the likelihood or probability score can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, treatment effectiveness assessment, etc.). For example, in one embodiment, if the likelihood or probability score exceeds a threshold, a doctor can prescribe an appropriate treatment. In some embodiments, a test report can be generated to provide the patient with their test results, including, for example, a probability score that the patient has a disease state (e.g., cancer), a disease type (e.g., cancer type), and / or a tissue of origin of the disease (e.g., a cancer tissue of origin). IX.A. Early Detection of Cancer
[0329] In some embodiments, the methods and / or classifiers of the present disclosure are used to detect the presence or absence of cancer in a subject suspected of having cancer. For example, a classifier (as described herein) can be used to determine the likelihood or probability score that a sample feature vector is from a subject with cancer.
[0330] In one embodiment, a probability score of greater than or equal to 60 can indicate that the subject has cancer. In yet other embodiments, a probability score of greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95 can indicate that the subject has cancer. In other embodiments, the probability score can indicate the severity of the disease. For example, a probability score of 80 can indicate a more severe form or a more advanced stage of the cancer compared to a score below 80 (e.g., a score of 70). Similarly, an increase in the probability score over time (e.g., at a second later time point) can indicate disease progression, or a decrease in the probability score over time (e.g., at a second later time point) can indicate treatment success.
[0331] In another embodiment, the cancer log-odds ratio of the test subject can be calculated by taking the logarithm of the ratio of the cancer probability to the non-cancerous probability (i.e., one minus the cancer probability), as described herein. According to this embodiment, a cancer log-odds ratio greater than 1 can indicate that the subject has cancer. In yet other embodiments, a cancer log-odds ratio greater than 1.2, greater than 1.3, greater than 1.4, greater than 1.5, greater than 1.7, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4 can indicate that the subject has cancer. In other embodiments, the cancer log-odds ratio can indicate the severity of the disease. For example, compared to a score lower than 2 (e.g., a score of 1), a cancer log-odds ratio greater than 2 can indicate a more severe form or more advanced stage of cancer. Similarly, a cancer log-odds ratio that increases over time (e.g., at a second later time point) can indicate disease progression, or a cancer log-odds ratio that decreases over time (e.g., at a second later time point) can indicate treatment success.
[0332] According to various aspects of the present disclosure, the methods and systems of the present disclosure can be trained to detect or classify a variety of cancer indications. For example, the methods, systems, and classifiers of the present disclosure can be used to detect the presence of one or more, two or more, three or more, five or more, or ten or more different types of cancer.
[0333] In some embodiments, the cancer is one or more of: head and neck cancer, liver / bile duct cancer, upper gastrointestinal cancer, pancreatic / gallbladder cancer; colorectal cancer, ovarian cancer, lung cancer, multiple myeloma, lymphoid tumors, melanoma, sarcoma, breast cancer, and uterine cancer. IX.B. Cancer and Treatment Monitoring
[0334] In certain embodiments, the first time point is before the cancer treatment (e.g., before resection or therapeutic intervention), the second time point is after the cancer treatment (e.g., after resection or therapeutic intervention), and the method is used to monitor the effectiveness of the treatment. For example, if the second likelihood or probability score decreases compared to the first likelihood or probability score, the treatment is considered successful. However, if the second likelihood or probability score increases compared to the first likelihood or probability score, the treatment is considered unsuccessful. In other embodiments, the first time point and the second time point are both before the cancer treatment (e.g., before resection or therapeutic intervention). In yet other embodiments, the first time point and the second time point are both after the cancer treatment (e.g., before resection or therapeutic intervention), and the method is used to monitor the effectiveness of the treatment or the loss of therapeutic effectiveness. In yet other embodiments, cfDNA samples may be obtained and analyzed from cancer patients at the first and second time points, for example, to monitor cancer progression, to determine whether the cancer is in remission (e.g., after treatment), to monitor or detect residual disease or disease recurrence, or to monitor treatment (e.g., therapeutic) effects.
[0335] Those skilled in the art will readily appreciate that a test sample can be obtained from a cancer patient at any desired set of time points and analyzed according to the methods disclosed herein to monitor the patient's cancer status. In some embodiments, the time interval between the first time point and the second time point ranges from about 15 minutes to about 30 years, such as about 30 minutes, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or about 24 hours, such as about 1, 2, 3, 4, 5, 10, 15, 20, 25 or about 30 days, or such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 months, or such as about 1, 1.5, 2, 2.5, 3, 3. 5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5 or about 30 years. In other embodiments, the frequency of obtaining test samples from patients may be: at least once every 3 months, at least once every 6 months, at least once a year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years. IX.C. Treatment
[0336] In yet another embodiment, the information obtained from any of the methods described herein (likelihood or probability scores) can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, assessment of treatment effectiveness, etc.). For example, in one embodiment, if the likelihood or probability score exceeds a threshold, the physician can prescribe an appropriate treatment (e.g., resection, radiation therapy, chemotherapy, and / or immunotherapy). In some embodiments, information such as the likelihood or probability score can be provided as a readout to the physician or subject.
[0337] Classifiers (as described herein) can be used to determine the likelihood or probability score of a sample feature vector from a subject suffering from cancer. In one embodiment, when the likelihood or probability exceeds a threshold value, a suitable treatment (e.g., resection or therapeutic intervention) is prescribed. For example, in one embodiment, if the likelihood or probability score is greater than or equal to 60, one or more suitable treatments are prescribed. In another embodiment, if the likelihood or probability score is greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95, one or more suitable treatments are prescribed. In other embodiments, the cancer log-odds ratio can indicate the effectiveness of a cancer treatment. For example, a cancer log-odds ratio increase over time (e.g., at the second time after treatment) can indicate that the treatment is ineffective. Similarly, a cancer log-odds ratio decrease over time (e.g., at the second time after treatment) can indicate that the treatment is successful. In another embodiment, if the cancer log-odds ratio is greater than 1, greater than 1.5, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4, one or more appropriate treatments are prescribed.
[0338] In some embodiments, treatment is one or more cancer therapeutic agents selected from the group consisting of chemotherapeutics, targeted cancer therapeutics, differentiation therapeutics, hormone therapeutics, and immunotherapeutics. For example, treatment can be one or more chemotherapeutics selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cell scaffold disruptors (taxanes), topoisomerase inhibitors, nuclear division inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, treatment is one or more targeted cancer therapeutics selected from the group consisting of signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, treatment is one or more differentiation therapeutics, including retinoids such as tretinoin, alitretinoin, and bexarotene. In some embodiments, treatment is one or more hormonal therapeutic agents selected from the group consisting of antiestrogens, aromatase inhibitors, progestins, estrogens, antiandrogens, and GnRH agonists or analogs. In one embodiment, treatment is one or more immunotherapeutics selected from the group consisting of monoclonal antibody therapy (such as rituximab (RITUXAN) and alemtuzumab (CAMPATH)), nonspecific immunotherapy and adjuvants (such as BCG, interleukin-2 (IL-2), and interferon-α), immunomodulatory drugs (e.g., thalidomide and lenalidomide (REVLIMID)). An experienced physician or oncologist can select an appropriate cancer therapeutic based on a variety of characteristics, such as tumor type, cancer stage, previous exposure to cancer treatments or therapeutic agents, and other characteristics of the cancer. X. Examples XA Example 1 – Whole Genome Bisulfite Sequencing (WBGS)
[0339] First CCGA sub-study: 7A to 7F The data shown in Figure 1 are from the first CCGA substudy, in which training data blood samples (N = 1785) were collected from individuals diagnosed with untreated cancer (including 20 tumor types and all stages of cancer) and healthy individuals without a cancer diagnosis (controls) for plasma cfDNA extraction. An additional set of blood samples (N = 1,010) was collected for validation. Unless otherwise specified, both cell-free DNA (cfDNA) and genomic DNA (gDNA) extracted from samples in the first CCGA substudy were subjected to whole-genome bisulfite sequencing.
[0340] During the classification process, the analysis system 200 considers the fragment methylation state as extracted from a mixture of potential methylation patterns. The analysis system 200 assigns a relative probability of the observed fragments being derived from a particular cancer tissue of origin.
[0341] More specifically, as described herein, a probabilistic model is fitted to sequence reads from multiple regions (or windows) of each cancer type (as well as non-cancer or healthy samples). In this case, a mixture model is used in which each mixture component is an independent site model (in which the methylation at each CpG is independent of the methylation at other CpGs). The model is fitted using maximum likelihood estimation to identify a parameter set that maximizes the overall log-likelihood of all fragments derived from a cancer type (or non-cancer).
[0342] For each region, for each cancer type pair (including non-cancer as a negative type), the best performing stratum was used to train a multinomial logistic regression classifier. For each sample (regardless of label), each region, each cancer type, each segment, the log-likelihood ratio was calculated as described above, and for each set of "stratum" values, the R 癌症类型 The number of fragments at each level. The quantized read counts at each level are binarized and used as features for training the classifier.
[0343] Finally, to generate predictions for unknown samples, if indicated, feature values are determined (as described above), and the generated features are used to create cancer and / or tissue of origin predictions using a trained multinomial logistic regression classifier.
[0344] Confusion Matrix Example: Figure 7A 、 Figure 7B and Figure 7C A confusion matrix indicating the accuracy of a classifier according to various embodiments is included. In some embodiments, the analysis system 200 uses a confusion matrix to determine the accuracy of a classifier. The confusion matrix includes information describing the success rate of the classifier in identifying each disease state.
[0345] like Figure 7A As shown, matrix 710 includes example performance of a classifier based on a multinomial model trained using a cfDNA sample set (without tissue samples). Matrix 720 includes example performance of a classifier based on a hybrid model trained by analysis system 200 using the same cfDNA sample set. Scores along the diagonal of the matrix indicate correct predictions, that is, the tissue of origin predicted for the fragment matches the true tissue of origin. Compared to the classifier based on the multinomial model as a baseline, the classifier based on the hybrid model has a higher overall accuracy in predicting the presence of cancer types, as shown in the matrix.
[0346] The samples of the training set can be screened based on one or more criteria (e.g., a particular level of specificity). For example, the training set includes samples that were determined to have cancer based on a specificity of 98% according to the m-score. For clarity, the remaining (e.g., 2%) non-cancer samples that were (erroneously) identified as having cancer are not shown in the confusion matrix.
[0347] like Figure 7B As shown, matrix 730 includes example performance of a classifier based on a hybrid model trained using a cross-validation training set of cfDNA samples (no tissue samples). Matrix 740 includes example performance of a classifier based on a hybrid model trained using a cross-validation training set of cfDNA samples and tissue samples.
[0348] like Figure 7C As shown, matrix 750 includes example performance of a classifier based on a hybrid model trained using a set of cfDNA samples (without tissue samples) from a clinical study titled "Circulating Cell-Free Genome Atlas Study" ("CCGA"). Matrix 740 includes example performance of a classifier based on a hybrid model trained using a set of cfDNA samples and tissue samples from the CCGA. The CCGA study is described using the following ClinicalTrial.gov identifier: NCT02889978 (https: / / www.clinicaltrials.gov / ct2 / show / NCT02889978). XB Example 2 – Classifying cancer using targeted bisulfite sequencing, an early breakthrough from the second CCGA substudy
[0349] Second CCGA sub-study: Figures 9A to 9B 、 FIG. 10A to FIG. 10B 、 Figure 11 and Figure 12The data shown in are from an early breakthrough of the second CCGA sub-study, in which training data blood samples (N=3,132) were collected from individuals diagnosed with untreated cancer (including 20 tumor types and all stages of cancer) and healthy individuals without cancer diagnosis (controls) for plasma cfDNA extraction. Another set of blood samples (N=1,354) was collected for validation. In some embodiments, where indicated, the training set also includes training data from tissue samples (i.e., gDNA). In order to determine the analysis population, the training data blood samples were screened based on several factors. For example, 105 samples were excluded because they were clinically unlocked; 11 samples were excluded based on eligibility criteria; 58 samples were excluded because their cancer or treatment status was not confirmed (unable to be evaluated); 4 untreated samples and 72 unevaluable tests (unable to be analyzed) were excluded; and 581 samples were retained for subsequent analysis. The results showed that the analysis population of 2,301 samples included 1,422 cancer samples and 879 non-cancer samples.
[0350] The participant demographics for the individuals in the sub-studies are shown in Table 1 below. Table 1: Demographics and Stage Distribution of Participants. The cancer and non-cancer groups were comparable with respect to age, race, sex, and body mass index (data not shown). *Includes anorectal cancer, bladder cancer, brain cancer, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastric cancer, head and neck cancer, hepatobiliary cancer, lung cancer, lymphoid neoplasms (chronic lymphocytic leukemia, lymphoma), multiple myeloma, myeloid neoplasms (acute myeloid leukemia, chronic myeloid leukemia), ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer, sarcoma, and uterine cancer. Thirty-eight participants with missing information on smoking status were excluded. Two participants with missing BMI values were excluded. §Invasive cancer only. Installment information is not available.
[0351] To identify cancer- and tissue-defined methylation signals, bisulfite sequencing assays were performed on extracted cfDNA, targeting the most informative regions in the methylome as identified based on GRAIL's proprietary whole-genome bisulfite sequencing assay and methylation database.
[0352] We used a methylation database that was interrogated for genome-wide fragment-level methylation patterns on 811 cancer cells representing 21 tumor types (accounting for 97% of SEER cancer incidence). To generate a methylation database of cancer-defined methylation signals, whole-genome bisulfite sequencing was performed on genomic DNA from formalin-fixed, paraffin-embedded (FFPE) tumor tissue and cells isolated from tumors. The methylation database was used for panel design and training to optimize classifier performance, as described in this article. Large methylation sequence databases for both cancer and non-cancer cancers were generated to select targets for a single test to classify multiple cancers with high specificity and identify tissue of origin.
[0353] Target selection and panel design: As described in this paper, the methylation sequence database from the CCGA study was used to select target genomic regions. Specifically, using the non-cancer distribution, the cfDNA sequences in the database were screened based on the p-value, and only fragments with p < 0.001 were retained. The selected cfDNA were further screened to retain only cfDNA that was at least 90% methylated or 90% unmethylated. Next, for each CpG site in the selected fragments, the number of cancer samples or non-cancer samples that included fragments overlapping with that CpG site was calculated. Specifically, P(cancer|overlapping fragments) was calculated for each CpG, and genomic sites with high P-values were selected as general cancer targets. By design, the noise of the selected fragments was very low (i.e., there were few overlapping non-cancer fragments).
[0354] To find cancer type-specific targets, a similar selection process was performed. CpG sites were ranked based on their information gain, comparing one cancer type to all other samples (i.e., non-cancer plus other cancer types). As described herein, cancer detection panels comprising probes targeting selected genomic regions are generated. Specifically, these detection panels are typically designed to detect the presence of cancer (i.e., distinguishing cancer from non-cancer) or specific cancer types (e.g., TOO). These detection panels include probe sets targeting each selected genomic region.
[0355] Probes are designed to overlap any CpG site included within the start / end range of any targeted region (eg, informative fragment).
[0356] Classification: During classification, analysis system 200 treats the fragment methylation state as extracted from a mixture of potential methylation patterns. Analysis system 200 assigns a relative probability of cancer origin to each observed fragment. For tissue-of-origin classification, analysis system 200 assigns a relative probability of origin to each observed fragment. Analysis system 200 combines the cancer and tissue-of-origin characteristics of the fragment in the targeted region to classify cancer from non-cancer and / or identify the tissue of origin. For binary cancer classification, analysis system 200 estimates sensitivity at 99% specificity.
[0357] More specifically, as described in Example VI.a, a probabilistic model is fitted to sequence reads from multiple regions (or windows) of each cancer type (as well as non-cancer or healthy samples), the identified features, and a trained multinomial logistic regression classifier. To generate predictions for unknown samples, feature values are determined (as described above), and the generated features are used to create cancer and / or tissue of origin predictions using a trained multinomial logistic regression classifier.
[0358] Figure 9A and Figure 9B The sensitivity of the tissue of origin classifier generated by the methods described in this disclosure is shown. Sensitivity is reported at 99% specificity, and 95% confidence intervals are indicated. Figure 9A Model predictions for a pre-specified list of cancers are presented. Figure 9B Model predictions for other cancers included in the CCGA study were presented. Demographic information alone (baseline modeling) correctly classified less than 5% of participants. Across a prespecified list of cancers (anorectal cancer, breast cancer [HR-negative], colorectal cancer, esophageal cancer, gastric cancer, head and neck cancer, hepatobiliary cancer, lung cancer, lymphoid neoplasms [chronic lymphocytic leukemia, lymphoma], multiple myeloma, ovarian cancer, pancreatic cancer), the overall sensitivity was 76.1% (95% CI: 73.1%-78.9%). Among early-stage (stage I-III) cancers in this cohort, the sensitivity was 68.8% (95% CI: 64.8%-72.6%). Across all cancer types and stages, the overall sensitivity was 55.1% (95% CI: 52.5%-57.7%). Among early-stage (stage I-III) cancers, the sensitivity was 43.8% (95% CI: 40.7%-46.8%).
[0359] Figure 10A and Figure 10BThe sensitivity of the tissue of origin classifier at different cancer stages is shown. As shown in the legend, for the pre-specified cancer of interest, the sensitivity by individual stage is reported as 99% specificity. The numbers in the boxes represent the total number of samples included in each stage. The 95% confidence interval is indicated. "Lymphoid tumors" include lymphomas (stages I-IV) and chronic lymphocytic leukemia (unstaged, classified as "NI").
[0360] Figure 11 A performance grid representing the accuracy of tissue of origin localization is shown. In stage I-IV samples, there is consistency between the true (x-axis) tissue of origin of each sample and the tissue of origin predicted (y-axis) using the tissue of origin classifier and the methylation database. The gradient legend corresponds to the proportion of predicted tissues of origin (y-axis) that are correct (x-axis). The analysis shows that the accuracy of tissue of origin localization (the correct proportion of all TOO predictions) is higher using the methylation database (p = 0.0066). This is consistent with the predictions for stage I-III: 89.9% (384 / 427), as further shown in Table 2. Table 2: Tissue-of-origin performance is improved when a methylation database is included. *P values were calculated using the Stuart-Maxwell test. An uncertain base call is defined as a sample that is detected as cancer but has no confident tissue of origin. Samples not base called by tissue of origin analysis were classified as non-cancer.
[0361] Ideally, an effective multicancer test should be able to simultaneously detect clinically significant cancers of all stages with high specificity (thus having a single, fixed, low false-positive rate) and accurately determine the tissue of origin. To demonstrate the potential of this approach, Figure 12 A summary of simultaneous detection (sensitivity reported at 99% specificity) and tissue of origin determination for a pre-specified list of cancer types at each stage is shown. Figure 12 The accuracy and sensitivity of the tissue-of-origin classifier for different cancer stages are demonstrated.
[0362] Figure 13A and Figure 13B Receiver operating characteristic (ROC) curves for the tissue of origin classifier are shown. The ROC curves show classifier performance at 99% specificity: 55% sensitivity for all cancers and 76% sensitivity for multiple cancers.
[0363] These data demonstrate that a classification method using targeted methylation signatures can simultaneously detect multiple cancer types at an early stage with a specificity (99%) suitable for population screening. This approach enables detection of multiple cancer types with a single, fixed, low false-positive rate. This approach also accurately localizes the tissue of origin, simplifying downstream diagnostic workup. Furthermore, incorporating data from a large methylation database improves classifier performance.
[0364] Taken together, this supports the potential clinical applicability of the methods described in this disclosure as early multi-cancer detection tests for a variety of clinically important cancer types. XC Example 3 – Classifying Cancer Using Targeted Bisulfite Sequencing (from the Complete Second CCGA Substudy)
[0365] Generation of Hybrid Model Classifiers: To maximize performance, the predictive cancer model described in this example was trained using sequence data from multiple samples of known cancer types and non-cancers from two CCGA sub-studies (CCGA1 and CCGA2), multiple tissue samples of known cancers from CCGA1, and multiple non-cancer samples from the STRIVE study (see Clinical Trail.gov identifier: NCT03085888 (http: / / clinicaltrials.gov / ct2 / show / NCT03085888)). Additional non-cancer training samples were obtained from the STRIVE study, a prospective, multicenter, observational cohort study designed to validate assays for the early detection of breast cancer and other invasive cancers, to train the classifier described herein. The known cancer types included in the CCGA sample set are as follows: breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, gastric cancer, and anorectal cancer. Thus, the model can be a multi-cancer model (or multi-cancer classifier) for detecting one or more, two or more, three or more, four or more, five or more, ten or more, or 20 or more different types of cancer. 4,841 participants in the CCGA study (2,836 with cancer; 2,005 without cancer) and 2,202 without cancer in the STRIVE study were included in this pre-specified analysis. Of these, 3,133 samples from CCGA were assigned to training (1,742 cancer samples; 1,391 non-cancer samples), and 1,354 samples were assigned to validation (740 cancer samples, 614 non-cancer samples); 1,587 samples from STRIVE were assigned to training, and 615 samples were assigned to validation. Participant disposition is indicated. Overall, 3,052 samples in the training set (1,531 cancer samples; 1,521 non-cancer samples) and 1,264 samples in the validation set (654 cancer samples; 610 non-cancer samples) were available for analysis and included in the pre-specified primary analysis population.Further details about the CCGA2 substudy and the analyses detailed in this example are described in an article in the Annals of Oncology, titled “Sensitive and specific multi-cancer detection and localization using methylation signatures in cell-free DNA,” published online on March 30, 2020 (https: / / www.annalsofoncology.org / article / S0923-7534(20)36058-0 / fulltext).
[0366] The classifier performance data shown below are reported for locked classifiers trained on cancer and non-cancer samples obtained from CCGA2 (a substudy of CCGA) and non-cancer samples obtained from STRIVE. The individuals in the CCGA2 substudy differed from the individuals in the CCGA1 substudy whose cfDNA was used to select the target gene set (as described in WO2019 / 195268, filed April 2, 2019, PCT / US2019 / 053509, filed September 27, 2019, and PCT / US2020 / 015082, filed January 24, 2020, which are incorporated herein by reference). From the CCGA2 study, blood samples were collected from individuals diagnosed with untreated cancer (including 20 tumor types and all stages of cancer) and healthy individuals without a cancer diagnosis (control group). For the STRIVE study, blood samples were collected from women within 28 days of their mammogram. Cell-free DNA (cfDNA) was extracted from each sample and treated with bisulfite to convert unmethylated cytosine to uracil. The bisulfite-treated cfDNA was enriched to obtain informative cfDNA molecules using hybridization probes designed to enrich for bisulfite-converted nucleic acids from each of multiple targeted genomic regions in three cancer panels: (1) Pan-Cancer Panel #4 as described and disclosed in WO 2019 / 195268 (labeled herein as Panel A); (2) Pan-Cancer Panel #5 as described and disclosed in WO 2019 / 195268 (labeled herein as Panel B); and (3) a large proprietary pan-cancer panel (Panel C, described below). The enriched bisulfite-converted nucleic acid molecules were sequenced on the Illumina platform (San Diego, California, USA) using paired-end sequencing technology to obtain a set of sequence reads for each training sample. The resulting read pairs were then aligned to the reference genome, assembled into fragments, and methylated and unmethylated CpG sites were identified. Feature extraction based on hybrid model
[0367] For each cancer type (including non-cancer), a probabilistic mixture model was trained and used to assign a probability to each fragment in each cancer sample and non-cancer sample based on the likelihood of observing the fragment in a given sample type. Fragment-level analysis
[0368] Briefly, for each sample type (cancer and non-cancer samples), for each region (where each region was used as is if less than 1 kb, otherwise subdivided into regions of 1 kb in length with 50% overlap (e.g., 500 base pairs overlap) between adjacent regions), a probabilistic model was fit to fragments derived from training samples of each cancer type and non-cancer. The probabilistic model trained for each sample type was a mixture model in which each of the three mixture components was an independent site model in which methylation at each CpG was considered independent of methylation at other CpGs. Fragments were excluded from the model if: their p-value (from the non-cancer Markov model) was greater than 0.01; they were marked as duplicates; the bag size of the fragment was greater than 1 (only for targeted methylation samples); they did not cover at least one CpG site; or if the fragment was greater than 1000 bases in length. Retained training fragments were assigned to a region if they overlapped with at least one CpG in that region. If a fragment overlaps with CpGs in multiple regions, the fragment will be assigned to all regions. Local Source Model
[0369] Each probabilistic model is fitted using maximum likelihood estimation to identify the set of parameters that maximizes the log-likelihood of all fragments from each sample type under the condition of regularization penalty. Specifically, in each classification region, a set of probabilistic models is trained, one for each training label (i.e., one for each cancer type and one for non-cancer). Each model takes the form of a Bernoulli mixture model with three components. The mathematical form is, Where n is the number of hybrid components, set to 3; m i ∈{0,1} is the observed methylation at position i of the fragment; f k is the score assigned to component k (where f k ≥0 and∑f k =1); and β ki is the methylation score of component k at CpG i. The product over i includes only those positions whose methylation status can be identified by sequencing. The parameters of each model {f k , β ki The maximum likelihood of} is estimated by using the rprop algorithm (such as that described in: Riedmiller M, Braun H. RPROP - A Fast Adaptive Learning Algorithm, Proceedings of the International Symposium on Computer and Information Science VII, 1992) to kiMaximize the total log-likelihood of a segment of a training label under the condition of a regularization penalty in the form of a beta distribution prior. Mathematically, the quantity being maximized is Here, r is the regularization strength, which is set to 1. Feature extraction
[0370] Once the probabilistic model is trained, a set of numerical features is calculated for each sample. Specifically, features are extracted for each fragment from each training sample, for each cancer type and non-cancer sample, and for each region. The extracted features are counts of anomalous fragments (i.e., informative fragments), which are defined as fragments whose log-likelihood under the first cancer model exceeds the log-likelihood under the second cancer model or the non-cancer model by at least a threshold level value. Anomalous fragments are counted separately for each genomic region, sample model (i.e., cancer type), and level (for levels 1, 2, 3, 4, 5, 6, 7, 8, and 9), resulting in 9 features per region for each sample type. In this way, each feature is defined by three attributes: genomic region; a "positive" cancer type label (excluding non-cancer); and a level value selected from the set {1, 2, 3, 4, 5, 6, 7, 8, 9}. The numerical value of each feature is defined as the number of fragments in the region such that where the probabilities are defined by equation (1) using the maximum likelihood estimated parameter values corresponding to “positive” cancer type (numerator of the logarithm) or non-cancer (denominator). Feature Ranking
[0371] For each set of paired features, the mutual information is used to rank these features based on their ability to distinguish between a first cancer type (which defines the log-likelihood model from which the feature is derived) and a second cancer type or non-cancer. Specifically, for each pair of unique class labels, two ranked feature lists are compiled: one list designates the first label as "positive" and the second label as "negative", while the other list swaps the positive / negative assignments (except for the "non-cancer" label, which is only allowed as a negative label). For each of these ranked lists, only features whose positive cancer type label (as shown in Equation (3)) matches the positive label under consideration are included in the ranking. For each such feature, the proportion of training samples with non-zero feature values is calculated for the positive label and the negative label, respectively. Features with a larger proportion in the positive label are ranked according to their mutual information relative to the pair of class labels.
[0372] The top 256 features in each pairwise comparison were identified and added to the final feature set for each cancer type and non-cancer. To avoid redundancy, if more than one feature was selected from the same positive type and genomic region (i.e., multiple negative types), only the feature with the lowest ranking (most informative) assigned to its cancer type pair was retained. If the rankings were the same, the feature with the higher rank value was selected. The features in the final feature set for each sample (cancer type and non-cancer) were binarized (any feature value greater than 0 was set to 1, so all features were either 0 or 1). Classifier training
[0373] The training samples were then split into different 5-fold cross-validation training sets, and a two-class classifier was trained for each fold, in each case training on 4 / 5 of the training samples and using the remaining 1 / 5 for validation.
[0374] In the first level of training, a binary (two-class) logistic regression model for detecting the presence of cancer was trained to distinguish between cancer samples (regardless of TOO) and non-cancer samples. When training this binary classifier, sample weights were assigned to male non-cancer samples to offset the gender imbalance in the training set. For each sample, the binary classifier outputs a prediction score indicating the likelihood of the presence or absence of cancer.
[0375] In the second level of training, a parallel multi-class logistic regression model for determining the tissue of origin of cancer is trained using TOO as the target label. Only cancer samples that score above the 95th percentile of non-cancer samples in the first level classifier are included in the training of the multi-class classifier. For each cancer sample used in training the multi-class classifier, the multi-class classifier outputs a predicted value for the classified cancer type, where each predicted value is the probability that a given sample has a certain cancer type. For example, the cancer classifier can return a cancer prediction for the test sample, including a prediction score for breast cancer, a prediction score for lung cancer, and / or a prediction score for the absence of cancer.
[0376] Both binary and multiclass classifiers were trained using mini-batch stochastic gradient descent, and in each case, training was stopped early when the performance of the validation fold (assessed by cross-entropy loss) began to decline. To make predictions for samples outside the training set, at each level, the scores assigned by the five cross-validated classifiers were averaged. Scores assigned to sex-inappropriate cancer types were set to zero, and the remaining values were renormalized to unity.
[0377] The scores assigned to the validation folds within the training set are retained for use in assigning cutoff values (thresholds) for certain performance metrics. Specifically, the probability scores assigned to the non-cancer samples in the training set are used to define a threshold corresponding to a particular level of specificity. For example, for a desired specificity target of 99.4%, the threshold is set to the 99.4th percentile of the cross-validated cancer detection probability scores assigned to the non-cancer samples in the training set. Training samples with a probability score exceeding the threshold are called cancer positive.
[0378] Subsequently, for each training sample determined to be cancer positive, a TOO or cancer type assessment was performed by a multi-class classifier. First, a multi-class logistic regression classifier assigned a set of probability scores to each sample, one score for each potential cancer type. Next, the confidence of these scores was assessed as the difference between the highest score and the second highest score assigned to each sample by the multi-class classifier. The cross-validated training set scores were then used to identify the lowest threshold so that 90% of the cancer samples in the training set whose first two scores differed by more than the threshold were assigned the correct TOO label as their highest score. In this way, the scores assigned to the validation folds during training were further used to determine a second threshold for distinguishing between confident and uncertain TOO base calls.
[0379] When predicting, samples whose scores from the binary (first-level) classifier are below a predefined specificity threshold are assigned a "non-cancer" label. For the remaining samples, samples whose first two TOO scores differ from each other in the second-level classifier below a second predefined threshold are assigned an "uncertain cancer" label. The remaining samples are assigned the cancer label that the TOO classifier has assigned the highest score. Classifier performance on the target genomic region detection panel
[0380] The discriminatory value of the target genomic regions of test panels AC was evaluated by testing the ability of the cancer classifier to detect cancer and 20 different cancer types based on the methylation status of these target genomic regions. For test panels AB, performance was evaluated on a training set of 1,531 cancer samples and 1,521 non-cancer samples used to train the classifier, as shown in Table 1. For test panel C, the performance of the classifier trained using the same 3,052 samples (1,531 cancer; 1,521 non-cancer) used for training on test panels AB was evaluated using 1,264 validation samples (654 cancer; 610 non-cancer). For each sample, a bait set of all target genomic regions included in test panels AC was used to enrich for differentially methylated cfDNA. The classifier was then constrained to provide a cancer determination based only on the methylation status of the target genomic regions in the evaluated list. The two-stage classifier embodiment includes a binary (two-class) logistic regression classifier model trained to distinguish cancer samples (regardless of TOO) from non-cancer for detecting the presence of cancer, and a second stage trained multi-class logistic regression classifier model for determining the tissue of origin of the cancer, the multi-class logistic regression classifier model being trained with TOO as the target label, as previously described in this example. Also as previously described, both classifier models were trained and validated using model-based feature extraction Table 1 - Cancer diagnosis for individuals whose cfDNA was used to train the classifier
[0381] Test groups A and B: Figure 26A and Figure 27A Figure 2 presents the results of the classifier performance analysis of detection groups A and B. In each figure, part A is a receiver operating characteristic curve (ROC) showing the true positive results and false positive results of cancer or non-cancer determination. The asymmetric shape of these ROC curves illustrates that the classifier is designed to minimize false positive results. For detection groups A and B, the area under the curve of these two detection groups is 0.83.
[0382] Cancer type (ie, TOO) was determined for all samples that tested positive for cancer using the classifier. Figure 26B and Figure 27B Included are confusion matrices indicating the accuracy of the TOO accuracy for detection groups A and B. The confusion matrices include information describing how successfully the classifier identified each cancer type and excluded uncertain cancer base calls.
[0383] like Figure 26B and Figure 27BAs shown, the TOO confusion matrix shows the performance of the multi-class logistic regression classifier, as described above. It describes the agreement between the actual (x-axis) tissue of origin of each sample and the tissue of origin predicted (y-axis) using the targeted methylation classifier. The scores along the diagonal of the matrix indicate correct predictions, that is, the tissue of origin predicted for the fragment matches the true tissue of origin. Figure 26B As shown, when the uncertain cancer base calls were excluded, the TOO accuracy of cancer detection panel A was about 90.8% (711 / 783). Figure 27B It was shown that when the uncertain cancer base calls were excluded, the TOO accuracy of test panel B was approximately 90.3% (705 / 781).
[0384] These classifier results are further summarized in Tables 2 to 3, which indicate that the accuracy of cancer detection and cancer type determination was made with a specificity of 0.990 (indicating a false positive rate of 1%). These results were divided by cancer stage. The results showed that cancer detection and cancer type determination were improved for samples from individuals with advanced cancer (e.g., stage III) compared to samples from individuals with early cancer (e.g., stage II). For all cancer stages (not categorized by stage), the accuracy of cancer type determination for test panels A and B (including uncertain cancer base calls) was approximately 89%.
[0385] Table 2. Accuracy of classification using genomic regions from test panel A. Data for cancer presence and cancer type are shown with percent accuracy at a specificity of 0.990, 95% confidence intervals in parentheses, and the number of correct assignments and totals in parentheses. Staging Cancer exists Cancer type I 20.4%[16.6-24.5](86 / 422) 71.8%[60.5-81.4](56 / 78) II 44.6%[39.6-49.7](173 / 388) 87.2%[81.1-91.9](143 / 164) III 81.5%[76.7-85.6](255 / 313) 90.5%[86.1-93.9](220 / 243) IV 90.9%[87.5-93.7](330 / 363) 93.3%[90-95.8](294 / 315) all 56.5%[54-59](866 / 1532) 89.1%[86.8-91.2](731 / 820)
[0386] Table 3. Accuracy of classification using genomic regions from test panel B. Data for cancer presence and cancer type are shown with percent accuracy at a specificity of 0.990, 95% confidence intervals in parentheses, and the number of correct assignments and totals in parentheses. Staging Cancer exists Cancer type I 19.9%[16.2-24](84 / 422) 72.7%[60.4-83](48 / 66) II 45.1%[40.1-50.2](175 / 388) 84.8%[78.2-90](134 / 158) III 81.2%[76.4-85.3](254 / 313) 91.3%[86.9-94.6](211 / 231) IV 90.9%[87.5-93.7](330 / 363) 93.2%[89.8-95.7](287 / 308) all 56.3%[53.7-58.8](862 / 1532) 89.2%[86.9-91.3](697 / 781)
[0387] Panel C: As described above, a third large proprietary pan-cancer panel was also tested. Panel C was designed based on WGBS data obtained from the first CCGA substudy, CCGA1, using the feature selection methods disclosed in PCT / US2019 / 053509, filed September 27, 2019, and PCT / US2020 / 015082, filed January 24, 2020 (incorporated herein by reference). The large proprietary targeted methylation panel covers 103,456 distinct regions (17.2 Mb) and 1,116,720 CpGs. Panel C included 363,033 CpGs in 68,059 regions (7.5 Mb) covered by probes targeting hypomethylated fragments; 585,181 CpGs in 28,521 regions (7.4 Mb) covered by probes targeting hypermethylated fragments; and 218,506 CpGs in 6,876 regions (2.3 Mb) targeting both types of fragments. Individual aberrant target regions contained 1 to 590 CpGs, with a median CpG count of 3 in hypomethylated target regions and 6 in hypermethylated target regions. CpGs are present in the following genomic regions: 193,818 (17%) are located in the 1-5 kbp region upstream of the transcription start site (TSS); 278,872 (24%) are located in promoters (<1 kbp upstream of the TSS); 500,996 (43%) are located in introns; 292,789 (25%) are located in exons; 247,752 (21%) are located in intron-exon boundaries; 134,144 (11%) are located in the 5'-untranslated region; 182,174 (16%) are located between genes; and the remaining 1,817 (<1%) are unannotated. Percentages are relative to the total number of CpGs and do not equal 100% because each CpG may receive multiple annotations due to overlap of genes and / or transcripts.
[0388] For this assessment, samples were divided into a training set (n=4,720) and an independent validation set (n=1,969). A total of 4,316 participants (training: 3,052 [1,531 cancer: stage I: 28%; stage II: 25%; stage III: 20%; stage IV: 24%; missing / unexpected: 3%; 1,521 non-cancer]; validation: 1,264 [654 cancer: stage I: 28%; stage II: 25%; stage III: 21%; stage IV: 23%; missing / unexpected: 3%; 610 non-cancer]) were available for analysis and included in the primary analysis population.
[0389] Figures 28 to 30 The classifier performance analysis results for the training set and the validation set are shown. Figure 28Figure A shows the specific results for the training and validation sets, and Figure B shows the sensitivity for prespecified cancers (a subset of 12 high-signal cancers based on results and mortality data from the first substudy (anal cancer, bladder cancer, colon / rectum cancer, esophageal cancer, head and neck cancer, liver / bile duct cancer, lung cancer, lymphoma, ovarian cancer, pancreatic cancer, plasma cell neoplasms, and gastric cancer)) and all cancer types (>20) in stages I to IV. Figure 28 Figure 29 shows the TOO confusion matrix for the training and validation sets, Figure 30 Sensitivity results for pre-specified cancer types are shown for the training and validation sets.
[0390] exist Figure 28 In the figure, sensitivity (y-axis) is reported for pre-specified cancer types (left inset) and all cancer types (right inset) by clinical stage (x-axis) for the training set (orange) and validation set (cyan). Tissue of origin accuracy (y-axis) is reported for pre-specified cancer types (left inset) and all cancer types (right inset) by clinical stage (x-axis) for the training set (orange) and validation set (cyan). Numbers indicate samples in the training|validation set.
[0391] like Figure 28As shown, the classifier achieved consistently high specificity between the cross-validation training set and the independent validation set (99.8% [95% CI: 99.4-99.9%] vs. 99.3% [98.3-99.8%], respectively; P = 0.095); this reflects a single, consistent false positive rate (FPR) of less than 1% across all 20 cancer types. For non-cancer samples from CCGA and STRIVE, the specificity of the validation set was similar (99.3% [97.4-99.9%] vs. 99.4% [97.9-99.9%], respectively), indicating that performance was not affected by the site or the selected sample. Sensitivity was consistent in the training and validation sets. Across all cancers, the sensitivity for stages I-III was 44.2% (95% CI: 41.3-47.2%) vs. 43.9% (39.4-48.5%), respectively (P = 1.000). For the pre-specified set of 12 high-signal cancers, the sensitivity for stages I-III was 69.8% (65.6-73.7%) and 67.3% (60.7-73.3%), respectively (P = 0.988). Similarly, the sensitivity for stages I-IV across all cancer types was 55.2% (52.7-57.7%) and 54.9% (51.0-58.8%), respectively (P = 0.897), and the sensitivity for pre-specified cancers was 77.9% (75.0-80.7%) and 76.4% (71.6-80.7%), respectively (P = 0.573).
[0392] In addition, if Figure 28 As shown, sensitivity increases with increasing disease stage. In validation, the sensitivity of pre-specified cancer types was 39% (27-52%) in stage I (n=62), 69% (56-80%) in stage II (n=62), 83% (75-90%) in stage III (n=102), and 92% (86-96%) in stage IV (n=130). Across all cancer types, sensitivity was 18% (13-25%) in stage I (n=185), 43% (35-51%) in stage II (n=166), 81% (73-87%) in stage III (n=134), and 93% (87-96%) in stage IV (n=148).
[0393] Figure 30 Performance in each tumor type is depicted. For individual cancer types with at least 50 samples, sensitivity at 99.8% specificity (training, orange) or 99.3% specificity (validation, cyan) is reported, with a 95% confidence interval. Clinical stage and the number of samples in training and validation are noted below the graph.
[0394] like Figure 28As shown, a pre-specified analysis of TOO accuracy (the proportion of all correct TOO predictions) found that among samples with cancer-like signatures in the validation set, TOO was predicted in 96% (344 / 359); of these, 93% (321 / 344) were accurate. Accuracy was consistent between the training and validation sets and across stages. The classifier distinguished between more than 20 cancer types included in the study, and performance was consistent across cancer types.
[0395] Figure 29 shows a confusion matrix representing the accuracy of tissue-of-origin positioning in (A) training set and (B) validation set. Describes the consistency between the actual (x-axis) tissue-of-origin of each sample and the tissue-of-origin predicted (y-axis) using the targeted methylation classifier. The color corresponds to the ratio of the predicted tissue-of-origin base calls. The included participants (training: n=844, validation: n=359) were cancer patients predicted to have cancer with 99.8% specificity (training set) or 99.3% specificity (validation set). In the training set, 95% (806 / 844) of the cases had been assigned tissue-of-origin base calls, and in the validation set, 96% (344 / 359) of the cases had been assigned tissue-of-origin base calls; In the training set, base calls in 92% (744 / 806) of the cases were correct, and in the validation set, base calls in 93% (321 / 344) of the cases were correct. XD Example 4 - Adjusting the Binary Classification Threshold
[0396] According to a general embodiment of binary cancer classification, an analysis system determines a cancer score for a test sample based on sequencing data (e.g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.) of the test sample. The analysis system compares the cancer score of the test sample with a binary threshold cutoff value that predicts whether the test sample is likely to have cancer. The binary threshold cutoff value can be adjusted using TOO thresholding based on one or more TOO subtype categories. The analysis system can further generate a feature vector for the test sample for determining a cancer prediction indicating one or more likely cancer types in a multi-class cancer classifier.
[0397] Figure 24AA confusion matrix illustrating the performance of a trained cancer classifier according to an example embodiment is shown. The cancer classifier was trained according to the principles described above. The TOO labels included: lymphoid tumors, lung cancer, kidney cancer, non-cancer, head and neck cancer, prostate cancer, breast cancer, upper gastrointestinal cancer, hepatobiliary cancer, colorectal cancer, cervical cancer, pancreatic and gallbladder cancer, uterine cancer, sarcoma, bladder and urothelial cancer, ovarian cancer, anorectal cancer, unknown origin, melanoma, multiple myeloma, medullary tumors, and thyroid cancer. Notably, the classification accuracy on the 1,151 samples considered in this holdout set was 89.1%.
[0398] Figure 24B The confusion matrix showing the performance of the trained cancer classifier for additional hematological cancer subtypes is shown. The cancer classifier was trained according to the above principles. Figure 24A In comparison, the TOO labels for hematological subtypes have been adjusted. Figure 24A In the context of hematologic subtypes, the subtypes include lymphoid neoplasms, multiple myeloma, and myeloid neoplasms. Figure 24B Among them, hematological subtypes include Hodgkin lymphoma (HL), aggressive NHL, indolent NHL, myeloid neoplasms, circulating lymphoma (or lymphoid neoplasms), and plasma cell neoplasms. Notably, the classification accuracy on 1,076 was 87.5%.
[0399] Figure 25A and Figure 25B A graph is shown showing the cancer prediction accuracy for multiple cancer types at different cancer stages. In this example, the cancer classifier is trained after pruning non-cancer samples according to process 1000 described above. The analysis system determines multiple TOO thresholds for hematological subtypes. The analysis system excludes non-cancer samples with at least one TOO probability equal to or higher than the corresponding TOO threshold for the hematological subtype. The graph shown shows the classification sensitivity of the following cancer types at different cancer stages: anorectal cancer, bladder and urothelial cancer, breast cancer, cervical cancer, colorectal cancer, head and neck cancer, hepatobiliary cancer, lung cancer, melanoma, ovarian cancer, pancreatic and gallbladder cancer, prostate cancer, kidney cancer, sarcoma, thyroid cancer, upper gastrointestinal cancer, and uterine cancer. The graph for each cancer type shows the prediction sensitivity of a first cancer classifier labeled "locked_v1_orgi" that does not use TOO thresholding and a second cancer classifier labeled "v2_custom" that uses TOO thresholding at each stage of the cancer type. Notably, for many cancer types, with more samples available for validation, the second cancer classifier achieved higher prediction accuracy while maintaining tight confidence intervals. Of particular note, many cancer types achieved higher prediction accuracy at stages I and II, suggesting that using TOO thresholding can improve the predictive potential of early-stage cancers. XI. Other Considerations
[0400] The above description of the disclosed embodiments has been presented for illustrative purposes; it is not intended to be exhaustive or to limit the invention to the precise form disclosed. Those skilled in the relevant art will appreciate that many modifications and variations can be made based on the above disclosure.
[0401] Certain portions of this specification describe embodiments of the present disclosure in terms of symbolic representations of algorithms and information operations. These algorithmic descriptions and representations are typically used by those skilled in the art of data processing to effectively convey the essence of their work to other persons skilled in the art. Although these operations are described functionally, computationally, or logically, they should be understood to be implemented by computer programs or equivalent circuits, microcode, etc. In addition, without loss of generality, it has been found that it is sometimes convenient to refer to these operational arrangements as modules. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.
[0402] Any steps, operations, or processes described herein may be performed or implemented using one or more hardware or software modules, alone or in combination with other devices. In some embodiments, the software modules are implemented by a computer program product comprising a non-transitory computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the steps, operations, or processes described.
[0403] Embodiments may also relate to products produced by the computing processes described herein. Such products may include information produced by the computing processes, wherein the information is stored on a tangible, non-transitory computer-readable storage medium, and may include any embodiment of a computer program product or other data combination described herein.
[0404] Finally, the language used in this specification is selected primarily for readability and instructional purposes, and is not selected to define or limit the subject matter of the present invention. Accordingly, it is intended that the scope of the present invention be limited not by this detailed description, but by any claims filed on an application based hereon. Accordingly, the embodiments disclosed herein are intended to illustrate, not to limit, the scope of the present invention, which is defined by the appended claims.
Claims
1. A method, characterized in that: The method comprises: obtaining a plurality of training samples from an individual, each training sample having one of a plurality of disease states and each training sample being associated with one of a plurality of labels of a covariate characteristic, wherein each training sample comprises at least 1,000 methylated sequence reads of cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; generating a first training sample set and a second training sample set by subdividing the plurality of training samples; For each training sample, a feature vector is generated based on the methylation sequence reads of the training sample in the following way: For each methylated sequence read in the training sample: applying each of a plurality of tissue models to the methylation sequence reads to determine a likelihood that the methylation sequence read indicates the presence of a disease state associated with the tissue model, assigning the methylated sequence read to one of the disease states having the highest likelihood output by the tissue models, and determining the feature vector based on the methylated sequence reads assigned to each disease state; Training a classifier using the feature vector of the first training sample set to generate a signal vector based on the input feature vector, wherein the signal vector includes a value for each disease state; Applying the cancer classifier to the feature vector of each training sample in the second training sample set to generate a cancer signature vector for each training sample in the second training sample set; and For each label in the plurality of labels of the covariate characteristic, a cutoff threshold for each disease state is determined based on the signal vector of the training sample with the label in the second training sample set.
2. The method according to claim 1, wherein: These methylation sequence reads were obtained using targeted methylation sequencing assays or whole-genome bisulfite sequencing assays.
3. The method according to any one of claims 1 to 2, characterized in that: Each training sample included at least 10,000 methylation sequence reads of cfDNA fragments.
4. The method according to any one of claims 1 to 3, characterized in that: The first training sample set and the second training sample set include similar proportions of training samples across the disease states.
5. The method according to any one of claims 1 to 4, characterized in that: The plurality of disease states includes a non-cancerous state and one or more cancerous states of one or more cancers of different origin.
6. The method according to claim 5, wherein: The one or more cancers of different origin include: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial cancer of the renal pelvis and ureter, renal cancer other than urothelial cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal squamous cell carcinoma, esophageal cancer other than squamous, gastric cancer, hepatobiliary cancer arising from hepatocellular cells, hepatobiliary cancer arising from cells other than hepatocellular cells, pancreatic cancer, head and neck cancer associated with human papillomavirus, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung cancer, lung squamous cell carcinoma, and lung cancer other than adenocarcinoma or small cell lung cancer, neuroendocrine cancer, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma and leukemia; in some embodiments, the cancer type is additionally selected from the group consisting of: brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma and gallbladder cancer.
7. The method according to any one of claims 1 to 6, characterized in that: The covariate characteristic is one of the following: age, smoker status, and biological sex.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: For each methylation sequence read, p-value screening is used to determine whether the methylation sequence read has an informative methylation pattern, wherein the feature vector of each training sample is generated based on the methylation sequence reads with informative methylation patterns.
9. The method according to any one of claims 1 to 8, characterized in that: Each of the multiple tissue models is trained in the following manner: obtaining a second plurality of training samples, each training sample having one of a plurality of disease states, wherein each training sample comprises methylated sequence reads of cfDNA fragments obtained from the individual; For each tissue model, generating a training dataset comprising at least 10,000 methylated sequence reads of training samples having a disease state associated with the tissue model; and Each tissue model is trained using the associated training dataset to predict the likelihood that a methylated sequence read indicates the presence of the associated disease state.
10. The method according to any one of claims 1 to 9, characterized in that: Each tissue model was one of the following: binomial, independent site, Markov, or mixture models.
11. The method according to any one of claims 1 to 9, characterized in that: The classifier is a machine learning model.
12. The method according to any one of claims 1 to 11, characterized in that: The method further comprises: obtaining a test sample from a test individual having an unknown disease state and associated with a first signature in the plurality of signatures characteristic of the covariate, wherein the test sample comprises methylated sequence reads of cfDNA fragments obtained from the test individual; A test feature vector is generated based on the methylation sequence reads of the test sample in the following way: For each methylated sequence read of the test sample: applying each of the plurality of tissue models to the methylation sequence reads to determine a likelihood that the methylation sequence read indicates the presence of a disease state associated with the tissue model, assigning the methylated sequence read to one of the disease states having the highest likelihood output by the tissue models, and determining the test feature vector based on the methylated sequence reads assigned to each disease state; applying the classifier to the test feature vector of the test to generate a signal vector for the test sample; and A positive disease signal for one or more of the disease states is detected by applying a cutoff threshold associated with the first label of the covariate characteristic to the signal vector of the test sample.
13. The method according to claim 12, wherein: Positive disease signatures for detecting one or more of these disease states include: For each disease state, a cutoff threshold for the disease state is applied to the values in the signal vector of the test sample corresponding to the disease state.
14. The method according to any one of claims 1 to 13, Its characteristics are: wherein each individual is further associated with one of a second plurality of labels characteristic of a second covariate; wherein determining the cutoff thresholds comprises determining for each combination of a label from the plurality of labels of the covariate characteristic and a label from the second plurality of labels of the second covariate characteristic; wherein the test sample is further associated with a second signature of the second covariate characteristic; and Wherein detecting a positive disease signal for one or more of the disease states comprises applying a cutoff threshold associated with a combination of the first signature of the covariate characteristic and the second signature of the second covariate characteristic.
15. The method according to any one of claims 12 to 14, characterized in that: The method further comprises Report the positive disease signal to the healthcare provider for additional diagnostic steps.
16. A method, characterized in that: The method comprises: obtaining a plurality of training samples from an individual, each training sample having one of a plurality of disease states and each training sample being associated with one of a plurality of labels of a covariate characteristic, wherein each training sample comprises at least 1,000 methylated sequence reads of cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; For each training sample, a feature vector is generated based on the methylation sequence reads of the training sample in the following way: For each methylated sequence read in the training sample: applying each of a plurality of tissue models to the methylation sequence reads to determine a likelihood that the methylation sequence read indicates the presence of a disease state associated with the tissue model, assigning the methylated sequence read to one of the disease states having the highest likelihood output by the tissue models, and determining the feature vector based on the methylated sequence reads assigned to each disease state; generating a training sample set for each of the plurality of labels of the covariate characteristic, comprising a feature vector of the training sample associated with the label of the covariate characteristic; For each label, a classifier is trained using the feature vector of the training sample set corresponding to the label to generate a signal vector based on the input feature vector, wherein the signal vector includes a value for each disease state.
17. The method according to claim 16, wherein: These methylation sequence reads were obtained using targeted methylation sequencing assays or whole-genome bisulfite sequencing assays.
18. The method according to any one of claims 16 to 17, characterized in that: Each training sample included at least 10,000 methylation sequence reads of cfDNA fragments.
19. The method according to any one of claims 16 to 18, characterized in that: The plurality of disease states includes a non-cancerous state and one or more cancerous states of one or more cancers of different origin.
20. The method of claim 19, wherein: The one or more cancers of different origin include: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial cancer of the renal pelvis and ureter, renal cancer other than urothelial cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal squamous cell carcinoma, esophageal cancer other than squamous, gastric cancer, hepatobiliary cancer arising from hepatocellular cells, hepatobiliary cancer arising from cells other than hepatocellular cells, pancreatic cancer, head and neck cancer associated with human papillomavirus, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung cancer, lung squamous cell carcinoma, and lung cancer other than adenocarcinoma or small cell lung cancer, neuroendocrine cancer, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma and leukemia; in some embodiments, the cancer type is additionally selected from the group consisting of: brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma and gallbladder cancer.
21. The method according to any one of claims 16 to 20, characterized in that: The covariate characteristic is one of the following: age, smoker status, and biological sex.
22. The method according to any one of claims 16 to 21, characterized in that: The method further comprises: For each methylation sequence read, p-value screening is used to determine whether the methylation sequence read has an informative methylation pattern, wherein the feature vector of each training sample is generated based on the methylation sequence reads with an informative methylation pattern.
23. The method according to any one of claims 16 to 22, characterized in that: Each of the multiple tissue models is trained in the following manner: obtaining a second plurality of training samples, each training sample having one of a plurality of disease states, wherein each training sample comprises methylated sequence reads of cfDNA fragments obtained from the individual; For each tissue model, generating a training dataset comprising at least 10,000 methylated sequence reads of training samples having a disease state associated with the tissue model; and Each tissue model is trained using the associated training dataset to predict the likelihood that a methylated sequence read indicates the presence of the associated disease state.
24. The method according to any one of claims 16 to 23, characterized in that: Each tissue model was one of the following: binomial, independent site, Markov, or mixture models.
25. The method according to any one of claims 16 to 24, characterized in that: Each classifier is a machine learning model.
26. The method according to any one of claims 16 to 25, characterized in that: The method further comprises: obtaining a test sample from a test individual having an unknown disease state and associated with a first signature in the plurality of signatures characteristic of the covariate, wherein the test sample comprises methylated sequence reads of cfDNA fragments obtained from the test individual; A test feature vector is generated based on the methylation sequence reads of the test sample in the following way: For each methylated sequence read of the test sample: applying each of the plurality of tissue models to the methylation sequence reads to determine a likelihood that the methylation sequence read indicates the presence of a disease state associated with the tissue model, assigning the methylated sequence read to one of the disease states having the highest likelihood output by the tissue models, and determining the test feature vector based on the methylated sequence reads assigned to each disease state; applying a classifier corresponding to the first label to the test feature vector of the test to generate a signal vector of the test sample; and A positive disease signature for one or more of the disease states is detected based on the signal vector of the test sample.
27. The method of claim 26, wherein: Positive disease signatures for detecting one or more of these disease states include: The disease state with the largest value in the signal vector is identified as the disease state with a detected positive disease signal.
28. The method according to any one of claims 16 to 27, Its characteristics are: wherein each individual is further associated with one of a second plurality of labels characteristic of a second covariate; Wherein, generating a plurality of training sample sets comprises: generating a training sample set for each combination of one label from the plurality of labels of the covariate characteristic and one label from the second plurality of labels of the second covariate characteristic; wherein training the classifiers comprises training a classifier for each combination using a feature vector from a training sample set corresponding to the combination; wherein the test sample is further associated with a second signature of the second covariate characteristic; and Wherein, applying the classifier includes applying a classifier trained for a combination of the first label and the second label.
29. The method according to any one of claims 12 to 14, characterized in that: The method further comprises Report the positive disease signal to the healthcare provider for additional diagnostic steps.
30. A method comprising: The method comprises: obtaining a plurality of training samples from an individual, each training sample having one of a plurality of disease states and each training sample being associated with one of a plurality of labels of a covariate characteristic, wherein each training sample comprises at least 1,000 methylated sequence reads of cell-free deoxyribonucleic acid (cfDNA) fragments obtained from the individual; For each training sample, a feature vector is generated based on the methylation sequence reads of the training sample in the following way: determining, for each of a plurality of genomic regions, a methylation feature value based on one or more methylated sequence reads of training samples that overlap with the genomic region, and Determining the feature vector based on the methylation feature values of the multiple genomic regions; generating a training sample set for each of the plurality of labels of the covariate characteristic, comprising a feature vector of the training sample associated with the label of the covariate characteristic; For each tag: determining a mutual information score for each genomic region based on a feature vector of a training sample set corresponding to the label; These genomic regions are ranked based on the mutual information score; selecting a feature set from the ranked genomic regions; modifying the feature vectors to include methylation feature values for the feature set; and A classifier is trained using the modified feature vector to generate a signal vector based on the input feature vector, wherein the signal vector includes a value for each disease state.
31. The method of claim 30, wherein: These methylation sequence reads were obtained using targeted methylation sequencing assays or whole-genome bisulfite sequencing assays.
32. The method according to any one of claims 30 to 31, characterized in that: Each training sample included at least 10,000 methylation sequence reads of cfDNA fragments.
33. The method according to any one of claims 30 to 32, characterized in that: Each genomic region includes one or more CpG sites.
34. The method according to any one of claims 30 to 33, wherein: The methylation feature value is the methylation density of the genomic region.
35. The method according to any one of claims 30 to 33, characterized in that: The method further comprises: For each methylation sequence read, p-value screening is used to determine whether the methylation sequence read has an informative methylation pattern, wherein the feature vector of each training sample is generated based on the methylation sequence reads with an informative methylation pattern.
36. The method of claim 35, wherein: The methylation feature value is the count of methylated sequence reads overlapping the genomic region.
37. The method according to any one of claims 30 to 36, characterized in that: The plurality of disease states includes a non-cancerous state and one or more cancerous states of one or more cancers of different origin.
38. The method of claim 37, wherein: The one or more cancers of different origin include: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urothelial cancer of the renal pelvis and ureter, renal cancer other than urothelial cancer, prostate cancer, anorectal cancer, colorectal cancer, esophageal squamous cell carcinoma, esophageal cancer other than squamous, gastric cancer, hepatobiliary cancer arising from hepatocellular cells, hepatobiliary cancer arising from cells other than hepatocellular cells, pancreatic cancer, head and neck cancer associated with human papillomavirus, head and neck cancer not associated with human papillomavirus, lung adenocarcinoma, small cell lung cancer, lung squamous cell carcinoma, and lung cancer other than adenocarcinoma or small cell lung cancer, neuroendocrine cancer, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma and leukemia; in some embodiments, the cancer type is additionally selected from the group consisting of: brain cancer, vulvar cancer, vaginal cancer, testicular cancer, pleural mesothelioma, peritoneal mesothelioma and gallbladder cancer.
39. The method according to any one of claims 30 to 37, wherein: The covariate characteristic is one of the following: age, smoker status, and biological sex.
40. The method according to any one of claims 30 to 38, wherein: The mutual information score for each genomic region is determined for each label by: determining a pairwise information gain for each pairwise combination of disease states based on the discriminatory ability of the genomic region to classify the two disease states; and The mutual information score is determined by combining the pairwise information gains of these pairwise disease state combinations.
41. The method according to any one of claims 30 to 39, wherein: The method further comprises: For each tag: Determining the coverage of each genomic region based on the sequencing depth of the training sample set corresponding to the label; and Based on the coverage being below a threshold, one or more genomic regions are excluded from selection for the feature set.
42. The method according to any one of claims 30 to 40, characterized in that: The method further comprises: For each tag: determining, for each genomic region, an activation score for the non-cancer disease state, the activation score indicating the number of training samples with the non-cancer disease state that exhibit activation of the genomic region; and One or more genomic regions are excluded from selection for the feature set based on an activation score above a threshold.
43. The method according to any one of claims 30 to 41, characterized in that: Determine the set of features selected for each label to optimize the precision of the classifier.
44. The method according to any one of claims 30 to 42, wherein: One feature set for one label is different from another feature set for another label.
45. The method of claim 43, wherein: The one feature set includes one or more features that are different from the other feature set.
46. The method of claim 43, wherein: The one feature set includes a different number of features than the other feature set.
47. The method according to any one of claims 30 to 45, characterized in that: Each classifier is a machine learning model.
48. The method according to any one of claims 16 to 25, wherein: The method further comprises: obtaining a test sample from a test individual having an unknown disease state and associated with a first signature in the plurality of signatures characteristic of the covariate, wherein the test sample comprises methylated sequence reads of cfDNA fragments obtained from the test individual; A test feature vector is generated based on the methylation sequence reads of the test sample in the following way: determining, for each of the plurality of genomic regions, a methylation signature value based on one or more methylated sequence reads of the test sample that overlap with the genomic region, and Determining the test feature vector based on the methylation feature values of the multiple genomic regions; applying a classifier corresponding to the first label to the test feature vector of the test to generate a signal vector of the test sample; and A positive disease signature for one or more of the disease states is detected based on the signal vector of the test sample.
49. The method of claim 47, wherein: Positive disease signatures for detecting one or more of these disease states include: The disease state with the largest value in the signal vector is identified as the disease state with a detected positive disease signal.
50. The method according to any one of claims 12 to 14, wherein: The method further comprises Report the positive disease signal to the healthcare provider for additional diagnostic steps.
51. A method for optimizing feature selection, characterized in that: The method comprises: obtaining a plurality of sequence reads from the sample via an analysis system; generating, via the analysis system, a plurality of features based on the plurality of sequence reads; generating, via the analysis system, one or more of an activation score, a mutual information score, or a coverage based on the plurality of sequence reads; determining, via the analysis system, a ranking of the plurality of features based on one or more activation scores, mutual information scores, or coverages; and A subset of features is selected from the plurality of features based on the ranking of the plurality of features, via the analysis system.
52. The method of claim 50, wherein: The sample includes a free nucleic acid sample.
53. The method according to any one of claims 50 to 51, characterized in that: The plurality of features includes sequence read counts of sequence reads exceeding a ratio threshold.
54. The method according to any one of claims 50 to 52, wherein: Generating the plurality of features includes determining methylation rates of a plurality of CpG sites in the plurality of sequence reads.
55. The method according to any one of claims 50 to 53, wherein: The activation score includes a non-cancer activation score.
56. The method according to any one of claims 50 to 54, wherein: A mutual information score among the mutual information scores corresponding to a feature among the plurality of features is inversely proportional to the rank of the feature.
57. The method according to any one of claims 50 to 55, wherein: Lower coverage corresponds to noisier features.
58. The method according to any one of claims 50 to 56, wherein: The method further includes training at least one classifier based on the subset of features, the at least one classifier being trained to predict the presence or absence of a disease, a type of disease, and / or a tissue of origin of a disease.
59. The method of claim 57, wherein: The at least one classifier is trained to determine the disease tissue of origin based on one or more characteristics.
60. The method of claim 58, wherein: The one or more characteristics include at least one of age, smoker status, and gender.
61. The method according to any one of claims 50 to 59, wherein: The method further comprises: determining a first activation score for positive prediction of disease; and Determine a second activation score for negative prediction of disease; Wherein the feature subset is selected at least in part by minimizing a difference between the first activation score and the second activation score.
62. The method according to any one of claims 50 to 60, wherein: The method further comprises: generating a training data set corresponding to the feature subset, where the training data is generated from sequence data labeled with disease type or tissue of origin or labeled as healthy; A machine learning model is trained using the generated training dataset, where the machine learning model is configured to predict disease status and type based on sequence data corresponding to the feature subset.
63. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores instructions that, when executed by a computer processor, cause the processor to perform the method of any one of claims 1 to 61.
64. A system, characterized in that: The system comprises: computer processors; and The non-transitory computer-readable storage medium of claim 62.
65. A therapeutic kit, characterized in that: The therapeutic kit comprises: A collection container for storing a biological sample including cfDNA fragments; Optionally, one or more reagents for isolating cfDNA fragments from the biological sample; Optionally, a sequencing panel comprising probes for targeting specific genomic regions; and The non-transitory computer-readable storage medium of claim 62.
Citation Information
Patent Citations
Anomalous fragment detection and classification
US12027237B2
Anomalous fragment detection and classification
US20190287652A1
Selective oxidation of 5-methylcytosine by TET-family proteins
WO2010037001A2
Composition and methods related to modification of 5-hydroxymethylcytosine (5-HMC)
WO2011127136A1
Methylation markers and targeted methylation probe panels
WO2019195268A2