Parallel cancer origin classification of organ type and oncobiological type
By training the CSO classifier in parallel, organ type and tumor biological type are predicted separately, which solves the problems of low detection rate and false positive rate in existing cancer detection, realizes early cancer detection and more refined cancer origin prediction, and supports more effective diagnosis and treatment options.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GRAIL INC
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-12
AI Technical Summary
Existing cancer detection methods are often specific to certain cancers, with low detection rates and high false positive rates, making it difficult to detect cancer in its early stages. Furthermore, they lack granular prediction of cancer origins, limiting the effectiveness of diagnosis and treatment.
A parallel-trained CSO classifier is used to predict the organ types and tumor biological types affected by cancer, respectively. Independent organ groups and tumor biological classifiers are trained using methylation sequencing data to avoid confusion, improve detection granularity, and avoid the need for multiple sequencing runs through parallel training.
It improves the accuracy and granularity of cancer detection, provides more detailed information on the origin of cancer, and helps healthcare professionals choose appropriate diagnostic and treatment options.
Smart Images

Figure CN122029607A_ABST
Abstract
Description
Background Technology
[0001] Cancer is a leading cause of death worldwide. Cancer mortality is exacerbated by the fact that cancer is often detected at a late stage, limiting the effectiveness of treatment options for long-term survival. Current detection methods are typically cancer-specific, meaning they screen for each type of cancer individually (breast cancer, lung cancer, colorectal cancer, prostate cancer, etc.). Therefore, each screening process is tailored to a specific cancer. For example, mammography is used to detect breast cancer, while colonoscopy or stool analysis helps detect colorectal cancer. Each different screening method generally does not cross-reference to other cancers. Furthermore, this screening method is hampered by low detection rates or high false-positive rates. Low detection rates often fail to detect early-stage cancers because they are just beginning to develop. High false-positive rates misdiagnose subjects without cancer as having a positive cancer status. Therefore, most screening tests are only practical when used to test subjects who are at high risk of developing the cancer being screened for or who have symptoms indicating the presence of suspected cancer. Consequently, most screening tests have limited ability to detect cancer in the general population.
[0002] Novel research indicates that aberrant DNA methylation occurs in many disease processes, including cancer. DNA methylation plays a role in regulating gene expression and defining tissue differentiation, cell identity, and / or embryonic lineage. Therefore, aberrant DNA methylation can cause problems in normal gene expression pathways or cell identity, leading to cancer or other diseases. For example, specific patterns of differentially methylated regions can be used as molecular markers for various disease states. The detection of these differentially methylated regions can be accomplished through sequencing analysis of cell-free DNA molecules. Typically, cell-free DNA molecules are DNA molecules produced in bodily fluids. These DNA molecules are typically released due to natural cell death, active release from healthy cells, or tumor-derived DNA molecules shed from tumor cells undergoing cell death. Nevertheless, even techniques for detecting differentially methylated regions face many challenges. Early cancer detection is particularly challenging due to the tiny ratio of tumor cells to non-cancer cells in a subject. The smallest ratios can be on the order of 1:1000, 1:10,000, or even 1:100,000. This presents a challenge in detecting small amounts of cancer “signals” amidst the “noise” of otherwise healthy conditions, especially when analyzing this signal with easily accessible sample conditions, such as blood draws to assess the presence of cancer signals in plasma (e.g., cell-free DNA).
[0003] Further challenges may arise when providing insights into cancers detected in subjects. For example, a multi-cancer detection test may only provide a binary prediction of whether a subject has cancer. This insight may limit healthcare professionals' ability to continue diagnosing cancer and / or treating the subject. Diagnostic testing and treatment options are often tailored to the specific organ groups affected and the biology of the tumor. Therefore, there is a need to increase the granularity of analytical predictions to better inform healthcare professionals about diagnostic testing options.
[0004] This disclosure addresses the challenges described above. The background description provided herein is intended to provide an overall overview of the information presented in this disclosure. Unless otherwise stated herein, the materials described in this section are not prior art to the claims of this application and will not be deemed prior art or a teaching of prior art by virtue of their inclusion in this section. Summary of the Invention
[0005] The invention disclosed herein provides improvements to cancer detection, diagnosis, and treatment, particularly in providing granularity for cancer origin (CSO) prediction. The invention covers training parallel CSO classifiers to predict, separately, the organ type and tumor biology type affected by cancer. Parallel training of the CSO classifiers divides CSO prediction analysis to avoid confounding organ type and tumor biology type predictions when only the origin of the cancer signal is typically predicted. In some instances, the CSO classifiers are trained using training datasets derived from the same training sample set. Training datasets are generated for each CSO classifier based, for example, methylation sequencing data for each sample and known CSO tags for the sample, representing the clinical truth currently obtained only after diagnostic examinations and cancer diagnoses. Training of the CSO classifiers can be parallelized, allowing each classifier to learn patterns in the methylation sequencing data separately and independently to distinguish between organ type and tumor biology type. Therefore, parallel-trained CSO classifiers provide additional granularity in CSO prediction, thus better informing examination steps that generate a cancer diagnosis after screening or early detection has already detected cancer signals, for example, in plasma. Furthermore, training a separate classifier with the same base data avoids the need for multiple sequencing measurements, thus improving the measurement process.
[0006] Clause 1. A method for training an independent parallel cancer signal origin (CSO) classifier, the method comprising: obtaining training samples derived from subjects having known cancer diagnoses, each training sample containing methylated sequence reads corresponding to nucleic acid fragments collected from biological samples from each subject, and each known cancer diagnosis including known organs or organ groups among a plurality of affected organs or organoids and known tumor biology among a plurality of tumor biology categories; generating feature vectors based on the methylated sequence reads for each training sample; generating a first training dataset containing the feature vectors of the training samples and the known organs or organoids; training an organ or organoid classifier with the first training dataset to predict organs or organoids from the plurality of organs or organoids based on input feature vectors; generating a second training dataset containing the feature vectors of the training samples and the known tumor biology categories; and training a tumor biology classifier with the second training dataset to predict tumor biology from the plurality of tumor biology categories based on input feature vectors.
[0007] Clause 2. The method as described in any of the preceding clauses further comprises: for each training sample, extracting the known organ or organ group and the known tumor biological category from the known cancer diagnosis and clinical information.
[0008] Clause 3. The method as described in any of the preceding clauses, wherein the feature vector is based at least in part on the methylation features of these methylated sequence reads.
[0009] Clause 4. The method as described in Clause 3, wherein such methylation features include: methylation density at one or more loci, density of hypermethylated sequence reads at one or more loci, density of hypomethylated sequence reads at one or more loci, count of methylated sequence reads identified as aberrantly methylated at one or more loci, or some combination thereof.
[0010] Clause 5. The method described in any of the preceding clauses, wherein generating the first training dataset includes excluding information about tumor biology.
[0011] Clause 6. The method described in any of the preceding clauses, wherein generating the second training dataset includes excluding information about the affected organ or group of organs.
[0012] Clause 7. The method as described in any of the preceding clauses further comprises: for each feature, determining an information gain in distinguishing these organs or organ groups; identifying discriminative features of the organ or organ group classifier based on these information gains; and modifying the feature vector of the first training set to consist of these discriminative features, wherein the modified feature vector is used to train the organ or organ group classifier.
[0013] Clause 8. The method as described in any of the preceding clauses further comprises: determining, for each feature, an information gain in distinguishing these tumor biological categories; identifying discriminative features of the tumor biological classifier based on these information gains; and modifying the feature vector of the second training set to consist of these discriminative features, wherein the modified feature vector is used to train the tumor biological classifier.
[0014] Clause 9. The method described in any of the preceding clauses, wherein the organ or organ group classifier or the tumor biology classifier is a machine learning model.
[0015] Clause 10. The method described in any of the preceding clauses further includes training the organ or organ group classifier and the tumor biology classifier in a parallel training process.
[0016] Clause 11. The method described in any of the preceding clauses further includes training the organ or organ group classifier prior to training the tumor biological classifier.
[0017] Clause 12. The method as described in Clause 11, wherein the output of the organ or organ group classifier is appended to the feature vector of the second training dataset before training the tumor biology classifier.
[0018] Clause 13. The method described in any of the preceding clauses further includes training the tumor biology classifier prior to training the organ or organ group classifier.
[0019] Clause 14. The method as described in Clause 13, wherein the output of the tumor biology classifier is appended to the feature vector of the first training dataset before training the organ or organ group classifier.
[0020] Clause 15. The method described in any of the preceding clauses, wherein these organs or groups of organs include: breast; prostate; lung; head or neck; anus; cervix; ovary or fallopian tube; uterus; bladder or urothelial lining; kidney; stomach or esophagus; liver or intrahepatic bile duct; pancreas, extrahepatic bile duct or gallbladder; colon or rectum; bone or soft tissue; skin; blood, lymphatic system or bone marrow; thyroid gland; obscure tissue; or some combination thereof.
[0021] Clause 16. As described in any of the preceding clauses, wherein these tumor biological categories include: lymphoid vegetations, medullary vegetations, plasma cell vegetations, neuroendocrine carcinomas or tumors, adenocarcinomas, squamous cell carcinomas that are not human papillomavirus-associated (HPV-associated), HPV-associated carcinomas, hepatocellular carcinomas, vegetations originating from Müllerian ducts, transitional cell carcinomas, mesenchymal tumors, melanocyte vegetations, mesothelial vegetations, other tumor biological categories, obscure tumor biological categories, or combinations thereof.
[0022] Clause 17. A method for predicting cancer signal origin (CSO), the method comprising: obtaining a test sample derived from a subject, the test sample containing methylated sequence reads corresponding to nucleic acid fragments collected from the subject; generating, for the test sample, a first feature vector based on methylated sequence reads associated with a first set of features identified as discriminative for organ or organoid classification; generating, for the test sample, a second feature vector based on methylated sequence reads associated with a second set of features identified as discriminative for tumor biological classification; and applying an organ or organoid classifier to these first feature vectors to predict, from multiple organs or organoids, the organoid associated with cancer in the test sample. The organ or organ group; a tumor biology classifier is applied to the second feature vector to predict the tumor biology of the cancer associated with the test sample from multiple tumor biology categories; wherein the organ or organ group classifier and the tumor biology classifier are independently trained on training samples derived from subjects with known cancer diagnoses, which include known organs or organ groups among multiple affected organs or organ groups and known tumor biology in multiple tumor biology categories, each training sample containing methylated sequence reads corresponding to nucleic acid fragments in biological samples collected from each subject; and information is provided for diagnostic testing to diagnose cancer based on the predicted organs or organ groups and the predicted tumor biology.
[0023] Clause 18. The method of Clause 17, wherein the organ or organ group classifier and the tumor biology classifier are trained by: generating a feature vector based on a methylated sequence read of the training sample for each training sample; generating a first training dataset containing the feature vectors of the training samples and the known organ or organ group of the known cancer diagnosis; training the organ or organ group classifier with the first training dataset to predict organs or organ groups from the plurality of organs or organ groups based on the input feature vectors; generating a second training dataset containing the feature vectors of the training samples and the known tumor biology category of the known cancer diagnosis; and training the tumor biology classifier with the second training dataset to predict tumor biology from the plurality of tumor biology categories based on the input feature vectors.
[0024] Clause 19. The method of any one of Clauses 17 to 18, wherein generating the first training dataset includes excluding information about tumor biology, and wherein generating the second training dataset includes excluding information about the affected organ or group of organs.
[0025] Clause 20. The method of any one of Clauses 17 to 19, further comprising: for each feature, determining an information gain in distinguishing these organs or groups of organs; identifying discriminative features of the organ or group of organs classifier based on these information gains; and modifying the feature vector of the first training set to consist of these discriminative features, wherein the modified feature vector is used to train the organ or group of organs classifier.
[0026] Clause 21. The method of any one of Clauses 17 to 20, further comprising: for each feature, determining an information gain in distinguishing these tumor biological categories; identifying discriminative features of the tumor biological classifier based on these information gains; and modifying the feature vector of the second training set to consist of these discriminative features, wherein the modified feature vector is used to train the tumor biological classifier.
[0027] Clause 22. The method of any one of Clauses 17 to 21, wherein the organ or organ group classifier or the tumor biology classifier is a machine learning model.
[0028] Clause 23. The method of any one of Clauses 17 to 22, further comprising training the organ or organ group classifier and the tumor biology classifier in parallel training.
[0029] Clause 24. The method of any one of Clauses 17 to 23 further includes training the organ or organ group classifier prior to training the tumor biological classifier.
[0030] Clause 25. The method of any one of Clauses 17 to 24, further comprising training the tumor biology classifier prior to training the organ or organ group classifier.
[0031] Clause 26. The method of any one of Clauses 17 to 25, further comprising: modifying the feature vector according to the discriminative features of the organ or organ group classifier to generate a first simplified feature vector before applying the organ or organ group classifier, such that the organ or organ group classifier is applied to the first simplified feature vector; and modifying the feature vector according to the discriminative features of the tumor biology classifier to generate a second simplified feature vector before applying the tumor biology classifier, such that the tumor biology classifier is applied to the second simplified feature vector.
[0032] Clause 27. The method of any one of Clauses 17 to 26, wherein providing information for diagnostic examinations of detected cancer signals includes identifying one or more diagnostic examination options based on the predicted organ or organ group, the predicted tumor biology, or some combination thereof.
[0033] Clause 28. The method as described in any one of Clauses 17 to 27, wherein these organs or groups of organs include: breast; prostate; lung; head or neck; anus; cervix; ovary or fallopian tube; uterus; bladder or urothelial lining; kidney; stomach or esophagus; liver or intrahepatic bile duct; pancreas, extrahepatic bile duct or gallbladder; colon or rectum; bone or soft tissue; skin; blood, lymphatic system or bone marrow; thyroid gland; indistinct tissue; or some combination thereof.
[0034] Clause 29. The method as described in any one of Clauses 17 to 28, wherein these tumor biological categories include: lymphoid vegetations, medullary vegetations, plasma cell vegetations, neuroendocrine carcinomas or tumors, adenocarcinomas, squamous cell carcinomas that are not human papillomavirus-associated (HPV-associated), HPV-associated carcinomas, hepatocellular carcinomas, vegetations originating from Müllerian ducts, transitional cell carcinomas, mesenchymal tumors, melanocyte vegetations, mesothelial vegetations, other tumor biological categories, obscure tumor biological categories, or combinations thereof.
[0035] Clause 30. The method of any one of Clauses 17 to 29, wherein the subject has previously been diagnosed with cancer of unknown origin, wherein providing information for the diagnostic examination includes providing information for the diagnostic examination to improve the diagnosis based on the predicted organ or organ group and the predicted tumor biology.
[0036] Clause 31. The method of any one of Clauses 17 to 30, wherein providing information for the diagnostic examination comprises: providing a report containing cancer signal detection readouts for the test sample, prediction of cancer signal origin, predicted organ or organ group, and predicted tumor biology.
[0037] Clause 32. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method as described in any one of Clauses 1-31.
[0038] Clause 33. A system comprising: one or more computer processors; and a non-transitory computer-readable storage medium as described in Clause 32.
[0039] Clause 34. A treatment kit comprising: a collection container for collecting a DNA sample from a subject; optionally, one or more reagents for isolating DNA fragments from the DNA sample; optionally, one or more probes targeting one or more loci identified as indicative of a cancer state; and a nontransitory computer-readable storage medium as described in Clause 32.
[0040] Clause 35. A method for providing a report of a test sample to a patient to assist in the patient's diagnostic examination, the report comprising cancer signal detection readout and cancer signal origin (CSO) prediction, the CSO prediction including a predicted organ or organoid group of the CSO and a predicted tumor biology of the CSO, the method comprising: obtaining a test sample derived from the patient, the test sample comprising methylated sequence readouts corresponding to nucleic acid fragments collected from the patient; generating a first feature vector for the test sample based on methylation information selected to provide information on cancer signals associated with the test sample; and generating a first feature vector for the test sample based on methylation information selected to provide information on cancer signals associated with the test sample. A second feature vector is generated based on methylated sequence reads associated with a first feature set identified as discriminative for organ or organoid group classification. For the test sample, a third feature vector is generated based on methylated sequence reads associated with the second feature set identified as discriminative for tumor biological classification. A cancer signal classifier is applied to the first feature vector to predict cancer signals associated with the test sample. An organ or organoid group classifier is applied to the second feature vector to predict the organ or organoid group associated with the cancer from multiple organs or organoid groups. A tumor biological classifier is applied to the third feature vector to... The tumor biology classifier predicts the tumor biology associated with a test sample from multiple tumor biology categories; wherein the cancer signal classifier is trained on training samples derived from multiple cancer-positive and cancer-negative subjects, each cancer-positive subject having a labeled cancer diagnosis and each cancer-negative subject known not to have cancer, and each training sample contains methylated sequence reads corresponding to nucleic acid fragments collected from biological samples from each subject; wherein the organ or organoid classifier and the tumor biology classifier are trained independently on training samples derived from subjects with known cancer diagnoses, which include affected multiple Known organs or organ groups in an organ or organoid group and known tumor biology in multiple tumor biology categories, each training sample containing methylated sequence reads corresponding to nucleic acid fragments in biological samples collected from each subject; a report of the test sample is generated based on the results from the cancer signal classifier, the organ or organ type classifier and the tumor biology classifier, the report containing the cancer signal detection readout and the cancer signal origin (CSO) prediction, the CSO prediction including the predicted organ or organoid group of the CSO and the predicted tumor biology of the CSO; and the report is provided to the patient or the patient's healthcare provider.
[0041] Clause 36. A method for providing a report of a test sample to a patient to assist in the patient's diagnostic examination, the report comprising cancer signal detection readout and cancer signal origin (CSO) prediction, the CSO prediction comprising a predicted organ or organoid group of the CSO and a predicted tumor biology of the CSO, wherein the CSO prediction is determined by: obtaining a test sample derived from the patient, the test sample comprising methylated sequence readouts corresponding to nucleic acid fragments collected from the patient; and generating a first feature vector for the test sample based on methylation information selected to provide information on cancer signals associated with the test sample; For the test sample, a second feature vector is generated based on methylated sequence reads associated with a first feature set identified as discriminative for organ or organoid group classification; for the test sample, a third feature vector is generated based on methylated sequence reads associated with a second feature set identified as discriminative for tumor biological classification; a cancer signal classifier is applied to the first feature vector to predict cancer signals associated with the test sample; an organ or organoid group classifier is applied to the second feature vector to predict the organ or organoid group associated with the cancer from multiple organs or organoid groups; and a tumor biological classifier is applied to the third feature vector. Vectors are used to predict the tumor biology associated with the test sample from multiple tumor biology categories; wherein the cancer signal classifier is trained on training samples derived from multiple cancer-positive and cancer-negative subjects, each cancer-positive subject having a labeled cancer diagnosis and each cancer-negative subject known not to have cancer, and each training sample contains methylated sequence reads corresponding to nucleic acid fragments collected from biological samples from each subject; wherein the organ or organoid classifier and the tumor biology classifier are trained independently on training samples derived from subjects with known cancer diagnoses, including those affected by cancer. The test sample contains known organs or organ groups from multiple organs or organoids and known tumor biology from multiple tumor biology categories. Each training sample contains methylated sequence reads corresponding to nucleic acid fragments collected from biological samples from each subject. A report on the test sample is generated based on the results from the cancer signal classifier, the organ or organ type classifier, and the tumor biology classifier. The report contains the cancer signal detection readout and the cancer signal origin (CSO) prediction, which includes the predicted organ or organoid and the predicted tumor biology of the CSO. The report is then provided to the patient or the patient's healthcare provider. Attached Figure Description
[0042] Figure 1 This is an exemplary flowchart illustrating the overall workflow for cancer classification of a sample according to one or more embodiments.
[0043] Figure 2A This is an exemplary flowchart illustrating the process of sequencing fragments of free (cf) DNA to obtain a methylation state vector according to one or more embodiments.
[0044] Figure 2B According to one or more embodiments Figure 2A The illustrated process involves sequencing fragments of cell-free (cf) DNA to obtain a methylation state vector.
[0045] Figure 3A This is an exemplary flowchart illustrating the process of generating a control group data structure to determine anomalous methylated fragments according to one or more embodiments.
[0046] Figure 3B This is an exemplary flowchart illustrating the process of determining fragments as anomalously methylated according to a control group data structure based on one or more embodiments.
[0047] Figure 4A This is an exemplary flowchart illustrating the process of training a cancer classifier according to one or more embodiments.
[0048] Figure 4B Examples of generating feature vectors according to one or more embodiments are shown, which are used to train a cancer classifier.
[0049] Figure 5A This is an exemplary flowchart illustrating the process of training a parallel Cancer Origin (CSO) classifier according to one or more embodiments.
[0050] Figure 5B This is an exemplary flowchart illustrating the process of deploying a parallel CSO classifier according to one or more embodiments.
[0051] Figure 6A An exemplary flowchart of an apparatus according to one or more embodiments for sequencing nucleic acid samples is shown.
[0052] Figure 6B This is an exemplary block diagram of an analysis system according to one or more embodiments.
[0053] Figure 7 Two confusion matrices are shown to demonstrate the prediction accuracy of a demonstration organ type classifier based on one or more example implementations.
[0054] Figure 8 Two confusion matrices are shown, illustrating the predictive accuracy of a tumor biological type classifier based on one or more example implementations.
[0055] These figures depict several embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods described herein may be employed without departing from the principles described herein. Detailed Implementation
[0056] Overview Early detection and classification of cancer is a crucial technology. Being able to detect cancer before it develops symptoms benefits all stakeholders, including patients, doctors, and loved ones. For patients, early cancer detection allows for a greater chance of a beneficial outcome; for doctors, it enables a deeper understanding of the context and status of disease progression and the availability of additional treatment options that could lead to a beneficial outcome; and for loved ones, it increases the likelihood of not losing friends and family due to the disease.
[0057] Early detection and classification of cancer can also be achieved while patients are examining symptoms that could indicate the presence of cancer. This process, which can be called the "diagnostic odyssey" today, causes stress and anxiety for patients and their loved ones through a variety of costly medical procedures with uncertainty that can sometimes last for months.
[0058] Recently, early cancer detection technologies have evolved towards analyzing gene fragments (e.g., DNA) from human samples to determine if any of those fragments originate from cancer cells. To enhance the benefits of early cancer detection, samples are typically liquid samples (e.g., blood, saliva, or urine) that are relatively easy to obtain from a person. These new technologies enable doctors to identify cancers in a patient's body that might otherwise be undetectable (e.g., using conventional screening methods). For example, consider the case of a person at high risk for breast cancer. Traditionally, this person would regularly visit their doctor for mammograms, which generate images of their breast tissue (e.g., taking X-ray images) that the doctor uses to identify cancerous tissue. Unfortunately, even with the highest resolution mammograms, doctors can only identify tumors as small as about one millimeter in size. This means the cancer has been present in the person for some time and has remained undiagnosed and untreated. Such definite limitations are common for most cancers—that is, most cancers can only be identified when they have grown large enough to be detected by some imaging technique. For many types of cancer, there is no screening paradigm like mammography available, and even regular doctor visits cannot detect the presence of growing and progressing cancers. For example, someone at high risk for breast cancer may also be at high risk for ovarian cancer, which cannot be detected by mammography or any other currently available image-based cancer screening method.
[0059] This problem is mitigated by analyzing gene fragments from a patient's blood (e.g., blood). For clarity, assume that once cancer cells form, DNA fragments detach and enter the bloodstream. This occurs when the number of cancer cells is very small, before they can be observed using imaging techniques. Therefore, a system that analyzes DNA fragments in the bloodstream (so-called "cell-free DNA" or "cfDNA") using appropriate methods can identify the presence of cancer in the body based on these detached cancer DNA fragments. More importantly, this system can identify the cancer before more traditional cancer detection techniques, especially for cancers for which traditional cancer screening methods are not available.
[0060] Cancer detection based on DNA fragment analysis can be achieved through next-generation sequencing (“NGS”) technology. Broadly speaking, NGS is a set of technologies that enable high-throughput sequencing of genetic material. As discussed in more detail here, NGS mainly consists of (1) sample preparation, (2) DNA sequencing, and (3) data analysis. Sample preparation includes the laboratory methods necessary to prepare the DNA fragments for sequencing, sequencing is the process of reading the ordered nucleotides in the sample, and data analysis includes processing and analyzing the genetic information in the sequencing data to identify the presence of cancer.
[0061] This analysis, used to identify the presence and type of cancer, is further relevant when: a patient has been diagnosed with cancer and additional information is needed for prognosis and treatment decisions; and when treatment has not yet or can no longer control cancer growth, and when it is used to identify residual disease, recurrence, or relapse in patients under treatment.
[0062] While these NGS steps may help enable early cancer detection, they also introduce complex and harmful problems to cancer detection. Therefore, any improvements to sample preparation, DNA sequencing, and / or data analysis (including preprocessing, algorithmic processing, and summarizing or presenting predictions or conclusions) will more broadly improve NGS technology, universal cancer detection techniques, early cancer detection, and ultimately, a cure.
[0063] For the sake of illustration, for example, (1) problems arising from sample preparation include DNA sample quality, sample contamination, fragmentation bias, and accurate indexing. Addressing these issues will provide better genetic data for cancer detection.
[0064] Similarly, (2) problems arising from sequencing include, for example, errors in accurate transcription of fragments (e.g., reading “cytosine (C)” as “adenine (A)”), incorrect or difficult fragment assembly and overlap, inconsistent coverage uniformity, trade-offs between sequencing depth and cost versus specificity, and insufficient sequencing length. Likewise, addressing any of these problems would provide improved genetic data for cancer detection.
[0065] (3) The challenges in data analysis are the most daunting and complex. The massive amounts of data generated by NGS sequencing present numerous challenges. Sequencing data from a single sample can range from hundreds of thousands (millions) of sequence reads, totaling terabytes of data. Training analytical models typically involves collecting and processing thousands (up to tens of thousands or more) of samples with identified and labeled clinical cancer states, affected organs or organoids, primary sites of cancer, and potential cancer biology. Analyzing such large amounts of data effectively and efficiently is computationally demanding. For example, analyzing NGS sequencing involves several baseline processing steps, such as aligning reads to each other, aligning and mapping reads to a reference genome, deduplicating duplicate reads, detecting sample contamination, identifying and identifying variant genes, identifying and identifying aberrant methylation sites or regions in an individual's genome, generating functional annotations, etc. Performing any of these processing steps on terabytes (or larger) of genetic data is computationally expensive, even for the most powerful computer architectures, and completely impossible for the normal human brain. Furthermore, due to the error-prone nature of sample preparation and sequence reading processes, a significant portion of the genetic data derived from these processes may be of low quality or unusable for cancer identification. For example, large amounts of genetic data may include contaminated samples, transcriptional errors, mismatched regions, overrepresented regions, and uninformative regions, potentially unsuitable for high-accuracy cancer detection. Identifying and processing low-quality genetic data from the vast amounts of genetic data obtained from NGS sequencing is procedurally and computationally extremely difficult, and impractical for the human brain to perform. Overall, any process that leads to more efficient processing of large-scale sequencing data in analytical models commonly used in early cancer detection via NGS will improve cancer detection using NGS sequencing. Moreover, such processes, as described and illustrated herein, are solutions designed to address the various inherent obstacles of NGS technology, and are therefore unconventional and innovative activities in this field.
[0066] As a further example, in (3) data analysis, accurately identifying informative DNA from NGS data to identify the presence of cancer is another inherently difficult task in the field of early cancer detection. To achieve effective detection, algorithms are being sought to compensate for errors such as those arising from sample preparation and sequencing, the large-scale genomic and methylation variations present in the population, and to overcome the large-scale data analysis challenges posed by NGS technology. That is, when designing one or more machine learning models or other computational processing algorithms to achieve early cancer detection based on next-generation sequencing technology, configurations must be made to account for the problems introduced by those technologies. Some of these technologies and models are discussed below, with further discussions of specific improvements to state-of-the-art technologies and models. Furthermore, such technologies represent an unconventional and innovative activity in the field of technological exploration.
[0067] A particular challenge arises in providing predictive insights to healthcare professionals based on patient-specific samples. A first approach to cancer prediction might involve providing healthcare professionals with a binary prediction of whether a test subject has a high probability of having cancer or not. While this insight is highly informative, it falls short in providing further insights into how best to process cancer signal detection for diagnosis while avoiding the so-called "diagnostic odyssey." A second approach might involve providing a multi-category prediction about the origin of a specific cancer signal, such as whether a patient's sample reflects one of several discrete cancer signal origins (often referred to as organ system or organ type). This insight improves upon the first approach but may lack the specificity of features surrounding the origin of the cancer signal, which can be highly informative for obtaining diagnostic testing options for a cancer diagnosis. Finally, a third approach might involve providing insights into both the origin of the cancer (including the affected organ type) and the tumor biology type of the cancer. This insight most comprehensively enables healthcare professionals to provide the optimal testing options for cancer diagnosis after a test predicts or detects a cancer signal. For example, a report generated by a system utilizing computational models or other algorithms trained to generate such insights can predict that a particular subject has a cancer with adenocarcinoma biology (tumor biology type) originating in gastric tissue (organ type). Reports generated based on cancer prediction can include information from all three cancer prediction methods: detected (or undetected) cancer signal, origin of the cancer signal (if detected), and additional predictive information such as organ type and / or tumor biology type. With this insight, healthcare professionals can better evaluate diagnostic testing options to tailor diagnostic tests accordingly for detected cancer signals. In some embodiments, the analysis system can further store a database of diagnostic steps or testing options for each combination of organ type and tumor biology type. The analysis system can provide healthcare professionals with diagnostic steps and testing recommendations based on the predicted organ type and tumor biology type.
[0068] IA Cancer Classification Workflow Figure 1 This is an exemplary flowchart illustrating an overall workflow 100 for cancer classification of samples according to one or more embodiments. The workflow 100 is performed by one or more entities, including healthcare professionals, sequencing devices, analysis systems, etc. The purpose of the workflow includes detecting and / or monitoring cancer in subjects. From a healthcare perspective, workflow 100 can be used to complement other existing cancer screening and early detection tools. Workflow 100 can be used to provide early cancer detection and / or routine cancer surveillance, minimal residual disease detection, prognosis, treatment prediction, or subtype information to better inform treatment planning for subjects diagnosed with cancer. The overall workflow 100 may include more than Figure 1 The steps shown have more / fewer steps.
[0069] Healthcare workers perform sample collection 110. Subjects undergoing screening, early detection, or cancer triage visit their healthcare workers. Healthcare workers collect samples for cancer triage. Examples of biological samples include, but are not limited to, tissue biopsies, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial effusion, or peritoneal fluid from the subject. Samples include genetic material belonging to the subject, which can be extracted and sequenced for screening, early detection, or cancer triage. Once collected, the sample is provided to the laboratory procedures and sequencing device. In addition to the sample, healthcare workers may collect other information relevant to the subject, such as biological sex, age, race, smoking status, other health indicators, any previous diagnoses, etc. The sequencing device performs sample sequencing 120. Laboratory clinicians may perform one or more processing steps on the sample to prepare it for sequencing. Once prepared, a laboratory technician loads the sample into the sequencing device. Figure 8 A and Figure 8 B further describes examples of devices used for sequencing. These sequencing devices typically extract and isolate nucleic acid fragments, which are then sequenced to determine the nucleotide sequences corresponding to those fragments. Sequencing may also include amplification of the nucleic acid material. Different sequencing processes include Sanger sequencing, fragment analysis, and next-generation sequencing. Sequencing can be whole-genome sequencing or targeted sequencing with a target genome. In the context of DNA methylation, bisulfite sequencing (e.g., in…) is used… Figure 2A and Figure 2B(As further described below) The methylation state can be determined by converting unmethylated cytosine at CpG sites to bisulfite. Sample sequencing 120 generates sequences of multiple nucleic acid fragments in the sample. In one or more embodiments, these sequences may include methylation state vectors, where each methylation state vector describes the methylation state of CpG sites on the fragment.
[0070] The analysis system performs pre-analysis processing 130. In Figure 8 Section B describes an example analysis system. Pre-analysis processing 130 may include, but is not limited to, deduplication of sequence reads, determination of coverage-related metrics, determination of sample contamination (including determination of WBC contamination), removal of contaminated fragments, and identification of base sequencing errors.
[0071] The analysis system performs one or more analyses 140. These analyses are statistical analyses or apply one or more trained models to at least predict the cancer status of the subject from whom the sample was sourced. Different genetic traits can be assessed and considered, such as methylation at CpG sites, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), other types of genetic mutations, etc. Analysis 140 may include contamination detection 142, identification of aberrant methylation 144 (e.g., in... Figure 3A and 3B (further description in the text), feature extraction 146 (e.g., in...) Figure 4A , Figure 4B , Figure 5A and Figure 5B (further description in the text) and cancer classification 148 to determine cancer prediction (e.g., in... Figure 4A , Figure 4B , Figure 5A and Figure 5B (Further description follows). Cancer classification typically requires input extracted features to determine cancer predictions. Cancer predictions can be labels or values. Labels can indicate a specific cancer state. Binary labels can indicate the presence or absence of a cancer signal; while multi-class labels can indicate one or more cancer signal origins from multiple potential cancer signal origins screened. Cancer signal origins can be segmented based on various characteristics of the cancer. For example, cancer signal origins can be segmented based on organ type and / or tumor biology type. In other instances, cancer signal origins can be further segmented based on progression (e.g., stage I, II, III, or IV), prognosis, predicted response to candidate treatments, or the presence of residual disease after or during treatment. Values can indicate the probability of a specific cancer state, such as the likelihood of having cancer and / or the likelihood of a specific cancer signal origin.
[0072] The analysis system returns prediction 150 to healthcare professionals. Healthcare professionals may use this cancer prediction to develop or adjust treatment plans. In some embodiments, the analysis system may return diagnostic test recommendations 160 to healthcare professionals and / or the subject, which identify one or more diagnostic steps that can examine test results and lead to a cancer diagnosis. In such embodiments, the analysis system may store a database that associates diagnostic test options with diagnostic test steps generally accepted by medical professionals as useful in the presence of the predicted origin of cancer signals. Treatment optimization is further described in Section VD. In some embodiments, the analysis system may utilize a cancer classification workflow for prognostic determination, treatment personalization, treatment evaluation, monitoring of cancer status, etc.
[0073] IB Methylation Overview According to this specification, cfDNA fragments from subjects are processed, for example by converting unmethylated cytosine to uracil, sequenced, and the sequence reads compared to a reference genome to identify the methylation status at specific CpG sites within the DNA fragment. Each CpG site can be methylated or unmethylated. Identification of aberrantly methylated fragments compared to healthy subjects can provide insights into the subject's cancer status. As is well known in the art, changes in DNA methylation (compared to healthy controls) can cause different effects, which may contribute to cancer. Various challenges arise in the identification of informatively methylated cfDNA fragments. First, determining whether a DNA fragment is informatively methylated requires comparison with a control group for the determination to be convincing; therefore, if the control group is small, the determination may lose credibility due to statistical variability caused by the small size of the control group. Furthermore, methylation status may vary within the control group, making it difficult to determine whether a subject's DNA fragment is informatively methylated. On the other hand, methylation of cytosine at one CpG site can have a causal effect on the methylation of subsequent CpG sites. Taking this dependency into account is itself a challenge.
[0074] Methylation typically occurs in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is converted into a methyl group, forming 5-methylcytosine. Specifically, methylation can occur at dinucleotides composed of cytosine and guanine, referred to herein as “CpG sites.” In other cases, methylation may occur at cytosine sites that are not part of a CpG site or at another nucleotide that is not cytosine; however, these cases are rare. In this disclosure, for clarity, methylation is discussed in reference to CpG sites. Abnormal DNA methylation can be identified as hypermethylation or hypomethylation, both of which can indicate a cancerous state. Throughout this disclosure, hypermethylation and hypomethylation of a DNA fragment can be characterized if the number of CpG sites exceeds a threshold, and more than a percentage of those CpG sites are methylated or unmethylated. In addition to hypermethylation or hypomethylation, the informative methylation state of DNA fragments can be further characterized by sequences of methylated and unmethylated CpGs that are not frequently observed in healthy populations, and can indicate the presence of cancer or the presence of one or more cancer signaling origins.
[0075] The principles described herein also apply to the detection of methylation in a non-CpG background, including non-cytosine methylation. In such embodiments, the laboratory wet assays used to detect methylation may differ from those described herein. Furthermore, the methylation state vector discussed herein may contain elements that are typically sites that have or have not undergone methylation (even if those sites are not specifically CpG sites). With this adjustment, the remaining processes described herein can remain unchanged; therefore, the inventive concepts described herein are also applicable to other forms of methylation.
[0076] IC Definition The term "cell-free nucleic acid" or "cfNA" refers to a fragment of nucleic acid that circulates in the body of a subject (e.g., in the blood) or is present in other bodily fluids (e.g., urine, cerebrospinal fluid) and originates from one or more healthy cells and / or one or more unhealthy cells (e.g., cancer cells). The term "cell-free DNA" or "cfDNA" refers to a fragment of deoxyribonucleic acid that is present in the extracellular space of the subject's body (e.g., in plasma, urine, sputum, cerebrospinal fluid, etc.).
[0077] The terms "genomic nucleic acid," "genomic DNA," or "gDNA" refer to nucleic acid or deoxyribonucleic acid molecules obtained from one or more cells. In many embodiments, gDNA can be extracted from healthy cells (e.g., non-tumor cells) or tumor cells (e.g., biopsy samples). In some embodiments, gDNA can be extracted from cells derived from a blood cell lineage (e.g., leukocytes).
[0078] The term “circulating tumor DNA” or “ctDNA” refers to nucleic acid fragments derived from tumor cells or other types of cancer cells that may be released into a subject’s bodily fluids (e.g., blood, sweat, urine, or saliva) due to biological processes such as apoptosis or necrosis of dying cells or by active release from surviving tumor cells.
[0079] The terms “DNA fragment,” “fragment,” or “DNA molecule” can generally refer to any deoxyribonucleic acid fragment, such as cfDNA, gDNA, ctDNA, etc.
[0080] The terms "informative fragment," "abnormally methylated fragment," or "fragment with an abnormal methylation pattern" refer to fragments with informative methylation at CpG sites. Abnormal methylation of a fragment can be determined using probabilistic models to assess the degree of surprise in observing fragment methylation patterns in control groups.
[0081] As used herein, the terms "about" or "approximately" can mean a specific value determined by those skilled in the art to be within an acceptable range of error, depending in part on how the value was measured or determined, for example, due to limitations of the measurement system. For example, according to practice in the art, "about" can mean within one or more standard deviations. "About" can mean a range of ±20%, ±10%, ±5%, or ±1% of a given value. The terms "about" or "approximately" can mean within the same order of magnitude, within 5 times, or within 2 times a value. When describing specific numerical values in this application and claims, unless otherwise stated, it should be assumed that the term "about" means within an acceptable range of error for that specific value. The term "about" can have the meaning commonly understood by those skilled in the art. The term "about" can mean ±10%. The term "about" can mean ±5%.
[0082] As used herein, the terms “biological sample,” “patient sample,” or “sample” refer to any sample taken from a subject that reflects a biological state relevant to that subject and includes gDNA or cell-free DNA. Samples can be liquid or solid (e.g., cell or tissue samples). Biological samples can be bodily fluids, such as blood, plasma, serum, urine, vaginal secretions, fluid in a hydrocele (e.g., hydrocele of the testis), vaginal douches, pleural fluid, pericardial fluid, peritoneal fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirated fluid from different parts of the body (e.g., thyroid, breast), etc. Biological samples can be fecal samples. Biological samples can include any tissue or material derived from a living or deceased subject. Biological samples can be free samples. Biological samples can contain nucleic acids (e.g., DNA or RNA) or fragments thereof.
[0083] The term "nucleic acid" can refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or any hybrid or fragment thereof. The nucleic acids in a sample can be free nucleic acids. In several embodiments, the majority of DNA in a biological sample where free DNA has been enriched (e.g., a plasma sample obtained by centrifugation) can be free (e.g., more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be free). Biological samples can be processed to physically disrupt tissue or cellular structures (e.g., centrifugation and / or cell lysis), thereby releasing intracellular components into a solution that may further contain enzymes, buffers, salts, detergents, etc., which can be used to prepare a sample for analysis.
[0084] As used herein, the terms “control,” “control sample,” “reference,” “reference sample,” “normal,” and “normal sample” describe samples taken from subjects who do not have a specific disease or are otherwise healthy. In one instance, the methods disclosed herein can be applied to a subject with a tumor, where the reference sample is a sample taken from the subject’s healthy tissue. The reference sample can be obtained from the subject or from a database. The reference can be, for example, a reference genome, which is used to compare nucleic acid fragment sequences obtained from sequencing of the subject’s sample. The reference genome can refer to a haploid or diploid genome to which nucleic acid fragment sequences from biological samples and body samples can be compared and aligned. An example of a body sample can be DNA from white blood cells obtained from the subject. For a haploid genome, each locus contains only one nucleotide. For a diploid genome, heterozygous loci can be identified; each heterozygous locus has two alleles, either of which can match the locus for alignment.
[0085] As used in this article, the terms “cancer” or “tumor” refer to an abnormal mass of tissue in which the growth exceeds and is not in harmony with the growth of normal tissue.
[0086] As used in this article, the phrase “healthy” means that the subject is in good physical condition. A healthy subject may be free of any malignant or non-malignant disease. A “healthy subject” may have other diseases or conditions unrelated to the condition being measured, which are not typically considered “healthy.”
[0087] As used herein, the term "methylation" refers to a modification of deoxyribonucleic acid (DNA) in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted into a methyl group, forming 5-methylcytosine. Specifically, methylation tends to occur at dinucleotides composed of cytosine and guanine, referred to herein as "CpG sites." In other cases, methylation may occur on cytosine that is not part of a CpG site or on another nucleotide that is not cytosine; however, these cases are rare. Abnormal cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which can indicate a cancer state. Abnormal DNA methylation (compared to healthy controls) can have different effects, potentially leading to cancer. The principles described herein also apply to the detection of methylation in both CpG and non-CpG backgrounds, including non-cytosine methylation. Furthermore, the methylation state vector may contain elements that are typically vectors of sites that have undergone or not undergone methylation (even if those sites are not specifically CpG sites).
[0088] As used interchangeably throughout this document, the term "methylated fragment" or "nucleic acid methylated fragment" refers to a sequence whose methylation status is determined by methylation sequencing of nucleic acids (e.g., nucleic acid molecules and / or nucleic acid fragments) across multiple CpG sites. Within a methylated fragment, the location and methylation status of each CpG site within the nucleic acid fragment are determined based on alignment of these sequence reads (e.g., obtained by sequencing these nucleic acids) with a reference genome. A nucleic acid methylated fragment contains the methylation status (e.g., a methylation status vector) of each CpG site across multiple CpG sites, specifying the location of the nucleic acid fragment in the reference genome (e.g., as specified using a CpG index based on the location of the first CpG site in the nucleic acid fragment or another similar indicator) and the number of CpG sites in the nucleic acid fragment. Based on methylation sequencing of nucleic acid molecules, CpG indexing can be used for alignment of sequence reads with a reference genome. As used herein, the term "CpG index" refers to a list (e.g., CpG 1, CpG 2, CpG 3, etc.) of each CpG site among multiple CpG sites in a reference genome (e.g., the human reference genome), and this list may be in electronic format. The CpG index further includes the corresponding genomic location of each corresponding CpG site in the CpG index within the corresponding reference genome. Therefore, each CpG site in each corresponding nucleic acid methylation fragment can be indexed to a specific location in the corresponding reference genome determined using the CpG index.
[0089] As used herein, the term "true positive" (TP) refers to a subject who has a disease. A "true positive" can refer to a subject who has a tumor, cancer, precancerous lesion (e.g., precancerous lesion), localized or metastatic cancer, or a non-malignant disease. A "true positive" can refer to a subject who has a condition and is identified as having that condition by the assays or methods disclosed herein, such as having cancer that will respond to treatment, having residual disease after or during treatment, or potentially experiencing events such as disease progression, recurrence, or even cancer-related death within a pre-specified timeframe. As used herein, the term "true negative" (TN) refers to a subject who does not have a disease or a detectable disease. A "true negative" can refer to a subject who does not have a disease or a detectable disease (e.g., tumor, cancer, precancerous lesion (e.g., precancerous lesion), localized or metastatic cancer, or a non-malignant disease) or is otherwise healthy. A "true negative" can refer to a subject who does not have a disease or a detectable disease, or is identified as not having that disease by the assays or methods disclosed herein.
[0090] As used herein, the term “reference genome” refers to any known, sequenced, or characterized genome, whether partial or complete, belonging to any organism or virus, that can be used to compare with identified sequences from a subject. Exemplary reference genomes for human subjects and many other organisms are provided by online genome browsers hosted by the National Center for Biotechnology Information (“NCBI”) or the University of California, Santa Cruz (UCSC). A “genome” refers to the complete genetic information of an organism or virus expressed as a nucleic acid sequence. As used herein, a reference sequence or reference genome is typically an assembled or partially assembled genome sequence from one or more subjects. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human subjects. A reference genome can be considered a representative instance of the gene set of a species. In some embodiments, a reference genome contains sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).
[0091] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence produced by any sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment ("single-end read"), and sometimes from both ends of the nucleic acid (e.g., paired-end read, paired-end read). In some embodiments, a sequence read (e.g., a single-end or paired-end read) can be generated from one or both strands of a target nucleic acid fragment. The length of the sequence read is typically related to the specific sequencing technology. For example, sequence reads provided by high-throughput methods can range in size from tens to hundreds of base pairs (bp). In some embodiments, the mean length, median length, or average length of these sequence reads is approximately 15 bp to 900 bp (e.g., approximately 20 bp, approximately 25 bp, approximately 30 bp, approximately 35 bp, approximately 40 bp, approximately 45 bp, approximately 50 bp, approximately 55 bp, approximately 60 bp, approximately 65 bp, approximately 70 bp, approximately 75 bp, approximately 80 bp, approximately 85 bp, approximately 90 bp, approximately 95 bp, approximately 100 bp, approximately 110 bp, approximately 120 bp, approximately 130 bp, approximately 140 bp, approximately 150 bp, approximately 200 bp, approximately 450 bp, approximately 300 bp, approximately 350 bp, approximately 400 bp, approximately 450 bp, or approximately 500 bp). In some embodiments, the mean, median, or average length of these sequence reads is approximately 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or longer. For example, nanopore sequencing can provide sequence reads ranging in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads with less variation in size; for example, most of these reads are less than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides from a portion of a nucleic acid fragment (e.g., about 20 to about 150), a string of nucleotides at one or both ends of a nucleic acid fragment, or nucleotides corresponding to the entire nucleic acid fragment. Sequence reads can be obtained in various ways, such as using sequencing technologies or using probes, for example, in hybridization arrays or capture probes, or amplification technologies, such as polymerase chain reaction (PCR) or linear or isothermal amplification using a single primer.
[0092] As used herein, the term "sequencing" generally refers to any and all biochemical processes that can be used to determine the sequence of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data may include all or part of the nucleotide bases in a nucleic acid molecule, such as a fragment of DNA.
[0093] As used herein, the terms "sequencing depth" and "coverage" are used interchangeably and refer to the number of times a shared sequence read corresponding to a unique nucleic acid target molecule aligned to a locus is aligned with that locus; for example, sequencing depth equals the number of unique nucleic acid target molecules covering that locus. A locus can be as small as a nucleotide, as large as a chromosome arm, or as large as an entire genome. Sequencing depth can be expressed as "Yx," such as 50x, 100x, etc., where "Y" refers to the number of times the locus is covered by a sequence corresponding to the target nucleic acid, for example, the number of times independent sequence information covering a specific locus is obtained. In some embodiments, the sequencing depth corresponds to the number of genomes that have been sequenced. Sequencing depth can also be applied to multiple loci or the entire genome, in which case Y can refer to the mean or average number of times a locus, a haploid genome, or the entire genome has been sequenced. When referring to average depth, the actual depth of different loci included in the dataset can cover a range of values. Ultra-deep sequencing can refer to a sequencing depth of at least 100x at a locus.
[0094] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of a assay or method to correctly identify the proportion of a population that actually has a certain disease. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects in a population who have cancer. In another instance, sensitivity can characterize the ability of a method to correctly identify one or more markers that indicate cancer.
[0095] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of a assay or method to correctly identify the proportion of a population that is truly free of a certain disease. For example, specificity can characterize the ability of a method to correctly identify the number of subjects within a population who do not have cancer. In another instance, specificity characterizes the ability of a method to correctly identify one or more markers that indicate cancer.
[0096] As used herein, the term "subject" means any living or non-living organism, including but not limited to humans (e.g., human males, human females, fetuses, pregnant women, children, etc.), non-human animals, plants, bacteria, fungi, or protozoa. Any human or non-human animal may be a subject, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovids (e.g., cattle), equines (e.g., horses), sheep and goats (e.g., sheep, goats), suidae (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), bears (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, a subject is a male or female at any stage (e.g., an adult male, an adult female, or a child). Subjects from whom samples are collected or who are treated using any of the methods or compositions described herein may be of any age and may be adults, infants, or children.
[0097] As used herein, the term "tissue" can refer to a group of cells aggregated together as a functional unit. More than one cell type can be found in a single tissue. Different types of tissues may be composed of different types of cells (e.g., hepatocytes, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother and fetus) or to healthy cells and tumor cells. The term "tissue" can generally refer to any group of cells found in the human body (e.g., heart tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some respects, the term "tissue" or "tissue type" can be used to refer to the tissue from which cell-free nucleic acids originate. In one instance, viral nucleic acid fragments may originate from blood tissue. In another instance, viral nucleic acid fragments may originate from tumor tissue.
[0098] As used herein, the term “genomic” refers to the characteristics of an organism’s genome. Examples of genomic characteristics include, but are not limited to: those characteristics associated with the primary nucleic acid sequences of all or part of the genome (e.g., the presence or absence of nucleotide polymorphisms, insertions or deletions, sequence rearrangements, mutation frequencies, etc.), copy numbers of one or more specific nucleotide sequences within the genome (e.g., copy number, allele frequency fractions, ploidy of a single chromosome or the entire genome, etc.), epigenetic states of all or part of the genome (e.g., covalent nucleic acid modifications such as methylation, histone modifications, nucleosome selection, etc.), and the expression profile of an organism’s genome (e.g., gene expression levels, isotype expression levels, gene expression ratios, etc.).
[0099] The terminology used herein is intended to describe a particular situation only and is not intended to be limiting. As used herein, the singular forms “a / an” and “the” are also intended to include the plural forms, unless the context clearly indicates otherwise. Furthermore, with regard to the scope of the terms “including,” “includes,” “having,” “has,” “with,” or variations thereof used in the detailed description and / or claims, such terms are intended to indicate inclusion in a manner similar to that of the term “comprising.”
[0100] ID Example Analysis System Figure 6A This is an exemplary flowchart of an apparatus for sequencing nucleic acid samples according to one or more embodiments. The illustrative flowchart includes apparatus such as a sequencer 620 and an analysis system 600. The sequencer 620 and the analysis system 600 may operate in series to perform one or more steps in the process.
[0101] In various embodiments, the sequencer 620 receives the enriched nucleic acid sample 610. For example... Figure 6A As shown, sequencer 620 may include a graphical user interface 625 that allows a user to interact with specific tasks (e.g., start or stop sequencing) and one or more loading stations 630 for loading sequencing cassettes containing the enriched fragment samples and / or buffers necessary for performing sequencing assays. Therefore, once the user of sequencer 620 has provided the necessary reagents and sequencing cassettes to the loading station 630 of sequencer 620, the user can start sequencing by interacting with the graphical user interface 625 of sequencer 620. Once started, sequencer 620 performs sequencing and outputs sequence reads of the enriched fragments in nucleic acid sample 610.
[0102] In some embodiments, sequencer 620 is communicatively coupled to analysis system 600. Analysis system 600 includes a number of computing devices for processing sequence reads to meet various application needs, such as assessing methylation status at one or more CpG sites, variant calling, or quality control. Sequencer 620 can provide sequence reads to analysis system 600 in BAM file format. Analysis system 600 can be communicatively coupled to sequencer 620 via wireless, wired, or a combination of wireless and wired communication technologies. Typically, analysis system 600 is configured with a processor and a non-transitory computer-readable storage medium storing computer instructions that, when executed by the processor, cause the processor to process sequence reads or perform one or more steps of any method or process disclosed herein.
[0103] In some embodiments, a sequence read can be aligned to a reference genome using methods known in the art to determine alignment location information. Alignment locations typically describe the start and end positions of a region in the reference genome, corresponding to the start and end nucleotide bases of a given sequence read. For methylation sequencing, the concept of alignment location information can be generalized to indicate the first and last CpG sites included in the sequence read, based on the alignment with the reference genome. Alignment location information can further indicate the methylation status and location of all CpG sites in a given sequence read. Regions in the reference genome may be associated with genes or fragments of genes; therefore, the analysis system 600 can label the sequence read with one or more genes aligned to it. In one embodiment, fragment length (or size) is determined by the start and end positions.
[0104] In various embodiments, such as when using paired-end sequencing, a sequence read consists of a pair of reads, denoted as R_1 and R_2. For example, the first read R_1 can be sequenced from one end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 can be sequenced from the other end of the same double-stranded DNA (dsDNA). Therefore, the nucleotide base pairs of the first read R_1 and the second read R_2 can be aligned with the nucleotide base pairs of a reference genome in a manner that is always consistent (e.g., in opposite directions). The alignment position information derived from the read pair R_1 and R_2 may include a start position in the reference genome corresponding to one end of the first read (e.g., R_1) and an end position in the reference genome corresponding to one end of the second read (e.g., R_2). In other words, the start and end positions in the reference genome can represent the possible positions of the nucleic acid fragments within the reference genome. Output files in SAM (Sequence Alignment Map) or BAM (Binary Alignment Map) format can be generated and output for further analysis.
[0105] Now for reference Figure 6B , Figure 6B This is a block diagram of an analysis system 600 for processing DNA samples according to one embodiment. The analysis system implements one or more computing devices for use in analyzing DNA samples. The analysis system 600 includes a sequence processor 640, a sequence database 645, a model database 655, a model 650, a parameter database 665, and a scoring engine 660. In some embodiments, the analysis system 600 performs some or all of the processes described throughout this disclosure.
[0106] Sequence processor 640 generates methylation status vectors for fragments of the sample. For each CpG site on a fragment, sequence processor 640 generates a methylation status vector for each fragment, indicating the fragment's location in the reference genome, the number of CpG sites within the fragment, and the methylation status of each CpG site in the fragment (methylated, unmethylated, or indeterminate). This process is performed through... Figure 2A The process 200 is implemented in the sequence processor 640. The sequence processor 640 can store the methylation state vectors of each fragment in the sequence database 645. The data in the sequence database 645 can be organized such that the methylation state vectors from the same sample are correlated with each other.
[0107] Furthermore, multiple different models 650 can be stored in the model database 655 or retrieved for use on test samples. In one instance, the model is a trained cancer classifier that uses feature vectors derived from informative fragments to determine cancer predictions for test samples. The training and use of the cancer classifier will be discussed further in Part III, “Cancer Classifiers for Cancer Determination.” The analysis system 600 can train one or more models 650 and store various trained parameters in the parameter database 665. The analysis system 600 stores the models 650 along with their functions in the model database 655.
[0108] During inference, the scoring engine 660 uses one or more models 650 to return outputs. The scoring engine 660 accesses the models 650 in the model database 655 and the trained parameters in the parameter database 665. For each model, the scoring engine receives inputs suitable for that model and computes the output based on the received inputs, parameters, and functions associated with the inputs and outputs in each model. In some use cases, the scoring engine 660 further computes metrics related to the confidence level of the computed outputs from that model. In other use cases, the scoring engine 660 computes additional intermediate values used in that model.
[0109] II. Sample Sequencing and Processing II.A. Generating a methylation state vector for DNA fragments According to one or more embodiments, Figure 2A This is an exemplary flowchart illustrating process 200 for sequencing fragments of cfDNA to obtain a methylation state vector. To analyze DNA methylation, the analysis system first obtains a sample 210 containing multiple cfDNA molecules from the subject. In another embodiment, process 200 can be applied to sequence other types of DNA molecules. Process 200 is... Figure 1 An example of sample sequencing 120.
[0110] The analytical system can isolate 210 individual cfDNA molecules from a sample. These cfDNA molecules can be processed 220 to convert unmethylated cytosine into uracil. In one embodiment, the method uses bisulfite treatment of the DNA, which converts unmethylated cytosine into uracil without converting methylated cytosine. For example, commercially available kits can be used for bisulfite conversion, such as EZ DNA Methylation. TM -Gold, EZ DNAMethylation TM -Direct or EZ DNA Methylation TM - The Lightning kit (available from Zymo Research, Irvine, CA). In another embodiment, an enzymatic reaction is used to complete the conversion of unmethylated cytosine to uracil. For example, this conversion can be performed using commercially available kits, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).
[0111] From the transformed cfDNA molecules, a sequencing library of 230 molecules can be prepared. During library preparation, a unique molecular identifier (UMI) can be added to the nucleic acid molecule via adapter ligation. For example The UMI (Universal Microarray) is attached to both ends of a DNA fragment (e.g., a DNA molecule lysed by physical shearing, enzymatic digestion, and / or chemical cleavage) during adapter ligation. The UMI can be a degenerate base pair, serving as a unique tag for identifying sequence reads derived from a specific DNA fragment. During PCR amplification following adapter ligation, the UMI is replicated along with the ligated DNA fragment. This provides a method for identifying sequence reads from the same original fragment in downstream analysis.
[0112] Optionally, multiple hybridization probes can be used to enrich cfDNA molecules or genomic regions in a sequencing library, which may provide information about cancer status. These hybridization probes are short oligonucleotides capable of hybridizing with specifically designated cfDNA molecules or target regions and enriching those fragments or regions for subsequent sequencing and analysis. Hybridization probes can be used to perform targeted, high-depth analysis on a set of designated CpG sites of interest to researchers. Hybridization probes can be tiled on one or more target sequences with coverage of 1X, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, or greater than 10X. For example, a hybridization probe tiled with 2X coverage contains overlapping probes, allowing each portion of the target sequence to hybridize with two independent probes. Hybridization probes can also be tiled on one or more target sequences with coverage less than 1X.
[0113] In one embodiment, hybridization probes are designed to enrich DNA molecules that have been treated (e.g., using bisulfite) to convert unmethylated cytosine into uracil. During enrichment, the hybridization probes (also referred to herein as “probes”) can be used to target and capture nucleic acid fragments that provide information about the presence or absence of cancer (or disease), cancer status, or cancer classification (e.g., cancer category or tissue of origin). These probes can be designed to anneal (or hybridize) with a target (complementary) strand of DNA. This target strand can be a “positive” strand (e.g., the strand transcribed into mRNA and subsequently translated into protein) or a complementary “negative” strand. The length of the probes can range from tens, hundreds, or thousands of base pairs. Probes can be designed based on a set of methylation sites. These probes can be designed based on a set of target genes to analyze specific mutations or target regions in the genome (e.g., in humans or other organisms) that are suspected to correspond to certain cancers or other types of diseases. Furthermore, these probes can cover overlapping portions of the target regions.
[0114] Once prepared, the sequencing library or a portion thereof can be sequenced 240 times to obtain multiple sequence reads. These sequence reads can be in a computer-readable digital format for processing and interpretation by computer software. These sequence reads may be aligned to a reference genome to determine alignment position information. This alignment position information can indicate the start and end positions of a region in the reference genome, corresponding to the start and end nucleotide bases of a given sequence read. The alignment position information may also include the sequence read length, which can be determined by the start and end positions. Regions in the reference genome may be associated with genes or segments of genes. A sequence read can consist of a pair of reads, represented as... R 1 and R 2. For example, the first reading paragraph R 1. Sequencing can be performed from the first end of the nucleic acid fragment, while the second read...R 2. Sequencing is performed from the second end of this nucleic acid fragment. Therefore, the first read... R 1 and the second reading segment R The nucleotide base pairs of 2 can be aligned in a manner consistent with the nucleotide base pairs of the reference genome (e.g., in the opposite direction). Derived from read pairs. R 1 and R The alignment location information for 2 may include a start position in the reference genome, which corresponds to the first read ( For example , R 1) at one end, and a termination position in the reference genome, which corresponds to the second read ( For example , R 2) One end. In other words, the start and end positions in the reference genome can represent the possible positions of nucleic acid fragments in the reference genome. Output files in SAM (Sequence Alignment Map) or BAM (Binary Alignment Map) format can be generated and output for further analysis, such as methylation status determination.
[0115] From these sequence reads, the analysis system determines the location and methylation status of 250 CpG sites per fragment based on alignment with a reference genome. The system generates a 260-fold methylation status vector for each fragment, specifying its location in the reference genome (e.g., as indicated by the location of the first CpG site in each fragment or another similar indicator), the number of CpG sites in the fragment, and the methylation status of each CpG site in the fragment (methylated (e.g., denoted as M), unmethylated (e.g., denoted as U), or indeterminate (e.g., denoted as I)). Observed states can be methylated or unmethylated, while unobserved states are indeterminate. Indeterminate methylation states may arise from sequencing errors and / or inconsistencies between the methylation states of the complementary strands of the DNA fragment. The methylation status vectors can be stored in transient or persistent computer memory for later use and processing. Furthermore, the analysis system can remove duplicate reads or duplicate methylation status vectors from individual samples. The analysis system can identify fragments with one or more CpG sites that have an indeterminate methylation state above a threshold number or percentage, and can exclude such fragments or selectively include such fragments, but construct a model that describes such indeterminate methylation states.
[0116] According to one or more embodiments, Figure 2B This is an example diagram that illustrates... Figure 2AThe procedure 200 involves sequencing cfDNA molecules to obtain a methylation state vector. For example, the analysis system receives cfDNA molecule 212, which in this example contains three CpG sites. As shown in the figure, the first and third CpG sites of cfDNA molecule 212 are methylated 214. During processing step 220, cfDNA molecule 212 is transformed to generate transformed cfDNA molecule 222. During processing 220, the cytosine at the unmethylated second CpG site is converted to uracil. However, the first and third CpG sites are not converted.
[0117] After transformation, a sequencing library 230 is prepared and sequenced 240 to generate sequence reads 242. The analysis system aligns sequence read 242 with a reference genome 244 250. The reference genome 244 provides context indicating the origin of the cfDNA fragment within the human genome. In this simplified example, the analysis system aligns sequence read 242 250 to associate the three CpG sites with CpG sites 23, 24, and 25 (arbitrary reference identifiers used for ease of description). This allows the analysis system to generate information about the methylation status of all CpG sites on the cfDNA molecule 212 and the locations of these CpG sites within the human genome. As shown in the figure, methylated CpG sites on sequence read 242 are read as cytosine. In this example, cytosine appears only at the first and third CpG sites of sequence read 242, suggesting that the first and third CpG sites in the original cfDNA molecule are methylated. The second CpG site can be read as thymine (U is converted to T during sequencing), indicating that the second CpG site in the original cfDNA molecule is unmethylated. With this information of methylation state and location, the analysis system generates a methylation state vector 252 for cfDNA fragment 212. In this example, the resulting methylation state vector 252 is... <M 23 U 24 M 25 > where M corresponds to methylated CpG sites, U corresponds to unmethylated CpG sites, and the subscript numbers correspond to the position of each CpG site in the reference genome.
[0118] One or more alternative sequencing methods can be used to obtain sequence reads from nucleic acids in biological samples. These sequencing methods can encompass any form of sequencing that can be used to obtain a number of sequence reads measured from nucleic acids (e.g., free nucleic acids), including but not limited to high-throughput sequencing systems such as Roche's 454 platform, Applied Biosystems' SOLID platform, Helicos' True Single Molecule DNA sequencing technology, Affymetrix's hybridization sequencing platform, Pacific Biosciences' Single Molecule Real-Time (SMRT) technology, 454 Life Sciences' synthetic sequencing platforms, Illumina / Solexa and Helicos Biosciences' synthetic sequencing platforms, and Applied Biosystems' ligation sequencing platforms. Life Technologies' ION TORRENT technology and nanopore sequencing can also be used to obtain sequence reads from nucleic acids (e.g., free nucleic acids) in biological samples. Synthetic sequencing and reversible terminator-based sequencing (e.g., Illumina Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 4500 (Illumina, San Diego, California)). Calif. can be used to extract sequence reads from cell-free nucleic acids obtained from biological samples acquired from training subjects to form a genotype dataset. Millions of cell-free nucleic acid (e.g., DNA) fragments can be sequenced in parallel. In one instance of this type of sequencing technology, a flow cell is used, which contains an optically clear slide with eight lanes for the subject, on the surface of which are bound to oligonucleotide anchors (e.g., linker primers). The cell-free nucleic acid sample may include a signal or tag that is easy to detect. The process of extracting sequence reads from cell-free nucleic acids obtained from biological samples may include obtaining quantitative information about the signal or tag using a variety of techniques, such as flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene chip analysis, microarrays, mass spectrometry, cellular fluorescence analysis, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch separation mode, electric field suspension, sequencing, and combinations thereof.
[0119] One or more sequencing methods may include whole-genome sequencing analysis. Whole-genome sequencing analysis may include a physical analysis that generates reads of the entire genome or a large portion of the genome, which can be used to identify large-scale variations such as copy number variations or copy number aberrations. This physical analysis may employ whole-genome sequencing technology or whole-exome sequencing technology. Whole-genome sequencing analysis may have an average sequencing depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, at least 20x, at least 30x, or at least 40x on the genome of the test subject. In some embodiments, the sequencing depth is approximately 30,000x. One or more sequencing methods may include targeted genome sequencing analysis. For the targeted genome, targeted genome sequencing analysis may have an average sequencing depth of at least 50,000x, at least 55,000x, at least 60,000x, or at least 70,000x. The targeted genome may contain 450 to 500 genes. The targeted genome can contain a range of 500 ± 5 genes, 500 ± 10 genes, or 500 ± 25 genes.
[0120] One or more sequencing methods may include paired-end sequencing. One or more sequencing methods may generate multiple sequence reads. These multiple sequence reads may have an average length range of 10 to 700, 50 to 400, or 100 to 300. One or more sequencing methods may include methylation sequencing analysis. Methylation sequencing may be i) whole-genome methylation sequencing or ii) targeted DNA methylation sequencing using multiple nucleic acid probes. For example, methylation sequencing is whole-genome bisulfite sequencing (e.g., WGBS). Methylation sequencing may be targeted DNA methylation sequencing using multiple nucleic acid probes that target the most informative regions of the methylome, may be a unique methylation database, or may be a previous whole-genome and targeted sequencing prototype analysis.
[0121] Methylation sequencing can detect one or more 5-methylcytosine (5mC) and / or 5-hydroxymethylcytosine (5hmC) in a corresponding nucleic acid methylation fragment. Methylation sequencing may involve converting one or more unmethylated cytosines, or one or more methylated cytosines, in a corresponding nucleic acid methylation fragment into one or more corresponding uracils. During methylation sequencing, one or more uracils can be detected as one or more corresponding thymines. The conversion of one or more unmethylated cytosines or one or more methylated cytosines can include chemical conversion, enzymatic conversion, or a combination thereof.
[0122] For example, bisulfite conversion involves converting cytosine to uracil without converting methylated cytosine (e.g., 5-methylcytosine or 5-mC). In some DNAs, approximately 95% of the cytosine may be unmethylated, and the resulting DNA fragment may include many uracils represented by thymine. Enzymatic conversion processes can be used to treat nucleic acids before sequencing and can be performed in various ways. One example of bisulfite-free conversion includes a bisulfite-free sequencing method with base-level resolution, TET-assisted pyridine-methylborane sequencing (TAPS), for the non-destructive and direct detection of 5-methylcytosine and 5-hydroxymethylcytosine without affecting unmodified cytosine. The methylation status of a CpG site within multiple corresponding CpG sites in a corresponding nucleic acid methylation fragment is determined by methylation sequencing; if methylation sequencing determines it is methylated, then the CpG site is methylated; if methylation sequencing determines it is unmethylated, then the CpG site is unmethylated.
[0123] Methylation sequencing analyses (e.g., WGBS and / or targeted methylation sequencing) can have average sequencing depths including, but not limited to, about 1,000x, 2,000x, 3,000x, 5,000x, 10,000x, 15,000x, 20,000x, or 30,000x. Methylation sequencing can have sequencing depths greater than 30,000x, for example, at least 40,000x or 50,000x. Whole-genome bisulfite sequencing methods can have average sequencing depths of 20x to 50x, while targeted methylation sequencing methods have average effective depths of 100x to 1,000x, where effective depth can be the equivalent coverage required for whole-genome bisulfite sequencing to obtain the same number of sequence reads obtained from targeted methylation sequencing. Methylation sequencing assays can include targeting 500 or more CpG sites, 1,000 or more CpG sites, 1,500 or more CpG sites, 2,000 or more CpG sites, 2,500 or more CpG sites, 3,000 or more CpG sites, 3,500 or more CpG sites, 4,000 or more CpG sites, 4,500 or more CpG sites, 5,000 or more CpG sites, 6,000 or more CpG sites, 7,000 or more CpG sites, 8,000 or more CpG sites. Probes with 9,000 or more CpG sites, 10,000 or more CpG sites, 20,000 or more CpG sites, 30,000 or more CpG sites, 40,000 or more CpG sites, 50,000 or more CpG sites, 60,000 or more CpG sites, 70,000 or more CpG sites, 80,000 or more CpG sites, 90,000 or more CpG sites, 100,000 or more CpG sites, 150,000 or more CpG sites, etc.
[0124] For further details on methylation sequencing (e.g., WGBS and / or targeted methylation sequencing), see, for example, U.S. Patent Application No. 16 / 719,902, filed December 18, 2019, entitled “Systems and Methods for Estimating Cell Source Fractions Using Methylation Information,” each of which is hereby incorporated by reference. Other methods for methylation sequencing, including those disclosed herein and / or any modifications, substitutions, or combinations thereof, can be used to obtain the methylation patterns of fragments. Methylation sequencing can be used to identify one or more methylation state vectors, for example, as described in U.S. Patent Application No. 16 / 352,602, filed March 13, 2019, entitled “Anomalous Fragment Detection and Classification,” or according to any technique disclosed in U.S. Patent Application No. 15 / 931,022, filed May 13, 2020, entitled “Model-Based Featurization and Classification,” each of which is hereby incorporated by reference.
[0125] Nucleic acid methylation sequencing and the resulting one or more methylation state vectors can be used to obtain multiple nucleic acid methylation fragments. Each group corresponding to multiple nucleic acid methylation fragments (e.g., for each corresponding genotype dataset) can contain more than 100 nucleic acid methylation fragments. The average number of nucleic acid methylation fragments can include 1,000 or more, 5,000 or more, 10,000 or more, 20,000 or more, 30,000 or more, 40,000 or more, or 50,000 or more. Within each group corresponding to multiple nucleic acid methylation fragments, the average number of nucleic acid methylation fragments can be between 10,000 and 50,000. Multiple methylated fragments can contain 1,000 or more, 10,000 or more, 100,000 or more, 1 million or more, 10 million or more, 100 million or more, 500 million or more, 1 billion or more, 2 billion or more, 3 billion or more, 4 billion or more, 5 billion or more, 6 billion or more, 7 billion or more, 8 billion or more, 9 billion or more, or 10 billion or more. The average length of multiple methylated fragments can be 140 to 480 nucleotides. The average length of multiple methylated fragments can be 100 to 600 nucleotides. The average length of multiple methylated fragments can be 50 to 750 nucleotides. The average length of multiple methylated fragments can be 30 to 1,000 nucleotides.
[0126] Further details regarding nucleic acid sequencing methods and methylation sequencing data are disclosed in U.S. Patent Application No. 17 / 191,914, filed March 4, 2021, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," which is hereby incorporated herein by reference in its entirety.
[0127] III. Cancer classification used to determine cancer Cancer classification involves extracting genetic features and applying one or more models to these extracted features to determine cancer predictions. The analysis system aggregates the extracted features into a feature vector, which can then be fed into a trained cancer prediction model to determine a cancer prediction based on the input feature vector. A cancer prediction may contain one or more labels and / or one or more values. One label may be binary, indicating the presence or absence of cancer in the test subject. Another label may be multi-class, indicating one or more specific cancer signal origins from multiple screened cancer signal origins. One value may indicate the probability of cancer presence. Another value may indicate the probability of cancer absence. Yet another value may additionally indicate an alternative prognosis for the cancer. For example, a value may quantify progress and / or potential response to cancer treatment.
[0128] In one or more embodiments, the feature vector input to the cancer classifier is determined based on a set of informative fragments (also known as “abnormal methylation” or “extreme methylation anomalous fragments” (UFXM)) identified from the test sample.
[0129] In some embodiments, the cancer classifier can be a machine learning model, which is a computational model containing multiple classification parameters and a function representing the relationship between an input feature vector and an output cancer prediction. The feature vector, along with the classification parameters, is input into the function to produce a cancer prediction. The machine learning model can be trained using training samples derived from subjects with known cancer diagnoses. The training samples can be divided into queues with different labels. For example, a queue of training samples can exist for each cancer signal origin.
[0130] Machine learning models can be trained using any combination of machine learning techniques, including but not limited to linear regression, logistic regression, support vector machines (SVM), k-nearest neighbors (KNN), decision trees (e.g., ID3, C4.5, CART), random forests, neural networks (e.g., multilayer perceptrons, convolutional neural networks, recurrent neural networks), Naive Bayes classifiers, gradient boosting machines (e.g., XGBoost, LightGBM, CatBoost), AdaBoost, k-means clustering, hierarchical clustering, DBSCAN (density-based spatial clustering with noise), Gaussian mixture models (GMM), PCA (principal component analysis), t-SNE (t-distributed random neighbor embedding), and UMAP (uniform manifold approximation and projection). Independent Component Analysis (ICA), autoencoders (e.g., variational autoencoders, denoising autoencoders), self-organizing maps (SOM), Q-learning, deep Q-networks (DQN), SARSA (state-action-reward-state-action), Monte Carlo methods, temporal difference learning (TD learning), policy gradients (e.g., reinforcement, actor-commentator), proximal policy optimization (PPO), dominant actor-commentator (A2C), soft actor-commentator (SAC), deep deterministic policy gradient (DDPG), hidden Markov models (HMM), conditional random fields (CRF), latent Dirichlet distribution (LDA), restricted Boltzmann machines (RBM), genetic algorithms and swarm intelligence algorithms (e.g., particle swarm optimization, ant colony optimization), expectation maximization (EM).
[0131] III.A. Identifying informative fragments The analysis system can use the methylation state vector of a sample to identify informative fragments of that sample. For each fragment in the sample, the analysis system can use the methylation state vector corresponding to that fragment to determine whether the fragment is informative. In some embodiments, the analysis system calculates a p-value score for each methylation state vector, which describes the probability of observing that methylation state vector or other less likely methylation state vectors in a healthy control group. The procedure for calculating the p-value score is further discussed in Section III.Ai below. p-value filtering.The analysis system can identify informative fragments as those with a p-value score below a threshold for their methylation state vectors. In some embodiments, the analysis system further labels fragments with at least a certain number of CpG sites and a methylation or unmethylation percentage exceeding a certain threshold as hypermethylated and hypomethylated fragments, respectively. Hypermethylated or hypomethylated fragments can also be referred to as extreme methylation anomalous fragments (UFXM). In other embodiments, the analysis system can implement various other probabilistic models to determine informative fragments. Examples of other probabilistic models include mixture models, deep probabilistic models, etc. In some embodiments, the analysis system can use any combination of the multiple processes described below to identify informative fragments. Based on the identified informative fragments, the analysis system can filter the methylation state vector set of a sample for use in other processes, such as training and deploying a cancer classifier.
[0132] III. Ai p value filtering In some embodiments, the analysis system calculates a p-value score for each methylation state vector compared to the methylation state vectors of fragments in a healthy control group. The p-value score describes the probability of observing a methylation state in the healthy control group that matches that methylation state vector or other less likely methylation state vectors. To determine if a DNA fragment is abnormally methylated, the analysis system can use a healthy control group where most fragments are normally methylated. When performing this probabilistic analysis to identify informative fragments, the determination needs to be compared with a control group comprising the healthy control group to be convincing. To enhance the robustness of the healthy control group, the analysis system can select a threshold number of healthy subjects to obtain samples including DNA fragments. The following... Figure 3A A method for generating a healthy control group data structure that can be used by the analysis system to calculate p-scores is described. Figure 3B A method for calculating p-value scores using the generated data structure is described.
[0133] Figure 3A This is a flowchart 300 describing the process of generating a healthy control group data structure according to an embodiment. To create the healthy control group data structure, the analysis system can receive multiple DNA fragments (e.g., cfDNA) from multiple healthy subjects. The analysis system can generate a 305-methylation state vector for each fragment, for example, through... Figure 2A The process is 200.
[0134] For each fragment's methylation state vector, the analysis system can subdivide it into 310 strings of CpG sites. In some embodiments, the analysis system subdivides the methylation state vector 310 times such that all resulting strings are shorter than a given length. For example, a methylation state vector of length 11 might be subdivided into strings of length 3 or less, resulting in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another instance, a methylation state vector of length 7 might be subdivided into strings of length 4 or less, resulting in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If a methylation state vector is shorter than or the same length as a specified string, then the methylation state vector might be transformed into a single string containing all CpG sites of that vector.
[0135] The analysis system counts the number of strings in the control group that have that methylation state probability, with the specified CpG site as the first CpG site in the string, for each possible CpG site and vector. For example, at a given CpG site and considering a string length of 3, there are 2^3 or 8 possible string configurations. At that given CpG site, for each of these 8 possible string configurations, the analysis system counts the number of times each methylation state vector probability occurs in the control group. Continuing this example, this might involve counting the following: for each starting CpG site in the reference genome, <M x M x+1 M x+2 >、 <M x M x+1 U x+2 >、. . .、 x U x+1 U x+2 The analysis system creates a 315-dimensional data structure to store statistical counts of each starting CpG site and the probability of a string.
[0136] Setting a maximum string length has several benefits. First, the size of the data structures created by the analysis system increases dramatically with the maximum string length. For example, a maximum string length of 4 means that at least 2^4 numbers need to be counted for each CpG site in a string of length 4. Increasing the maximum string length to 5 means that there are another 2^4 or 16 numbers to count for each CpG site, doubling the number of numbers to count (and the required computer memory) compared to the previous string length. Reducing the string size can help keep the creation and performance of the data structures (e.g., for subsequent access, as described below) reasonable in terms of computation and storage. Second, a statistical consideration for limiting the maximum string length could be to avoid overfitting downstream models that use string counting. If long strings of CpG sites do not have a strong biological impact on outcomes (e.g., predictions of predictive anomalies in cancer), then calculating probabilities based on large strings of CpG sites could be problematic because it would use a large amount of potentially unavailable data, making it too sparse for the model to function properly. For example, calculating the probability of an anomalous / cancer condition based on the top 100 CpG sites could be done using a count of 100 strings in a data structure, ideally some of which would perfectly match these top 100 methylation states. However, if only sparse counts of 100 strings are available, the data might be insufficient to determine whether a given 100-string in a test sample is anomalous.
[0137] According to one embodiment, Figure 3B This is a flowchart describing process 330 for identifying aberrant methylation fragments from a subject. In process 330, the analysis system generates a methylation state vector 340 from the subject's cfDNA fragment, for example, through... Figure 2A The process is as follows: 200. The analysis system can process each methylation state vector as follows.
[0138] For a given methylation state vector, the analysis system enumerates all possibilities of methylation state vectors with the same starting CpG sites and the same length (i.e., the set of CpG sites) as that given methylation state vector. Since each methylation state is typically either methylated or unmethylated, there are effectively two possible states at each CpG site. Therefore, the number of distinct possibilities for a methylation state vector can depend on powers of 2, such that a methylation state vector of length n will be associated with 2... n The probability of a methylation state vector. For a methylation state vector containing one or more CpG sites in an uncertain state, the analysis system may only consider those CpG sites with observed states to enumerate the probability of a methylation state vector.
[0139] The analysis system accesses the healthy control group data structure to calculate the probability of observing each methylation state vector for 350 possible occurrences, given the identified initial CpG sites and methylation state vector lengths. In some embodiments, Markov chain probabilities are used to model the joint probability calculation when calculating the probability of observing a given possibility. The Markov model can be trained, at least in part, based on the evaluation of the methylation state of each CpG site in the corresponding fragments (e.g., nucleic acid methylated fragments) of nucleic acid methylated fragments with corresponding CpG sites in the healthy non-cancer cohort dataset. For example, a Markov model (e.g., a hidden Markov model or HMM) is used to determine the probability that a certain methylation state (containing, for example, "M" or "U") sequence can be observed for one of the nucleic acid methylated fragments, given a set of probabilities (which, for each state in the sequence, determine the likelihood of observing the next state in the sequence). This set of probabilities can be obtained by training an HMM. Such training may involve computing statistical parameters (e.g., the probability that a first state can transition to a second state (transition probability) and / or the probability that a given methylation state can be observed at the corresponding CpG site (emission probability)) given an initial training dataset of observed methylation state sequences (e.g., methylation patterns). The HMM can be trained using supervised training (e.g., using samples where both the base sequence and observed states are known) and / or unsupervised training (e.g., Viterbi learning, maximum likelihood estimation, expectation-maximization training, and / or Baum-Welch training). In other embodiments, computational methods other than Markov chain probabilities are used to determine the probability of observing each methylation state vector. For example, such computational methods may include a learned representation. The p-value threshold may be between 0.01 and 0.10, or between 0.03 and 0.06. The p-value threshold may be 0.05. The p-value threshold may be less than 0.01, less than 0.001, or less than 0.0001.
[0140] The analysis system uses the calculated probability of each possibility to calculate a 355 p-value score for the methylation state vector. In some embodiments, this includes identifying the calculated probabilities corresponding to the possibilities matching the methylation state vector in question. Specifically, this could be possibilities having the same set of CpG sites as the methylation state vector or similarly having the same starting CpG sites and length as the methylation state vector. The analysis system can sum the calculated probabilities of any possibilities whose probabilities are less than or equal to the identified probabilities to generate the p-value score.
[0141] This p-value can represent the probability of observing a fragment's methylation state vector or other less likely methylation state vectors in a healthy control group. Therefore, a low p-value score typically corresponds to a methylation state vector that is rare in healthy subjects and causes the fragment to be labeled as aberrantly methylated relative to healthy controls. A high p-value score typically correlates with methylation state vectors expected to be present in relatively healthy subjects. For example, if the healthy control group is non-cancer, a low p-value could indicate that the fragment is aberrantly methylated relative to that non-cancer group, and thus may indicate the presence of cancer in the test subject.
[0142] As described above, the analysis system can calculate a p-value score for each of multiple methylation state vectors, each representing a cfDNA fragment in the test sample. To identify which fragments are aberrantly methylated, the analysis system may filter a set of 365 methylation state vectors based on their p-value scores. In some embodiments, screening is performed by comparing the p-value scores to a threshold and retaining only those fragments below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001, 0.0001, or similar orders of magnitude.
[0143] Based on the example results of process 300, the analysis system can generate a median (range) of 2,800 (1,500–12,000) fragments with aberrant methylation patterns in participants without cancer during training, and a median (range) of 3,000 (1,200–420,000) fragments with aberrant methylation patterns in participants with cancer during training. These selected fragment sets with aberrant methylation patterns can be used for downstream analyses as described in Sections III.B and III.C below.
[0144] In some embodiments, the analysis system uses a 360-degree sliding window to determine the probabilities of methylation state vectors and calculate p-values. The analysis system only needs to enumerate probabilities and calculate p-values for windows containing consecutive CpG sites, where the length of the window (CpG sites) is shorter than at least some segments (otherwise, the window would be meaningless), without needing to enumerate probabilities and calculate p-values for all methylation state vectors. The window length may be static, user-defined, dynamic, or otherwise selected.
[0145] When calculating the p-value for a methylation state vector larger than the window size, the window can be used to identify a set of consecutive CpG sites within the vector that are within the window, starting from the first CpG site in the vector. The analysis system can calculate a p-value score for the window that includes the first CpG site. The system can then "slide" the window to the second CpG site in the vector and calculate another p-value score for the second window. Therefore, for window size... l and methylation vector length mEach methylation state vector can generate m-l+1 Each p-value score. After calculating the p-value for each part of the vector, the lowest p-value score across all sliding windows can be used as the overall p-value score for that methylation state vector. In other embodiments, the analysis system aggregates the p-value scores of the methylation state vectors to generate an overall p-value score.
[0146] Using a sliding window can help reduce the number of possible methylation state vectors that need to be enumerated and the corresponding probability calculations that would otherwise be required. For example, a fragment might have up to 54 CpG sites. The analysis system could use, for instance, a window of size 5 to perform 50 p-value calculations for each of the 50 windows representing the fragment's methylation state vectors, instead of calculating the probabilities of 2^54 (approximately 1.8 × 10^16) possibilities to generate a single p-value. Each of these 50 calculations can enumerate 2^5 (32) possible methylation state vectors, totaling 50 × 2^5 (1.6 × 10^3) probability calculations. This significantly reduces the computations required without noticeably affecting the accurate identification of informative fragments.
[0147] In embodiments involving uncertain states, the analysis system can calculate a p-value score summing the CpG sites with uncertain states in the methylation state vector of a fragment. The analysis system can identify all possibilities that correspond to all methylation states in the methylation state vector except for the uncertain states. The analysis system can assign a probability to the methylation state vector as the sum of the probabilities of these identified possibilities. For example, the analysis system can calculate the methylation state vector...<M1,I2,U3> The probability of , as the methylation state vector<M1,M2,U3> and<M1,U2,U3> The sum of probabilities for the possible methylation states is calculated because the methylation states of CpG sites 1 and 3 have been observed and are consistent with the methylation states of the fragment at CpG sites 1 and 3. This method of summing probabilities for CpG sites with uncertain states can be used to calculate probabilities of up to 2^i possibilities, where i represents the number of uncertain states in the methylation state vector. In another embodiment, a dynamic programming algorithm can be used to calculate the probabilities of the methylation state vector with one or more uncertain states. Advantageously, dynamic programming algorithms run in linear computation time.
[0148] In some embodiments, caching at least some computations may further reduce the computational load of probabilities and / or p-value scores. For example, the analysis system may cache the probability calculations of methylation state vectors (or windows thereof) in transient or persistent memory. If other fragments have the same CpG sites, caching the probabilities of these possibilities can allow for efficient computation of p-value scores without recalculating the probabilities of potential possibilities. Equivalently, the analysis system can compute a p-value score for each possibility of a methylation state vector associated with a set of CpG sites in the vector (or window thereof). The analysis system can cache these p-value scores to determine the p-value scores of other fragments that include the same CpG sites. Overall, the p-value scores of the possibilities of methylation state vectors with the same CpG sites can be used to determine the p-value scores of different possibilities under the same set of CpG sites.
[0149] Before training a region model or cancer classifier, one or more nucleic acid methylation fragments can be screened. Screening can involve removing each corresponding nucleic acid methylation fragment from a plurality of corresponding nucleic acid methylation fragments that does not meet one or more selection criteria (e.g., below or above one selection criterion). One or more selection criteria can include a p-value threshold. The output p-value of a corresponding nucleic acid methylation fragment can be determined, at least in part, by comparing the corresponding methylation pattern of the corresponding nucleic acid methylation fragment with the distribution of corresponding methylation patterns of nucleic acid methylation fragments with corresponding multiple CpG sites in a healthy non-cancer cohort dataset.
[0150] Screening multiple nucleic acid methylated fragments can involve removing each corresponding nucleic acid methylated fragment that does not meet a p-value threshold. The screener is applied to the methylation pattern of each corresponding nucleic acid methylated fragment using the methylation patterns observed in the first batch of multiple nucleic acid methylated fragments. Each corresponding methylation pattern of each corresponding nucleic acid methylated fragment (e.g., fragment 1, ..., fragment N) can contain one or more corresponding methylation sites (e.g., CpG sites), identified by a methylation site identifier and a corresponding methylation pattern, represented as a sequence containing multiple 1s and multiple 0s, where each "1" represents a methylated CpG site in these one or more CpG sites, and each "0" represents an unmethylated CpG site in these one or more CpG sites. The methylation patterns observed in the first batch of multiple nucleic acid methylated fragments can be used to construct a methylation state distribution for these CpG site states, which are collectively represented by the first batch of multiple nucleic acid methylated fragments (e.g., CpG site A, CpG site B, ..., CpG site ZZZ). Further details regarding the processing of nucleic acid methylation fragments are disclosed in U.S. Patent Application No. 17 / 191,914, filed March 4, 2021, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," which is hereby incorporated herein by reference in its entirety.
[0151] When the aberrant methylation score of a corresponding nucleic acid methylation fragment is below an aberrant methylation score threshold, the fragment may not meet one or more selection criteria. In this case, the aberrant methylation score can be determined by a hybrid model. For example, a hybrid model can detect aberrant methylation patterns in a nucleic acid methylation fragment by determining the probability of a methylation state vector (e.g., methylation pattern) based on the number of possible methylation state vectors of the same length at the same corresponding genomic location. This process can be performed by generating multiple possible methylation states for a vector of a specified length at each genomic location in the reference genome. Using multiple possible methylation states, the total number of possible methylation states can be determined, and then the probability of each predicted methylation state at the genomic location can be determined. Then, by matching the sample nucleic acid methylation fragment with the predicted (e.g., possible) methylation states and retrieving the calculated probability of the predicted methylation state, the probability of the sample nucleic acid methylation fragment corresponding to the genomic location within the reference genome can be determined. The aberrant methylation score can then be calculated based on the probability of the sample nucleic acid methylation fragment.
[0152] When the number of residues in a corresponding nucleic acid methylation fragment is below a threshold number, the nucleic acid methylation fragment fails to meet one or more selection criteria. The threshold number of residues can be between 10 and 50, 50 and 100, 100 and 150, or more than 150. The threshold number of residues can also be a fixed value between 20 and 90. When the number of CpG sites in a corresponding nucleic acid methylation fragment is below a threshold number, the nucleic acid methylation fragment may fail to meet one or more selection criteria. The threshold number of CpG sites can be 4, 5, 6, 7, 8, 9, or 10. When the genomic start and genomic end positions of a corresponding nucleic acid methylation fragment indicate that the number of nucleotides represented by the nucleic acid methylation fragment in the human genome reference sequence is below a threshold number, the nucleic acid methylation fragment fails to meet one or more selection criteria.
[0153] Screening can remove nucleic acid methylation fragments from a plurality of corresponding nucleic acid methylation fragments that have the same corresponding methylation pattern, the same corresponding genomic start position, and the same genomic end position as another nucleic acid methylation fragment in the plurality of corresponding nucleic acid methylation fragments. This screening step can remove completely duplicated redundant fragments, including PCR duplicates in some cases. Screening can remove nucleic acid methylation fragments that have the same corresponding genomic start and end positions as another nucleic acid methylation fragment in the plurality of corresponding nucleic acid methylation fragments and have fewer than a threshold number of different methylation states. The threshold number of different methylation states used to retain a nucleic acid methylation fragment can be 1, 2, 3, 4, 5, or more than 5. For example, if a first nucleic acid methylation fragment has the same corresponding genomic start and end positions as a second nucleic acid methylation fragment, but has at least 1, at least 2, at least 3, at least 4, or at least 5 different methylation states at the corresponding CpG sites (e.g., aligned with the reference genome), then the first nucleic acid methylation fragment is retained. For example, if the first nucleic acid methylation fragment has the same methylation state vector (e.g., methylation pattern) as the second nucleic acid methylation fragment, but has different corresponding genomic start and end positions, then the first nucleic acid methylation fragment is also retained.
[0154] Screening can remove detection artifacts from multiple nucleic acid methylation fragments. Removal of detection artifacts may include removing sequence reads obtained from sequencing hybridization probes and / or sequence reads obtained from sequences that did not undergo bisulfite conversion. Screening can also remove contaminants (e.g., contaminants generated during sequencing, nucleic acid isolation, and / or sample preparation).
[0155] Screening can be based on mutual information of corresponding methylation fragments and cancer status in multiple training subjects, removing a subset of methylated fragments from a pool of methylated fragments. For example, mutual information can measure the degree of interdependence between two simultaneously sampled conditions of interest. Mutual information can be determined by selecting an independent set of CpG sites (e.g., within all or part of nucleic acid methylation fragments) from one or more datasets and comparing the probabilities of the methylation status of that set of CpG sites between two sample groups (e.g., genotype datasets, biological samples, and / or subsets and / or groups of subjects). The mutual information score can represent the probability of the methylation pattern of the first condition versus the second condition at a corresponding region within the corresponding frame of the sliding window, thus indicating the discriminative power of that region. Similarly, the mutual information score can be calculated for each region in each frame of the sliding window as the sliding window moves forward over the selected set of CpG sites and / or the selected genomic region. Further details regarding mutual information screening are disclosed in U.S. Patent Application 17 / 119,606, filed December 11, 2020, entitled “Cancer Classification Using Patched Convolutional Neural Networks,” which is hereby incorporated herein by reference in its entirety.
[0156] III.A.ii. Hypermethylated and hypomethylated fragments In some embodiments, the analysis system 370 identifies hypomethylated or hypermethylated fragments from the screened set as informative fragments. The analysis system identifies hypermethylated fragments that have more than a threshold number of CpG sites and that more than a threshold percentage of these CpG sites are methylated. The analysis system identifies hypomethylated fragments that have more than a threshold number of CpG sites and that more than a threshold percentage of these CpG sites are unmethylated. Instance thresholds for fragment (or CpG site) length include more than 3, 4, 5, 6, 7, 8, 9, 10, etc. Instance thresholds for the percentage of methylation or unmethylation include more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50%-100%.
[0157] III.B. Training of the Cancer Classifier Figure 4A This is a flowchart describing a process 400 for training a cancer classifier according to an embodiment. The analysis system acquires more than 410 training samples, each with an informative set of fragments and a cancer signal origin label. The multiple training samples can include any combination of samples from healthy subjects with an overall label of "non-cancer"; samples from subjects with an overall label of "cancer"; or samples with specific labels (e.g., "breast cancer," "lung cancer," etc.). Training samples from subjects targeting a specific cancer signal origin can be referred to as a cohort of that cancer signal origin or a cancer signal origin cohort.
[0158] The analysis system uses an informative set of training samples to determine a 420-feature vector for each training sample. The system can calculate an anomaly score for each CpG site in the initial CpG site set. This initial CpG site set can be all CpG sites in the human genome or a portion thereof—its order of magnitude can be 10-1. 4 10 5 10 6 10 7 10 8 In one embodiment, the analysis system uses a binary score to define anomaly scores for feature vectors based on the presence or absence of informative fragments in a set of informative fragments covering CpG sites. In another embodiment, the analysis system defines anomaly scores based on the count of informative fragments overlapping with CpG sites. In one instance, the analysis system might use a ternary score, assigning a first score to the absence of informative fragments, a second score to the presence of a few informative fragments, and a third score to the presence of multiple informative fragments. For example, the analysis system counts 5 informative fragments in samples overlapping with CpG sites and calculates anomaly scores based on this count of 5.
[0159] Once all anomaly scores have been determined for the training samples, the analysis system can define the feature vector as a vector containing multiple elements, each element being one of the anomaly scores associated with one of the CpG sites in the initial set. The analysis system can then normalize the anomaly scores of the feature vector based on the coverage of the samples. Here, coverage can refer to the median or average sequencing depth over all CpG sites covered by the initial set of CpG sites used in the classifier (or based on an informative fragment set of given training samples).
[0160] For example, now refer to Figure 4B The figure shows matrix 422 of the training feature vectors. In this example, the analysis system has identified CpG sites [K] 426 to be considered when generating the feature vectors for the cancer classifier. The analysis system selects training samples [N] 424. The analysis system determines a first anomaly score 428 for the first arbitrary CpG site [k1], which will be used in the feature vectors of the training samples [n1]. The analysis system examines each informative fragment in the informative fragment set. If the analysis system identifies at least one informative fragment that includes the first CpG site, the analysis system sets the first anomaly score 428 of the first CpG site to 1, as shown below. Figure 4BAs shown. Considering a second arbitrary CpG site [k2], the analysis system similarly examines at least one informative fragment in the informative fragment set that includes the second CpG site [k2]. If the analysis system does not find any such informative fragment including the second CpG site, the analysis system determines the second anomaly score 429 of the second CpG site [k2] to be 0, as... Figure 4B As shown. Once the analysis system has determined all the anomaly scores of the initial CpG site set, the analysis system determines a feature vector for the first training sample [n1]. This feature vector includes the anomaly scores, where the first anomaly score 428 of the first CpG site [k1] is 1, the second anomaly score 429 of the second CpG site [k2] is 0, and the subsequent anomaly scores are added to form the feature vector [1, 0, …].
[0161] Other methods for sample characterization can be found in the following documents: U.S. Patent Application No. 15 / 931,022, entitled "Model-Based Featureization and Classification"; U.S. Patent Application No. 16 / 579,805, entitled "Mixture Model for Targeted Sequencing"; U.S. Patent Application No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification"; and U.S. Patent Application No. 16 / 723,716, entitled "Source of Origin Deconvolution Based on Methylation Fragments in Cell-Free DNA Samples"; all of these documents are incorporated herein by reference in their entirety. Various characterization methods can generate different features to be included in the feature vector of the sample.
[0162] In some embodiments, each classifier may be trained on a different training feature vector matrix. For example, one classifier may be trained on a first matrix covering a first set of genomic regions, while another classifier may be trained on a second matrix covering a second, different set of genomic regions. In another instance, one classifier may be trained on a first matrix having features determined according to one featureization method, while another classifier may be trained on a second matrix having features determined according to another featureization method.
[0163] The analysis system may further limit the CpG sites considered for use in the cancer classifier. Based on the feature vectors of the training samples, the analysis system calculates an information gain of 430 for each CpG site in the initial CpG site set. According to step 420, each training sample has a feature vector that may contain anomaly scores for all CpG sites in the initial CpG site set, which may include up to all CpG sites in the human genome. However, some CpG sites in the initial CpG site set may be less informative than others in distinguishing between multiple cancer signaling origins, or the information they provide may overlap with that of other CpG sites.
[0164] In one embodiment, the analysis system calculates an information gain of 430 for each CpG site in each cancer signal origin and initial set to determine whether to include that CpG site in the classifier. Information gains are calculated for training samples given a cancer signal origin, compared to all other training samples. For example, two random variables, “informative fragment” (“IF”) and “cancer signal” (“CS”), are used. In one embodiment, IF is a binary variable indicating whether an informative fragment overlapping with a given CpG site exists in a given sample (as determined above with an anomalous score / eigenvector). CS is a random variable indicating whether the cancer signal has a specific origin, such as a specific organ or organoid group, or a specific cancer biology. The analysis system calculates the mutual information relative to CS given an IF. That is, if the existence of an informative fragment overlapping with a specific CpG site is known, the number of information bits about that cancer signal origin can be obtained. In practice, for a first cancer signal origin, the analysis system calculates pairwise mutual information gains with each of the other cancer signal origins and sums the mutual information gains of all other cancer signal origins.
[0165] For a given cancer signal origin, the analysis system can use this information to rank these sites according to their cancer-specificity. This process can be repeated for all cancer signal origins considered. If a particular region is typically aberrantly methylated in training samples for a given cancer but not in training samples for other cancer signal origins or healthy training samples, then for that given cancer signal origin, the CpG sites whose informative fragments overlap can have high information gain. The ranked CpG sites for each cancer signal origin are then added (selected) 440 to a selected set of CpG sites using a greedy algorithm, based on their ranking in the cancer classifier.
[0166] In another embodiment, the analysis system may consider other selection criteria for selecting informative CpG sites to be used in the cancer classifier. One selection criterion could be that the distance between a selected CpG site and other selected CpG sites is greater than a threshold. For example, a selected CpG site is more than a threshold number of base pairs (e.g., 100 base pairs) away from any other selected CpG site, such that CpG sites within the threshold distance are not also selected for consideration in the cancer classifier.
[0167] In one embodiment, the analysis system may modify the feature vectors of 450 training samples as needed, based on a selected set of CpG sites from an initial set. For example, the analysis system may truncate the feature vectors to remove outlier scores corresponding to CpG sites not in the selected set of CpG sites.
[0168] Using the feature vectors of the training samples, the analysis system can train a cancer classifier in any of a number of ways. The feature vectors may correspond to the initial CpG site set in step 420 or the selected CpG site set in step 450. In one embodiment, the analysis system trains a binary cancer classifier (460) based on the feature vectors of the training samples to distinguish between cancer and non-cancer samples. In this way, the training samples used by the analysis system include non-cancer samples from healthy subjects and cancer samples from subjects. Each training sample may have one of two labels: "cancer" or "non-cancer." In this embodiment, the classifier outputs a cancer prediction that indicates the likelihood of the presence or absence of cancer.
[0169] In another embodiment, the analysis system trains a 470-class cancer classifier to distinguish multiple cancer signaling origins (also known as CSO labels). A CSO can include one or more organs, organ groups, cancer biological categories, cell lineages of cancer cells, or cancer drivers such as viral states. For this purpose, the analysis system can use a cohort of cancer signaling origins, and may or may not include a non-cancer cohort. In this multi-class embodiment, the cancer classifier is trained to determine a cancer prediction (or, more specifically, a CSO prediction) that includes a predicted value for each of the cancer signaling origins being classified. The predicted values may correspond to the likelihood that a given training sample (or, during inference, the test sample) has each of the cancer signaling origins. In one implementation, the predicted value score is between 0 and 100, where the sum of the predicted values equals 100. For example, the cancer classifier returns a cancer prediction that includes predicted values for the affected organ or organ group of breast cancer, lung cancer, and non-cancer origins. For example, the classifier may return a cancer prediction that predicts the test sample's signaling origin has a 65% likelihood of breast cancer, a 25% likelihood of lung cancer, and a 10% likelihood of non-cancer origins. The analysis system can further evaluate the predicted values to generate predictions of the presence of one or more cancers in the sample, also known as CSO predictions indicating one or more CSO tags (e.g., the first CSO prediction with the highest predicted value, the second CSO prediction with the second highest predicted value, etc.). Continuing with the example above and given these percentages, in this example, given that the CSO for breast cancer has the highest likelihood, the system might determine that the origin of the cancer signal in the sample is breast cancer.
[0170] In both embodiments, the analysis system trains the cancer classifier by inputting a training sample set containing its feature vectors into the cancer classifier and adjusting the classification parameters so that the classifier's function accurately associates the training feature vectors with their corresponding labels. The analysis system may group the training samples into one or more sets for iterative batch training of the cancer classifier. After inputting all training sample sets including its training feature vectors and adjusting the classification parameters, the cancer classifier can be sufficiently trained to label test samples based on their feature vectors within a certain error bound. The analysis system can train the cancer classifier using any of many methods. For example, a binary cancer classifier could be an L2-regularized logistic regression classifier trained using a logarithmic loss function. Another example is a multinomial logistic regression classifier. In practice, other techniques can be used to train both types of cancer classifiers. These techniques are numerous and include kernel methods, random forest classifiers, mixture models, autoencoder models, machine learning algorithms (such as multilayer neural networks), etc. In some embodiments, such a model may include 1,000 or more parameters, 2,000 or more parameters, 3,000 or more parameters, 4,000 or more parameters, 5,000 or more parameters, 10,000 or more parameters, 15,000 or more parameters, 20,000 or more parameters, 25,000 or more parameters, 50,000 or more parameters, 100,000 or more parameters, 150,000 or more parameters, 200,000 or more parameters, 250,000 or more parameters, 300,000 or more parameters, 350,000 or more parameters, 400,000 or more parameters, 450,000 or more parameters, or 500,000 or more parameters.
[0171] Classifiers can include logistic regression, neural network, support vector machine, naive Bayes, nearest neighbor, reinforcement tree, random forest, decision tree, multinomial logistic regression, linear model, or linear regression.
[0172] III.C. Deployment of Cancer Classifiers During the use of the cancer classifier, the analysis system can acquire test samples from subjects whose cancer signaling origin is unknown. The analysis system may use any combination of procedures 200 and 330 to process the test samples, which consist of DNA molecules, to obtain an informative fragment set. Based on similar principles discussed in procedure 400, the analysis system can determine a test feature vector for use by the cancer classifier. The analysis system can calculate anomaly scores for each of the multiple CpG sites used by the cancer classifier. For example, the cancer classifier receives an input feature vector including anomaly scores for 1,000 selected CpG sites. Therefore, based on the informative fragment set, the analysis system can determine a test feature vector including anomaly scores for those 1,000 selected CpG sites. The analysis system can calculate the anomaly scores in the same manner as with training samples. In some embodiments, the analysis system defines the anomaly score as a binary score based on the presence of hypermethylated or hypomethylated fragments in the informative fragment set covering CpG sites.
[0173] The analysis system can then input the test feature vector into a cancer classifier. The cancer classifier's function, based on the classification parameters trained in process 400 and the test feature vector, can then generate a cancer prediction. In a first approach, the cancer prediction can be binary, selected from a group consisting of "cancer" or "non-cancer"; in a second approach, the cancer prediction is selected from a group of potential cancer signaling origins that can identify the affected organ or organ group and simultaneously identify the presence of potential cancer biology (such as histological type, cancer cell lineage of origin, or cancer drivers (such as oncogenic viral infection)) and "non-cancer". In another embodiment, the cancer prediction has a prediction value for each of these multiple cancer signaling origins. Furthermore, the analysis system can determine that the cancer signal of the test sample is most likely to have one of the cancer signaling origins. Continuing the example above, if the cancer prediction for a test sample is 65% likely to originate from the breast organ, 25% likely to originate from the lung organ, and 10% likely to be non-cancer, the analysis system might determine that the test sample is most likely to have cancer originating from the breast. In another example, where the cancer prediction is binary, with a 60% probability of non-cancer and a 40% probability of cancer, the analysis system determines that the test sample is most likely to lack a cancer signal. In yet another embodiment, to identify whether a test subject has a cancer signal or the origin of that cancer signal, the most probable cancer prediction may still be compared to a threshold (e.g., 40%, 50%, 60%, 70%). If the cancer prediction with the highest probability does not exceed this threshold, the analysis system may return an indeterminate result or report the presence of a cancer signal without reporting the origin of the cancer signal. In another embodiment, the analysis system concatenates the cancer classifier trained in step 460 of process 400 with one or more additional cancer classifiers trained in step 470 of process 400. The analysis system can input the test feature vector into the cancer classifier trained as a binary classifier in step 460 of process 400. The analysis system can receive the output of cancer predictions. The cancer predictions can be binary regarding whether the test subject is likely to have or not have cancer and whether a cancer signal is detected in the sample. In other embodiments, the cancer predictions include predicted values describing cancer likelihood and non-cancer likelihood. For example, the cancer prediction has 85% cancer prediction value and 15% non-cancer prediction value. The analysis system may determine that the test subject is likely to have cancer. Once the analysis system determines that the test subject is likely to have cancer, the analysis system can input the test feature vector into a multi-class cancer classifier trained to distinguish different cancer signal origins. The multi-class cancer classifier can receive the test feature vector and return one or more cancer signal predictions for the origin of multiple cancer signals in one or more cancer signal origin categories. For example, a multi-class cancer classifier provides a cancer prediction that specifies the test subject is most likely to have a cancer signal originating from a predicted organ or group of organs (such as ovarian cancer or fallopian tube cancer), while also providing a cancer signal origin prediction that has cells of the Miller lineage as the origin. In another implementation, the multi-class cancer classifier provides a predicted value for each cancer signal origin in multiple cancer signal origins and cancer signal origin classifiers. For example, the cancer prediction may include, for organs in an organ group, a cancer signal origin of 40% for the head and neck, a cancer signal origin of 20% for the anus, and a cancer signal origin of 20% for the cervix, and in parallel, the following predictions: 80% of cancer signals originate from cancers associated with human papillomavirus (HPV) infection, 10% from squamous cell carcinoma, and 10% from cancers originating from Miller lineage cells.
[0174] According to a common embodiment of binary cancer classification, the analysis system can determine a cancer score for a test sample based on its sequencing data (e.g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.). The analysis system can compare the cancer score of the test sample with a binary threshold that predicts whether the test sample is likely to have cancer. The binary threshold cutoff can be tuned using CSO thresholding based on one or more CSO predictions. The analysis system can further generate a feature vector of the test sample for use in a multi-class cancer classifier to determine a cancer signal origin prediction that indicates one or more possible organs or groups of organs as cancer signal origins, and in parallel indicates one or more histological types, cell lineages, or oncogenic drivers as cancer signal origins.
[0175] A classifier can be used to determine the disease status of a test subject (e.g., a subject with an unknown disease status). The method may include acquiring a test genome data construct in electronic format (e.g., single-timepoint test data) that includes the value of each of multiple genomic features corresponding to multiple nucleic acid fragments from a biological sample obtained from the test subject. The method may then include applying the test genome data construct to a test classifier to determine the status of the test subject's disease condition. The test subject may not have been previously diagnosed with this disease condition.
[0176] The classifier may be a time classifier that uses at least (i) a first test genome data construct generated from a first biological sample obtained from the test subject at a first time point and (ii) a second test genome data construct generated from a second biological sample obtained from the test subject at a second time point.
[0177] A trained classifier can be used to determine the disease status of a test subject (e.g., a subject with an unknown disease status). In this case, the method may include acquiring a test time-series dataset of the test subject in electronic format, wherein: for each corresponding time point among a plurality of time points, the test time-series dataset includes a corresponding test genotype data construct, which includes values of multiple genotype features of corresponding multiple nucleic acid fragments in a corresponding biological sample obtained from the test subject at that corresponding time point; for each pair of corresponding consecutive time points among the plurality of time points, the test time-series dataset includes an indication of the time length between the pair of corresponding consecutive time points. The method may then include applying the test genotype data construct to a test classifier to determine the status of the test subject's disease condition. The test subject may not have been previously diagnosed with this disease condition.
[0178] IV. Parallel Cancer Origin Classifier The analysis system can implement a cancer signal origin classifier to predict cancer signal origin characteristics. The cancer signal origin classifier can be based on... Figure 4A An example of a trained cancer classifier. Cancer can be characterized based on its cancer signaling origin features. For example, cancer signaling origins can include the organ or organ group primarily affected by the cancer and / or the tumor biology type indicating how the cancer behaves, or what its cellular lineage or oncogenic drivers are. Cancer signaling origins can be further analyzed based on the primarily affected cell types. Each type (organ or organ group, cell type, and tumor biology type) can be an independent variable and obtain independent predictions; for example, an organ or organ group can cross with one of many tumor biology signaling origins, etc.
[0179] The results of a cancer origin classifier can be used to inform examination procedures to examine the origin of detected cancer signals in order to obtain a diagnosis of cancer or a confidence assessment that the cancer signal detection was a false positive and the patient does not have cancer. In some embodiments, the results of the cancer origin classifier are provided to healthcare professionals to inform them of diagnostic examination options for the detected cancer signals. In other embodiments, the analysis system may identify one or more diagnostic examination options based on the results of the cancer origin classifier to recommend to healthcare professionals. The results of the cancer origin classifier can provide additional details for cancer prediction, thereby allowing diagnostic examinations to be better tailored to the subject's cancer.
[0180] IV.A. Organ type and tumor biological type In one or more embodiments, organ type and tumor biology type can be orthogonal. For example, any organ or organ group signaling origin can be paired with many tumor biology signaling origins. In other embodiments, additional cancer signaling origin features (e.g., cell type) may be used.
[0181] In one or more embodiments, the organ or group of organ signal origins may include: breast; prostate; lung; head or neck; anus; cervix; ovary or fallopian tube; uterus; bladder or urothelial tissue; kidney; stomach or esophagus; liver or intrahepatic bile duct; pancreas, extrahepatic bile duct or gallbladder; colon or rectum; bone or soft tissue; skin; blood, lymphatic system or bone marrow; thyroid gland; fuzzy or absent cancer signal origins; or some combinations thereof. In one or more embodiments, some organ types may be further subdivided; for example, a stomach or esophagus signal origin may be divided into a signal originating from the stomach and a signal originating from the esophagus.
[0182] In one or more embodiments, cancers with a signaling origin in the breast include cancers such as invasive ductal breast cancer, non-specific type (NST) breast cancer, invasive lobular breast cancer, or some combination thereof. Examples of cancers that may have different signaling origins include sarcomas (with signaling origins in bone or soft tissue) or lymphomas (with signaling origins in blood, lymphatic system, or bone marrow) reported in the breast, any invasive skin cancer of the breast (with signaling origins in the skin), Paget's disease, or some combination thereof.
[0183] In one or more embodiments, the cancer signaling origin of the prostate includes cancers such as invasive ductal prostatic adenocarcinoma, invasive acinar adenocarcinoma of the prostate, small cell prostatic carcinoma, or some combination thereof. Examples of cancers that may have different cancer signaling origins include sarcomas (with cancer signaling origins in bone or soft tissue) or lymphomas (with cancer signaling origins in bone, lymphatic system, or bone marrow) reported in the prostate.
[0184] In one or more embodiments, the origin of cancer signals in the lung includes cancers such as lung adenocarcinoma, lung squamous cell carcinoma, non-small cell lung cancer (NSCLC NOS) not otherwise specified, small cell lung cancer (SCLC), lung carcinoid, or combinations thereof. Examples of cancers that may have different origins of cancer signals include sarcomas or lymphomas reported in the lung, or combinations thereof.
[0185] In one or more embodiments, the cancer signaling origin in the head or neck includes cancers such as oropharyngeal human papillomavirus-associated (HPV-associated) squamous cell carcinoma, laryngeal HPV-negative squamous cell carcinoma, salivary gland adenocarcinoma, or some combination thereof. Examples of cancers that may have different cancer signaling origins include sarcomas or lymphomas reported in the head and neck region, skin cancers reported in the head and neck region, or some combination thereof.
[0186] In one or more embodiments, the cancer signaling origin of the anus may include cancers such as anal HPV-associated squamous cell carcinoma, anal glandular carcinoma, HPV-positive squamous cell carcinoma reported in the primary rectum, or combinations thereof. Examples of cancers that may have different cancer signaling origins include skin cancers reported in the anal region, extramammary Paget's disease in the region, squamous cell carcinomas reported in the colon (even if HPV positive), sarcomas or lymphomas reported in the anal region, or combinations thereof.
[0187] In one or more embodiments, the cancer signal originating from the cervix may include cancer, such as cervical HPV-associated squamous cell carcinoma, cervical HPV-associated adenocarcinoma, cervical neuroendocrine carcinoma, cervical non-HPV-associated adenocarcinoma, or some combination thereof.
[0188] In one or more embodiments, the origin of cancer signals in the ovary or fallopian tube can include cancers such as fallopian tube-derived serous cystadenocarcinoma of the ovary, endometrioid ovarian cancer, malignant Müllerian mixed carcinoma of the ovary, clear cell carcinoma of the ovary, small cell carcinoma of the ovary, malignant Brenner tumor, or combinations thereof. Examples of cancers that may have different origins of cancer signals include germ cell tumors, sarcomas or lymphomas reported in the ovarian or peritoneal region, mesotheliomas in the peritoneal region, or combinations thereof.
[0189] In one or more embodiments, the origin of cancer signals in the uterus can include cancers such as endometrial cancer, uterine carcinosarcoma, high-grade serous cystadenocarcinoma reported as the primary site in the uterus, or combinations thereof. Examples of cancers with different origins of cancer signals include germ cell tumors, uterine sarcomas, endometrial stromal tumors, or combinations thereof.
[0190] In one or more embodiments, the origin of cancer signaling in bladder cancer or urothelial carcinoma can include cancers such as bladder adenocarcinoma, bladder transitional cell carcinoma, urothelial carcinoma of the ureter or renal pelvis, transitional or renal cell carcinoma reported in the kidney, urothelial carcinoma of the renal pelvis, ureteral cancer, small cell carcinoma of the bladder, adenocarcinoma NOS in the renal pelvis, or some combinations thereof. Examples of cancers that may have different origins of cancer signaling include sarcomas or lymphomas reported in the bladder, ureter, or renal pelvis, urothelial carcinoma of the kidney, or some combinations thereof.
[0191] In one or more embodiments, the origin of cancer signaling in the kidney includes renal cell carcinoma, renal carcinoid, renal adenocarcinoma (NOS), or some combination thereof. Examples of cancers that may have different origins of cancer signaling include sarcomas or lymphomas reported in the kidney, transitional or urothelial carcinomas reported in the kidney, or some combination thereof.
[0192] In one or more embodiments, the origin of cancer signaling in the stomach or esophagus may include gastric adenocarcinoma, esophageal adenocarcinoma, esophageal squamous cell carcinoma, small cell gastric carcinoma, gastric carcinoid, or combinations thereof. Examples of cancers that may have different origins of cancer signaling include gastrointestinal stromal tumors (GIST), gastric mucosa-associated lymphoid tissue (MALT) lymphoma, small bowel cancer, or combinations thereof.
[0193] In one or more embodiments, the origin of cancer signals in the liver or intrahepatic bile ducts includes cancers such as hepatocellular carcinoma, intrahepatic cholangiocarcinoma, small hepatocellular carcinoma, hepatic carcinoid, or combinations thereof. Examples of cancers that may have different origins of cancer signals include hilar cholangiocarcinoma, reported liver sarcomas or lymphomas, or combinations thereof.
[0194] In one or more embodiments, cancer signaling originating from the pancreas, extrahepatic bile duct, or gallbladder includes cancers such as pancreatic ductal adenocarcinoma, gallbladder adenocarcinoma, extrahepatic bile duct carcinoma, cystic duct bile duct carcinoma, pancreatic neuroendocrine tumors, or combinations thereof. Examples of cancers that may have different cancer signaling origins include sarcomas or lymphomas reported in the pancreas, gallbladder, or bile ducts, intrahepatic bile duct carcinoma, or combinations thereof.
[0195] In one or more embodiments, the origin of cancer signals in the colon or rectum can include cancers such as colorectal adenocarcinoma, colonic signet ring cell carcinoma, colonic large cell neuroendocrine carcinoma, colonic carcinoid, appendiceal adenocarcinoma, appendiceal carcinoid, colonic squamous cell carcinoma, or combinations thereof. Examples of cancers that may have different origins of cancer signals include HPV-positive squamous cell carcinoma reported in the rectum, sarcomas or lymphomas reported in the colon or rectum, small intestinal adenocarcinoma or carcinoid, or combinations thereof.
[0196] In one or more embodiments, the bone or soft tissue from which the cancer signal originates may include cancers such as uterine leiomyosarcoma, malignant solitary fibroma, osteosarcoma, gastrointestinal stromal tumor, (cutaneous) malignant fibrous histiocytoma, malignant hemangiopericytoma of the brain, or some combination thereof. Examples of cancers that may have different cancer signal origins include myeloid sarcoma.
[0197] In one or more embodiments, the origin of cancer signals in the skin can include cancers such as melanoma of the extremities, melanoma of the head and neck region, Merkel cell carcinoma, papillary carcinoma of the fingers and toes, or combinations thereof. Examples of cancers that may have different origins of cancer signals include basal cell carcinoma of the skin (unless metastatic), squamous cell carcinoma of the skin (unmetastatic), lymphoma of the skin, or combinations thereof.
[0198] In one or more embodiments, cancer signaling origins in the blood, lymphatic system, or bone marrow include cancers such as lymphomas (including gastrointestinal CNS lymphomas and MALT lymphomas), lymphoid leukemias, myeloid leukemias, multiple myelomas, or plasma cell myelomas, or combinations thereof. Examples of cancers with different cancer signaling origins include hematologic precursor diseases.
[0199] In one or more embodiments, the cancer signaling origin of the thyroid gland can include cancers such as medullary thyroid carcinoma, papillary thyroid carcinoma, or combinations thereof. Examples of cancers that may have different cancer signaling origins include sarcomas or lymphomas reported in the thyroid gland.
[0200] In one or more embodiments, there may be no cancer signal or a vague cancer signal origin assigned to a cancer, such as mesothelioma, small bowel cancer, penile cancer, vulvar cancer or vaginal cancer, clinically unknown primary cancer, multiple primary cancers, brain and spinal cord cancer (excluding sarcoma and lymphoma), or some combination thereof.
[0201] In one or more embodiments, the cancer signaling origin categories in tumor biology include categories identifying tumor cells or cell lineages of origin, tumor histological types, or carcinogenic drivers. Cancer signaling origins identifying cell lineages include lymphoid vegetations, medullary vegetations, plasma cell vegetations, neuroendocrine carcinomas or tumors, vegetations originating from Müllerian ducts, mesenchymal tumors, melanocyte vegetations, and mesothelial vegetations. Cancer signaling origins identifying histological types include adenocarcinoma, squamous cell carcinoma (non-HPV-related), hepatocellular carcinoma, and transitional cell carcinoma. Cancer signaling origins identifying carcinogenic drivers include HPV-related carcinomas. Cancer signaling origins identifying any other tumor biology may also exist, while some cases may have no tumor biology or have a vague tumor biology, or be assigned some combination thereof.
[0202] In one or more embodiments, cancer signaling origin lymphoid vegetations include cancers such as Hodgkin and non-Hodgkin lymphoma, T-cell lymphoma, B-cell lymphoblastic lymphoma (BLL), small lymphocytic lymphoma (SLL), primary cutaneous follicular center lymphoma, precursor B and T-cell lymphoblastic leukemia, gastric mucosa-associated lymphoid tissue lymphoma, or some combinations thereof. Examples of cancers that can be excluded include malignant pre-hematologic disorders. Although plasma cells differentiate from B lymphocytes, all malignant vegetations of plasma cell origin have alternative cancer signaling origins as plasma cell vegetations.
[0203] In one or more embodiments, the cancer signaling origin of the myeloid vegetation includes cancers such as acute myeloid leukemia (AML), chronic myeloid leukemia (CML), myelodysplastic syndrome (MDS), malignant mastocytosis, myeloid sarcoma, acute erythroid leukemia, or some combination thereof. Examples of cancers with different cancer signaling origins may include malignant pre-hematologic disorders.
[0204] In one or more embodiments, the cancer signaling origin of plasma cell vegetations includes multiple myeloma, plasma cell myeloma, or combinations thereof. Examples of cancers that can be excluded include malignant pre-hematologic disorders.
[0205] In one or more embodiments, the cancer signal originating from neuroendocrine carcinoma includes cancers such as SCLC, small cell prostate carcinoma, large cell colonic neuroendocrine carcinoma, medullary thyroid carcinoma, typical and atypical carcinoids (functional or non-functional), neuroendocrine carcinoma NOS, pancreatic neuroendocrine tumors, collisional tumors or complex small cell carcinomas of SCLC and lung adenocarcinoma, adenocarcinoma with neuroendocrine differentiation, mixed neuroendocrine-nonneurocrine vegetations containing low-grade neuroendocrine components, or some combination thereof.
[0206] In one or more embodiments, the cancer signaling origin of adenocarcinoma includes cancers such as colorectal adenocarcinoma, bronchoalveolar or minimally invasive lung adenocarcinoma, bladder adenocarcinoma (if not reported as transitional cell carcinoma), gastric signet ring cell carcinoma, pancreatic ductal adenocarcinoma, acinar prostate cancer, cholangiocarcinoma of intrahepatic or extrahepatic bile ducts, thyroid follicular carcinoma, salivary gland adenocarcinoma, or some combination thereof. Examples of cancers that may have different cancer signaling origins include HPV-positive adenocarcinoma (whose cancer signaling origin is HPV-associated carcinoma), bladder transitional cell carcinoma, and adenocarcinoma derived from cells of embryonic origin in the Müllerian duct. Mixed or collisional tumors between adenocarcinoma and high-grade neuroendocrine carcinomas with a cancer signaling origin neuroendocrine tumor or carcinoma, including subtypes of a lineage in other class specifications, or adenocarcinomas with a cancer signaling origin representing that lineage.
[0207] In one or more embodiments, squamous cell carcinoma (non-HPV-related) with cancer signaling origins includes cancers such as keratinized and non-keratinized squamous cell carcinoma of the bronchi or lungs, HPV-negative squamous cell carcinoma of the larynx, verrucous carcinoma of the urothelial tract, squamous cell carcinoma of the esophagus, basal cell carcinoma of the salivary glands, or combinations thereof. Examples of cancers with different cancer signaling origins include HPV-positive squamous cell carcinomas because their cancer signaling origin is HPV-related.
[0208] In one or more embodiments, HPV-related cancers with cancer signaling origins include cancers such as HPV-positive oropharyngeal squamous cell carcinoma, cervical HPV-positive SCC, anal HPV-positive SCC, cervical HPV-positive adenocarcinoma, or combinations thereof. Examples of cancers that may have different cancer signaling origins include laryngeal HPV-negative squamous cell carcinoma, squamous cell carcinoma with no clinical or molecular evidence of HPV status in the H&N region, squamous cell carcinoma with a definitively negative HPV test result in the cervix or anus, squamous cell carcinoma of the skin in the anal region, non-HPV-related cancers in patients with active HPV infection, or combinations thereof.
[0209] In one or more embodiments, the cancer signaling origin of hepatocellular carcinoma includes cancers such as hepatocellular carcinoma and all its subtypes. Examples of cancers that may have different cancer signaling origins include intrahepatic cholangiocarcinoma, cancers in the liver not identified as hepatocellular carcinoma, or some combination thereof.
[0210] In one or more embodiments, the cancer signaling origin vegetations originating from the Müllerian ducts include cancers such as ovarian serous cystadenocarcinoma, endometrioid adenocarcinoma of the ovary or uterus, clear cell carcinoma of the ovary or uterus, malignant mixed Müllerian duct tumors, uterine carcinosarcoma, cervical HPV-negative adenocarcinoma, or combinations thereof. Examples of cancers that may have different cancer signaling origins include cervical HPV-positive squamous cell carcinoma, cancers reported in the ovary or uterus from different cell lineages (e.g., uterine leiomyosarcoma), germ cell carcinoma, or combinations thereof.
[0211] In one or more embodiments, the cancer signaling origin of transitional cell carcinoma includes cancers such as bladder transitional cell carcinoma, renal pelvis urothelial carcinoma, ureteral urothelial carcinoma, urothelial carcinoma or transitional cell carcinoma reported in the kidney, or some combination thereof. Examples of cancers that may have different cancer signaling origins include bladder adenocarcinoma (if not reported as transitional cell carcinoma), kidney or renal pelvis cancer not reported as urothelial carcinoma or transitional cell carcinoma, or some combination thereof.
[0212] In one or more embodiments, the cancer signaling origin of the mesenchymal tumor includes cancers such as sarcomas in muscle or connective tissue, uterine leiomyosarcoma, malignant solitary fibroma, osteosarcoma, gastrointestinal stromal tumors, malignant fibrous histiocytoma of the skin, malignant hemangiopericytoma of the brain, or combinations thereof. Examples of cancers that may have different cancer signaling origins include myeloid sarcoma or other myeloid vegetation cancers.
[0213] In one or more embodiments, the cancer signaling origin of melanocyte vegetations includes cancers such as cutaneous melanoma, mucosal melanoma of the head and neck region, conjunctival and uveal melanoma, amelanoma, or combinations thereof. Examples of cancers that may have different cancer signaling origins include Merkel cell carcinoma, sweat gland carcinoma, basal cell carcinoma of the skin, squamous cell carcinoma of the skin, or combinations thereof.
[0214] In one or more embodiments, the cancer signaling origin of the mesothelial vegetation includes cancers such as pleural mesothelioma, peritoneal mesothelioma, or some combination thereof. Examples of cancers that may have different cancer signaling origins include pleural carcinomas that are not of mesothelial origin.
[0215] In one or more embodiments, other categories of cancer signal origin include cancers such as germ cell tumors, anaplastic carcinomas, pleomorphic carcinomas, medullary carcinomas (if not thyroid cancer), undifferentiated carcinomas, mucoepidermoid carcinomas, astrocytomas, glioblastomas, lymphoepithelial carcinomas, carcinosarcomas (if not of Müllerian duct origin), or some combination thereof.
[0216] In one or more embodiments, some cases may not have a cancer signaling origin in cancer biology, or may have an ambiguous cancer signaling origin. Cancers in this group include cancer NOS, malignant vegetation NOS, NSCLC NOS, mixed hepatocellular carcinoma and cholangiocarcinoma, adenosquamous carcinoma, biphenotypic leukemia, invasive ductal carcinoma mixed with other types of cancer, cancers with mixed subtypes, or collisional tumors. However, collisional tumors of SCLC and lung adenocarcinoma, or mixed small cell carcinomas, may have a cancer signaling origin in neuroendocrine tumors or mixed tumors of carcinoma, adenocarcinoma, or squamous cell carcinoma, and low-grade neuroendocrine tumors may have a cancer signaling origin in the epithelial component, and mixed tumors with only one malignant component (e.g., adenosarcoma) may have a target cancer signaling origin in that malignant component, other mixed tumors, or some combination thereof.
[0217] IV.B. Training of Parallel CSO Classifiers The analysis system trains a parallel CSO classifier. Training typically requires using training data to build a predictive model that accurately predicts the origin of cancer signals. Training data can be derived from training samples collected from subjects with different cancer diagnoses. In some embodiments, a parallel CSO classifier is deployed in response to the detection of a cancer signal in a sample. In such embodiments, the training data for the parallel CSO classifier can be derived from cancer subjects with different cancer signal origins. Each cancer signal origin can be characterized by CSO features of an organ or group of organs and a tumor biological type. In one or more embodiments, when training the CSO classifier, the analysis system can obtain two or more training datasets, wherein at least one of the training datasets includes CSO features for training one type of CSO classifier. For example, the analysis system can receive a general training dataset that can be used to train any CSO classifier, and a second specific training dataset that can be used to train one type of CSO classifier. In another instance, the analysis system can utilize a first training dataset specific to one type of CSO classifier and a second training dataset specific to another type of CSO classifier. The training data can be used to generate multiple training sets, each used to train a CSO classifier. During training, the analysis system can cross-validate the trained CSO classifier to verify its predictive accuracy. The advantage of trained CSO classifiers is that they learn patterns in the training data that indicate various CSO categories. In particular, the parallel training process increases the granularity of CSO predictions by predicting organs or organ groups individually from tumor biology. This increased granularity provides deeper insights when cancer signals are detected in a sample, informing diagnostic steps for cancer assessment.
[0218] Figure 5A This is an exemplary flowchart illustrating process 500 for training a parallel Cancer Origin (CSO) classifier according to one or more embodiments. An analysis system may execute process 500. In other embodiments, another computing device or system may execute any, some, or all of the steps of process 500. Figure 5A In the embodiment shown, the analysis system trains two parallel CSO classifiers, one for predicting CSO organs or organ groups and the other for predicting CSO tumor biological categories.
[0219] The analysis system obtains 505 cancer samples for training a parallel CSO classifier. Each cancer sample is derived from a subject with a positive diagnosis of cancer. Cancer samples may include methylation sequencing data (e.g., obtained via WGBS or targeted methylation assays) and CSO tags, which include known organs or organ groups and known tumor biological categories. In other embodiments, each cancer sample may include additional CSO tags for other CSO features. In some embodiments, the analysis system receives a cancer diagnosis for each cancer sample. Based on the cancer diagnosis, the analysis system can resolve the CSO tags for various CSO features. For example, for a cancer sample diagnosed with HPV-positive adenocarcinoma of the cervix, the analysis system may resolve that the cancer signal origin of the organ or organ group is the cervix and the tumor biological CSO category is HPV-associated cancer.
[0220] For each cancer sample, the analysis system generates a 510 feature vector based on methylation sequencing data. The feature vector can include multiple methylation features based on the methylation sequencing data. Methylation features can be characteristics of the sequencing data associated with the methylation of fragments in the sequencing data. For example, one type of methylation feature could be the methylation density at a specific locus. Methylation density is the percentage of sequence reads that are methylated at a specific locus. As another example, another type of methylation feature could be the density of highly methylated sequence reads at a specific locus or the density of highly unmethylated sequence reads at a specific locus. In yet another example, another type of methylation feature could be the count of sequence reads overlapping with a specific locus identified as anomalously methylated (e.g., as shown in the original text). Figure 3A and Figure 3B (As described in the text). A locus may contain one or more CpG sites. For example, a locus may be a single CpG site, while a second locus may be a series of adjacent CpG sites, i.e., a CpG region.
[0221] In one or more embodiments, the feature vector may include methylation features spanning target loci. In other embodiments, the analysis system may perform feature selection to identify particularly informative methylation features for CSO classification. This feature selection may utilize the mutual information gain of each methylation feature to distinguish different categories classified by the classifier. Features may be ranked and selected accordingly based on information gain. Discriminative features are features that facilitate classification among different labels. The analysis system may evaluate the discriminative power of features by calculating information gain as a measure of the correlation between features and labels; that is, features with high information gain are more relevant to labels than features with low information gain. The analysis system may identify discriminative features based on the calculated information gain. In one or more embodiments, the analysis system may select discriminative features from all features as those with information gain above a threshold, or those with information gain above a certain percentile.
[0222] In some embodiments, the number of features considered is 100 or more features, 200 or more features, 300 or more features, 400 or more features, 500 or more features, 600 or more features, 700 or more features, 800 or more features, 900 or more features, 1,000 or more features, 1,500 or more features, 2,000 or more features, 2,500 or more features, 3,000 or more features, 3,500 or more features, 4,000 or more features, or 4,500 features. More than 5,000 features, 6,000 features, 7,000 features, 8,000 features, 9,000 features, 10,000 features, 15,000 features, 20,000 features, 25,000 features, 30,000 features, 35,000 features, 40,000 features, 50,000 features, 75,000 features, or 100,000 features.
[0223] The analysis system generates a first training dataset of 515, which includes feature vectors of cancer samples and known organs or organ groups as CSO categories. The first training dataset may exclude tumor biology categories or other CSO features assigned in parallel to each training case. As described above, in some embodiments, the analysis system may perform feature selection to identify discriminative features for use in training organ or organ group classifiers. During feature selection, the analysis system may modify the feature vectors to include only the selected features. This modification of the feature vectors results in a reduction of the training dataset, thereby improving computational efficiency. In some embodiments, the training dataset is stored as a data table (or data array), where each feature vector and known organ type is an entry in the data table.
[0224] The analysis system trains a 520 organ type classifier using a first training dataset to predict organ or organ group CSO categories based on input feature vectors. The analysis system can train the organ or organ group classifier into a computer model containing multiple parameters and a function that associates the input feature vector with the predicted organ or organ group CSO category. In one or more embodiments, the analysis system can train the organ or organ group classifier into a machine learning model implementing one or more machine learning algorithms. In a further embodiment, the analysis system can train the organ or organ group classifier into a neural network containing interconnected layers of nodes. To perform training, the analysis system can input (multiple batches) feature vectors into the organ type classifier while simultaneously tuning the parameters of the organ type classifier to direct the classifier's predictions toward the known organ or organ group CSO category for each cancer sample. The analysis system can cross-validate the trained organ or organ group classifier to evaluate the predictive accuracy of the organ type classifier.
[0225] In some embodiments, the number of parameters in the organ type classifier is 1,000 or more, 1,500 or more, 2,000 or more, 2,500 or more, 3,000 or more, 3,500 or more, 4,000 or more, 4,500 or more, 5,000 or more, 6,000 or more, 7,000 or more, 8,000 or more, 9,000 or more, 10,000 or more, 15,000 or more, 20,000 or more, 25,000 or more, 30,000 or more, 35,000 or more, 40,000 or more, 50,000 or more, 75,000 or more. More parameters, 100,000 or more parameters, 200,000 or more parameters, 300,000 or more parameters, 400,000 or more parameters, 500,000 or more parameters, 600,000 or more parameters, 700,000 or more parameters, 800,000 or more parameters, 900,000 or more parameters, 1,000,000 or more parameters Number, 2,000,000 or more parameters, 3,000,000 or more parameters, 4,000,000 or more parameters, 5,000,000 or more parameters, 6,000,000 or more parameters, 7,000,000 or more parameters, 8,000,000 or more parameters, 9,000,000 or more parameters, or 10,000,000 or more parameters.
[0226] The analysis system generates a second training dataset of 525, which includes feature vectors of cancer samples and known tumor biology CSO categories. The second training dataset may exclude organ groups or other CSO features. As described above, in some embodiments, the analysis system may perform feature selection to identify discriminative features used in training the tumor biology CSO classifier. During feature selection, the analysis system may modify the feature vectors to include only the selected features. This modification of the feature vectors results in a reduction of the training dataset, thereby improving computational efficiency. In some embodiments, the training dataset is stored as a data table (or data array), where each feature vector and known organ group, as well as each cancer biology CSO, are entries in the data table. The first and second training datasets are different. Although CSO categories may contain genetic or epigenetic data from the same samples, the CSO categories differ between the first and second training datasets. Furthermore, feature selection can modify the feature vectors of cancer samples to be different between the first and second training datasets.
[0227] The analysis system trains a 530-classifier for tumor biology CSOs using a second training dataset to predict tumor biology CSO categories based on input feature vectors. The analysis system can train the tumor biology CSO classifier into a computer model containing multiple parameters and a function that associates the input feature vector with the predicted tumor biology CSO category. In one or more embodiments, the analysis system can train the tumor biology CSO classifier into a machine learning model implementing one or more machine learning algorithms. In a further embodiment, the analysis system can train the tumor biology CSO classifier into a neural network containing interconnected layers of nodes. To perform training, the analysis system can input (multiple batches) of feature vectors into the tumor biology CSO classifier while simultaneously tuning the parameters of the tumor biology CSO classifier to direct the predictions of the tumor biology type classifier towards the known tumor biology CSO category for each cancer sample. The analysis system can cross-validate the trained tumor biology CSO classifier to evaluate the predictive accuracy of the tumor biology type classifier.
[0228] In some embodiments, the number of parameters in the tumor biology CSO classifier is 1,000 or more, 1,500 or more, 2,000 or more, 2,500 or more, 3,000 or more, 3,500 or more, 4,000 or more, 4,500 or more, 5,000 or more, 6,000 or more, 7,000 or more, 8,000 or more, 9,000 or more, 10,000 or more, 15,000 or more, 20,000 or more, 25,000 or more, 30,000 or more, 35,000 or more, 40,000 or more, 50,000 or more, 75,000 or more. One or more parameters, 100,000 or more parameters, 200,000 or more parameters, 300,000 or more parameters, 400,000 or more parameters, 500,000 or more parameters, 600,000 or more parameters, 700,000 or more parameters, 800,000 or more parameters, 900,000 or more parameters, 1,000,000 or more parameters Parameters, 2,000,000 or more, 3,000,000 or more, 4,000,000 or more, 5,000,000 or more, 6,000,000 or more, 7,000,000 or more, 8,000,000 or more, 9,000,000 or more, or 10,000,000 or more.
[0229] In one or more embodiments, the analysis system trains an organ or organoid group classifier and a tumor biology classifier in parallel. Parallel training means training two models separately, such that the two models do not share parameters. Parallel training may also require utilizing the same base sequencing data, as both models are trained using the same base methylation sequencing data from cancer samples. However, in embodiments, the features used to train the organ or organoid group classifier may differ from those used for the tumor biology classifier. Therefore, when training each type of classifier, the analysis system modifies the methylation sequencing data to target the features of the specific type of classifier being trained, i.e., generating two distinct derived training datasets.
[0230] In other embodiments, the analysis system may sequentially train an organ or organ group classifier and a tumor biology classifier. In sequential training, the result of the first classifier can be appended to a feature vector as input to a second classifier. In such embodiments, the second classifier is trained using the appended result of the first classifier or a known CSO category. In other embodiments, the analysis system may train multiple CSO classifiers based on the predictions of the first CSO classifier for the first CSO. For example, the first classifier may be an organ or organ group classifier. For each organ or organ group, the analysis system trains a separate tumor biology classifier. Thus, for example, for 15 organs or organ groups, the analysis system trains 15 tumor biology classifiers. Each tumor biology classifier is trained using cancer samples with the same organ or organ group.
[0231] The advantages of training organ or organome classifiers and tumor biology classifiers separately are multifaceted. First, compared to a single CSO classifier with mixed categories or CSO classes, predicting organs or organomes and tumor biology (and / or any other CSO features) provides granularity for CSO prediction. This increased granularity better informs diagnostic testing options, which can efficiently approve safety checks for cases where cancer signals are detected. Second, the analysis system extracts two distinct training datasets from the same sequencing dataset. This extraction compresses the assay process while still adding the aforementioned granularity to CSO prediction. Third, parallel training of the CSO classifiers allows each classifier to infer different patterns used to distinguish CSO labels. All these advantages equate to technological improvements in both the assay process and CSO prediction analysis. Furthermore, these advantages improve CSO prediction, thereby better informing treatment options and potentially leading to improved treatment outcomes.
[0232] IV.B. Deployment of Parallel CSO Classifiers The analysis system deploys a trained CSO classifier. The CSO classifier can be based on the above... Figure 5A The process described in section 500 involves training. The analysis system can deploy the trained CSO classifier to output a CSO prediction that includes one or more CSO features of the sample. For example, the CSO prediction may include organs or organoids and tumor biology. The analysis system can deploy the trained CSO classifier on samples predicted to have cancer or known to have cancer.
[0233] Figure 5B This is an exemplary flowchart illustrating process 540 of a parallel CSO classifier deployment according to one or more embodiments. An analysis system can execute process 540. In other embodiments, another computing device or system can execute any, some, or all of the steps of process 540. Figure 5AIn the embodiment shown, the analysis system deploys two parallel CSO classifiers, one for predicting CSO organs or organ groups and the other for predicting CSO tumor biology.
[0234] The analysis system yielded 545 test samples containing methylation sequencing data. In one or more embodiments, the test samples may have an unknown cancer state. In other embodiments, the test samples may have a positive cancer diagnosis. The methylation sequencing data may be derived from the sequencing of biological samples containing nucleic acid fragments (e.g., via WGBS or targeted methylation assays).
[0235] The analysis system generated feature vectors for 550 test samples based on methylation sequencing data. Methylation features can be compared with those described above. Figure 5A The features described in the text are the same.
[0236] In some embodiments, the analysis system applies a 555 cancer classifier to predict the cancer status of a test sample. In such embodiments, the test sample may have an unknown cancer status. Therefore, a cancer classifier (e.g., as by...) can be applied. Figure 4A The process (400) is trained to determine cancer predictions. Cancer predictions can indicate the cancer status of subjects in derived test samples.
[0237] The analysis system applies a 560 organ type classifier to predict the organ or organ group of a test sample. In some embodiments, the analysis system may modify a feature vector based on selected features from the organ or organ group classifier to generate a first simplified feature vector. The analysis system inputs the feature vector (or the first simplified feature vector) into an organ type classifier, which outputs an organ or organ group prediction for the test sample based on the input feature vector. In one or more embodiments, the organ or organ group prediction identifies one of a plurality of organs or organ groups classified as a major origin of cancer.
[0238] The analysis system applies a 565 tumor biology classifier to predict the tumor biology of a test sample. In some embodiments, the analysis system can modify a feature vector based on selected features of the tumor biology classifier to generate a second simplified feature vector. The analysis system inputs the feature vector (or the second simplified feature vector) into the tumor biology classifier, which outputs a tumor biology prediction for the test sample based on the input feature vector. In one or more embodiments, the tumor biology prediction identifies one of a plurality of tumor biology categories.
[0239] In some embodiments, the analysis system can apply CSO classifiers one after another. In such embodiments, a subsequent CSO classifier can combine feature vectors with further input from the predictions made by the previous CSO classifier to output a subsequent prediction. For example, an organ type classifier can output a CSO prediction for an organ or organ group. The analysis system can input feature vectors and the predicted organ type into a tumor biology classifier to predict the tumor biology type. In a contrasting example, a tumor biology classifier can output a CSO prediction for the tumor biology of a sample. The analysis system can then input feature vectors (e.g., which may be specific to an organ or organ group classifier) and the predicted tumor biology into an organ or organ group classifier to predict the organ or organ group of the tumor.
[0240] The analysis system can integrate 570 predictions from CSO classifiers. It can receive CSO predictions from organ or organogroup classifiers and CSO predictions from tumor biology classifiers. In this step, these two predictions are compared to a list of combined CSO predictions, which inform safe and efficient diagnostic testing for detected cancer signals. The analysis system can determine whether organ or organogroup predictions and cancer biology predictions are reported.
[0241] The analysis system reports 575 the results selected in step 570, along with optional diagnostic test recommendations. The results may include cancer signal detection (e.g., output by a cancer classifier) and / or the selected CSO prediction (e.g., output by a CSO classifier). In some embodiments, the results reported by the analysis system may include cancer signal detection, a basic CSO prediction (e.g., organ system), and additional predictive information from the CSO classifier, such as organ type and / or tumor biology. In some embodiments, the analysis system may identify one or more diagnostic test options based on cancer signal detection and / or CSO prediction. Diagnostic test options may be associated with various combinations of CSO predictions in a database.
[0242] Using the results and recommendations for optional diagnostic tests, healthcare professionals can consult with the subject to determine a diagnostic testing plan. In some embodiments, the diagnostic testing plan may include one or more options recommended by the analysis system. In one or more embodiments, the analysis system may assist in ordering subsequent diagnostic tests (e.g., in response to instructions from healthcare professionals).
[0243] V. Application In some embodiments, the methods, analysis systems, and / or classifiers of the present invention can be used to detect the presence of cancer, monitor cancer progression or recurrence, monitor treatment response or effectiveness, determine the presence of or monitor minimal residual disease (MRD), or any combination thereof. For example, as described herein, the classifier can be used to generate a probability score (e.g., from 0 to 100) that describes the likelihood that a test feature vector originates from a subject with cancer. In some embodiments, the probability score is compared to a threshold probability to determine whether the subject has a detectable cancer signal.
[0244] In one or more embodiments, probability or likelihood scores can be assessed at multiple different time points (e.g., before or after treatment) to monitor disease progression or to monitor treatment effectiveness (e.g., treatment outcome). For example, the monitored individual can undergo repeated sampling. Biological samples extracted from the individual can be sequenced, for example, using one or more sequencing devices, to measure sequencing data from the samples. The analysis system can perform various analyses, including any cancer classification analysis described elsewhere. In the context of monitoring progression, repeated sample collection, sequencing, and analysis can track changes in cancer signals in an individual over time. If the cancer signal increases over time, the analysis system can infer that the tumor is malignant and is progressing. If the cancer signal is stable, the analysis system can infer that the tumor is benign or stationary. In the context of treatment evaluation, the analysis system can collect samples, sequence the samples, and analyze the samples to predict cancer signals before treatment compared to after treatment. If the cancer signal decreases after treatment administration, the analysis system can infer treatment success. The analysis system can further compare changes in cancer signals across populations to assess inter-individual efficacy. In one or more embodiments, when significant cancer signals are detected, the analysis system may transmit a notification to an individual (e.g., for viewing on their personal computing device) to seek medical attention for diagnostic examinations.
[0245] In other embodiments, probability or likelihood scores can be used to make or influence clinical decisions (e.g., diagnostic tests for cancer signal detection in cancer diagnosis, treatment options, treatment efficacy assessments, etc.). For example, in one embodiment, if the probability score exceeds a threshold, a physician may prescribe appropriate treatment. In other embodiments, the prediction outcome may include a CSO prediction, which includes one or more predicted CSO categories. For example, a CSO prediction may include an organ or organ group, tumor biology, cell of origin, or some combination thereof. Such outcomes (or combinations of outcomes) can inform diagnostic testing options, treatment choices, and follow-up strategies between healthcare professionals and subjects.
[0246] VA Cancer Detection In some embodiments, the method and / or classifier of the present invention are used to detect the presence or absence of cancer in a subject who is not suspected of having cancer. For example, the classifier (e.g., as described in Part III above) can be used to determine a cancer prediction that describes the probability that a test feature vector originates from a subject with cancer.
[0247] In one embodiment, cancer prediction is the probability (i.e., a binary classification) that a test sample has cancer (e.g., a score between 0 and 100). Therefore, the analysis system may determine a threshold for determining whether a test subject has cancer. For example, a cancer prediction greater than or equal to 60 may indicate that the subject has cancer. In further embodiments, a cancer prediction greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95 indicates that the subject has cancer. In other embodiments, cancer prediction may indicate the severity of the disease. For example, a cancer prediction of 80 compared to a cancer prediction below 80 (e.g., a probability score of 70) may indicate a more severe form or later stage of cancer. Similarly, an increase in cancer prediction over time (e.g., determined by classifying test feature vectors from multiple samples collected from the same subject at two or more time points) may indicate disease progression, or a decrease in cancer prediction over time may indicate treatment success.
[0248] In another embodiment, cancer prediction comprises a plurality of predicted values, each of the multiple different cancer signaling origins being classified (i.e., multi-class classification) having a predicted value (e.g., a score between 0 and 100). The predicted value may correspond to the probability that a given training sample (corresponding to the training sample during inference) possesses each cancer signaling origin. The analysis system can identify the cancer signaling origin with the highest predicted value and indicate that the test subject is likely to have cancer originating from that signaling origin. In other embodiments, the analysis system further compares the highest predicted value to a threshold (e.g., 50, 55, 60, 65, 70, 75, 80, 85, etc.) to determine that the test subject is highly likely to have cancer originating from that signaling origin. In other embodiments, the predicted value may also indicate the severity of the disease. For example, a predicted value greater than 80 compared to a predicted value of 60 may indicate a more severe form or later stage of the cancer. Similarly, an increase in the predicted value over time (e.g., determined by classifying test feature vectors from multiple samples collected from the same subject at two or more time points) may indicate disease progression, or a decrease in the predicted value over time may indicate treatment success.
[0249] In some embodiments, the method and / or classifier of the present invention are used to determine a cancer origin prediction for a subject suspected of having cancer. For example, one or more CSO classifiers (e.g., as described in Section IV above) may be used to determine a cancer origin prediction to aid in diagnostic examination options.
[0250] According to various aspects of the present invention, the methods and systems of the present invention can be trained to detect or classify multiple signs of cancer. For example, the methods, systems, and classifiers of the present invention can be used to detect one or more, two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different types of cancer.
[0251] Examples of cancers that can be detected using the methods, systems, and classifiers of this invention include carcinoma, lymphoma, blastoma, sarcoma, and leukemia or lymphoma. More specific examples of this type of cancer include, but are not limited to, squamous cell carcinoma (e.g., epithelial squamous cell carcinoma), skin cancer, melanoma, lung cancer (including small cell lung cancer, non-small cell lung cancer (“NSCLC”), lung adenocarcinoma, and lung squamous cell carcinoma), peritoneal cancer, gastric (or stomach) cancer (including gastrointestinal cancer), pancreatic cancer (e.g., pancreatic ductal adenocarcinoma), cervical cancer, ovarian cancer (e.g., high-grade serous ovarian cancer), liver cancer (e.g., hepatocellular carcinoma (HCC)), hepatotum, liver tumor, bladder cancer (e.g., urothelial bladder cancer), testicular (germ cell tumor) cancer, breast cancer (e.g., HER2-positive, HER2-negative, and triple-negative breast cancer), brain cancer (e.g., astrocytoma, glioma (e.g., glioblastoma)), colon cancer, rectal cancer, colorectal cancer, endometrial cancer or uterine cancer, salivary gland cancer, kidney cancer or renal cancer (e.g., renal cell carcinoma, nephroblastoma, or Wilms' disease). Tumors, prostate cancer, vulvar cancer, thyroid cancer, anal canal cancer, penile cancer, head and neck cancer, esophageal cancer, and nasopharyngeal carcinoma (NPC). Other examples of cancers include, but are not limited to, retinoblastoma, theca cell tumor, ovarian-testicular blastoma, hematologic malignancies, including but not limited to non-Hodgkin's lymphoma (NHL), multiple myeloma and acute hematologic malignancies, endometriosis, fibrosarcoma, choriocarcinoma, laryngeal cancer, Kaposi's sarcoma, schwannoma, oligodendroglioma, neuroblastoma, rhabdomyosarcoma, osteosarcoma, leiomyosarcoma, and urinary tract cancers.
[0252] In some embodiments, cancer is one or more of the following, or any combination thereof: anorectal cancer, bladder cancer, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, stomach cancer, head and neck cancer, hepatobiliary cancer, leukemia, lung cancer, lymphoma, melanoma, multiple myeloma, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer, thyroid cancer, and uterine cancer.
[0253] In some embodiments, one or more cancers can be “high-signal” cancers (defined as cancers with a 5-year cancer-specific mortality rate greater than 50%), such as anorectal cancer, colorectal cancer, esophageal cancer, head and neck cancer, hepatobiliary cancer, lung cancer, ovarian cancer, and pancreatic cancer, as well as lymphoma and multiple myeloma. High-signal cancers tend to be more aggressive and typically have above-average concentrations of free nucleic acids in test samples obtained from patients.
[0254] VB Parallel Cancer Origin Classification In some embodiments, the method and / or classifier of the present invention are used to determine multiple individual CSO predictions for a sample. For example, a CSO classifier trained in parallel (e.g., as described in Section IV above) can be used to determine CSO predictions describing the organ or group of organs, tumor biological category, cell type, or some combination thereof that describe the origin of the cancer signal observed in the sample. Figure 7 and Figure 8 In the example implementation of the results shown, the organ or organ group classifier and the tumor biology classifier can be based on Figure 5A The process was trained using 500, and based on... Figure 5B The process 540 was deployed. In an instance demonstrating the feasibility of training classifiers separately for organs and organoids from the same cases and with the same genetic or epigenetic information, as well as for the origin of cancer biological signals, the training population consisted of 2,496 evaluable study participants with invasive cancer and 2,591 participants without cancer. With 99.4% specificity, both classifiers consistently detected 1,401 cancer cases. The first classifier for organs or organoids had a 91.8% [90.2–93.2] (1,286 / 1,401) CSO prediction accuracy. The second classifier for tumor biology was trained independently of the first classifier, with the tumor biology classifier prediction accuracy corrected to 86.2% [84.2–87.9] (1,207 / 1,401). Of the 1,401 cases detected, 97 had clinical information that did not allow for explicit designation of a target CSO for classifier training or assessment of prediction accuracy (e.g., "cancer not otherwise specified"), and these were counted as incorrect. Then, both classifiers were used to predict in a retained population of 1,583 cancer patients and 521 cancer-free patients. Cancer signals were detected in 763 cases. The first classifier for organs or organ groups had 87.9% [85.4–90.2] (671 / 763) CSO correct predictions. The second classifier for tumor biology had 84.3% [81.5–86.8] (643 / 763) CSO correct predictions.
[0255] Figure 7Two confusion matrices are presented to demonstrate the prediction accuracy of a demonstrative organ or organoid group classifier according to one or more example implementations. The top confusion matrix 710 shows the prediction accuracy from cross-validation with the training cohort. The bottom confusion matrix 720 shows the prediction accuracy for the retained set. For both confusion matrices, the x-axis represents the known CSO category (e.g., an organ or organoid group representing clinical truth), and the y-axis represents the predicted CSO category. Notably, the prediction accuracy is consistent in both exemplary results.
[0256] Figure 8 Two confusion matrices illustrating the predictive accuracy of a tumor biology classifier according to one or more example implementations are presented. The top confusion matrix 810 shows the predictive accuracy from cross-validation with the training cohort. The bottom confusion matrix 820 shows the predictive accuracy for the retained set. For both confusion matrices, the x-axis represents the known CSO category (e.g., representing clinical truth in tumor biology), and the y-axis represents the predicted CSO category. Notably, the predictive accuracy is consistent in both exemplary results.
[0257] VC Cancer and Treatment Monitoring In some embodiments, cancer predictions can be assessed at multiple different time points (e.g., before or after treatment) to determine patient prognosis, predict response to candidate treatments, monitor disease progression, or monitor treatment effectiveness (e.g., treatment efficacy). For example, the method included in this invention involves obtaining a first sample (e.g., a first plasma cfDNA sample) from a cancer patient at a first time point, determining a first cancer prediction therefrom (as described herein), and obtaining a second test sample (e.g., a second plasma cfDNA sample) from the same cancer patient at a second time point, determining a second cancer prediction therefrom (as described herein).
[0258] In some embodiments, the first time point is before cancer treatment (e.g., before resection or treatment intervention), and the second time point is after cancer treatment (e.g., after resection or treatment intervention), and a classifier is used to monitor the effectiveness of the treatment. For example, if the second cancer prediction decreases compared to the first cancer prediction, then the treatment is considered successful. However, if the second cancer prediction increases compared to the first cancer prediction, then the treatment is considered unsuccessful. In other embodiments, both the first and second time points are before cancer treatment (e.g., before resection or treatment intervention). In still other embodiments, both the first and second time points are after cancer treatment (e.g., after resection or treatment intervention). In yet still other embodiments, cfDNA samples may be obtained from and analyzed from the cancer patient at the first and second time points, for example, to monitor cancer progression, determine whether the cancer is in remission (e.g., post-treatment), to monitor or detect residual disease or disease recurrence, or to monitor the effectiveness of treatment (e.g., therapeutic).
[0259] Those skilled in the art will readily understand that test samples can be obtained from cancer patients at any desired set of time points and analyzed according to the method of the present invention to monitor the patient's cancer status. In some embodiments, the duration of the interval between the first and second time points ranges from about 15 minutes to about 30 years, such as about 30 minutes, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or about 24 hours, such as about 1, 2, 3, 4, 5, 10, 15, 20, 25 or about 50 days, or such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 months, or such as about 1, 1.5, 2, 2.5, 3, 3. 5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5 or approximately 30 years. In other embodiments, the frequency of obtaining test samples from patients may be: at least once every 5 months, at least once every 6 months, at least once a year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years.
[0260] VD processing In yet another embodiment, cancer prediction can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, treatment effectiveness evaluation, etc.). For example, in one embodiment, if the cancer prediction (e.g., for cancer or for a specific cancer signal origin) exceeds a threshold, a physician can prescribe appropriate treatment (e.g., surgical resection, radiation therapy, chemotherapy, and / or immunotherapy).
[0261] A classifier (as described herein) can be used to determine a sample feature vector based on a cancer prediction for a subject with cancer. In one embodiment, when the cancer prediction exceeds a threshold, an appropriate treatment (e.g., surgical resection or therapeutic intervention) is prescribed. For example, in one embodiment, if the cancer prediction is greater than or equal to 60, one or more appropriate treatments are prescribed. In another embodiment, if the cancer prediction is greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95, one or more appropriate treatments are prescribed. In other embodiments, the cancer prediction may indicate the severity of the disease. An appropriate treatment matching the severity of the disease can then be prescribed.
[0262] In some embodiments, the treatment is one or more cancer therapeutic agents selected from the group consisting of chemotherapeutic agents, targeted cancer therapeutic agents, differentiation therapeutic agents, hormone therapeutic agents, and immunotherapeutic agents. For example, the treatment may be one or more chemotherapeutic agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, scaffold disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapeutic agents selected from the group consisting of signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteasome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapeutic agents, including retinoids such as retinoic acid, levamisole, and bexarotin. In some embodiments, the treatment is one or more hormonal therapeutic agents selected from the group consisting of: anti-estrogens, aromatase inhibitors, progestins, estrogens, anti-androgens, and GnRH agonists or analogues. In one embodiment, the treatment is one or more immunotherapeutic agents selected from the group consisting of: monoclonal antibody therapies (such as rituximab (Rituxan) and alemtuzumab (CAMPATH)), nonspecific immunotherapies and adjuvants (such as BCG, interleukin-2 (IL-2), and interferon-α), and immunomodulatory drugs (e.g., thalidomide and lenalidomide (Revlimid)). Experienced physicians or oncologists can select appropriate cancer therapeutic agents based on a variety of characteristics, such as tumor type, cancer stage, previous cancer treatments or agents, and other characteristics of the cancer.
[0263] VE kit implementation method This document also discloses kits for performing the methods described above, including those related to cancer classifiers. The kit may include one or more collection containers for collecting samples containing genetic material from a subject. Samples may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. Such kits may include reagents for isolating nucleic acids from the sample. The reagents may further include reagents for sequencing the nucleic acids, including buffers and detection reagents. In one or more embodiments, the kit may include one or more sequencing units comprising probes for targeting specific genomic regions, specific mutations, specific genetic variations, or combinations thereof. In other embodiments, samples collected via the kit are provided to a sequencing laboratory that can use the sequencing units to sequence the nucleic acids in the sample. WBC contamination detection can be applied to various kit configurations to minimize WBC contamination that may originate from components of the kit. For example, experiments comparing the types of collection containers can be run. WBC contamination can be evaluated and compared between types of collection containers to identify the optimal type that minimizes WBC contamination.
[0264] The kit may further include instructions for use of the reagents contained in the kit. For example, the kit may include instructions for collecting samples and extracting nucleic acids from test samples. Exemplary instructions may include the order of adding reagents, the centrifugation speed for isolating nucleic acids from test samples, how to amplify nucleic acids, how to sequence nucleic acids, or any combination thereof. The instructions may further explain how to operate the computing device, as analysis system 200, to perform the steps of any of the described methods.
[0265] In addition to the components described above, the kit may also include a computer-readable storage medium containing computer software for performing the various methods described herein. These instructions may exist in one form as printed information on a suitable medium or substrate (e.g., one or more sheets of paper on which the information is printed), in the kit packaging, or in a packaging insert. Another form would be a computer-readable medium, such as a floppy disk, CD, hard disk, or network data storage, where the instructions are stored in the form of computer code. Yet another means may be a URL or QR code, which can be used to access information at a remote site via the Internet.
[0266] VI. Other considerations The above detailed description of the embodiments refers to the accompanying drawings, which illustrate specific embodiments of this disclosure. Other embodiments with different structures and operations do not depart from the scope of this disclosure. The terms "invention" and the like are used to refer to certain specific instances of the many alternative aspects or embodiments of the applicant's invention shown in this specification, and their use or absence is not intended to limit the scope of the applicant's invention or the scope of the claims.
[0267] Embodiments of the present invention may also relate to an apparatus for performing the operations described herein. This apparatus may be specifically constructed for the desired purpose and / or may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory, tangible, computer-readable storage medium, or in any type of medium suitable for storing electronic instructions, which may be coupled to a computer system bus. Furthermore, any computing system mentioned herein may include a single processor, or may employ a multiprocessor architecture to enhance computing power.
[0268] Any step, operation, or process described herein, performed by the analysis system, may be performed individually by one or more hardware or software modules of the device, or in combination with other computing devices. In one embodiment, the software module is implemented as a computer program product comprising a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or processes.
Claims
1. A method for training an independent parallel cancer signal origin classifier, characterized in that: The method includes: Training samples derived from subjects with known cancer diagnoses were obtained, each training sample containing methylated sequence reads corresponding to nucleic acid fragments collected from biological samples from each subject, and each known cancer diagnosis included known organs or organ groups among multiple affected organs or organ groups and known tumor biology among multiple tumor biology categories. For each training sample, a feature vector is generated based on these methylated sequence reads; Generate a first training dataset, which contains the feature vectors of these training samples and the known organ or group of organs; The first training dataset is used to train an organ or organ group classifier to predict organs or organ groups from the plurality of organs or organ groups based on the input feature vector. Generate a second training dataset, which contains the feature vectors of these training samples and these known tumor biological categories; and A tumor biology classifier is trained using the second training dataset to predict tumor biology from the multiple tumor biology categories based on the input feature vector.
2. The method according to any of the preceding claims, characterized in that: The method further includes: For each training sample, the known organ or organ group and the known tumor biological category are extracted from the known cancer diagnosis and clinical information.
3. The method according to any of the preceding claims, characterized in that: The feature vector is based at least in part on the methylation features of these methylated sequence reads.
4. The method as described in claim 3, characterized in that: These methylation features include: methylation density at one or more loci, density of hypermethylated sequence reads at one or more loci, density of hypomethylated sequence reads at one or more loci, count of methylated sequence reads identified as aberrantly methylated at one or more loci, or some combination thereof.
5. The method according to any of the preceding claims, characterized in that: Generating the first training dataset includes excluding information about tumor biology.
6. The method according to any of the preceding claims, characterized in that: Generating the second training dataset includes excluding information about the affected organs or groups of organs.
7. The method according to any of the preceding claims, characterized in that: The method further includes: For each feature, determine the information gain in distinguishing these organs or groups of organs; Based on these information gains, the discriminative features of the organ or organ group classifier are identified; and The feature vectors of the first training set are modified to consist of these discriminative features, wherein the modified feature vectors are used to train the organ or organ group classifier.
8. The method according to any of the preceding claims, characterized in that: The method further includes: For each feature, determine the information gain in distinguishing these tumor biological categories; Based on these information gains, the discriminative features of the tumor biological classifier are identified; and The feature vectors of the second training set are modified to consist of these discriminative features, wherein the modified feature vectors are used to train the tumor biology classifier.
9. The method according to any of the preceding claims, characterized in that: The organ or organ group classifier or the tumor biology classifier is a machine learning model.
10. The method according to any of the preceding claims, characterized in that: The method further includes training the organ or organ group classifier and the tumor biology classifier in a parallel training process.
11. The method according to any of the preceding claims, characterized in that: The method further includes training the organ or organ group classifier before training the tumor biological classifier.
12. The method as described in claim 11, characterized in that: Before training the tumor biology classifier, the output of the organ or organ group classifier is appended to the feature vector of the second training dataset.
13. The method according to any of the preceding claims, characterized in that: The method further includes training the tumor biology classifier before training the organ or organ group classifier.
14. The method as described in claim 13, characterized in that: Before training the organ or organ group classifier, the output of the tumor biology classifier is appended to the feature vector of the first training dataset.
15. The method according to any of the preceding claims, characterized in that: These organs or groups of organs include: breast; prostate; lung; head or neck; anus; cervix; ovary or fallopian tube; uterus; bladder or urothelial tissue; kidney; stomach or esophagus; liver or intrahepatic bile duct; pancreas, extrahepatic bile duct or gallbladder; colon or rectum; bone or soft tissue; skin; blood, lymphatic system or bone marrow; thyroid gland; indistinct tissue; or some combination thereof.
16. The method according to any of the preceding claims, characterized in that: These tumor biology categories include: lymphoid vegetations, medullary vegetations, plasma cell vegetations, neuroendocrine carcinomas or tumors, adenocarcinomas, squamous cell carcinomas that are not human papillomavirus-associated, human papillomavirus-associated carcinomas, hepatocellular carcinomas, vegetations originating from Müllerian ducts, transitional cell carcinomas, mesenchymal tumors, melanocyte vegetations, mesothelial vegetations, other tumor biology, obscure tumor biology, or some combination thereof.
17. A method for predicting the origin of cancer signals, characterized in that: The method includes: Obtain a test sample derived from a subject, the test sample containing methylated sequence reads corresponding to nucleic acid fragments collected from biological samples from the subject; For the test sample, a first feature vector is generated based on methylated sequence reads associated with a first feature set that is identified as having discriminative significance for the classification of organs or organ groups; For the test sample, a second feature vector is generated based on methylated sequence reads associated with a second feature set that is identified as having discriminative significance for tumor biological classification; An organ or organ group classifier is applied to these first feature vectors to predict the organ or organ group of cancer associated with the test sample from multiple organs or organ groups. A tumor biology classifier is applied to the second feature vector to predict the tumor biology of the cancer associated with the test sample from multiple tumor biology categories; The organ or organoid group classifier and the tumor biology classifier are independently trained on training samples derived from subjects with known cancer diagnoses, which include known organs or organoid groups among multiple affected organs or organoid groups and known tumor biology among multiple tumor biology categories. Each training sample contains a methylated sequence read corresponding to a nucleic acid fragment in a biological sample collected from each subject. The diagnostic tests are based on the predicted organ or organ group and the predicted tumor biology to provide information for diagnosing cancer.
18. The method as described in claim 17, characterized in that: The organ or organ group classifier and the tumor biology classifier are trained by the following: For each training sample, a feature vector is generated based on the methylation sequence reads of the training sample; Generate a first training dataset, which contains the feature vectors of these training samples and the known organ or group of organs for the known cancer diagnosis; The first training dataset is used to train an organ or organ group classifier to predict organs or organ groups from the plurality of organs or organ groups based on the input feature vector. Generate a second training dataset, which contains the feature vectors of these training samples and the known tumor biological categories of the known cancer diagnoses; as well as A tumor biology classifier is trained using the second training dataset to predict tumor biology from the multiple tumor biology categories based on the input feature vector.
19. The method according to any one of claims 17 to 18, characterized in that: Generating the first training dataset includes excluding information about tumor biology, and generating the second training dataset includes excluding information about the affected organ or organ group.
20. The method according to any one of claims 17 to 19, characterized in that: The method further includes: For each feature, determine the information gain in distinguishing these organs or groups of organs; Based on these information gains, the discriminative features of the organ or organ group classifier are identified; and The feature vectors of the first training set are modified to consist of these discriminative features, wherein the modified feature vectors are used to train the organ or organ group classifier.
21. The method according to any one of claims 17 to 20, characterized in that: The method further includes: For each feature, determine the information gain in distinguishing these tumor biological categories; Based on these information gains, the discriminative features of the tumor biological classifier are identified; and The feature vectors of the second training set are modified to consist of these discriminative features, wherein the modified feature vectors are used to train the tumor biology classifier.
22. The method according to any one of claims 17 to 21, characterized in that: The organ or organ group classifier or the tumor biology classifier is a machine learning model.
23. The method according to any one of claims 17 to 22, characterized in that: The method further includes training the organ or organ group classifier and the tumor biology classifier in a parallel training process.
24. The method according to any one of claims 17 to 23, characterized in that: The method further includes training the organ or organ group classifier before training the tumor biological classifier.
25. The method according to any one of claims 17 to 24, characterized in that: The method further includes training the tumor biology classifier before training the organ or organ group classifier.
26. The method according to any one of claims 17 to 25, characterized in that: The method further includes: Before applying the organ or organ group classifier, the feature vector is modified based on the discriminative features of the organ or organ group classifier to generate a first simplified feature vector, such that the organ or organ group classifier is applied to the first simplified feature vector; and Before applying the tumor biology classifier, the feature vector is modified according to the discriminative features of the tumor biology classifier to generate a second simplified feature vector, so that the tumor biology classifier is applied to the second simplified feature vector.
27. The method according to any one of claims 17 to 26, characterized in that: Information provided for diagnostic testing of detected cancer signals includes identifying one or more diagnostic test options based on the predicted organ or group of organs, the predicted tumor biology, or some combination thereof.
28. The method according to any one of claims 17 to 27, characterized in that: These organs or groups of organs include: breast; prostate; lung; head or neck; anus; cervix; ovary or fallopian tube; uterus; bladder or urothelial tissue; kidney; stomach or esophagus; liver or intrahepatic bile duct; pancreas, extrahepatic bile duct or gallbladder; colon or rectum; bone or soft tissue; skin; blood, lymphatic system or bone marrow; thyroid gland; indistinct tissue; or some combination thereof.
29. The method according to any one of claims 17 to 28, characterized in that: These tumor biology categories include: lymphoid vegetations, medullary vegetations, plasma cell vegetations, neuroendocrine carcinomas or tumors, adenocarcinomas, squamous cell carcinomas that are not related to human papillomavirus (HPV), HPV-associated carcinomas, hepatocellular carcinomas, vegetations originating from Müllerian ducts, transitional cell carcinomas, mesenchymal tumors, melanocyte vegetations, mesothelial vegetations, other tumor biology, obscure tumor biology, or some combination thereof.
30. The method according to any one of claims 17 to 29, characterized in that: The subject was previously diagnosed with cancer of unknown origin, wherein providing information for the diagnostic examination includes providing information for the diagnostic examination to refine the diagnosis based on the predicted organ or organ group and the predicted tumor biology.
31. The method according to any one of claims 17 to 30, characterized in that: Information provided for the diagnostic examination includes: A report is provided that includes cancer signal detection readouts for the test sample, prediction of cancer signal origin, predicted organ or organoid group, and predicted tumor biology.
32. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method as described in any one of claims 1 to 31.
33. A system, characterized in that: The system includes: One or more computer processors; and The non-transitory computer-readable storage medium as described in claim 32.
34. A treatment kit, characterized in that: The treatment kit contains: A collection container for collecting DNA samples from a subject; Optionally, one or more reagents are used to separate DNA fragments from the DNA sample; Optionally, one or more probes are used to target one or more loci identified as indicating a cancer state. as well as The non-transitory computer-readable storage medium as described in claim 32.
35. A method for providing a report of test samples to a patient to assist in the patient's diagnostic examination, characterized in that: The report includes cancer signal detection readout and cancer signal origin prediction, wherein the cancer signal origin prediction includes the predicted organ or organ group of the cancer signal origin and the predicted tumor biology of the cancer signal origin, and the method includes: Obtain a test sample derived from a patient, the test sample containing methylated sequence reads corresponding to nucleic acid fragments from biological samples collected from the patient; For the test sample, a first feature vector is generated based on methylation information, wherein the methylation information is selected to provide information on cancer signals associated with the test sample; For the test sample, a second feature vector is generated based on methylated sequence reads associated with a first feature set that is identified as having discriminative significance for the classification of organs or organ groups; For the test sample, a third feature vector is generated based on methylated sequence reads associated with a second feature set that is identified as having discriminative significance for tumor biological classification. A cancer signal classifier is applied to the first feature vector to predict cancer signals associated with the test sample; An organ or organ group classifier is applied to the second feature vector to predict the organ or organ group associated with the cancer from multiple organs or organ groups; A tumor biology classifier is applied to the third feature vector to predict the tumor biology of the cancer associated with the test sample from multiple tumor biology categories; The cancer signal classifier is trained on training samples derived from multiple cancer-positive and cancer-negative subjects, each cancer-positive subject having a labeled cancer diagnosis and each cancer-negative subject being known not to have cancer, and each training sample containing a methylated sequence read corresponding to a nucleic acid fragment from a biological sample collected from each subject. The organ or organ group classifier and the tumor biology classifier are trained independently on training samples derived from subjects with known cancer diagnoses, which include known organs or organ groups among multiple affected organs or organ groups and known tumor biology among multiple tumor biology categories, and each training sample contains a methylated sequence read corresponding to a nucleic acid fragment in a biological sample collected from each subject. A report on the test sample is generated based on the results from the cancer signal classifier, the organ or organ type classifier, and the tumor biology classifier. The report includes the cancer signal detection readout and the predicted origin of the cancer signal, the predicted origin of the cancer signal including the predicted organ or organ group of the cancer signal origin and the predicted tumor biology of the cancer signal origin; and The report shall be provided to the patient or the patient's healthcare provider.
36. A method for providing a report of test samples to a patient to assist in the patient's diagnostic examination, characterized in that: The report includes cancer signal detection readout and cancer signal origin prediction, wherein the cancer signal origin prediction includes the predicted organ or organ group of the cancer signal origin and the predicted tumor biology of the cancer signal origin, wherein the cancer signal origin prediction is determined by the following: Obtain a test sample derived from a patient, the test sample containing methylated sequence reads corresponding to nucleic acid fragments from biological samples collected from the patient; For the test sample, a first feature vector is generated based on methylation information, wherein the methylation information is selected to provide information on cancer signals associated with the test sample; For the test sample, a second feature vector is generated based on methylated sequence reads associated with a first feature set that is identified as having discriminative significance for the classification of organs or organ groups; For the test sample, a third feature vector is generated based on methylated sequence reads associated with a second feature set that is identified as having discriminative significance for tumor biological classification. A cancer signal classifier is applied to the first feature vector to predict cancer signals associated with the test sample; An organ or organ group classifier is applied to the second feature vector to predict the organ or organ group associated with the cancer from multiple organs or organ groups; A tumor biology classifier is applied to the third feature vector to predict the tumor biology of the cancer associated with the test sample from multiple tumor biology categories; The cancer signal classifier is trained on training samples derived from multiple cancer-positive and cancer-negative subjects, each cancer-positive subject having a labeled cancer diagnosis and each cancer-negative subject being known not to have cancer, and each training sample containing a methylated sequence read corresponding to a nucleic acid fragment from a biological sample collected from each subject. The organ or organ group classifier and the tumor biology classifier are trained independently on training samples derived from subjects with known cancer diagnoses, which include known organs or organ groups among multiple affected organs or organ groups and known tumor biology among multiple tumor biology categories, and each training sample contains a methylated sequence read corresponding to a nucleic acid fragment in a biological sample collected from each subject. A report on the test sample is generated based on the results from the cancer signal classifier, the organ or organ type classifier, and the tumor biology classifier. The report includes the cancer signal detection readout and the predicted origin of the cancer signal, the prediction including the predicted organ or organ group of origin and the predicted tumor biology of the origin. The report shall be provided to the patient or the patient's healthcare provider.