Systems and methods for detecting disease subtypes

The system uses targeted methylation sequencing and machine learning to analyze cfDNA for accurate cancer subtype identification and resistance mechanism detection, enhancing early cancer detection and treatment strategies.

JP2025528064APending Publication Date: 2025-08-26GRAIL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025505598
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-30
Filing Date
2023-07-31
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Current methods for analyzing methylation sequencing data from free DNA are inadequate for accurately detecting and monitoring disease states, particularly in early stages of cancer, and do not effectively identify cancer subtypes or detect resistance mechanisms during treatment.

Method used

A system and method using targeted methylation sequencing and machine learning models to analyze methylation data from nucleic acid samples, enabling the identification of disease subtypes and detection of resistance mechanisms in cancer by processing methylation patterns in cfDNA.

Benefits of technology

Enables accurate detection of cancer subtypes and resistance mechanisms, allowing for personalized treatment strategies and improved patient prognosis through blood-based assays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528064000001_ABST
    Figure 2025528064000001_ABST
Patent Text Reader

Abstract

Systems and methods for detecting disease state subtypes and determining the occurrence of resistance mechanisms in a disease are disclosed. One method may include receiving, at an input component of the system, a set of sequence reads associated with a nucleic acid sample, generating methylation data through analysis of the set of sequence reads using a processor of the system, and analyzing the methylation data using the processor to identify the disease state subtype. Another method may include obtaining methylation data from a targeted methylation sequencing assay, applying the methylation data to a trained machine learning model, and receiving an output indicating whether MRD is present in the test subject and / or whether a resistance mechanism has occurred with the disease.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 370,014, filed August 1, 2022, U.S. Provisional Patent Application No. 63 / 382,664, filed November 7, 2022, and U.S. Provisional Patent Application No. 63 / 385,590, filed November 30, 2022, each of which is incorporated by reference herein in its entirety.

[0002] The present disclosure relates generally to the field of model-based feature quantification and classifiers for predicting disease states and subtypes from nucleic acid samples. [Background technology]

[0003] It has been observed that deoxyribonucleic acid (DNA) methylation plays an important role in gene regulation and expression. Abnormal DNA methylation is involved in many disease processes, including cancer. Advances in research and technology have led to the development of novel techniques for detecting various disease states at early stages. For example, DNA methylation profiling using methylation sequencing (e.g., whole-genome bisulfite sequencing (WGBS)) is increasingly recognized as a valuable diagnostic tool for the detection, diagnosis, and / or monitoring of cancer. More specifically, specific patterns of differentially methylated regions and / or allele-specific methylation patterns can be useful as molecular markers for non-invasive diagnosis using blood-free (cf) DNA. However, there remains a need in the art for improved methods for analyzing methylation sequencing data from free DNA for the detection, diagnosis, and / or monitoring of diseases such as cancer. Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure is directed to addressing one or more of the above problems. The background discussion provided herein is intended to generally present the context of the present disclosure. Unless otherwise noted herein, the material discussed in this section is not prior art to the claims of this application and is not admitted as prior art or an indication of prior art by inclusion in this section. [Means for solving the problem]

[0005] In summary, one aspect provides a method for detecting a subtype of a disease state using a system, the method including receiving, at an input component of the system, a set of sequence reads associated with a nucleic acid sample; generating methylation data through analysis of the set of sequence reads using a processor of the system; and analyzing the methylation data using the processor to identify a subtype of the disease state.

[0006] Another aspect provides a method for training a machine learning model to detect the occurrence of resistance mechanisms in cancer, the method comprising: obtaining a set of training data from a source, the training data comprising methylation data derived from a targeted methylation sequencing assay; following the obtaining, annotating the set of training data by assigning a histological label to each of the training data in the set; applying the annotated set of training data to the machine learning model; and optimizing a pattern recognition capability of an algorithm associated with the machine learning model based on the applying.

[0007] Yet another aspect provides a method for detecting the emergence of resistance mechanisms in cancer during treatment using a computer system and an associated trained machine learning model, the method comprising: receiving methylation data derived from a targeted methylation sequencing assay from a biological sample associated with a test subject; and, following the receiving, applying the methylation data to a trained machine learning model; and, following the applying, receiving output from the trained machine learning model, the output including: A) a first indication of whether minimal residual disease is present in the test subject following administration of a cancer treatment; and B) in response to the first indication providing a finding that minimal residual disease is present in the test subject, a second indication that at least a portion of the cancer cells within the minimal residual disease have transformed from a first cancer type to a second cancer type as a result of the emergence of the resistance mechanism.

[0008] The foregoing is a summary and, as such, may contain simplifications, generalizations, and omissions of detail; thus, those skilled in the art will appreciate that the summary is illustrative only and is not intended to be in any way limiting.

[0009] For a better understanding of the embodiments, together with other and further features and advantages thereof, reference is made to the following description taken in conjunction with the accompanying drawings, the scope of which is pointed out in the appended claims.

[0010] The file of this patent contains at least one drawing / photograph executed in color. Copies of this patent with color drawing(s) / photograph(s) will be provided by the Office upon request and payment of the necessary fee.

[0011] The accompanying drawings, which are incorporated herein and constitute the present invention, illustrate several embodiments and, together with the description, serve to explain the principles of the present disclosure. [Brief explanation of the drawings]

[0012] [Figure 1A] 1 illustrates an exemplary computer system for performing the techniques described herein. [Figure 1B]1 illustrates an exemplary software platform for implementing the techniques described herein. [Figure 2] 1 illustrates a confusion matrix according to one aspect of the present disclosure. [Figure 3] 1 shows a table according to one aspect of the present disclosure. [Figure 4] 10 shows another table according to an aspect of the present disclosure. [Figure 5] 1 illustrates a heat map according to an aspect of the present disclosure. [Figure 6A] 1 shows a data plot according to one aspect of the present disclosure. [Figure 6B] 1 illustrates a confusion matrix according to one aspect of the present disclosure. [Figure 7A] 1 illustrates a heat map according to an aspect of the present disclosure. [Figure 7B] 10 illustrates another heat map according to an aspect of the present disclosure. [Figure 8] 1 illustrates a bar graph according to an aspect of the present disclosure. [Figure 9A] 1 shows a gene body diagram of a transcription factor according to one embodiment of the present disclosure. [Figure 9B] FIG. 1 shows another gene body diagram of a transcription factor motif according to one embodiment of the present disclosure. [Figure 10A] 1 illustrates an inspection process flow according to one aspect of the present disclosure. [Figure 10B] 10 illustrates another inspection process flow according to an aspect of the present disclosure. [Figure 11] 1 shows a diagram of how β values ​​are calculated, according to one aspect of the present disclosure. [Figure 12] 10 illustrates another heat map according to an aspect of the present disclosure. [Figure 13] 1 illustrates a clustering graph pattern according to one aspect of the present disclosure. [Figure 14A] 1 shows a line graph of transcription factors according to one embodiment of the present disclosure. [Figure 14B] 1 shows a heat map of transcription factors according to one embodiment of the present disclosure. [Figure 15] 1 shows another string plot graph of transcription factors according to one embodiment of the present disclosure. [Figure 16]1 shows another string plot graph of transcription factors according to one embodiment of the present disclosure. [Figure 17] 1 shows another string plot graph of transcription factors according to one embodiment of the present disclosure. [Figure 18] 1 shows another string plot graph of transcription factors according to one embodiment of the present disclosure. [Figure 19] 10 illustrates another heat map according to an aspect of the present disclosure. [Figure 20] 1 illustrates a clustering graph pattern according to one aspect of the present disclosure. [Figure 21] 1 shows a graph of sample size considerations according to one aspect of the present disclosure. [Figure 22] 1 illustrates exemplary information according to one aspect of the present disclosure. [Figure 23] 1 illustrates exemplary information according to one aspect of the present disclosure. [Figure 24] 1 illustrates an exemplary method for training a machine learning model to determine whether a resistance mechanism has occurred in a cancer, according to one aspect of the present disclosure. [Figure 25] 1 illustrates an exemplary method utilizing a trained machine learning model to determine whether a resistance mechanism has occurred in a cancer, according to one aspect of the present disclosure. [Figure 26] 1 illustrates a confusion matrix according to one aspect of the present disclosure. [Figure 27] 1 illustrates an exemplary method for training a machine learning model to generate a prognostic score from patterns of DNA methylation data, according to one embodiment of the present disclosure. [Figure 28] 1 illustrates exemplary information according to one aspect of the present disclosure. [Figure 29] FIG. 1 shows an exemplary method for generating a final patient prognosis score by combining a machine-generated prognosis score from a trained machine learning model with a ctDNA score, according to one embodiment of the present disclosure. [Figure 30] 1 shows a graph presenting patient prognosis scores generated from different approaches, according to one aspect of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] The terms used hereinafter may be interpreted in their broadest possible sense, even when used in conjunction with the detailed description of specific embodiments of the present disclosure. Indeed, certain terms may even be emphasized below, but any terms intended to be interpreted in any restrictive manner are so expressly and specifically defined in this detailed description section. Both the general description above and the detailed description below are merely exemplary and explanatory and do not limit the features recited in the claims.

[0014] In this disclosure, the term "based on" means "based at least in part on." The singular forms "a," "an," and "the" include the plural unless the context dictates otherwise. The term "exemplary" is used to mean "example" rather than "ideal." The terms "comprises," "comprising," "includes," and "including," or variations thereof, are intended to encompass a non-exclusive inclusion, such that a process, method, or product that includes a list of elements does not necessarily include only those elements, but may include other elements not expressly listed or inherent in such process, method, article, or apparatus. Relative terms, such as "about," "approximately," "approximately," and "generally," are used to indicate that a variation of ±10% of the stated or understood value is possible. Additionally, the term "between," when used in describing a range of values, is intended to include the minimum and maximum values ​​set forth herein. The use of the term "or" in the claims and this specification is used to mean "and / or" unless clearly indicated to refer to alternatives only or unless the alternatives are mutually exclusive, but this disclosure supports a definition that refers to alternatives only and "and / or." As used herein, "another" can mean at least a second or more.

[0015] As used herein, the term "user" generally encompasses any person or entity, such as a researcher and / or medical professional (e.g., a physician), who may desire information, problem resolution, or any other type of interaction with a provider of the systems and methods described herein (e.g., via an application interface resident on their electronic device, etc.). The terms "electronic application" or "application" may be used interchangeably with other terms, such as "program," and generally encompass software configured to interact with, modify, override, supplement, or cooperate with other software.

[0016] As used herein, a "machine learning model" generally encompasses instructions, data, and / or a model configured to receive an input and apply one or more weights, biases, classifications, or analyses to the input to generate an output. The output may include, for example, a prediction, suggestion, or recommendation associated with the input, a dynamic action performed by the system, or any other suitable type of output-based analysis. Machine learning models are generally trained using training data, e.g., empirical data and / or samples of input data, that are fed to the model to establish, adjust, or change one or more aspects of the model, e.g., weights, biases, criteria for forming classifications, or clusters. Aspects of a machine learning model may operate on inputs linearly, in parallel, via a network (e.g., a neural network), or via any suitable configuration.

[0017] DNA methylation has long been regarded as a hallmark of cancer and holds great promise for early-stage cancer detection. In particular, through the use of targeted whole-genome bisulfite sequencing ("WGBS") in conjunction with machine learning techniques and associated processing power, methylated DNA sequences can be efficiently read and aberrantly methylated sequences (i.e., sequences that may be indicative of cancer) can be identified. Thus, targeted DNA methylation assays performed on free DNA ("cfDNA") fragments may be capable of detecting multiple cancers across all stages, including early stages when treatment may be more effective.

[0018] GRAIL's Multi-Cancer Early Detection (MCED) test uses targeted methylation assays from plasma samples with trained classifiers to detect the presence of invasive cancer and identify the source of the cancer signal (i.e., "tissue of origin") in the patient's body. Additionally, GRAIL's post-diagnostic program provides a non-tissue-informed estimate of circulating methylated variant allele frequency (mVaF) (i.e., circulating tumor fraction) to quantify disease burden and detect minimal residual disease (MRD). GRAIL has also demonstrated the ability to distinguish different cancer subtypes from cfDNA. More specifically, techniques have been developed that can broadly distinguish cancers by their origin, and more specifically, separate solid cell cancers into histological types or subtypes, such as adenocarcinoma, HPV-associated squamous cell carcinoma, non-HPV-associated squamous cell carcinoma, adenocarcinoma of Müllerian origin, and transitional cell carcinoma. GRAIL can further distinguish type II ovarian and uterine high-grade serous carcinoma (more aggressive subtypes) from type I ovarian endometrioid and uterine endometrioid carcinoma (less aggressive subtypes) using cfDNA methylation patterns.

[0019] Embodiments of the present application may utilize existing systems and processes to identify cancer subtypes, as further described herein. Embodiments of the present disclosure are generally drawn to the use of methylation data to define molecular subtypes of cancer. The present disclosure provides an in-depth case study to define subtypes of small cell lung cancer (SCLC) as an example of cancer.

[0020] SCLC is an aggressive form of lung cancer often associated with a poor prognosis. In contrast to non-small cell lung cancer (NSCLC), where therapies targeting genomic alterations in diversely expressed genes have been successfully deployed, SCLC has not benefited from targeted therapies, in part because SCLC malignancies are almost exclusively caused by loss-of-function mutations in the tumor suppressor genes TP53 and RB1.

[0021] Subtypes of SCLC have been identified using gene expression profiling, with four major subtypes defined based on the expression levels of three key transcription factors: ASCL1, NEUROD1, and POU2F3.

[0022] Surgical resection is not frequently performed for patients with advanced SCLC, and SCLC generally goes undetected until late stages. As a result, access to tumor tissue may be limited. However, blood-based methylation assays may offer advantages in research and development and clinical settings. Methylation-based approaches in cfDNA have been attempted to classify lung cancer into broader types, i.e., NSCLC (and its subtypes: lung adenocarcinoma, lung squamous cell carcinoma, and large cell carcinoma) versus SCLC. However, such approaches may not identify how to further subdivide SCLC into its constituent subtypes. Furthermore, such approaches focus on methylation levels derived from quantitative methylation-specific PCR in only four prespecified genes (APC, HOXA9, RARB2, and RASSF1A), which do not include the three SCLC subtypes defined in recent literature: transcription factors.

[0023] Accordingly, some aspects of the present disclosure are directed to the use of single-target blood-based methylation assays for early cancer detection, quantification of patient disease burden for treatment, and subclassification of cancers, such as SCLC. Identifying cancer molecular subtypes can be useful, for example, in understanding cancer, determining patient prognosis, or guiding treatment decisions. More specifically, machine learning techniques can be employed on targeted methylation data from blood-derived cfDNA to identify molecular cancer subtypes and predict the susceptibility of individual cancers to therapeutic intervention. Embodiments disclosed herein propose learning methylation signatures of cancer subtypes across different and potentially many more genomic loci than traditionally covered. Furthermore, cancer subclassification with information for therapy selection can be performed using the same blood draw and sample used to detect the presence of invasive cancer in a screening setting. While some embodiments of the present disclosure focus on SCLC, the methods described herein can be used to identify various molecular types of cancer using methylation data.

[0024] Another situation addressed by some embodiments of the present disclosure relates to the detection of resistance mechanisms that arise in some cancers during treatment. More specifically, a variety of targeted therapies exist and are available for the treatment of many common types of cancer, such as adenocarcinoma. For example, epidermal growth factor receptor (EGFR)-mutated lung adenocarcinoma is potentially treated with EGFR inhibition, and prostate cancer is potentially treated with hormone deprivation therapy or androgen receptor signaling inhibitors.

[0025] Selective pressures associated with cancer treatment can lead to the development of specific resistance mechanisms. For example, lung adenocarcinomas treated with EGFR inhibition may exhibit resistance after a period of time, e.g., on the order of approximately 12 months. As another example, prostate cancers may evade treatment and transform into castration-resistant and potentially metastatic cancers. One mechanism for acquiring treatment resistance is tumor transdifferentiation, i.e., the transformation of the original adenocarcinoma into small cell neuroendocrine carcinoma. After transformation, neuroendocrine carcinomas may no longer respond to the initial therapy, necessitating the selection and implementation of a new treatment regimen. This type of treatment evasion has been reported in approximately 15–20% of late-stage prostate cancers and approximately 5–14% of adenocarcinomas resistant to EGFR inhibition.

[0026] Traditionally, transdifferentiation has typically only been detected through re-biopsy of metastatic, recurrent, or growing primary tumors after initial treatment failure (i.e., observed by tumor recurrence). Re-biopsy can be a clinically intensive process, and by the time a re-biopsy is performed, it may generally be too late to successfully adjust treatment. A solution has been proposed to detect transdifferentiation from ctDNA using genomic profiling in conjunction with targeted small variant assays. However, there are often too many genomic abnormalities shared between the original adenocarcinoma and transformed small cell carcinoma to discern a clear distinction between each genomic variant. For example, mutations in the p53 and Rb1 pathways, which are nearly universal in small cell carcinoma, are also found in adenocarcinoma and are not specific to neuroendocrine transformation.

[0027] Treatment monitoring often involves disease burden assessment or detection of MRD, which corresponds to the small number of cancer cells that may remain in the body after treatment (e.g., tumor removal, therapy administration, etc.). During or after cancer treatment, any remaining cancer cells may become active and begin to proliferate, potentially leading to disease recurrence. Thus, detection of MRD may indicate that treatment is not fully effective or that treatment was incomplete. MRD may exist, for example, because certain cancer cells have become resistant to the drugs used, for example, through transdifferentiation. MRD may be detected, for example, via liquid biopsy. However, even with available liquid biopsy assays to detect transdifferentiation to small cell carcinoma, secondary assays and analyses are required to obtain information about MRD and disease burden, and possibly transdifferentiation, for disease management, which may be time-consuming and / or cumbersome.

[0028] In light of the above, therefore, some aspects of the present disclosure are directed to utilizing a single targeted liquid biopsy (e.g., blood-based) methylation assay for both MRD assessments simultaneously, as described above, to perform parallel disease monitoring and identify cancer recurrence and / or the development of resistance mechanisms, e.g., via transdifferentiation. In this regard, machine learning techniques may be employed to recognize and distinguish cancer types, e.g., adenocarcinoma and neuroendocrine carcinoma, by training models to recognize cancer signal of origin (CSO) based on tissue-type labeling rather than anatomical site-based labeling. One or more downstream actions, e.g., alternative treatment suggestions and administration of different therapeutic regimens, may then be performed.

[0029] Another situation addressed by some embodiments of the present disclosure relates to patient prognosis. More specifically, patient prognosis may provide an estimate of the course and / or outcome of cancer by assessing the patient's risk of death and / or disease progression of the cancer. Knowledge of a patient's prognosis may be valuable in clinical trial design and / or therapy selection, for example, to prescribe more aggressive treatments to patients with a worse prognosis, and in contrast, to prescribe more conservative treatments to patients with a less severe prognosis.

[0030] Prognostic biomarkers are desirable for the management of subjects identified as having cancer. Understanding a subject's prognosis at the cancer detection stage, particularly for early cancer detection, can influence one or more treatment regimens (including surgical or non-surgical), follow-up screening schedules, or other aspects of cancer and patient management. For example, in the case of stage I NSCLC, prognostic biomarkers can identify subjects with a poorer prognosis and therefore a higher risk profile, which can then justify neoadjuvant therapy for those subjects. A wide range of biomarkers, including proteomic and genomic biomarkers, are available for patient prognosis. For example, tumor fraction (mVAF) is a biomarker that can be an important prognostic feature. More specifically, increased tumor fraction can be associated with many biological phenomena known to indicate poor patient prognosis, such as increased tumor size, the presence of tumor-involved lymph nodes, the presence of distant metastases, increased metabolic and mitotic activity of tumor cells, and tumor invasion of various biological structures (e.g., blood vessels, lymphatic vessels, adjacent structures, etc.).

[0031] Existing prognostic information obtained from liquid biopsies condenses the presence of ctDNA into a single metric or ctDNA score, namely, tumor fraction, mVAF, or ctDNA concentration. While this score can provide valuable insight into a subject's disease status, it is largely binary in nature; for example, a large tumor fraction corresponds to a poor prognosis, while a small tumor fraction corresponds to a good prognosis. However, studies have shown that some subjects with a relatively large tumor fraction may still perform relatively well, even better than other patients with a relatively low tumor fraction. Therefore, even when ctDNA scores are similar, it may be advantageous to utilize additional patient-specific information to better capture the broad heterogeneity of cancer and its impact on prognosis.

[0032] Therefore, in view of the above, some aspects of the present disclosure are directed to utilizing DNA methylation features to improve the accuracy of subject prognosis. More specifically, in one embodiment, a supervised machine learning classifier may be trained with DNA methylation data labeled with known outcomes for each subject (e.g., each label provides an indication of subject survival after X years, along with an indication of disease progression, etc.). A prognostic score may be obtained from the classifier and then combined with the ctDNA score to generate a final prognostic score for the subject. This final prognostic score may provide a more accurate indication of subject prognosis than using cTF, VAF, or ctDNA concentrations alone.

[0033] More specifically, the circulating tumor fraction (cTF) or methylated variant allele frequency (mVAF) measures the fraction of DNA molecules in a plasma sample (cTF) that have methylation signals that are approximately unique to tumor cells (mVAF). Circulating tumor DNA molecules carry methylation signals that contain distinct methylation patterns that characterize tumors, independent of the total amount captured by cTF or mVAF, and thus can enable prognosis. Classifiers can be trained to learn and recognize methylation signals that indicate a good or bad prognosis. Thus, once two independent prognostic metrics (i.e., the amount of ctDNA (cTF or mVAF) and the methylation pattern associated with ctDNA) are obtained, these metrics can be combined into one final metric (e.g., by multiplying the two metrics together).

[0034] The subject matter of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof and which show, by way of illustration, certain exemplary embodiments. Any embodiment or implementation described herein as "exemplary" should not be construed as preferred or advantageous over other embodiments or implementations, for example, but rather is intended to reflect or indicate that the embodiment is an "exemplary" embodiment. The subject matter may be embodied in a variety of different forms, and thus, the subject matter encompassed or claimed should not be construed as limited to any exemplary embodiment described herein, which exemplary embodiment is provided merely as an example. Likewise, a fairly broad scope of subject matter recited or encompassed in the claims is intended. Among other things, for example, the subject matter may be embodied as a method, device, component, or system. Thus, embodiments may take the form of, for example, hardware, software, firmware, or any combination thereof. Therefore, the following detailed description is not intended to be taken in a limiting sense.

[0035] Throughout this specification and the claims, terms may have a particular meaning beyond that expressly stated, that is suggested or implied by the context. Similarly, the phrases "in one embodiment" or "in some embodiments" used herein do not necessarily refer to the same embodiment, and the phrase "in another embodiment" used herein does not necessarily refer to different embodiments. For example, it is intended that claimed subject matter include, in whole or in part, any combination of the example embodiments.

[0036] 1A illustrates an exemplary system for identifying cancer subtypes using a targeted methylation assay. The exemplary system 100 includes a data collection component 10, a database 20, and a device data intelligence component 30, which are connected to each other via a network 40. Alternatively or additionally, one or more of the components may be locally connected to another component, for example, through a wired connection, without relying on a network connection. The concept is illustrated using sequencing data of free nucleic acids. However, those skilled in the art will understand that the method is equally applicable to sequencing data or non-sequencing data of other substances.

[0037] As disclosed herein, the data collection component 10 may include a device or machine capable of generating sequencing data. In some embodiments, the data collection component 10 may include a sequencing machine or equipment for generating nucleic acid sequence data of a biological sample using a sequencing machine. Any suitable biological sample may be used. In some embodiments, the biological sample is cell-based, e.g., one or more tissues. In some embodiments, the biological sample is a sample containing free nucleic acid fragments. Examples of biological samples include, but are not limited to, a blood sample, a serum sample, a plasma sample, a urine sample, a saliva sample, etc.

[0038] Examples of sequencing data can include, but are not limited to, sequence read data of a target genomic location, partial or whole genome sequencing data of a genome represented by nucleic acid fragments in a free sample or a cell-based sample, partial or whole genome sequencing data that includes one or more epigenetic modifications (e.g., methylation), or combinations thereof.

[0039] Data acquired by data collection component 10 may be transferred to database 20 via network 40. In some embodiments, the collected data may be analyzed by data intelligence component 30 via a local or network connection. Figure 1B shows exemplary functional modules that may be implemented to perform the tasks of data intelligence component 30.

[0040] 1B illustrates an exemplary computer system 110 for processing sequencing data. The exemplary embodiment 110 achieves such functionality by implementing, on one or more computer devices, a user input / output (I / O) module 120, a memory or database 130, a data processing module 140, a data analysis module 150, a classification module 160, a network communication module 170, and any other functional modules (e.g., error correction or compensation module, data compression module, etc.) that may be necessary to perform a particular task. As disclosed herein, the user I / O module 120 may further include an input sub-module, such as a keyboard, and an output sub-module, such as a display (e.g., a printer, monitor, or touchpad). In some embodiments, all functionality is performed by one computer system. In some embodiments, functionality is performed by two or more computers.

[0041] Also, as disclosed herein, a particular task may be performed by implementing one or more functional modules. Notably, each of the listed modules may itself include multiple sub-modules. For example, data processing module 140 may include a sub-module for data quality assessment (e.g., discarding very short sequence reads or sequence reads containing obvious errors), a sub-module for normalizing the number of sequence reads aligned to different regions of a reference genome, a sub-module for compensating / correcting for GC bias, etc.

[0042] In some embodiments, a user may use I / O module 120 to manipulate data available on the local device or data that can be obtained either from a remote service device or another user device via a network connection. For example, I / O module 120 may enable a user to perform data analysis via a graphical user interface (GUI), e.g., via a keyboard, mouse, or touchpad. In some embodiments, a user may manipulate data via voice control. In some embodiments, user authentication may be required before the user is granted access to the requested data.

[0043] In some embodiments, user I / O module 120 may be used to manage various functional modules. For example, while an existing data processing session is in progress, a user may request input data via user I / O module 120. The user may do so by selecting menu options or by directly typing commands without interrupting the existing process.

[0044] A user may use any type of input to direct and control the processing and analysis of data via I / O module 120 as disclosed herein.

[0045] In some embodiments, system 110 further includes memory or database 130. In some embodiments, database 130 includes a local database that may be accessed through user I / O module 120. In some embodiments, database 130 includes a remote database that may be accessed by user I / O module 120 over a network connection. In some embodiments, database 130 is a local database that stores data retrieved from another device (e.g., a user device or a server). In some embodiments, memory or database 130 may store data retrieved in real time from an internet search.

[0046] In some embodiments, database 130 may exchange data with one or more other functional modules, including, but not limited to, a data collection module (not shown), a data processing module 140, a data analysis module 150, a classification module 160, a network communication module 170, etc.

[0047] In some embodiments, database 130 may be a database local to other functional modules. In some embodiments, database 130 may be a remote database that may be accessed by other functional modules via a wired or wireless network connection (e.g., via network communications module 170). In some embodiments, database 130 may include a local portion and a remote portion.

[0048] In some embodiments, system 110 includes a data processing module 140. Data processing module 140 may receive real-time data from I / O module 120 or database 130. In some embodiments, data processing module 140 may perform standard data processing algorithms, such as one or more of noise reduction, signal enhancement, sequence read count normalization, GC bias correction, etc. In some embodiments, data processing module 140 may identify global or local systematic errors. For example, sequencing data may be aligned to regions within a reference genome. The number of sequence reads aligned to different genomic regions may vary within the same subject. The number of sequence reads aligned to the same genomic region may vary from subject to subject. Depending on these differences, differences observed, particularly in healthy subjects (i.e., including all species of organisms, not just humans), may be attributable to systematic errors rather than to an association with one or more pathological conditions. For example, if sequencing data corresponding to a particular genomic region exhibits a wide range of variation among healthy subjects, data processing module 140 may classify the particular genomic region as a high-noise region and exclude the corresponding data from further analysis. In some embodiments, identification and processing of possible systematic errors may be performed by data analysis module 140, as described below.

[0049] In some embodiments, the system 110 includes a data analysis module 150. In some embodiments, the data analysis module 150 includes identifying and processing systematic errors in the sequencing data, as described with respect to the data processing module 140.

[0050] In some embodiments, the system 110 includes a classification module 160 that analyzes data from test subjects whose status with respect to a medical condition is unknown and subsequently classifies the unknown test subjects based on the likelihood that the subject fits into a particular category. In some embodiments, the one or more parameters include a binomial probability score calculated based on a logistic regression analysis. As disclosed herein, the binomial probability score may correspond to the likelihood that the subject has a particular medical condition, such as cancer. For example, a score above a predefined threshold may indicate the likelihood that the subject has a particular subtype of cancer. In some embodiments, the one or more parameters may include a sequencing data distribution pattern that correlates with the presence of a particular cancer type or subtype. Subjects with a pattern similar to the cancer pattern may be diagnosed as suffering from that type of cancer. In some embodiments, the sequencing data distribution pattern may be jointly identified as two or more specific types of cancer, thus allowing for further classification of the unknown subject.

[0051] As disclosed herein, the network communication module 170 may be used to facilitate communication between user devices, one or more databases, and any other suitable systems or devices through a wired or wireless network connection. Any communication protocol / device may be used, including, without limitation, an Ethernet connection, a network card (wireless or wired), an infrared communication device, a wireless communication device and / or chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication facility, etc.), near field communication (NFC), Zigbee communication, radio frequency (RF) or radio frequency identification (RFID) communication, a PLC protocol, 3G / 4G / 5G / LTE-based communication, etc. For example, a user device having a user interface platform that processes / analyzes low-coverage sequencing data may communicate with another user device having the same platform, a regular user device (e.g., a regular smartphone) that does not have the same platform, a remote server, a physical device in a remote IoT local network, a wearable device, a user device communicatively connected to a remote server, etc.

[0052] The functional modules described herein are provided as examples. It will be understood that different functional modules can be combined to create different utilities. It will also be understood that additional functional modules or sub-modules may be created to implement specific utilities.

[0053] Figures 2-6 below show how a trained cancer classifier uses cfDNA methylation data to distinguish cancer subtypes.

[0054] Referring now to Figure 2, a confusion matrix is ​​provided presenting data analyzed by a machine learning model (i.e., a trained cancer classifier) ​​for samples identified as having a particular histological type or cancer subtype. In one embodiment, a classifier using cfDNA methylation as input was trained to distinguish cancer cell lineages (epithelial, neural, lymphoid, myeloid, plasma, mesenchymal, melanocyte, neuroendocrine, and germ cell) and further distinguish between all epithelial cancers (the most common types of solid cancers) and their histological types: adenocarcinoma (ADC), HPV-associated squamous cell carcinoma (HPV), non-HPV-associated squamous cell carcinoma (SCC), adenocarcinoma of Müllerian origin (Müllerian), and transitional cell carcinoma. The classifier was evaluated in cross-validation on a set of 2,920 plasma samples from cancer patients and 3,704 plasma samples from cancer-free patients. In 1,619 samples, the plasma samples contained detectable levels of circulating tumor DNA (ctDNA).

[0055] In the confusion matrix shown in Figure 2, each row contains samples from one ground truth occurrence series and cancer histology type. Each column represents samples with detectable ctDNA that this trained classifier identified as having this histology type or cancer subtype. Samples counted on the diagonal represent instances where the classifier correctly identified cancer by its embryological origin and histology type based on methylation signals in cfDNA.

[0056] Referring now to FIG. 3, a table listing the accuracy (number of cases detected and number with correct cancer signal origin result) broken down by each ground truth class is presented.

[0057] Figures 3 and 4 show a CSO accuracy of 84% and 86.3% for identifying neuroendocrine cancers and distinguishing them from epithelial cancers. In the case of lung cancer, this distinction allows the classifier to reliably separate non-small cell lung cancer (NSCLC) from small cell lung cancer (SCLC) using only plasma samples obtained from blood draws, without the need to obtain tissue samples from the lung cancer.

[0058] Referring now to Figure 4, we present a table listing the accuracies of the classifiers discussed in Figure 2. For each class, we list the frequency with which the class returned by the classifier was correct and the frequency with which samples actually had this class. Numbers are presented once for all cases detected by the classifier (cso_precision) and once restricted to cases that were actually invasive cancer (cso_precision_tp). In this latter example, results from samples where the patient did not have cancer (false positive detections) are not included.

[0059] Referring now to Figure 5, a heatmap showing classifier results for ovarian and uterine cancers with different histological types and different anatomical locations of primary cancer is presented. In one embodiment, the classifier was trained using classes defined by anatomical origin and further trained on two distinct subtypes of ovarian and uterine cancer. More specifically, Class 1 ovarian cancers have their cell of origin in the epithelial cells of the ovary. These are typically indolent cancers with a relatively good prognosis. Class 2 ovarian cancers are cancers whose cell of origin likely resides in the fallopian tube. These neoplasms of the fallopian tube disseminate tumor cells very early (perhaps from a neoplastic growth of only a few hundred cells) and colonize the ovaries. These cancers are typically malignant and are almost exclusively detected when they have already metastasized to the ovaries. These tumor metastases from fallopian tube origin are also found in the lining of the rectum and peritoneum. Together, Type II ovarian cancers and these similar cancers of fallopian tube origin are characterized as high-grade serous carcinomas.

[0060] In the classifier shown in Figure 5, high-grade serous carcinoma was trained as one class (i.e., named hgs_pelvis), and ovarian and uterine epithelial grade 1 or endometrial cancer was trained as a second class (i.e., named ovary_uterus). Figure 5 shows the classifier results for all ovarian and uterine cancers. More specifically, the heatmap color varies from low classifier signal (blue) to high classifier signal (yellow). The classifier results for each sample are displayed as column labels, and the scores are shown as row labels. Type II high-grade serous carcinoma can be distinguished from type I low-grade epithelial carcinoma. One of the cases was misclassified as pancreatic or gallbladder cancer, and not all of the plasma samples had detectable ctDNA (classifier result non_cancer).

[0061] Referring now collectively to Figures 6A and 6B, we provide data plots and confusion matrices showing the classifier results from Figure 5 separated by ground truth sequence labels: type II high-grade serous ovarian cancer, type I epithelial ovarian cancer, high-grade serous uterine cancer, and epithelial endometrial uterine cancer. Type II high-grade serous ovarian cancer is identified with high confidence by this classifier (20 cases had ctDNA detected and correctly classified, 2 cases had no detectable ctDNA, and 2 cases had detectable ctDNA but were misclassified as type I cancer).

[0062] Some or all of the processes identified in Figures 2-6 above may also be applicable to distinguishing SCLC subtypes, as further described herein and illustrated in Figures 7-26. Additionally, the entire disclosures of commonly owned U.S. Patent Application Publication Nos. 2020 / 0365229 and 2021 / 0313006, excluding any definitions, subject matter disclaimers, or disclaimers, are incorporated herein by reference, except to the extent the incorporated material contradicts the express disclosure herein, in which case the terms in this disclosure shall control. Some or all of U.S. Patent Application Publication Nos. 2020 / 0365229 and 2021 / 0313006 may also be applicable to distinguishing cancer subtypes.

[0063] Referring now to Figures 7A and 7B, we provide heat maps of the different known subtypes of SCLC previously documented in the literature. These subtypes are currently defined by tissue expression signatures, more specifically, by the expression levels of three transcription factors, namely, three genes: ASCL1, NEUROD1, and POU2F3. The expression level of a fourth transcription factor, YAP1, may also be associated with subtype. The SCLC subtypes in Figures 7A and 7B were identified using RNA sequencing on tissue samples and cell lines. Cancer tissue samples are often unavailable for the management of patients diagnosed with SCLC. Therefore, it is desirable to identify these and other subtypes using only plasma samples obtained using a simple blood draw. Furthermore, examination of the heat maps in Figures 7A and 7B indicates that the expression levels of these genes are generally mutually exclusive in the tissue and cell line data presented. However, data generated from tissue samples may not reflect the intratumor heterogeneity that can be detected from plasma samples, which may provide a more complete understanding of SCLC subtype heterogeneity and its therapeutic implications.

[0064] Referring now to Figure 8, a bar graph is provided showing the different types of samples obtained for analysis of methylation signatures used to distinguish SCLC subtypes. More specifically, the samples were collected as part of the Circulating Cell-Free Genome Atlas (CCGA) study, in which 15,000 participants with and without cancer participated in a prospective observational case-control study for training and validation of the Galleri classifier. Fifty-five patients with SCLC were selected for the analytical process described in this invention, separated by patient type and clinical stage, as shown in Figure 8, and 57 non-cancer samples were matched by age, sex, and smoking status.

[0065] Referring now to Figures 9A and 9B, hypothetical gene body diagrams of transcription factor gene bodies and motifs are provided. One such hypothetical transcription factor could be, for example, ASCL1. While no information currently exists about the response of sampled patients (e.g., as shown in Figure 8) to therapy, methylation patterns in different regions of the genome that are biologically relevant to the subtype-defining transcription factors can be examined. More specifically, subtype-defining transcription factors encode proteins that bind to specific genomic sequences (i.e., motifs) that can be found throughout the genome. Through the investigations described herein, stratification of regions biologically relevant to known subtype biomarkers can be demonstrated. In particular, using methylation data, substructures can be observed in the methylation signal that allow for the generation of the subtypes. One or all of these candidate subtypes may ultimately prove to be predictive of treatment response or prognosis of clinical outcome.

[0066] Referring now to FIG. 10A, a general flowchart of the cfDNA methylation testing process is provided. First, a sample, e.g., a blood sample, may be collected from a subject. The cfDNA in the sample may be evaluated through the use of bisulfite conversion, which may distinguish between methylated and unmethylated cytosines in the cfDNA. Library preparation may then be performed through suitable means, and an enrichment step may be performed. Sequencing may then be performed, followed by demultiplexing and alignment, and then methylation calling. Based on the methylation data, classification of cancer type, histology type, or subtype may be performed. Finally, quality control and reporting to a healthcare professional or patient may be performed. Embodiments of the present disclosure may provide more data for classification or may output additional data during the classification step.

[0067] A flow chart of how pilot analysis of methylation data can be incorporated into the current cfDNA methylation testing process is provided in FIG. 10B. As shown in FIG. 10B, following methylation calling, the methylation data can be analyzed to identify cancer subtypes, and analysis can be performed based on the identified cancer subtypes. This information can be used to refine the classification step or to broaden the range of clinical or biological predictions that can be provided by the classifier. This methylation analysis step, referred to as the methylation toolbox in FIG. 10B, can enable exploration of methylation signature substructures in samples. In some embodiments, analysis of methylation signature substructures can be performed very rapidly on thousands of samples.

[0068] Referring now to Figure 11, a diagram of how a methylation beta data value or methylation-like beta value is calculated is provided. The beta value refers to the proportion of molecules identified at a particular CpG site that are methylated. This proportion can be calculated by dividing the number of methylated molecules by the total number of molecules. For illustrative purposes, red represents methylated molecules and blue represents unmethylated molecules.

[0069] 12, which provides a heatmap 1200 illustrating methylation beta values ​​per CpG in specific regions of gene bodies and promoters of SCLC-associated transcription factors. The CpGs selected are those in the gene body and promoter regions of SCLC subtypes that define the transcription factors, although alternative regions may be selected, including regions identified as potential transcription factor binding sites, regions or individual CpG loci identified as differentially methylated between known transcriptionally defined SCLC subtypes, regions or individual CpG loci identified as differentially methylated between responders and non-responders to a particular treatment, and / or regions or individual CpG loci that correlate with any of the above definitions.

[0070] Referring to heatmap 1200, CpGs are shown in rows, while samples are shown in columns. The rows are annotated to indicate the gene in which each CpG is located and whether it is in the gene body or promoter. Examination of heatmap 1200 may reveal "blockiness" in areas of the heatmap structure where samples are hypermethylated or hypomethylated. For example, heatmap 1200 reveals that sample 122's cluster has hypermethylation along the gene body 1222 and promoter region 1224 of NEUROD1 1226. As another example, sample 124 has hypermethylation along the gene body 1242 and promoter region 1244 of ASCL1 1246. As yet another example, sample 126 has hypomethylation along the gene body 1262 of NEUROD1 1264.

[0071] Embodiments of the present disclosure aim to use the β values ​​in such a matrix as feature values ​​and the feature values ​​in a set of labeled training samples to create a cancer subtype classifier - in this case, an SCLC subtype classifier - to create a machine learning model (e.g., a penalized regression model, a support vector machine, a shallow neural network, etc.) that distinguishes between different SCLC subtypes defined either by traditional expression profiles or by response to treatment.

[0072] In addition to the β value as a feature value for this classifier, the machine learning model of this embodiment may also utilize the total count of methylated and unmethylated molecules in each region or CpG locus and / or the count of aberrantly methylated molecules in each region of CpG loci. Thus, a single run of GRAIL's current targeted methylation assay can be applied for either early cancer detection or disease burden monitoring, and can also be used in parallel for the detection of cancer subtypes, e.g., SCLC subtypes, to inform prognosis, inform treatment decisions, or identify the occurrence of different treatment-resistant subtypes.

[0073] Referring now to Figures 13A-13F, dimensionality reduction was performed to reduce the number of variables in the data. The charts shown in Figures 13A-13F show how the samples from Figure 12 cluster after dimensionality reduction was performed. In Figures 13A-13F, principal component analysis (PCA) dimensionality reduction was performed on the data, and in Figures 13C and 13D, uniform manifold approximation and projection (UMAP) nonlinear dimensionality reduction was further performed on the data. The colors of the participant samples in Figures 13A, 13B, and 13E correspond to the three cluster colors assigned below the dendrogram at the top of Figure 12, allowing visualization of where the participant samples shown in Figure 12 map onto the two-dimensional charts in Figures 13A, 13B, 13E, and 13F. In Figure 12, there were roughly six tree clusters identified based on the methylation signatures of the samples. These six tree clusters represent subtypes imposed on the SCLC data based on the methylation signature of each sample; therefore, Figures 13A-13F can show whether the subtypes inferred using the methylation signatures correspond to significant differences in participant samples or to other clinical or patient information, such as gender or smoking status, that are not indicative of cancer subtype.

[0074] Figure 13A shows that participant samples from the three cluster rows in Figure 12 appear to be grouped into similar clusters in the two-dimensional space defined by Principal Component 1 (PC1, 73.1% explained variance) and Principal Component 2 (PC2, 6.36% explained variance). Figures 13C and 13D show the dataset stratified by gender and smoking status, respectively. Given that the data points for each participant sample appear to be uniformly distributed across clusters by both gender and smoking status, Figures 13C and 13D can indicate that the clusters shown in the charts do not appear to be due to these other covariates, which can introduce spurious signals when assessing cancer data. Thus, SCLC subtypes identified using methylation status can be attributed to significant differences in methylation signatures across samples, as opposed to confounding variables. Figure 13E again shows that the same tree clusters—or subtypes—tend to cluster together. Figure 13F generally shows sample clustering by cancer status, although there are some outliers. The "Stage" row at the top of Figure 12 indicates the stage of lung cancer, with blue samples indicating non-cancer samples and increasingly dark purple samples indicating increasingly advanced cancer stages. The colors of these samples in Figure 12 correspond to the data points shown in Figure 13F. Cancer samples found to cluster with non-cancer samples in Figure 13F may have a low tumor fraction, and therefore it may not be possible to see a large signal in these samples, which may result in them clustering with the non-cancer samples as opposed to the cancer samples.

[0075] Referring now to FIG. 14A, a line graph 1400 is provided based on the heat map 1200 data shown in FIG. 12. Such a line graph 1400 can provide a clearer visual indication of the genomic context of the methylation status of each cluster of samples. In this graph, the data presented from FIG. 12 is restricted to one specific gene, namely, NEUROD1. β values ​​are plotted on the y-axis of the line graph 1400, and genomic locations are plotted on the x-axis of the line graph 1400. Generally, a larger β value indicates more methylation of the DNA fragment in this region. Thus, as a representative example of the information that can be obtained from such a graph, one may look at sample cluster 4, which corresponds to the cluster of sample 122 discussed above in FIG. 12. It can be seen in the line graph 142 that the samples associated with sample cluster 4 are all hypermethylated at NEUROD1, which is consistent with what is presented in FIG. 12. It can also be seen that NEUROD1 is hypermethylated both away from and close to the transcription start site ("TSS").

[0076] Referring now to Figure 14B, we present a heatmap previously documented in the literature. See https: / / pubmed.ncbi.nlm.nih.gov / 33482121 / . This heatmap also shows beta values, with more yellow regions indicating more methylation and more blue regions indicating less methylation. This heatmap, derived from cell lines, demonstrates that subtypes defined by the expression levels of different transcription factors are associated with methylation levels in cell lines at different distances from the TSS of the gene NEUROD1. The heatmap data show methylation beta values ​​for expression-defined SCLC subtypes termed "ASCL1 only," "NEUROD1 only," and "both," defined by the expression of either the gene ASCL1 or NEUROD1 or both, and these subtypes are indicated by the dark bands at the top of the heatmap in Figure 14B. ASCL1-only samples in the heatmap tend to have hypermethylation both near the TSS, approximately 200 base pairs away, and further from the TSS, approximately 1,500 base pairs away. This pattern of hypermethylation relative to the NEUROD1 TSS in the ASCL1-only samples in heatmap 14B is also seen in samples from cluster 4, which show hypermethylation both 200 base pairs from the NEUROD1 TSS and 1,500 base pairs from the NEUROD1 TSS, as shown in Figure 14A. This correspondence in the data may suggest that the subtype identified in sample cluster 4 may correspond to the ASCL1-high subtype. Thus, this correlation may indicate that it is possible to identify cancer subtypes using methylation signatures as described herein. Figure 14A also shows that there may be additional layers of stratification that can be learned by looking at methylation patterns in the NEUROD1 gene body; for example, samples from cluster 3 in red show intermediate-range beta values ​​more than 5,000 base pairs from the NEUROD1 TSS, which are not discernible in Figure 14B. Thus, using targeted methylation assays, it may be possible to analyze methylation signatures as described herein to identify subtypes previously defined in the literature using conventional methods and to reproduce these patterns using methylation signatures.Embodiments of the present disclosure may utilize targeted methylation assays and pattern recognition to recreate existing subtypes or discover new subtypes that may be useful, for example, in identifying cancer, determining prognosis, or informing treatment options.

[0077] A similar pattern is shown in the line graphs of Figures 15-18, each looking at a different region of the genome. Referring now to Figure 15, there is provided a line graph 1500 of methylation values ​​within the gene body of the transcription factor ASCL1 based on data from the heatmap shown in Figure 12. Tree cluster 1 shown in Figure 15 is separated from the rest of the samples and corresponds to sample 124 shown in Figure 12.

[0078] 16-18, we provide line graphs of methylation values ​​within the gene body of POU2F3 (i.e., 1600-1800) based on the heatmap data shown in Figure 12. As discussed above with reference to Figures 14A and 15, Figures 16-18 call out the methylation patterns seen in different genomic regions.

[0079] Referring now to FIG. 19, another heatmap 1900 showing methylation beta values ​​is provided. Using CpGs or binding sites within transcription factor motifs as a starting point, a similar analysis to that described above (e.g., with reference to FIG. 12) can be performed. More specifically, instead of looking at CpGs in the gene bodies of these four transcription factors, we can look at CpGs within regions identified as harboring motifs or binding sites for SCLC-associated transcription factors, e.g., ASCL1, NEUROD1, and POU2F3. In this analysis, motifs were identified from HOCOMOCO [cite], their locations in the genome were identified using PWMTools [cite], and CpGs within 50 base pairs on either side of these motifs (referred to as "motif windows") were identified. Furthermore, a β2-term model was used to compare methylation values ​​from SCLC patients with those from non-cancer patients to identify differential methylation, and CpGs from these motif locations were filtered to require significant differential methylation. CpGs that were present in more than one motif window or that overlapped a CTCF motif window were also removed. In examining Figure 19, structure can also be seen in the heatmap, which structure differs from the structure observed from the heatmap shown in Figure 12. In other words, because an alternative set of genomic regions was used, the clusters learned from this heatmap may differ from those learned in Figure 12 and may provide different information than those learned in Figure 12.

[0080] 20A-20F, several charts are provided that show how the samples from FIG. 19 were clustered, similar to the cluster charts shown in FIGS. 13A-13F.

[0081] Referring now to Figure 21, a graph 2100 of sample size considerations is provided. In one embodiment, twice as many plasma samples should be allocated as are needed to build a classifier from tissue data (fewer may be sufficient for SCLC).

[0082] Referring now to Figure 22, a diagram illustrating how targeted methylation probes can improve testing efficiency is provided. More specifically, cfDNA fragments can have different methylation states, for example, methylated and unmethylated. Two types of probes, i.e., hyperprobes and hypoprobes, can be designed to target cfDNA fragments with the same methylation status at CpGs. However, this does not imply that all captured cfDNA is completely methylated or unmethylated. In one embodiment, each type of probe can be selected for one type of methylation state (after bisulfite conversion). For example, hyperprobes can be selected for methylated fragments, while hypoprobes can be selected for unmethylated fragments.

[0083] Referring now to Figure 23, an additional illustration is provided showing how targeted methylation probes can improve testing efficiency. More specifically, as shown in Figure 22, cfDNA fragments can have different methylation states. These fragments can be designed as binary targets or semi-binary targets. Regarding the former, binary targets can be targeted by both types of probes (i.e., hyperprobes and hypoprobes) described above. Conversely, regarding the latter, semi-binary targets can be targeted by only one type of probe (e.g., hyperprobes for methylated fragments and hypoprobes for unmethylated fragments).

[0084] As discussed above, some embodiments of the present disclosure may be relevant to detecting resistance mechanisms developed by some cancers during treatment. In this case, methylation assays may be used to identify transdifferentiation. To do so, machine learning techniques may be employed to recognize and distinguish cancer types, such as lung cancer and neuroendocrine cancer, by training models to recognize cancer signal of origin (CSO) based on tissue type labeling rather than anatomical site-based labeling.

[0085] Generally, after cancer is detected in a patient (e.g., via the use of one or more cancer detection techniques), appropriate treatment can be prescribed. Following initiation of treatment, routine testing for cancer (i.e., in the form of MRD in disease monitoring) can be performed at predetermined intervals, the length of which can depend on one or more factors (e.g., the severity of the cancer, the type or intensity of treatment, biological characteristics associated with the subject, etc.). Suitable time intervals can be on the order of weeks, months, or years, e.g., every 2-6 months (e.g., 2, 3, 4, 5, or 6 months), annually, every 2 years, etc. If cancer recurrence is detected, one or more contrast-enhanced scans can be performed (e.g., bone scans, whole-body scans, etc.), and a new course of treatment can be determined based on the results of the scans. Additionally or alternatively, identifying that transdifferentiation has occurred can provide an additional indicator metric of failure of a prescribed primary or current treatment. In this manner, targeted methylation assays can be used for disease monitoring. Furthermore, if transdifferentiation is detected and a second-line or further therapy is initiated, targeted methylation assays may continue to be used at regular intervals to determine whether the disease burden changes over time. Detecting changes in disease burden may help determine whether the second-line therapy is working and / or whether or when the second-line therapy fails.

[0086] Referring now to FIG. 24, machine learning models can be trained to determine whether MRD is present in a biological sample, and if so, whether cancer cells associated with the MRD have developed resistance mechanisms to the treatment.

[0087] In step 2405, a set of training data may be obtained from a source. In one embodiment, the source may be an accessible database, and the training data may include DNA methylation data derived from targeted methylation sequencing assays performed on biological samples obtained from training subjects. In one embodiment, the accessible database may be constantly or periodically updated with new training data.

[0088] In step 2410, the set of training data may be annotated so that each of the training data in the set is assigned a tissue label, as opposed to an anatomical site-based label. Cancer histology is traditionally determined in clinical practice by pathological review of tissue or other specimens. The training data may be labeled to indicate whether the DNA methylation data was obtained from a source having a given type of cancer, e.g., lung cancer, prostate cancer, etc. The label may include age, sex, race, or other information about the individual source of DNA methylation data (e.g., the patient's responsiveness to a particular course of treatment, etc.).

[0089] In step 2415, the annotated training data may be applied to a machine learning model. In one embodiment, the machine learning model may be almost any type of supervised, unsupervised, or hybrid machine learning model selected by the user.

[0090] In step 2420, application of the annotated set of training data to the machine learning model may optimize the pattern recognition capabilities of the algorithm associated with the machine learning model. More specifically, the algorithm of the machine learning model may be trained to: A) identify whether MRD is present in the test subject (i.e., by determining whether any cancer signal is present in the biological sample, or by imaging or otherwise), and, if present, B) determine whether cancer cells associated with MRD have developed a resistance mechanism to a previously administered or ongoing treatment. In this regard, the algorithm may be trained to recognize whether transdifferentiation has occurred by identifying that at least some of the cancer cells have transformed from a first cancer type (e.g., adenocarcinoma) to a second cancer type (e.g., small cell neuroendocrine carcinoma) based on the tissue label. In an optional embodiment, the machine learning model may be further trained to output new treatment recommendations in response to the development of a resistance mechanism and / or the identification of the type of resistance mechanism. In one embodiment, the new treatment recommendations may include suggesting a new drug therapy to administer to the patient, suggesting a surgical procedure to perform on the patient, adjusting the patient's liquid biopsy collection (e.g., blood draw), etc.

[0091] In one embodiment, the new treatment recommendation may take into account how the test subject may respond to various types of potential treatments. More specifically, the machine learning model may be able to identify biological patterns in the test subject's blood draw and predict the extent to which the test subject may respond to a potential type of treatment (e.g., by utilizing data procured through one or more dedicated training phases that identify how well other test subjects with similar biological patterns respond to various types of treatments). In one embodiment, the new treatment recommendation may include a recommendation for a disease monitoring schedule based on a determination of the extent to which the subject may respond to a potential treatment (e.g., if the test subject is expected to respond well to treatment, a six-month detection interval is suggested, whereas if the test subject is expected to respond poorly to treatment, a two-month detection interval is suggested, etc.).

[0092] Referring now to Figure 25, a trained machine learning model can be used to determine whether MRD is present, and if so, whether the cancer cells have developed resistance to ongoing treatment.

[0093] In step 2510, methylation data derived from a targeted methylation sequencing assay performed on a biological sample obtained from the test subject may be obtained. In step 2515, the methylation data may be applied to a trained machine learning model (e.g., the machine learning model discussed above with reference to FIG. 24). The methylation data may be processed by an algorithm of the trained machine learning model, and an output result may be received. More specifically, in step 2520, the machine learning model may determine whether MRD is present in the test subject. Such a determination may be facilitated by identifying whether any cancer signals are present in the methylation data. In response to determining in step 2520 that no cancer signals are present in the methylation data, the machine learning model may output a result indicating that cancer was not detected in step 2525. Alternatively, in response to determining that MRD is present in step 2520, the machine learning model may further determine in step 2530 whether a resistance mechanism has been developed by the cancer cells in response to ongoing treatment. This determination may be facilitated by identifying whether at least a portion of the detected cancer cells have transformed from a first cancer type (i.e., the cancer type that was first identified and for which the current treatment was originally prescribed) to another cancer type. The trained machine learning model algorithm may be capable of performing this identification based on a training dataset annotated with tissue labels for each cancer type. More specifically, when predicting tissue type from detected ctDNA in a liquid biopsy sample, such as a blood plasma sample, a robust separation between adenocarcinoma and neuroendocrine carcinoma may be identified. In response to determining in step 2530 that the cancer cells associated with MRD are equivalent to the first cancer type, the machine learning model may output a result in step 2535 indicating that the first cancer type has recurred in the test subject. Alternatively, in response to determining in step 2530 that at least a portion of the cancer cells associated with MRD have transformed from the first cancer type to a second cancer type, the machine learning model may output a result in step 2540 indicating that the cancer cells of the first cancer type have developed resistance to ongoing treatment.In an optional embodiment, the trained machine learning model may further output a treatment recommendation based on a determination that the cancer has developed resistance to the treatment. For example, the treatment recommendation may suggest an alternative therapy that may be administered to the test subject, a potential surgery that may be performed on the test subject, an adjustment to the test subject's liquid biopsy collection (e.g., blood draw) schedule, etc.

[0094] Referring now to FIG. 26 , a confusion matrix 2600 generated from an exemplary application of methylation data derived from a set of test subjects to a trained machine learning model is presented. In one embodiment, the machine learning model was trained by utilizing tissue cancer types as CSO labels. For clarity, only samples in which cancer signals were detected are shown. Column labels correspond to ground truth classes of tissue types, and row labels indicate classifier results. For interpretation purposes, "adc" corresponds to adenocarcinoma, "hpv" corresponds to HPV-positive squamous cell carcinoma, "scc" corresponds to HPV-negative squamous cell carcinoma, and "epithelial_nos" corresponds to epithelial carcinoma with unspecified histology. Examining the confusion matrix can reveal the accuracy of the machine learning model's prediction results. The confusion matrix indicates how frequently the classifier generates calls from an individual's blood draw that are consistent with the diagnosis determined for that individual based on clinical testing. Experimental results indicate that targeted methylation sequencing assays can be used for disease monitoring, and can be used in parallel with disease burden quantification or MRD detection and cancer transdifferentiation detection.

[0095] Referring now to FIG. 27 , a machine learning model may be trained to generate a subject prognosis score. In step 2705, a set of training data may be obtained from a source. In one embodiment, the source may be an accessible database, and the training data may include DNA methylation data derived from targeted methylation sequencing assays performed on biological samples, e.g., liquid samples such as blood samples, obtained from one or more training subjects. In one embodiment, the accessible database may be continuously or periodically updated with new training data. In one embodiment, the subject methylation data on which the final prognosis score is based may be the same methylation data utilized for other cancer classification determinations. More specifically, both prognosis and cancer classification determinations may be performed from the same blood sample, thereby avoiding the subject undergoing additional liquid sample collections, such as additional blood draws.

[0096] Aspects of the embodiments may employ filtering of cfDNA fragments based on methylation status. For example, embodiments of the present disclosure may identify regions in the genome where the presence of certain aberrantly methylated fragments can separate cancer subjects with a good prognosis from cancer subjects with a poor prognosis, or separate cancer subjects from non-cancer subjects. In one embodiment, the DNA methylation data utilized to train a machine learning model may include only aberrantly methylated cfDNA fragments. More specifically, DNA fragments may be included in subsequent training and scoring if their methylation patterns are unlikely to be observed among cfDNA fragments found in healthy, non-cancer study participants.

[0097] In addition to the above, the training data may be based on substantially the same methylation feature set as that used in the Galleri classifier (i.e., the feature set may represent aberrantly methylated CpG sites identified for cancer detection or for distinguishing between different cancer signal origins). For example, the P_prognostic_CSO labeling method and the P_prognostic_Galleri_fragment labeling method may both utilize methylation patterns typically used to distinguish between different cancer types. More specifically, each of the above methods utilizes the same set of methylation sites (i.e., CpG sites in the genome relied upon for cancer detection and / or cancer signal origin) that were used to train the Galleri classifier. Alternatively, the training data may be based on a different feature set corresponding to methylation sites that are deemed particularly informative for the subject prognosis. For example, the P_prognostic_fragment method may be a ground-up feature identification in which a methylation panel is interrogated looking for unique methylation sites that may be particularly associated with (i.e., more correlated with) the subject prognosis. There may or may not be overlap between traditional Galleri methylation patterns and methylation patterns identified as being particularly relevant to determining a subject's prognosis.

[0098] In step 2710, the set of training data may be annotated such that each of the training data in the set (i.e., each set of DNA methylation data associated with a training subject) is assigned a label indicating the outcome the disease had for the training subject. In one embodiment, possible labels may include subject death after X years, disease progression in the subject within X years, disease progression and subject death within X years, subject metastasis or recurrence after X years, and subject disease-free and alive after X years. For example, with reference to FIG. 28, multiple sample labels are provided for use in the set of training data. The training labels in set 2805 provide an indication of whether the training subject is alive or dead after a predetermined number of years and / or is cancer-free, has plateaued, or has progressed (e.g., the cancer continues to grow or begins to grow again even after surgery or treatment).

[0099]

[00110] Referring back to Figure 27, in step 2715, the annotated training data may be applied to a machine learning model. In one embodiment, the machine learning model may be nearly any type of supervised, unsupervised, or hybrid machine learning model selected by a user. In one embodiment, application of the annotated training data may optimize the pattern recognition capabilities of an algorithm associated with the machine learning model to learn which methylation patterns in cancer subjects are indicative of subject prognosis. For example, the machine learning model may learn that the presence of a group of aberrantly methylated CpG sites (Group A) is associated with a poor subject prognosis (e.g., disease progression and death within two years), while a different group of aberrantly methylated CpG sites (Group B) is associated with a more favorable subject prognosis (no disease progression and subject survival after two years).

[0100] Referring now to Figure 29, a method for generating a final prognostic score for a patient from a combination of the model-generated prognostic score and the ctDNA score is disclosed.

[0101] In step 2905, methylation data derived from a targeted methylation sequencing assay performed on a biological sample obtained from a test subject may be obtained. The biological sample may be a liquid biological sample, such as a blood sample. In step 2910, the methylation data derived from the targeted methylation sequencing assay may be applied to a trained machine learning model (e.g., the machine learning model discussed above with reference to FIG. 24). The methylation data may be processed by the trained machine learning model's algorithm, and in 2915, an output result in the form of a prognostic score may be received. In one embodiment, the prognostic score may have a value between 0 and 1, with 0 indicating a methylation pattern significantly similar to that of patients in a good prognosis training population and 1 indicating a methylation pattern significantly similar to that of patients in a poor prognosis training population. In step 2920, a ctDNA score may be identified that represents the proportion of cfDNA in the biological sample derived from tumor rather than non-cancerous tissue (i.e., the tumor fraction). In one embodiment, the tumor fraction may be estimated using one or more computational tools and / or methods known in the art. In one embodiment, the ctDNA score may be obtained by taking the log10 of the calculated tumor fraction value, i.e., taking log10(mVAF). In step 2925, a final patient prognosis score may be generated by combining the machine learning-generated prognosis score received in step 2915 with the ctDNA score identified in step 2920. In one embodiment, the ctDNA score and the machine learning-generated prognosis score may be combined in one or more different ways. For example, one non-limiting way of combining the ctDNA score and the machine learning-generated prognosis score may be by multiplying the two metrics together. Other ways of combining the ctDNA score and the machine learning-generated prognosis score include by adding the two metrics together or by combining the two metrics through almost any other mathematical means.

[0102] In one embodiment, the final patient prognosis score may be used as the basis for one or more downstream actions and / or decisions. For example, the range of values ​​that the final prognosis score falls within may provide an indication of how good or poor the prognosis is. For example, a good prognosis indication may be provided if the final patient prognosis score falls within a first range of values ​​(e.g., 0 to 0.3), a fair prognosis indication may be provided if the final patient prognosis score falls within a second range of values ​​(e.g., 0.3 to 0.6), and a poor prognosis indication may be provided if the final patient prognosis score falls within a third range of values ​​(e.g., 0.6 to 1). These ranges are merely exemplary, and different groupings or different numbers of groups may be utilized according to clinically relevant ranges. In other examples, a threshold may be established and the final patient prognosis score may be compared to the threshold. A value below the threshold may indicate a good prognosis score, while a value above the threshold may indicate a poor prognosis indication.

[0103] Additionally or alternatively, various recommendations (e.g., recommendations for the type of clinical study to participate in, recommendations for disease management processes, recommendations for follow-up regimens during / after treatment, etc.) may be dynamically provided to the subject, clinician, and / or clinical trial supervisor based on the range of values ​​that the final patient prognosis score falls into or whether the value is above or below a threshold. For example, if the final patient prognosis score falls within the third value range described above or exceeds a threshold, a recommendation may be made suggesting that the subject participate in a more aggressive clinical trial in an attempt to improve the poor predicted prognosis. In other embodiments, a more aggressive treatment protocol may be recommended or a more frequent follow-up schedule during / after treatment may be recommended (e.g., more frequent tests, scans, biopsy collections (e.g., blood draws), or other follow-up regimens may be recommended). A recommendation may also be made suggesting a change to the subject's post-treatment biopsy collection (e.g., blood draw) schedule, such that blood draws may be performed more frequently. In yet another aspect, if a patient identified as having a poorer prognosis participates in a clinical trial, changes may be made to the clinical trial participation or resulting data analysis so that the prognosis is considered as a relevant variable to potentially control for.

[0104] As another example, if the final patient prognosis score falls within the first value range described above or below a threshold, a less aggressive treatment protocol may be recommended or a less frequent follow-up regimen during / after treatment may be recommended. For example, more frequent tests, scans, biopsy collection (e.g., blood draws), or other follow-up regimens may be recommended. Recommendations may be made suggesting changes to the subject's post-treatment biopsy collection (e.g., blood draw) schedule to draw blood less frequently. In yet other aspects, if a patient identified as having a better prognosis participates in a clinical trial, changes may be made to the clinical trial participation or resulting data analysis so that the prognosis is considered as a relevant variable to potentially control for.

[0105] In addition to the above, the final patient prognosis score can be utilized in a variety of ways in clinical trial settings. More specifically, knowledge of a patient's prognosis can be valuable in the context of clinical trial design for patient stratification and inclusion criteria to increase event rates and event power. For example, given a particular clinical trial, the patient population can be divided into distinct subgroups based on the prognostic indication derived from the final patient prognosis score.

[0106] 30 , a graph 3000 illustrating multiple patient prognostic predictions for a cancer type (i.e., squamous cell carcinoma of the lung) is presented. In graph 3000, multiple sets of prognostic predictions are presented with different levels of specificity, with each set including prognostic predictions generated using a different technique. In this discussion, two particular prognostic prediction metrics (3005, 3010) within one set 3015 of prognostic predictions in graph 3000 may be examined. More specifically, the associated set 3015 may include an indication of a first prognostic prediction 3005 corresponding to only the tumor fraction value (i.e., mVAF) and an indication of another prognostic prediction 3010 (i.e., P_prognostic_Galleri_fragment) corresponding to a combination of a machine-generated prognostic score and a ctDNA score, as previously disclosed above. Inspection of graph 3000 shows that the prognostic predictions 3010 associated with P_prognostic_Galleri_fragment generally have greater specificity and sensitivity than the tumor fraction values ​​3005, an indication that the former provide a better patient outcome than the mVAF values ​​alone. An additional observation that can be made from graph 3000 is that the prognostic predictions 3010 associated with P_prognostic_Galleri_fragment have similar specificity to the P_prognostic_CSO metric and P_prognostic_fragment in each set, but outperform both in sensitivity.

[0107] In some embodiments, the disclosed methods, systems, and / or classifiers can be used to detect the presence (or absence) of cancer, monitor the progression or recurrence of cancer, monitor the response or effectiveness of treatment, monitor the transdifferentiation of cancer, determine or monitor the presence of minimal residual disease (MRD), quantify disease burden, generate a patient prognosis score, or any combination thereof. In some embodiments, the disclosed systems and / or classifiers can be used to identify the histological type or molecular subtype of cancer. For example, the disclosed systems and / or classifiers can be used to identify cancer as any of the following cancer types: epithelial, neural, lymphoid, myeloid, plasma, mesenchymal, melanocyte, neuroendocrine, and germ cell, and, for all epithelial cancers, further distinguish between histological types: adenocarcinoma (ADC), HPV-associated squamous cell carcinoma (HPV), non-HPV-associated squamous cell (SCC), adenocarcinoma of Müllerian origin (Müllerian), and transitional cell carcinoma. In some embodiments, subtypes of known cancer types, such as type I ovarian epithelial cancer and type II high-grade serous ovarian cancer, or different molecular subtypes of SCLC that can be distinguished by key transcription factors, can be identified. In some embodiments, a test report can be generated to provide the patient with test results including, for example, a probability score that the patient has a disease state (e.g., cancer), the type of disease (e.g., cancer type), and / or histologic type or molecular subtype (i.e., when generating the report, a CSO label can be generated for certain people based on cofactors such as gender, smoking status, etc.). In some embodiments, the disclosed methods and / or classifiers are used to identify cancer types, histologic types, or molecular subtypes in patients suspected of having been diagnosed with cancer. In some embodiments, this identification of cancer types, histologic types, or molecular subtypes can use the same blood draws, plasma samples, and sequencing results used to detect cancer in subjects participating in cancer early detection or screening programs. According to aspects of the present disclosure, the disclosed methods and systems can be trained to detect or classify multiple indications of cancer.For example, the disclosed methods, systems, and classifiers can be used to detect the presence of one or more, two or more, three or more, five or more, or ten or more different cancer types, histological types, or molecular subtypes. In some embodiments, the cancer is one or more of adenocarcinoma, HPV-associated squamous cell carcinoma, non-HPV-associated squamous cell, type I adenocarcinoma of Müllerian origin, type II carcinoma of Müllerian origin, transitional cell carcinoma, or neuroendocrine carcinoma such as small cell carcinoma with molecular subtypes distinguished by different transcription factors.

[0108] Datasets of sequence data obtained from a cancer patient over any desired set of time points can be generated and analyzed according to the methods of the present disclosure to monitor the patient's cancer status. In some embodiments, a first time point is before cancer treatment (e.g., before resection surgery or intervention) and a second time point is after cancer treatment (e.g., after resection surgery or intervention), and the method is utilized to monitor the effectiveness of the treatment. In other embodiments, the first and second time points are both before cancer treatment (e.g., before resection surgery or intervention). In yet other embodiments, the first and second time points are both after cancer treatment (e.g., before resection surgery or intervention), and the method is utilized to monitor the effectiveness of the treatment or loss of effectiveness of the treatment. In yet another embodiment, a dataset of sequence data obtained from a cancer patient at a first and second time point may be generated and analyzed to, for example, monitor cancer progression, determine whether the cancer has gone into remission (e.g., after treatment), monitor or detect residual disease or disease recurrence, monitor the efficacy of a treatment (e.g., therapy), or monitor whether the cancer has changed histological type or molecular subtype while developing resistance to cancer treatment.

[0109] In some embodiments, the first and second time points are separated by an amount of time ranging from about 15 minutes to about 30 years, e.g., about 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9 hours, 10 hours, 11 hours, 12 hours, 13 hours, 14 hours, 15 hours, 16 hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hours, or about 24 hours, e.g., about 1 day, 2 days, 3 days, 4 days, 5 days, 10 days, 15 days, 20 days, 25 days, or about 30 days, or e.g., about 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, or 12 months, or e.g., about 1 year, 1.5 years, 2 years, 2.5 years, 3 years, 3.5 years, 4 years, 4.5 years, 5 years, 5.5 years, 6 years, 6.5 years, 7 years, 7.5 years, 8 years, 8.5 years, 9 years, 9.5 years, 10 years, 10. 5 years, 11 years, 11.5 years, 12 years, 12.5 years, 13 years, 13.5 years, 14 years, 14.5 years, 15 years, 15.5 years, 16 years, 16.5 years, 17 years, 17.5 20 years, 21 years, 21.5 years, 22 years, 22.5 years, 23 years, 23.5 years, 24 years, 24.5 years, 25 years, 25.5 years, 26 years, 26.5 years, 27 years, 27.5 years, 28 years, 28.5 years, 29 years, 29.5 years, or about 30 years, etc. In other embodiments, datasets of sequence data obtained from patients can be generated at least once every three months, at least once every six months, at least once a year, at least once every two years, at least once every three years, at least once every four years, or at least once every five years.

[0110] In yet another embodiment, information obtained from any of the methods described herein can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, assessment of treatment efficacy, etc.). For example, in one embodiment, a physician can prescribe an appropriate treatment (e.g., resective surgery, radiation therapy, chemotherapy, and / or immunotherapy) based on the use of the dataset in the classification process. In some embodiments, information such as a classification based on the dataset can be provided to a physician or subject as a readout. In some embodiments, a classification based on the dataset can indicate the efficacy of a cancer treatment.

[0111] In some embodiments, the treatment is one or more cancer therapy agents selected from the group including chemotherapy agents, targeted cancer therapy agents, differentiation therapy agents, hormonal therapy agents, and immunotherapy agents. For example, the treatment can be one or more chemotherapy agents selected from the group including alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeletal disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapy agents selected from the group including signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteasome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapy agents, including retinoids such as tretinoin, alitretinoin, and bexarotene. In some embodiments, the therapy is one or more hormone therapy agents selected from the group including antiestrogens, aromatase inhibitors, progestins, estrogens, antiandrogens, and GnRH agonists. In one embodiment, the therapy is one or more immunotherapy agents selected from the group including monoclonal antibody therapy, e.g., rituximab (Rituxan) and alemtuzumab (Campus), non-specific immunotherapy and adjuvants, e.g., BCG, interleukin-2 (IL-2), and interferon alpha, immunomodulatory agents, e.g., thalidomide and lenalidomide (Revlimid). Appropriate cancer therapy agents can be selected based on characteristics such as cancer type, histological type, molecular subtype, cancer stage, previous cancer treatments or therapeutic agents, and other characteristics of the cancer.

[0112] In addition to standard desktops or servers, any computer system capable of the required storage and processing needs is fully within the scope of this disclosure that may be suitable for practicing embodiments of the present disclosure. This may include tablet devices, smartphones, PIN pad devices, and any other computing device, whether mobile or even distributed across a network (i.e., cloud-based).

[0113] Unless otherwise indicated, and as will be apparent from the discussion that follows, throughout this specification discussion utilizing terms such as "processing," "calculating," "calculating," "determining," "analyzing," and the like will be understood to refer to operations and / or processes of a computer or computing system or similar electronic computing device that manipulates and / or transforms data represented as physical quantities, such as electronic quantities, into other data similarly represented as physical quantities.

[0114] Similarly, the term "processor" may refer to any device or portion of a device that processes electronic data, e.g., from registers and / or memory, and converts the electronic data into other electronic data that may be stored, e.g., in registers and / or memory. A "computer," "computing machine," "computing platform," "computing device," or "server" may include one or more processors.

[0115] According to various embodiments of the present disclosure, the methods described herein may be implemented by a software program executable by a computer system. Furthermore, in one exemplary, non-limiting embodiment, the implementation may include distributed processing, component / object distributed processing, and parallel payment. Alternatively, a virtual computer system process may be constructed to implement one or more of the methods or functions as described herein.

[0116] Although this specification describes components and functions that may be implemented in particular embodiments with reference to particular standards and protocols, the present disclosure is not limited to such standards and protocols. For example, standards for Internet and other packet-switched network transmissions (e.g., TCP / IP, UDP / IP, HTML, HTTP, etc.) represent examples of the state of the art. Such standards are periodically superseded by faster or more efficient equivalents having essentially the same functionality. Accordingly, replaced standards and protocols having the same or similar functionality as those disclosed herein are considered equivalents thereof.

[0117] It will be understood that the steps of the discussed methods are, in one embodiment, performed by a suitable processor (or processors) of a processing (i.e., computer) system executing instructions (computer-readable code) stored in storage. It will also be understood that the disclosed embodiments are not limited to any particular implementation or programming technique, and that the disclosed embodiments may be implemented using any suitable technique for implementing the functions described herein. The disclosed embodiments are not limited to any particular programming language or operating system.

[0118] In the foregoing description of exemplary embodiments, it should be understood that various features of the embodiments are sometimes grouped together in a single embodiment, figure, or description thereof to streamline the disclosure and facilitate understanding of one or more of the various inventive aspects. This method of disclosure, however, should not be interpreted as reflecting an intention that the claimed embodiments require more features than are expressly recited in each claim. Rather, the following claims reflect that inventive aspects lie in less than all features of a single foregoing disclosed embodiment. Thus, the claims following the Detailed Description are hereby expressly incorporated into this Detailed Description, with each claim standing on its own as a separate embodiment.

[0119] Furthermore, as will be understood by those skilled in the art, while some embodiments described herein include some features but not other features included in other embodiments, combinations of features from different embodiments are intended to be within the scope of the present disclosure and to form different embodiments. For example, in the following claims, any of the claimed embodiments may be used in any combination.

[0120] Furthermore, some of the embodiments are described herein as methods or combinations of elements of methods that can be implemented by a processor of a computer system or other means for performing a function. Thus, a processor with the necessary instructions for carrying out such a method or element of a method forms a means for carrying out the method or element of a method. Furthermore, a described element of an apparatus embodiment is an example of a means for carrying out a function, where the element is performed by the element to perform the function.

[0121] In the description provided herein, numerous specific details are set forth. However, it will be understood that embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been described in detail so as not to obscure an understanding of this description.

[0122] Similarly, it should be noted that the term coupled, when used in the claims, should not be construed as being limited to only direct connections. The terms "coupled" and "connected," along with their derivatives, may be used. It should be understood that these terms are not intended as synonyms for each other. Thus, the scope of the phrase device A coupled to device B should not be limited to devices or systems in which the output of device A is directly connected to the input of device B. It means that there is a path between the output of A and the input of B, which may be a path that includes other devices or means. "Coupled" can mean that two or more elements are in either direct physical or electronic contact, or that two or more elements are not in direct contact with each other, but yet still cooperate or interact with each other.

[0123] Thus, while what are considered to be preferred embodiments of the present disclosure have been described, those skilled in the art will recognize that other and further modifications may be made without departing from the spirit of the present disclosure, and that all such changes and modifications are intended to be within the scope of the present disclosure. For example, any formulas above merely represent procedures that may be used. Functions may be added or deleted from block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted from methods described within the scope of the present disclosure.

[0124] The subject matter disclosed above should be considered illustrative and not limiting, and the appended claims are intended to encompass all such modifications, improvements, and other embodiments that fall within the true spirit and scope of the present disclosure. Accordingly, to the maximum extent permitted by law, the scope of the present disclosure should be determined by the broadest possible interpretation of the following claims and their equivalents, and not be limited or restricted by the above detailed description. While various embodiments of the present disclosure have been described, it will be apparent to those skilled in the art that many more embodiments and implementations are possible within the scope of the present disclosure. Accordingly, the present disclosure should not be limited except in light of the appended claims and their equivalents.

Claims

1. 1. A method of detecting a subtype of a disease state using a system comprising: receiving, at an input component of the system, a set of sequence reads associated with a nucleic acid sample; generating methylation data through analysis of the set of sequence reads using a processor of the system; analyzing the methylation data using the processor to identify the subtype of the disease state; A method comprising:

2. applying, using the processor, information associated with the identified subtypes as training inputs to a disease state classifier; utilizing the disease state classifier trained with the information associated with the identified subtypes on one or more subsequent sets of nucleic acid samples; The method of claim 1 further comprising:

3. The method of claim 1 , wherein the subtype of the disease state comprises a source of cancer cells.

4. 10. The method of claim 1, wherein the subtype of the disease state comprises a histological subtype.

5. The method of claim 1 , wherein the subtype of the disease state comprises a molecular subtype.

6. 6. The method of claim 5, wherein the molecular subtypes are previously defined based on protein expression identified using cancer tissue samples.

7. 6. The method of claim 5, wherein the molecular subtypes are previously defined based on gene expression identified using cancer tissue samples.

8. 6. The method of claim 5, wherein the molecular subtypes are previously defined based on genomic aberrations identified using cancer tissue samples.

9. The method of claim 2 , wherein the molecular subtype information is defined and trained based on the results of different treatments.

10. The method of claim 2 , wherein the molecular subtype information is defined and trained based on a prognosis of the subject's cancer progression.

11. 3. The method of claim 2, wherein the molecular subtype information is defined and trained based on the subject's prognosis of cancer recurrence.

12. 1. A method for training a machine learning model to detect the occurrence of resistance mechanisms in cancer, comprising: obtaining a set of training data from a source, the training data comprising methylation data derived from a targeted methylation sequencing assay; subsequent to said obtaining, annotating the set of training data by assigning a histological label to each of the training data in the set; applying the annotated set of training data to the machine learning model; optimizing the pattern recognition capabilities of an algorithm associated with the machine learning model based on the applying; and A method comprising:

13. The method of claim 12 , wherein the source is a plurality of training subjects.

14. 13. The method of claim 12, wherein the histological label is associated with one of adenocarcinoma and small cell neuroendocrine carcinoma.

15. Optimizing the pattern recognition capabilities of the algorithm includes: The machine learning model includes: A) determining whether minimal residual disease is present in a test set of methylation data; B) in response to determining that the minimal residual disease is present in the test set, determining whether at least a portion of the cancer cells within the minimal residual disease have transformed from a first cancer type to a second cancer type as a result of the development of the resistance mechanism; 13. The method of claim 12, comprising causing:

16. Optimizing the pattern recognition capabilities of the algorithm includes: The machine learning model includes: C) in response to determining that the minimal residual disease is present and that at least some of the cancer cells within the minimal residual disease have transformed from the first cancer type to the second cancer type, suggesting a treatment recommendation directed toward the second cancer type.

17. 1. A method for detecting the emergence of resistance mechanisms in cancer during treatment using a computer system and associated trained machine learning model, comprising: receiving methylation data derived from a targeted methylation sequencing assay from a biological sample associated with the test subject; subsequent to said receiving, applying said methylation data to said trained machine learning model; Following the applying, receiving an output from the trained learning model, the output comprising: A) a first indication of whether minimal residual disease is present in the test subject following administration of the treatment for the cancer; and B) in response to said first indication providing a finding that said minimal residual disease is present in said test subject, a second indication that at least a portion of cancer cells within said minimal residual disease have transformed from a first cancer type to a second cancer type as a result of said development of said resistance mechanism. receiving, A method comprising:

18. 18. The method of claim 17, wherein the resistance mechanism is transdifferentiation.

19. 18. The method of claim 17, wherein the first cancer type is adenocarcinoma and the second cancer type is small cell neuroendocrine carcinoma.

20. The output is C) in response to the first instruction providing the finding that minimal residual disease is present in the test subject and the second instruction providing a separate finding that the at least some of the cancer cells within the minimal residual disease have transformed from the first cancer type to the second carcinoma, further comprising a third instruction for a treatment recommendation directed toward the second cancer type.

21. 1. A method of training a machine learning model to generate a patient prognosis score, comprising: obtaining a set of training data from a source, the training data comprising methylation data derived from a targeted methylation sequencing assay; subsequent to said obtaining, annotating the set of training data by assigning a known patient outcome to each of the training data in the set; applying the set of annotated training data to the machine learning model; optimizing the patient outcome prediction ability of the machine learning model and associated algorithm based on the applying; A method comprising:

22. 1. A method for determining a final prognostic score for a test subject, comprising: receiving methylation data derived from a targeted methylation sequencing assay from a biological sample associated with the test subject; subsequent to said receiving, applying said methylation data to said trained machine learning model; subsequent to the applying, receiving an output from the trained learning model, the output including a first prognostic score for the test subject; and identifying a ctDNA score based on the biological sample; combining the first prognostic score and the ctDNA score together; generating the final prognostic score for the test subject based on said combining; A method comprising:

23. Identifying the ctDNA score comprises: determining a tumor incidence value associated with said biological sample; calculating the log10 of said tumor fraction value; 23. The method of claim 22, comprising:

24. 24. The method of claim 23, wherein said combining comprises multiplying said loglO of said tumor fraction value by said first prognostic score.

25. 23. The method of claim 22, wherein the final prognostic score is a value between 0 and 1.